Seu agent dorme? OpenAI lança always-on agents. Custa quanto?
OpenAI lança Dots (always-on agents). Seu agent dorme? Always-on = customer expectation agora. Como competir (e custo).
Equipe OpenClaw · Time de Engenharia & Produto
A Equipe OpenClaw é formada por engenheiros, designers e especialistas em IA dedicados a construir a melhor plataforma de agentes conversacionais para negócios brasileiros. Combinamos expertise…
Seu agent dorme? OpenAI lança always-on agents. Custa quanto?
Você é founder de SaaS.
Seu SaaS tem agent de IA (WhatsApp, atendimento ao cliente, automação de vendas).
Current agent architecture:
Your agent today (request-response): │ ├─ How it works: │ ├─ Customer sends message: "Oi, qual o status do meu pedido?" │ ├─ Your server wakes up agent (spins up container) │ ├─ Agent processes request (LLM API call) │ ├─ Agent sends response ("Your order is...") │ ├─ Server goes to sleep (shuts down container) │ └─ Repeat for next customer │ ├─ Issues with this approach: │ ├─ Latency: 2-3 seconds per response (server startup time) │ ├─ Cost: Cheap (pay only when customer messages) │ ├─ Capabilities: Limited (agent has no context between messages) │ ├─ Proactivity: ZERO (can't initiate conversations) │ ├─ Experience: Feels slow (customer notices lag) │ └─ Status: Works, but feels dated │ └─ Your assumption: "This is how all agents work. It's fine."
Then you read (October 2026): Headline: "OpenAI launches Dots: Always-on agents (available 24/7)" │ What changed: ├─ Old agent (request-response): │ ├─ Active: Only when customer sends message │ ├─ Availability: Reactive (waits for input) │ ├─ Response time: 2-3 seconds (server startup) │ ├─ Context: Limited (resets between messages) │ ├─ Proactivity: ZERO (never reaches out first) │ └─ Customer perception: "Agent is slow. Feels like waiting on hold." │ ├─ New agent (always-on, OpenAI Dots): │ ├─ Active: Always running (24/7, never sleeps) │ ├─ Availability: Proactive (can reach out to customers) │ ├─ Response time: Instant (<100ms, no startup) │ ├─ Context: Rich (remembers everything across conversations) │ ├─ Proactivity: YES (can send reminders, follow-ups, alerts) │ └─ Customer perception: "This agent is always there. Amazing." │ ├─ Implication: │ ├─ Market shift: From request-response → Always-on │ ├─ Expectation: Customers now expect 24/7 availability │ ├─ Competition: Your old agent architecture is now dated │ ├─ Pressure: You need to upgrade (or lose to competitors) │ ├─ Cost: Always-on is more expensive (always running) │ └─ Question: Can you afford to upgrade? Should you? │ └─ Your realization: ├─ My agent is reactive (customers must reach out) ├─ OpenAI agents are proactive (reach out to customers) ├─ Gap is now visible (customer experience suffers) ├─ I need to decide: Upgrade to always-on? Or stay reactive? ├─ Cost matters (budget is limited) └─ Timeline: Competitors will upgrade soon (window is closing)
The Always-On Agent Paradigm Shift: Request-Response is Dead
Always-on agents fundamentally change how AI assistants work (and how much they cost).
Architecture comparison: Request-Response vs. Always-On
Request-Response Agent (traditional): │ ├─ Flow: │ ├─ 00:00 - Agent sleeps (not running) │ ├─ 09:15 - Customer sends message ("Hi, help me") │ ├─ 09:15:500ms - Server wakes up agent (starts container) │ ├─ 09:15:1s - Agent loads model weights (LLM into memory) │ ├─ 09:15:1.5s - Agent processes request (generates response) │ ├─ 09:15:2s - Agent sends response ("Hello, how can I help?") │ ├─ 09:15:2.1s - Server shuts down agent (stops container) │ ├─ 09:15:2.1s - 09:20 - Agent sleeps (not running) │ ├─ 09:20 - Customer sends follow-up ("What about pricing?") │ ├─ 09:20:500ms - Server wakes up agent AGAIN (repeats cycle) │ └─ Total latency per message: 2-3 seconds (server startup) │ ├─ Characteristics: │ ├─ Stateless: Agent has no memory between messages (context reset) │ ├─ Reactive: Only responds to customer input (never reaches out) │ ├─ Cheap: Pay only when processing (minutes = dollars) │ ├─ Limited: Can't do background tasks (scheduling, reminders) │ └─ Predictable: Costs scale linearly with usage │ ├─ Example use cases: │ ├─ FAQs (answer questions) │ ├─ Basic support (triage tickets) │ ├─ Transaction lookup (order status) │ └─ One-off help (not ongoing relationships) │ └─ Customer experience: ├─ "Agent is responsive" ✓ (if message comes quickly) ├─ "Agent remembers me" ✗ (context resets) ├─ "Agent is always there" ✗ (only when I message) ├─ "Agent proactively helps" ✗ (no outbound reach) └─ Overall: Acceptable, but feels reactive
Always-On Agent (OpenAI Dots model): │ ├─ Flow: │ ├─ 00:00 - Agent starts (initialized, running) │ ├─ 00:00 - 09:15 - Agent waits (but is active in memory) │ ├─ 09:15 - Customer sends message ("Hi, help me") │ ├─ 09:15:050ms - Agent processes (no startup latency!) │ ├─ 09:15:100ms - Agent sends response (instant) │ ├─ 09:15:100ms - Agent stays active (ready for next message) │ ├─ 09:20 - Customer sends follow-up ("What about pricing?") │ ├─ 09:20:050ms - Agent processes (same latency, no startup) │ ├─ 09:20:100ms - Agent sends response (instant again) │ ├─ 09:25 - Agent is idle (but still running, remembering context) │ ├─ 14:00 - Agent proactively sends message ("Hey, your order shipped!") │ │ └─ (Agent can reach out without customer trigger) │ └─ Total latency per message: 50-100ms (no startup overhead) │ ├─ Characteristics: │ ├─ Stateful: Agent has persistent memory (context across all interactions) │ ├─ Proactive: Can reach out to customers (send reminders, alerts, offers) │ ├─ Expensive: Always running (always consuming resources) │ ├─ Powerful: Can do background tasks (scheduling, analyzing, learning) │ └─ Unpredictable: Costs scale with features, not just usage │ ├─ Example use cases: │ ├─ Sales follow-ups (proactive outreach) │ ├─ Relationship management (remember customer preferences) │ ├─ Predictive support (reach out before customer asks) │ ├─ Ongoing tasks (scheduling, reminders, automation) │ └─ Complex workflows (multi-step processes over time) │ └─ Customer experience: ├─ "Agent is responsive" ✓ (instant, <100ms) ├─ "Agent remembers me" ✓ (persistent memory) ├─ "Agent is always there" ✓ (always running) ├─ "Agent proactively helps" ✓ (sends reminders, alerts) └─ Overall: Amazing, feels like personal assistant
Key differences summary: ├─ Latency: Request-response (2-3s) vs Always-on (50-100ms) = 30x faster ├─ Memory: Request-response (resets) vs Always-on (persistent) = context preserved ├─ Proactivity: Request-response (zero) vs Always-on (unlimited) = new capabilities ├─ Cost: Request-response (pay-per-use) vs Always-on (always-on) = 10-100x more expensive ├─ Complexity: Request-response (simple) vs Always-on (complex) = more engineering └─ Experience: Request-response (good) vs Always-on (exceptional) = clear winner for UX
Market Shift: Always-On Becomes The Expectation
OpenAI launching Dots signals the market is moving (fast) toward always-on agents.
Timeline: How always-on becomes standard
Now (October 2026): ├─ OpenAI: Launches Dots (always-on agents available) ├─ Early adopters: Start building always-on agents ├─ Competitive advantage: HUGE (most agents still request-response) ├─ Customer expectation: Still mixed (some expect 24/7, others don't mind lag) ├─ Market signal: Always-on is becoming table-stakes └─ Window: 3-6 months (early adopter advantage)
6 months (April 2027): ├─ Competition: Major SaaS companies launch always-on agents ├─ Adoption: 20-30% of AI agents are now always-on ├─ Expectation: High-touch SaaS (CRM, sales, support) REQUIRES always-on ├─ Advantage: FADING (competitors also launched) ├─ Pressure: Request-response agents start looking dated └─ Decision: Upgrade now or be perceived as behind
12 months (October 2027): ├─ Standard: Always-on is table-stakes for AI agents ├─ Legacy agents: Request-response agents are "budget" option ├─ Perception: "My SaaS uses old request-response" = cheap/outdated ├─ Adoption: 50%+ of agents are always-on ├─ Expectation: Customers expect instant <100ms response time + proactive features └─ Cost: Everyone paying more for always-on (unavoidable)
24 months (October 2028): ├─ Default: New agents are always-on (request-response is "legacy") ├─ Perception: Request-response agents are like Web 1.0 (feels ancient) ├─ Market: Only budget SaaS uses request-response ├─ Adoption: 80%+ of agents are always-on ├─ Competitive: Always-on is no longer differentiator (table-stakes) └─ New differentiator: AI reasoning, personalization, autonomy
Implication for your SaaS: ├─ If you upgrade to always-on NOW: │ ├─ Advantage: 6-12 month window (before competitors) │ ├─ Positioning: "State-of-the-art agent architecture" │ ├─ Customer perception: "This SaaS is modern" │ ├─ Conversion impact: +15-30% (better experience) │ └─ Cost: Premium (always-on is expensive) │ ├─ If you wait 6-12 months: │ ├─ Advantage: GONE (competitors already there) │ ├─ Positioning: "We also have always-on now" (me-too) │ ├─ Customer perception: "Finally, you caught up" │ ├─ Conversion impact: No change (everyone has it) │ └─ Cost: Premium (still expensive, but no differentiator) │ ├─ If you never upgrade: │ ├─ Advantage: NEGATIVE (falling behind) │ ├─ Positioning: "Our agent uses legacy architecture" │ ├─ Customer perception: "This SaaS is cheap/outdated" │ ├─ Conversion impact: -20-30% (vs competitors) │ ├─ Churn risk: HIGH (customers leave for modern alternatives) │ └─ Outcome: Forced to upgrade eventually (but at cost) │ └─ Recommendation: Strategic decision required NOW
Cost Breakdown: What Always-On Agents Actually Cost
Always-on is 10-100x more expensive than request-response (but ROI can justify it).
Cost comparison: Request-Response vs. Always-On
Scenario: Mid-sized SaaS with 10,000 customers
Request-Response Agent (current setup): │ ├─ Usage assumption: │ ├─ Active users per day: 5,000 │ ├─ Messages per user per day: 3 │ ├─ Total messages per day: 15,000 │ ├─ Avg processing time per message: 2 seconds │ ├─ Total compute time per day: 30,000 seconds (8.3 hours) │ └─ Assumption: Pay per second of compute │ ├─ Cost breakdown: │ ├─ LLM API (OpenAI GPT): R$ 0.50 per message = R$ 7.500/day │ ├─ Compute (server startup + processing): │ │ ├─ Container startup: 500ms per message = 7,500 seconds = R$ 1.500/day │ │ ├─ Model inference: 1.5s per message = 22,500 seconds = R$ 4.500/day │ │ └─ Subtotal: R$ 6.000/day │ │ │ ├─ Storage (conversation history): R$ 500/day │ ├─ Monitoring + logging: R$ 300/day │ ├─ Bandwidth: R$ 200/day │ └─ Total per day: R$ 14.500 │ ├─ Monthly cost: │ ├─ 14.500 × 30 days = R$ 435.000/month │ ├─ Per customer: R$ 435.000 / 10.000 = R$ 43.50/customer/month │ └─ As % of revenue (R$ 100/customer): 43.5% (if price is R$ 100) │ └─ Status: Cheap, but slow (2-3 second latency)
Always-On Agent (OpenAI Dots model): │ ├─ Architecture assumption: │ ├─ Agent instance: Always running (never stops) │ ├─ Deployment: Per customer or shared (depends on design) │ ├─ Memory: Persistent (context stored in vector DB) │ ├─ Background tasks: Continuous (monitoring, scheduling, proactivity) │ └─ Assumption: Pay per instance + storage + compute │ ├─ Cost breakdown (scenario 1: Per-customer instance): │ ├─ Agent instance (always running): │ │ ├─ GPU/compute per instance: R$ 50/day (keeping model in memory) │ │ ├─ Number of instances: 10.000 (one per customer) │ │ ├─ Total: R$ 500.000/day (!!) │ │ └─ Status: PROHIBITIVELY EXPENSIVE │ │ │ ├─ Cost breakdown (scenario 2: Shared instances, load-balanced): │ ├─ Agent instances (shared, load-balanced): │ │ ├─ Replicas needed: 50 (to handle 5,000 concurrent users) │ │ ├─ Compute per instance: R$ 50/day │ │ ├─ Total: R$ 2.500/day (way better) │ │ │ │ │ ├─ Vector DB (persistent memory): │ │ │ ├─ Storage: 10.000 customers × 1GB context = 10TB │ │ │ ├─ Cost: R$ 10.000/day (or R$ 300K/month for Pinecone/Weaviate) │ │ │ └─ Operations: R$ 2.000/day │ │ │ │ │ ├─ LLM API (same as before): R$ 7.500/day │ │ ├─ Inference on edge (optimized): R$ 2.000/day (vs R$ 4.500 before) │ │ ├─ Background tasks (proactivity): R$ 3.000/day │ │ ├─ Monitoring + logging (increased): R$ 2.000/day │ │ ├─ Bandwidth (higher): R$ 1.000/day │ │ └─ Total per day: R$ 28.000 │ │ │ ├─ Monthly cost: │ │ ├─ 28.000 × 30 days = R$ 840.000/month │ │ ├─ Per customer: R$ 840.000 / 10.000 = R$ 84/customer/month │ │ └─ As % of revenue (R$ 100/customer): 84% (if price is R$ 100) │ │ │ └─ Status: Expensive (2x request-response), but AMAZING experience │ └─ Cost multiplier: Always-on is ~2x more expensive (scenario 2) └─ But: Scenario 1 (per-customer) is 100x more expensive (not viable)
ROI Analysis: Is always-on worth it? │ ├─ Additional cost per month: R$ 840K - R$ 435K = R$ 405K/month ├─ Additional cost per customer: R$ 84 - R$ 43.50 = R$ 40.50/month │ ├─ Expected benefits of always-on: │ ├─ Conversion improvement: +15-30% (better experience) │ ├─ Churn reduction: -10-20% (customers stay longer) │ ├─ NPS improvement: +2-3 points (much better satisfaction) │ ├─ Upsell: +5-10% (proactive features drive upgrades) │ └─ Time-to-value: Faster (instant responses help adoption) │ ├─ Revenue impact (assuming 20% conversion boost): │ ├─ Current revenue: 10.000 customers × R$ 100 = R$ 1.000.000/month │ ├─ With always-on (+20% conversion): R$ 1.200.000/month │ ├─ Additional revenue: R$ 200.000/month │ ├─ Additional cost: R$ 405.000/month │ ├─ Net impact: -R$ 205.000/month (NEGATIVE?!) │ └─ Implication: At +20% boost, always-on doesn't pay off │ ├─ Break-even analysis (what boost is needed?): │ ├─ Additional cost: R$ 405K/month │ ├─ Revenue per customer: R$ 100/month │ ├─ Customers needed to cover cost: 405K / 100 = 4.050 new customers │ ├─ Current customer base: 10.000 │ ├─ Growth rate needed: 4.050 / 10.000 = +40.5% (break-even) │ ├─ Is +40% realistic? Only if: │ │ ├─ Conversion improves significantly (30%+) │ │ ├─ AND churn improves (15%+ reduction) │ │ ├─ AND upsell increases (10%+ boost) │ │ └─ Combined: ~40-50% growth possible │ │ │ └─ Verdict: Always-on pays off IF you can hit 40%+ growth from it │ ├─ Strategic considerations: │ ├─ Timing: Early adopter advantage (3-6 month window) │ ├─ Positioning: "State-of-the-art" vs "budget option" │ ├─ Price increase: Can you charge more for always-on? │ │ ├─ Current price: R$ 100/month │ │ ├─ Premium tier with always-on: R$ 150/month │ │ ├─ Premium customers: 30% of base = 3.000 customers │ │ ├─ Additional revenue: 3.000 × R$ 50 = R$ 150K/month │ │ ├─ Cost covered? Partial (150K of 405K) │ │ └─ Remaining gap: R$ 255K (need 40% of base to upgrade) │ │ │ └─ Recommendation: Only viable if you can 1) Price higher, 2) Grow 40%+, or 3) Reduce other costs │ └─ Bottom line: ├─ Always-on is NOT automatically profitable ├─ But: Early-adopter advantage might justify it (premium positioning) ├─ Strategy: Launch "premium always-on" tier (higher price) ├─ This way: Only profitable customers get always-on (covers costs) └─ Timeline: Decide in next 4-6 weeks (before competitors launch)
Strategic Options: How to Compete With Always-On
You don't have to go full always-on (there are hybrid approaches).
Option 1: Full Always-On Agent
✓ Pros: ├─ Best customer experience (instant response, proactive) ├─ Early adopter advantage (if you launch first) ├─ Positioning: "State-of-the-art AI agent" ├─ Capabilities: Full suite (everything is possible) └─ Long-term: Table-stakes in 12 months
✗ Cons: ├─ Cost: 2-3x more expensive ├─ Complexity: Much harder to build + maintain ├─ Infrastructure: Need persistent DB + real-time systems ├─ Risk: Needs careful scaling (costs can spiral) └─ ROI: Only works if you can 1) Price higher or 2) Grow 40%+
$ Cost: R$ 840K+/month (depends on scale) Timeline: 4-8 weeks to launch (if building on OpenAI APIs) Recommendation: Only if you can Premium price + expect 40%+ growth
Option 2: Hybrid Agent (Always-On for Premium Tier)
✓ Pros: ├─ Cost: Only pay for premium customers (lower total cost) ├─ Positioning: Premium feature (justifies higher price) ├─ Profitability: Premium customers cover costs ├─ Low risk: Test with small customer segment first └─ Revenue: Premium tier increases ARPU
✗ Cons: ├─ Complexity: Two agent architectures (must maintain both) ├─ Perception: "Budget tier" uses old tech (might hurt positioning) ├─ Migration: Customers might hesitate to upgrade └─ Cannibalization: Some customers might downgrade to budget
$ Cost: R$ 200-300K+/month (only for premium tier) Timeline: 4-6 weeks to launch Recommendation: BEST OPTION (balanced cost + positioning)
Option 3: Stay Request-Response (For Now)
✓ Pros: ├─ Cost: Continue current spend (no increase) ├─ Simplicity: No architectural changes needed ├─ Stability: Works, proven, stable └─ Time: Focus on other features
✗ Cons: ├─ Competitive: Competitors will upgrade (you'll fall behind) ├─ Perception: "Budget option" once everyone upgrades ├─ Experience: 2-3 second latency (feels slow vs always-on) ├─ Churn risk: Customers leave for modern competitors └─ Timeline: Must upgrade in 12 months anyway (forced)
$ Cost: R$ 435K/month (no change) Timeline: 0 weeks (no changes) Recommendation: RISKY (delay = competitive disadvantage)
Recommendation: Strategic Path Forward
Best move: Launch hybrid (always-on premium tier) in next 6 weeks.
6-week launch plan: Always-on premium tier
Week 1-2: Planning + Design ├─ Define premium tier (what features = premium?) ├─ Design always-on architecture (persistent memory, proactivity) ├─ Set premium pricing (R$ 150-200/month vs R$ 100 base) ├─ Identify early customers (5-10 beta testers) └─ Timeline: 2 weeks
Week 3-4: Development ├─ Build always-on backend (persistent agent instance) ├─ Integrate vector DB (memory storage) ├─ Add proactive features (reminders, alerts, follow-ups) ├─ Test with beta customers └─ Timeline: 2 weeks
Week 5-6: Launch + Monitor ├─ Launch premium tier (limited availability) ├─ Announce: "New always-on agent tier" ├─ Marketing push (early adopter positioning) ├─ Monitor costs + performance ├─ Iterate based on feedback └─ Timeline: 2 weeks
Results (6 weeks out): ├─ Launch date: Mid-November 2026 ├─ Competitive advantage: 3-6 month window (before competitors) ├─ Premium adoption: Target 20-30% of base (2.000-3.000 customers) ├─ Premium revenue: R$ 300K-450K/month ├─ Additional costs: R$ 200-300K/month (only premium instances) ├─ Net impact: +R$ 50-150K/month (POSITIVE) ├─ Positioning: "State-of-the-art agent architecture" (premium tier only) └─ Long-term: Prepared when always-on becomes table-stakes (12 months out)
Next Steps: Always-On Agent Strategy Assessment
At OpenClaw, we help SaaS companies plan always-on agent upgrades (architecture + cost modeling + launch strategy):
- Always-on feasibility assessment (can your SaaS support this cost?)
- Architecture design (persistent agents, vector DB, real-time systems)
- Cost modeling (what will it actually cost at your scale?)
- Pricing strategy (how to charge for always-on premium tier?)
- Launch plan (6-8 week deployment timeline)
- Competitive positioning (how to market "always-on")
Get a free always-on readiness assessment: Schedule 30 minutes with our AI agent architect. We'll assess your current agent architecture (request-response baseline), forecast costs of always-on upgrade (realistic budget?), model revenue impact (+15-40% conversion possible?), design hybrid premium tier (profit-positive day 1?), create 6-week launch plan (launch before competitors?), and help you decide: Premium tier (recommended) or full always-on (risky) or stay request-response (falling behind)?
[Book your free always-on readiness assessment] → [Button: Schedule 30-Minute Call]
FAQ
Q: Always-on custa 2x mais? Mas meu budget é apertado. Vale a pena mesmo?
A: NÃO universal. Só vale se: (1) Você consegue Premium price (R$ 150-200 vs R$ 100 base), (2) Customers esperam 24/7 (se não esperam, eles não vão pagar extra), (3) Seu LTV permite (se customer lifetime value é baixo, custo de sempre-on não paga). Recomendação: Faça model ROI PRIMEIRO. Se sair negativo, fica com request-response (por enquanto).
Q: Meu agent é request-response. Se eu não upgradar, meu SaaS morre?
A: NÃO morre, mas sofre. Timeline: (1) Próximos 6 meses = Você tá OK (early adopters vão upgradar, maioria fica request-response), (2) 12 meses = Competitors têm always-on, clientes começam comparar, (3) 18+ meses = Always-on é expectation, você tá atrasado (perda de market share). Recomendação: Não precisa fazer NOW, mas prepare strategy para 6-12 meses (não deixe pra última hora).
Q: Dá pra fazer um "Always-On Light" (mais barato que full always-on)?
A: SIM! Opções: (1) Always-on só pra horários de pico (economiza 30-40% de custo), (2) Always-on em shared instances (divide custo entre múltiplos clientes, reduz 50-70%), (3) Always-on com feature-gating (proatividade limitada = menos processamento = mais barato). Recomendação: Explore opção 2 (shared instances) = melhor cost-benefit.
Publicado em 29 de setembro de 2026