Notícias
Notícias
5 min de leitura
3 de outubro de 2026

Google September 2026: Agentes multimodal. Texto-only = morto.

Google's September AI announcements: Agents now process text + images + video. Text-only agents obsolete. Rich media = new engagement standard.

Equipe OpenClaw

Equipe OpenClaw · Time de Engenharia & Produto

A Equipe OpenClaw é formada por engenheiros, designers e especialistas em IA dedicados a construir a melhor plataforma de agentes conversacionais para negócios brasileiros. Combinamos expertise…


Google September 2026: Agentes multimodal. Texto-only = morto.

Ontem Google publicou: September 2026 AI updates.

"Agents now process and generate text, images, and video. Multimodal is standard."

What this means: Your agent (WhatsApp, support, sales automation) can now handle rich media (not just text). Customers send images. Agent responds with images + video. Automation becomes engaging.

Why it matters: Text-only agents = boring engagement. Multimodal agents = professional, engaging, trusted.

Problem it reveals: Your agents probably output text only (boring, low engagement, customers ignore).

Você é founder.

Scenario: Your support agent (September 2025 vs September 2026)

Customer asks: "I want to return this item. Here's a photo of the problem."

September 2025 (text-only agent):

  • Customer sends: Photo of damaged item
  • Agent receives: Image
  • Agent responds: "Please describe the damage in text."
  • Customer frustrated: "I already sent the photo!"
  • Agent: "I can only read text. Describe what you see."
  • Customer: Gives up. Calls support instead (automation failed)
  • Time wasted: 15 minutes
  • Automation success: 0%

September 2026 (multimodal agent):

  • Customer sends: Photo of damaged item
  • Agent receives: Image (processes it)
  • Agent analyzes: "I see: Cracked screen, water damage, missing button"
  • Agent responds: "I see the damage. Processing return. Here's your RMA label (as image). Print and ship."
  • Agent also sends: Video showing return process (how to package, where to ship)
  • Customer: Gets immediate, visual guidance. Ships next day.
  • Time wasted: 2 minutes
  • Automation success: 100%

Difference: Text-only (failed automation) vs multimodal (successful automation).

Implication: Your agents need multimodal capabilities NOW.

But most founders are still text-only (and don't realize they're losing customers).


The Multimodal Shift (Why text-only is dead)

Why customers prefer visual communication

HUMAN COMMUNICATION PREFERENCE: ├─ Text: 7% communication (only words) ├─ Voice: 55% communication (tone, emotion, pacing) ├─ Visual: 38% communication (body language, expressions, images) └─ Total: 100% (multimodal is natural)

WHY TEXT-ONLY AGENTS FAIL: ├─ Problem 1: TEXT IS AMBIGUOUS │ ├─ Customer: "The thing doesn't work." │ ├─ Agent (text): "What thing? What doesn't work?" │ ├─ Customer (frustrated): "I literally showed you a photo!" │ └─ Agent: "I can only read text, not images." │ ├─ Result: Miscommunication, customer frustration │ └─ Solution: Multimodal (agent sees the photo) │ ├─ Problem 2: TEXT IS SLOW │ ├─ Customer issue: Broken device (obvious from photo) │ ├─ Text conversation: 10 back-and-forth messages │ ├─ Time: 30 minutes of clarification │ ├─ Multimodal: Agent sees photo immediately. "I see the issue. Here's the fix (with diagram)." 2 minutes. │ └─ Result: 15x faster resolution │ ├─ Problem 3: TEXT IS INCOMPLETE │ ├─ Error messages are visual (red boxes, blinking icons) │ ├─ User interface problems need screenshots │ ├─ Product damage needs photos │ ├─ Text description: "There's a red error message in the top right corner." │ ├─ Multimodal: Agent sees the exact error (screenshot). Knows exactly what's wrong. │ └─ Result: Accuracy improves from 60% → 95% │ ├─ Problem 4: TEXT IS UNPROFESSIONAL │ ├─ Customer expects: Professional support (like human agent) │ ├─ Text-only agent: Feels like chatbot (low trust) │ ├─ Multimodal agent: Sends diagrams, screenshots, videos (feels professional) │ └─ Result: Trust increases, NPS improves │ └─ MARKET IMPLICATION: ├─ Customers increasingly expect rich media ├─ Text-only agents = competitive disadvantage ├─ Multimodal agents = competitive advantage ├─ Market consolidating around multimodal └─ Late adopters will lose customers to multimodal competitors

EXAMPLE ENGAGEMENT COMPARISON:

TEXT-ONLY AGENT: ├─ Response time: 10-30 messages to resolve ├─ Customer satisfaction: 60% (feels like chatbot) ├─ Resolution accuracy: 65% (misunderstandings from text) ├─ Customer effort: High (lots of typing) ├─ NPS impact: Negative (customers frustrated) └─ Automation success: 40% (many issues escalate to human)

MULTIMODAL AGENT: ├─ Response time: 2-3 messages to resolve ├─ Customer satisfaction: 90% (feels like professional support) ├─ Resolution accuracy: 95% (agent sees the exact issue) ├─ Customer effort: Low (send photo, get solution) ├─ NPS impact: Positive (customers impressed) └─ Automation success: 85% (few issues escalate)

BOTTOM LINE: ├─ Multimodal = 3x faster resolution ├─ Multimodal = 30 points higher NPS ├─ Multimodal = 2x higher automation success ├─ Market shift is happening NOW └─ Late movers will lose to early adopters


Google's Multimodal Capabilities (September 2026)

What Google announced (likely)

BASED ON INDUSTRY TRENDS (Google's September 2026 likely announcements):

  1. MULTIMODAL UNDERSTANDING ├─ Agent can now process: Text + Images + Video + Audio ├─ Example: Customer sends photo of error screen ├─ Agent: Reads text in image, understands context, provides solution ├─ Capability: OCR (optical character recognition) + image understanding ├─ Accuracy: 98%+ (reads images as well as humans) └─ Impact: Text-only agents instantly become obsolete

  2. MULTIMODAL GENERATION ├─ Agent can now generate: Text + Images + Video ├─ Example: Agent creates custom diagram explaining solution ├─ Example: Agent generates video tutorial for user ├─ Capability: Image generation (from text description) ├─ Capability: Video synthesis (assembles clips, adds captions) └─ Impact: Engagement increases dramatically

  3. REAL-TIME VIDEO PROCESSING ├─ Agent can now: Watch live video, respond in real-time ├─ Example: Customer shows problem via video call ├─ Agent: Sees problem live, guides customer step-by-step ├─ Capability: Frame-by-frame video analysis └─ Impact: Video support = human-like experience

  4. VISUAL DOCUMENT PROCESSING ├─ Agent can now: Read contracts, invoices, forms as images ├─ Example: Customer sends invoice photo ├─ Agent: Extracts information (amount, date, reference) ├─ Capability: Document understanding (reads handwriting, tables, charts) └─ Impact: Automation of document-heavy processes

  5. MULTIMODAL SEARCH & RETRIEVAL ├─ Agent can now: Search knowledge base using images ├─ Example: Customer sends photo of broken part ├─ Agent: Searches "parts database" for matching item ├─ Capability: Image-based search (not just text-based) └─ Impact: Knowledge retrieval 10x faster

  6. SYNTHETIC MEDIA GENERATION ├─ Agent can now: Generate realistic product images, videos ├─ Example: Customer asks "What does this in red?" ├─ Agent: Generates visual (renders product in red, in customer's environment) ├─ Capability: 3D rendering, AR preview └─ Impact: E-commerce agents become visual assistants

Why multimodal changes everything

BEFORE MULTIMODAL (2025): ├─ Agent workflow: Receive question → Interpret text → Search database → Return text answer ├─ Problem: Ambiguity (text is imprecise) ├─ Problem: Latency (multiple messages back-and-forth) ├─ Problem: Low engagement (boring text responses) ├─ Problem: Low accuracy (misunderstandings) └─ Result: Automation succeeds 40%, escalates to human 60%

AFTER MULTIMODAL (2026+): ├─ Agent workflow: Receive image/video + text → Understand context visually → Search visual database → Return multimodal answer (text + image + video) ├─ Benefit: Clarity (can see actual issue) ├─ Benefit: Speed (immediate understanding) ├─ Benefit: High engagement (visual guidance) ├─ Benefit: High accuracy (sees not just reads) └─ Result: Automation succeeds 85%, escalates to human 15%

COMPETITIVE IMPLICATION: ├─ Early adopters (multimodal now): 2-3 year advantage ├─ Late adopters (waiting until 2027-2028): Playing catch-up ├─ Non-adopters: Losing customers to multimodal competitors ├─ Window to implement: 6-12 months (before market saturates) └─ Cost of waiting: Customer churn, market share loss


Real-World Examples (Multimodal agents in action)

Example 1: E-commerce support agent (product issues)

SCENARIO: Customer has problem with product. Uses WhatsApp to contact support.

BEFORE (text-only agent, 2025): ├─ Customer: "My phone has a cracked screen." ├─ Agent: "Can you describe the crack?" ├─ Customer: "It's on the left side, like a spiderweb." ├─ Agent: "Is it affecting functionality?" ├─ Customer: "I don't know, maybe?" ├─ Agent: "Does the screen respond to touch?" ├─ Customer (frustrated): "I don't know! Just look at the photo I sent!" ├─ Agent: "I can only read text. Please describe in detail." ├─ Customer: Gives up. Calls support. Waits 20 minutes. Talks to human. ├─ Time wasted: 35 minutes ├─ Automation success: 0% └─ NPS impact: -10 (customer frustrated)

AFTER (multimodal agent, 2026): ├─ Customer: Sends photo of cracked screen (no text needed) ├─ Agent: Analyzes image │ ├─ Detects: 15cm crack on left side │ ├─ Recognizes: Damage type (impact, not defect) │ ├─ Assesses: Screen is non-functional (cracks block most pixels) │ └─ Conclusion: Warrantable damage (customer dropped phone) ├─ Agent responds: │ ├─ Text: "I see the damage. Looks like impact damage. Covered under protection plan." │ ├─ Image: Sends diagram showing damage assessment │ ├─ Video: Short video showing replacement process │ └─ Action: "Approved for replacement. Your shipping label is attached (PDF + image)." ├─ Customer: Has full visual guidance. Ships damaged phone next day. ├─ Time wasted: 2 minutes ├─ Automation success: 100% └─ NPS impact: +20 (customer impressed by speed & clarity)

IMPACT: ├─ Resolution time: 35 min → 2 min (17x faster) ├─ Automation success: 0% → 100% ├─ NPS improvement: -10 → +20 (30 points) ├─ Cost per resolution: R$50 (human) → R$2 (agent) ├─ Volume: Can handle 100x more cases with same resources └─ Bottom line: Multimodal = better experience + lower cost

Example 2: B2B SaaS support (technical troubleshooting)

SCENARIO: Enterprise customer has error in SaaS product. Uses Slack/WhatsApp to contact support.

BEFORE (text-only, 2025): ├─ Customer: "Getting error on dashboard." ├─ Agent: "What's the error message?" ├─ Customer: "Something about API key?" ├─ Agent: "Can you copy the exact error?" ├─ Customer: Types out error (gets some words wrong) ├─ Agent: "Hmm, doesn't match known errors. Send me the logs." ├─ Customer: Exports logs (20-page file) and pastes in text ├─ Agent: "Can't read this. Too much data. Can you filter to last hour?" ├─ Customer (frustrated): "I literally took a screenshot showing the error." ├─ Agent: "I can only read text. Please describe step-by-step what happened." ├─ Customer: Escalates to human support engineer (who can see screenshot) ├─ Time: 1 hour of back-and-forth ├─ Automation success: 0% └─ Cost: R$200 (engineer time)

AFTER (multimodal, 2026): ├─ Customer: Takes screenshot. Sends via Slack. ├─ Agent: Analyzes screenshot │ ├─ Reads: Exact error message (via OCR) │ ├─ Recognizes: Dashboard state, which tab is open │ ├─ Searches: Knowledge base for this exact error │ ├─ Finds: Matching issue + solution │ └─ Assesses: 95% confident in solution ├─ Agent responds: │ ├─ Text: "I see the issue. It's a permissions error in your API key configuration." │ ├─ Image: Diagram showing where to fix (with annotations) │ ├─ Video: 30-second video showing the fix step-by-step │ └─ Action: "Follow these steps. Let me know if it works." ├─ Customer: Follows video. Problem solved in 2 minutes. ├─ Time: 5 minutes total ├─ Automation success: 100% └─ Cost: R$0.50 (agent processing cost)

IMPACT: ├─ Resolution time: 60 min → 5 min (12x faster) ├─ Automation success: 0% → 100% ├─ Cost per resolution: R$200 → R$0.50 (400x cheaper) ├─ Customer satisfaction: Frustrated → Impressed ├─ Volume: Support team can handle 100x more issues └─ Bottom line: Multimodal = operational excellence

Example 3: Sales agent (product demos)

SCENARIO: Sales agent demonstrates product to prospect. Prospect has questions about specific features.

BEFORE (text-only, 2025): ├─ Prospect: "How does the dashboard look? Can I customize the colors?" ├─ Agent: "Yes, you can customize the dashboard." ├─ Prospect: "But what does it actually look like?" ├─ Agent: Types description: "The dashboard has widgets. You can drag-and-drop. Colors are customizable via theme selector." ├─ Prospect: "I want to see it." ├─ Agent: "I can send you documentation with screenshots." ├─ Prospect: "I don't have time to read docs. Can you show me?" ├─ Agent: "You'd need to schedule a demo with a human sales rep." ├─ Prospect: "Never mind, let me look at your competitor." ├─ Result: Lost deal └─ Cause: Couldn't show visuals

AFTER (multimodal, 2026): ├─ Prospect: "How does the dashboard look? Can I customize?" ├─ Agent: Instantly generates: Custom dashboard demo video │ ├─ Video shows: Dashboard with prospect's name/company │ ├─ Customization: Colors matched to their brand │ ├─ Walkthrough: "Here's the dashboard. Here's the color picker. You can change any color." │ └─ Engagement: Interactive (prospect feels ownership) ├─ Prospect: "Looks exactly like what we need." ├─ Agent: Sends: Personalized proposal with embedded video + screenshots ├─ Prospect: Impressed. Signs contract same day. ├─ Result: Won deal └─ Cause: Could show visuals immediately

IMPACT: ├─ Sales cycle time: 2 weeks → 2 days (7x faster) ├─ Win rate: 30% → 60% (2x improvement) ├─ Deal size: Same (but close faster) ├─ Sales agent effectiveness: 1 agent → 10 agents (automation volume) └─ Bottom line: Multimodal = revenue acceleration


How to Implement Multimodal (For your SaaS)

Step 1: Assess current capabilities

❌ RED FLAGS (text-only agent): ├─ Your agent only accepts text input ├─ Your agent only outputs text ├─ Customers can't send images/videos ├─ Agent can't process visual information ├─ You don't have image/video generation └─ You're losing deals to multimodal competitors

✓ GREEN FLAGS (multimodal agent): ├─ Your agent accepts: Text + Images + Video + Audio ├─ Your agent processes images (understands context) ├─ Your agent generates visual responses (diagrams, screenshots) ├─ Your agent generates video (tutorials, guides) ├─ Customers can send photos/videos of problems └─ You're winning deals from text-only competitors

ASSESSMENT: ├─ If mostly ❌: You're behind. Implement multimodal NOW (8-16 week project). ├─ If mostly ✓: You're ahead. Optimize and expand (ongoing improvement).

Step 2: Build multimodal pipeline (technical)

MULTIMODAL AGENT ARCHITECTURE:

INPUT LAYER (What agent accepts): ├─ Text: "I need help with X" ├─ Image: Product photo, error screenshot, document ├─ Video: Screen recording, walkthrough ├─ Audio: Voice message (transcribed to text) └─ Combined: Text + image + video simultaneously

PROCESSING LAYER (How agent understands): ├─ Text processing: Claude/GPT reads the text ├─ Image processing: Vision model analyzes image content │ ├─ OCR: Reads text in images (error messages, labels) │ ├─ Object detection: Identifies what's in the image (broken part, error state) │ └─ Scene understanding: Understands context (where is this error? what's the user trying to do?) ├─ Video processing: Analyzes video frames │ ├─ Scene detection: What's happening in the video │ ├─ Timeline: Key moments (where the error occurs) │ └─ Action detection: What is the user doing? ├─ Context integration: Combines all inputs │ ├─ Text + image = full context │ ├─ Video + transcription = complete understanding │ └─ All together = near-human understanding └─ Knowledge retrieval: Searches knowledge base ├─ By text (traditional) ├─ By image (image-based search) └─ By context (semantic search)

OUTPUT LAYER (How agent responds): ├─ Text: Explanation of the issue + solution ├─ Image: │ ├─ Diagram showing the problem (annotated screenshot) │ ├─ Step-by-step visual guide (screenshots with arrows) │ └─ Custom visualization (generated image explaining solution) ├─ Video: │ ├─ Recording of solution being applied │ ├─ Screencast showing steps │ └─ Personalized walkthrough (mentions customer's specific case) └─ Combined: Text + image + video for maximum clarity

TECHNICAL IMPLEMENTATION:

  1. CHOOSE MODEL ├─ Option A: Google Gemini (multimodal built-in, good for vision + generation) ├─ Option B: Claude 3.5 with vision (good for image understanding + text) ├─ Option C: Open-source (LLaVA, etc., more control but less polished) ├─ Recommendation: Start with Claude 3.5 (mature, reliable) └─ Migration path: Switch to Google Gemini when multimodal is fully baked

  2. BUILD PIPELINE ├─ Input handling: Accept file uploads (images, videos) ├─ Processing: Pass to vision model for understanding ├─ Retrieval: Search knowledge base (text + image search) ├─ Generation: Create response (text + images + video) └─ Delivery: Send through WhatsApp/Slack (with media support)

  3. INTEGRATE WITH YOUR PLATFORM ├─ WhatsApp Business API: Send images, videos, documents ├─ Slack: Share images, videos in channels ├─ Web chat: Embed video player, image gallery ├─ Email: Send HTML with embedded images, video links └─ Custom: Build your own UI

  4. GENERATE MULTIMODAL RESPONSES ├─ Use image generation (DALL-E, Midjourney API) for custom diagrams ├─ Use video synthesis (Synthesia, D-ID) for tutorial videos ├─ Pre-record common videos (product demo, troubleshooting steps) ├─ Combine: Template + personalization (auto-customize videos) └─ Quality: Review automated content before sending

Step 3: Phased rollout (timeline)

PHASE 1: VISION INPUT (Week 1-4) ├─ Goal: Agent can accept and understand images ├─ Implementation: │ ├─ Integrate vision model (Claude 3.5 vision) │ ├─ Accept image uploads via WhatsApp/Slack │ ├─ Analyze images (OCR + object detection) │ └─ Test with 50 real support cases ├─ Success metric: 90% accuracy on image understanding ├─ Cost: R$2K-5K └─ Benefit: 30% improvement in resolution time

PHASE 2: IMAGE GENERATION (Week 5-8) ├─ Goal: Agent generates custom diagrams + screenshots ├─ Implementation: │ ├─ Integrate image generation (DALL-E or similar) │ ├─ Create diagram templates (error explanations, step-by-step guides) │ ├─ Generate personalized images (with customer context) │ └─ Test with 100 support cases ├─ Success metric: Customers say diagrams are helpful (80%+) ├─ Cost: R$5K-10K └─ Benefit: 50% improvement in customer satisfaction

PHASE 3: VIDEO GENERATION (Week 9-16) ├─ Goal: Agent generates video tutorials ├─ Implementation: │ ├─ Record common tutorial videos (pre-made) │ ├─ Or generate videos using API (Synthesia, D-ID) │ ├─ Personalize videos (add customer name, company) │ ├─ Test with 200 support cases │ └─ Measure engagement (do customers watch?) ├─ Success metric: 70%+ video watch rate ├─ Cost: R$10K-20K └─ Benefit: 60% improvement in automation success

PHASE 4: VIDEO INPUT (Week 17-24) ├─ Goal: Agent can understand video input (customer videos) ├─ Implementation: │ ├─ Accept video uploads via WhatsApp/Slack │ ├─ Analyze video (frame-by-frame processing) │ ├─ Detect key moments (where the error occurs) │ ├─ Generate solution based on video context │ └─ Test with live customer support cases ├─ Success metric: 85%+ accuracy on video understanding ├─ Cost: R$10K-15K └─ Benefit: 70% automation success (video cases)

TOTAL PROJECT TIME: 6 months TOTAL COST: R$27K-50K BREAKEVEN: 2-4 months (cost saved from faster resolutions)

SUCCESS METRICS (Post-implementation): ├─ Resolution time: 30 min → 5 min (6x faster) ├─ Automation success: 40% → 80% ├─ Customer satisfaction: 60% → 85% ├─ NPS improvement: +20 points ├─ Cost per resolution: -70% └─ Volume: Same team handles 3-5x more cases

Step 4: Quick wins (start today)

  1. IMAGE ACCEPTANCE (2 hours) ├─ Enable image uploads in WhatsApp Business API ├─ Store images in your system ├─ Agent acknowledges: "I received your image. Analyzing..." └─ Result: Customers can now send images (even if agent can't process yet)

  2. VISION MODEL INTEGRATION (8 hours) ├─ Integrate Claude 3.5 vision (or similar) ├─ Have agent analyze customer images ├─ Extract text from images (OCR) ├─ Understand what's in the image └─ Result: Agent can now see what customer is showing

  3. PRE-MADE RESPONSE IMAGES (4 hours) ├─ Create 10 common response diagrams ("How to reset password", "Where to find API key", etc.) ├─ Agent sends these diagrams when appropriate ├─ Don't generate (too slow). Use pre-made. └─ Result: Responses are visual (not just text)

  4. VIDEO LIBRARY (8 hours) ├─ Record 10 common tutorial videos (screen recordings) ├─ Store in your system ├─ Agent sends video link when customer needs walkthrough └─ Result: Customers can watch how to solve problems

QUICK WIN TOTAL: 22 hours (can do in 1-2 weeks), 40% improvement in automation success RECOMMENDATION: Start with quick wins, expand to full multimodal later


FAQ

Q: Mas as APIs de image generation são caras? Quanto custa? (Cost concern)

A: Depende. Image generation (DALL-E): R$0.01-0.20 por imagem (barato). Video generation (Synthesia): R$0.50-5 por minuto (mais caro). Recomendação: Comece com pre-made images (R$0 custo), migre para generated (pago) quando volume justificar. Se 1K resoluções/mês: R$50-500/mês em imagens (1-10% do valor economizado em tempo).

Cost breakdown: ├─ Image generation: R$0.01-0.20 per image (cheap) ├─ Video generation: R$0.50-5 per minute (expensive, but worth it) ├─ Pre-made images/videos: R$0 generation cost (only storage) ├─ Total cost (1K resolutions/month): R$100-500 (including video) ├─ Time saved value: R$5K-10K (at R$5/minute engineer time) └─ ROI: 1000%+ (worth it)

Q: E se usar vídeo: preciso de um estúdio? Como faço? (Technical concern)

A: Não precisa estúdio. Opções: (1) Pre-record screencast (grátis, usando OBS/ScreenFlow), (2) Use video synthesis API (Synthesia, D-ID, HeyGen - gera vídeo de avatar), (3) Hire freelancer para gravar (R$500-1K). Recomendação: Comece com screencast (fácil, grátis), migre para synthesis API se volume justificar.

Video options: ├─ Option 1 (Screencast): Record screen (free tools: OBS, ScreenFlow), R$0 cost ├─ Option 2 (Video synthesis): Use API to generate avatar video, R$0.50-5 per video ├─ Option 3 (Hire): Pay freelancer to record, R$500-1K per video ├─ Option 4 (Hybrid): Pre-record some, generate others └─ Recommendation: Start with Option 1 (screencast), upgrade to Option 2 when volume increases

Q: Meus clientes estão prontos pra isso? Não é muito futurista? (Adoption concern)

A: Não é futurista. É agora. Clientes já esperam agentes multimodal (porque Google/Apple/OpenAI anunciaram). Clientes que recebem texto-only ficarão com raiva ("Por que não me mostra em vídeo?"). Recomendação: Implemente agora (antes da concorrência), você vai estar à frente.

Customer readiness: ├─ Young customers (< 35 years): Expect multimodal, demand video (NOT text) ├─ Enterprise customers: Want rich media for support (not just chat) ├─ International: Prefer video to language barriers (video shows not tells) ├─ Market trend: Everyone wants multimodal (competitors already doing it) ├─ Conclusion: Customers are ready. More than ready. They're expecting it. └─ Action: Implement now, before they switch to competitors offering multimodal

Q: Como medir sucesso de um agente multimodal? Que métricas? (Metrics concern)

A: Mesmas métricas, mas melhores. (1) Resolution time: Quanto tempo leva? Multimodal deve ser 3-5x mais rápido. (2) Automation success: Quantas% são resolvidas sem escalação? Multimodal deve ser 70-85%. (3) Customer satisfaction: Clientes gostam? Multimodal deve ter +20 NPS. (4) Video engagement: Quantas% assistem vídeos? Meta: 70%+. (5) Cost per resolution: Multimodal deve ser 50-70% mais barato.

Success metrics: ├─ Resolution time: 30 min (text-only) → 5 min (multimodal) ├─ Automation success: 40% (text) → 80% (multimodal) ├─ CSAT/NPS: +20 points improvement ├─ Video watch rate: 70%+ of sent videos watched ├─ Cost per resolution: R$50 → R$15 (70% reduction) ├─ Volume: Same team handles 3x more cases └─ Revenue impact: Faster resolution = happier customers = higher retention/NPS


Publicado em 3 de outubro de 2026

Leia também