Notícias
Notícias
5 min de leitura
11 de setembro de 2026

Seu agente WhatsApp é só texto (Amazon: video search agora obrigatório)

Amazon Bedrock: agente busca em vídeos/imagens (Marengo 3.0). Seu agente é só texto? Vision é futuro.

Equipe OpenClaw

Equipe OpenClaw · Time de Engenharia & Produto

A Equipe OpenClaw é formada por engenheiros, designers e especialistas em IA dedicados a construir a melhor plataforma de agentes conversacionais para negócios brasileiros. Combinamos expertise…


Seu agente WhatsApp é só texto (Amazon: video search agora obrigatório)

Você é founder/CEO de SaaS.

Seu SaaS: agente IA em WhatsApp (atendimento, vendas, suporte).

Seu agente: Busca informações em BASE DE TEXTO (documentos, políticas, FAQs).

Ontem: Amazon Bedrock announced Marengo 3.0 (video + image search, GA).

What Marengo 3.0 does (the breakthrough):

  • Video search (find specific moments in hours of footage, by meaning)
  • Image search (search photos, diagrams, screenshots, visual content)
  • Natural language queries ("show me when customer complained" vs time-codes)
  • Multimodal embeddings (understand video + image + text together)
  • Production-ready (GA in Bedrock Knowledge Base, not beta)
  • Scale (process hours of video, thousands of images efficiently)
  • Integration (agentes can query video/image knowledge base directly)

Queries Marengo 3.0 handles (examples):

  • "Show me product defects in manufacturing footage"
  • "Find invoice with customer ID 12345 in scanned documents"
  • "When did customer complain about shipping?"
  • "Show me tutorial video for fixing error code XYZ"
  • "Find all contracts mentioning penalty clause"
  • "Locate product review screenshots with negative sentiment"

Your assumption (WRONG):

  • "My agente searches text, that's sufficient (vision is nice-to-have)"
  • "Video/image search is advanced (not needed for basic support)"
  • "My customers don't upload videos (they just ask questions)"
  • "Text-based knowledge base covers everything (no gaps)"
  • "Competitors won't have vision for months (I have time)"

Your reality (Amazon just proved otherwise):

  • Multimodal agentes (vision + text) are now production-standard (Sept 2026, AWS)
    • Problem: Your agente is text-only (can't search visual content)
    • Evidence: Amazon Bedrock integrated Marengo 3.0 (GA, production)
    • Impact: Agentes that see videos/images answer better (richer context)
    • Customer expectation: "Can my agente see my attachment?"
    • Your agente: "No, only text" (customer frustrated)
    • Competitive signal: Competitors adding vision (will win deals)
    • Timeline: In 6-12 months, vision will be table-stakes
    • Implication: If you don't add vision, you'll lose to multimodal competitors

Why text-only agentes are losing (the vision gap)

The knowledge gap (what text-only agentes miss)

Scenario: Customer support (insurance claims)

Customer: "Meu sinistro não foi aprovado. Por quê?"

=== TEXT-ONLY AGENTE (current) === Agent knowledge base: ├─ Policy documents (text) ├─ Claims guidelines (text) ├─ FAQ (text) ├─ Email conversations (text) └─ Missing: Actual claim form (PDF image), damage photos, video documentation

Agent response: ├─ "Vou verificar sua política..." ├─ Agent searches text: "Você tem cobertura para danos elétricos." ├─ Agent: "Sua reclamação foi rejeitada porque... [generic reason]" ├─ Reality: Agent never saw the actual damage (photo), can't verify ├─ Customer frustrated: "Mas a foto mostra dano claro!" ├─ Result: Customer disputes decision, escalates to human └─ Cost: Manual investigation, customer churn

=== MULTIMODAL AGENTE (Marengo 3.0) === Agent knowledge base: ├─ Policy documents (text) ├─ Claims guidelines (text) ├─ FAQ (text) ├─ Email conversations (text) ├─ Claim forms (images, searchable by meaning) ├─ Damage photos (searchable: "what kind of damage?", "where located?") ├─ Video documentation (searchable: "when did damage occur?", "progression?") └─ Complete context

Agent response: ├─ Customer uploads: Claim form (PDF), damage photo, video ├─ Agent searches multimodal KB: │ ├─ Text: "Cobertura para danos elétricos" │ ├─ Image: Recognizes damage type from photo │ ├─ Video: Understands damage progression │ └─ Combined: Full context (policy + visual evidence) ├─ Agent: "Análise completa: Sua reclamação é elegível. Danificação é clara. Aprovando reembolso de R$ 5.000." ├─ Reality: Agent saw actual evidence, can make confident decision ├─ Customer satisfied: "Agente viu a foto, entendeu tudo!" ├─ Result: Instant approval, no escalation, customer loyal └─ Benefit: Faster resolution, higher customer satisfaction, lower cost

The resolution speed gap (vision = faster decisions)

Comparison: Text-only vs Multimodal agente

=== TEXT-ONLY AGENTE === Customer: "Meu pedido chegou quebrado" Agent: "Entendo. Qual é o problema exatamente?" Customer: "O produto está rachado" Agent: "E qual é a marca/modelo?" Customer: "iPhone 15 Pro" Agent: "Pode enviar foto?" Customer: [uploads photo] Agent: "Recebida, mas não consigo analisar. Vou marcar para análise manual." Escalation: Human reviews photo (2-4 horas) Decision: Refund approved Total time: 4-8 hours

=== MULTIMODAL AGENTE (Marengo 3.0) === Customer: "Meu pedido chegou quebrado" Customer: [uploads photo immediately] Agent: [searches image KB: "broken phone screen", "damage type", "repairable?"] Agent: "Vejo a rachadura no canto. Cobertura total. Processando refund." Decision: Refund approved instantly Total time: 30 seconds

=== BUSINESS IMPACT === Text-only: 4-8 horas de latência Multimodal: 30 segundos de latência Difference: 480-960x faster

Customer satisfaction: ├─ Text-only: Customer frustrated (waited hours) ├─ Multimodal: Customer delighted (instant resolution) └─ Churn impact: Multimodal wins loyalty

Cost impact: ├─ Text-only: 1 human agent per 5 escalations (expensive) ├─ Multimodal: 0 escalations (agente handles autonomously) └─ Savings: 80% reduction in support labor

The context gap (vision = smarter decisions)

Scenario: E-commerce returns (what's wrong with product?)

=== TEXT-ONLY AGENTE === Customer: "Quero devolver o sapato" Agent: "Por quê?" Customer: "Não combina com meu estilo" Agent: "Entendido. Processando devolução." Agent's context: Only text description ("didn't match style") Problem: ├─ Agent can't assess actual condition (worn? damaged? new?) ├─ Agent processes return blindly (assumes customer is honest) ├─ Warehouse receives: Could be new, could be trashed ├─ Inventory value: Unknown (can't resell if damaged) └─ Financial loss: Can't predict

=== MULTIMODAL AGENTE (Marengo 3.0) === Customer: "Quero devolver o sapato" Customer: [uploads 4 photos of shoe] Agent: [searches image KB: "shoe condition", "wear pattern", "damage assessment"] Agent's context: ├─ Photo 1: Sole pattern (worn or new?) ├─ Photo 2: Overall condition (scratches, stains?) ├─ Photo 3: Heel wear (heavy use?) ├─ Photo 4: Side view (cosmetic damage?) Agent decision: ├─ "Shoe shows light wear. Resaleable. Refund full amount, send to warehouse B." ├─ OR "Shoe heavily worn. Non-resaleable. Refund 50%, mark for disposal." ├─ OR "Shoe unworn, perfect condition. Refund, resell as new." Inventory optimization: ├─ Visual assessment automates routing (new stock, B-grade stock, disposal) ├─ Reduces manual inspection (warehouse savings) ├─ Improves resale value prediction (financial accuracy) └─ Scale: 1000s of returns/day, each with instant visual assessment


How Marengo 3.0 works (technical architecture)

Multimodal embedding model (understanding video + image + text)

Marengo 3.0 capabilities:

  1. Video understanding ├─ Input: Raw video file (hours of footage) ├─ Processing: Frame extraction + scene understanding ├─ Embedding: Convert video to searchable vectors (by meaning) ├─ Query: Natural language ("find moment when customer arrives") ├─ Output: Specific timestamp + relevant scene └─ Example: "Show me the 2:15 minute mark (customer interaction)"

  2. Image understanding ├─ Input: Photos, diagrams, screenshots, document scans ├─ Processing: Visual content recognition (what's in image?) ├─ Embedding: Convert image to semantic vectors ├─ Query: Natural language ("find invoice with amount > R$ 1000") ├─ Output: Matching image + highlighted section └─ Example: "Found invoice #12345 with amount R$ 1.500"

  3. Text understanding ├─ Input: Documents, emails, chat logs (traditional) ├─ Processing: Standard NLP (text extraction) ├─ Embedding: Convert text to vectors ├─ Query: Natural language ("what's our refund policy?") ├─ Output: Relevant text passages └─ Example: "Refund available within 30 days of purchase"

  4. Cross-modal search ├─ Query: "Show me damage in customer video" ├─ Searches: VIDEO knowledge (finds matching scene) ├─ Also searches: IMAGE knowledge (similar damage photos) ├─ Also searches: TEXT knowledge (policy on damage coverage) ├─ Output: Video timestamp + similar photos + relevant policy └─ Result: Agent has complete multi-source context

=== INTEGRATION WITH AGENTES ===

Agent workflow (with Marengo 3.0):

  1. Customer sends message + attachment (image/video) └─ "Meu produto está com defeito" + photo

  2. Agent receives + processes ├─ Extracts text from message ├─ Sends image to Bedrock Knowledge Base ├─ Marengo 3.0 creates multimodal embedding └─ Searches KB (text + image + video knowledge)

  3. Agent gets results ├─ Text matches: Similar defect descriptions from KB ├─ Image matches: Customer photos vs previous defects (visual similarity) ├─ Video matches: Tutorial videos showing same issue ├─ Policy matches: Coverage details for this defect type └─ Combined: Rich context for decision

  4. Agent responds with confidence ├─ "Vi sua foto. Defeit tipo X. Cobertura: Sim. Refund: R$ 500." ├─ Based on: Visual evidence (foto) + historical data (KB images) + policy (text) ├─ Explanation: Evidence-based (customer can see logic) └─ Satisfaction: High (agent understood visually)

  5. Customer outcome ├─ Instant resolution (no escalation needed) ├─ High confidence (agent saw evidence) ├─ Loyalty: Customer appreciates visual understanding └─ Cost: Saved 1 human agent interaction

Knowledge base evolution (text → multimodal)

Bedrock Knowledge Base version timeline:

=== Before Marengo 3.0 (text-only) === Knowledge base structure: ├─ Documents (PDF, docx, txt) ├─ Indexed by: Text content only ├─ Searchable by: Keywords + semantic meaning (text) ├─ Limitations: │ ├─ Photos in PDFs: Not searchable (treated as images, not understood) │ ├─ Embedded videos: Not supported │ ├─ Customer attachments (images): Not indexed │ └─ Visual information: Locked away (invisible to agente) └─ Agent knowledge: Text-only (incomplete)

=== After Marengo 3.0 (multimodal) === Knowledge base structure: ├─ Documents (PDF, docx, txt) + photos + videos ├─ Indexed by: Text + image + video content (simultaneous) ├─ Searchable by: Keywords + semantic meaning + visual concepts ├─ Capabilities: │ ├─ Photos in PDFs: Fully searchable ("find invoices with amount > X") │ ├─ Embedded videos: Searchable by scene ("find when product fails") │ ├─ Customer attachments (images): Auto-indexed, searchable │ └─ Visual information: Fully accessible to agente └─ Agent knowledge: Text + visual (complete)

=== MIGRATION PATH ===

Step 1: Set up Bedrock + Marengo 3.0 ├─ Cost: Included in Bedrock pricing (no extra cost) ├─ Timeline: 1-2 weeks └─ Effort: Low (configuration, not development)

Step 2: Upload knowledge base ├─ Documents: Transfer existing text KB (no change) ├─ Photos: Upload customer photos, product images, invoices ├─ Videos: Upload product demos, training videos, customer videos ├─ Cost: Storage (minimal, Bedrock is S3-backed) └─ Timeline: Depends on content volume (1-4 weeks)

Step 3: Integrate with agente ├─ Update agente: Use multimodal search (not text-only search) ├─ Prompts: Tell agente to analyze images + videos when customer uploads ├─ Integration: Connect agente to Bedrock KB (already built, Bedrock handles) ├─ Code: Minimal changes (API calls same, results are richer) └─ Timeline: 1-2 weeks

Step 4: Test + deploy ├─ Red team: Try queries with images/videos (does search work?) ├─ Measure: Agent accuracy improved? Resolution time decreased? ├─ Monitor: Track metrics (satisfaction, escalation rate, cost) └─ Timeline: 1-2 weeks

=== TOTAL IMPLEMENTATION === Cost: R$ 0-20K (mostly labor, Bedrock pricing is usage-based) Timeline: 4-8 weeks to full deployment Benefit: Agents become multimodal (3-10x better at visual questions) ROI: High (faster resolutions, fewer escalations, happier customers)


Real-world use cases (where vision changes everything)

E-commerce returns (visual condition assessment)

Use case: Returns processing agente

Current (text-only): ├─ Customer: "Defective shoe" ├─ Agent: "Send it back, we'll inspect." ├─ Result: Guess if shoe is returnable (expensive)

With Marengo 3.0 (vision-enabled): ├─ Customer uploads photos + video ├─ Agent: "Shoe shows R$ 400 condition (resaleable)" ├─ Agent: "Approve refund, route to refurbished inventory" ├─ Result: Precise value assessment, efficient routing

Business impact: ├─ Cost savings: 40-60% lower refund processing cost (automation) ├─ Accuracy: 95%+ condition assessment (vs 60% human guess) ├─ Speed: 30 seconds (vs 2 hours manual inspection) └─ Customer satisfaction: Instant decision (vs wait for inspection)

Insurance claims (damage assessment)

Use case: Claims agente (auto insurance)

Current (text-only): ├─ Customer: "Damaged in accident" ├─ Agent: "Send documentation" ├─ Customer emails: Photos + police report (text) ├─ Agent: "Forwarding to adjuster (2-3 days)" ├─ Result: Slow (customer waits)

With Marengo 3.0 (vision-enabled): ├─ Customer uploads: Photos + video of damage ├─ Agent searches image KB: Similar damage assessments + payouts ├─ Agent: "Damage type: rear bumper impact. Typical payout: R$ 3.000-5.000" ├─ Agent: "Pre-approval: Estimate R$ 4.500. Full adjuster review to follow." ├─ Result: Instant pre-approval (customer knows expected payout)

Business impact: ├─ Customer satisfaction: Instant feedback (vs 2-3 day wait) ├─ Operational efficiency: 70% of claims pre-approved automatically ├─ Fraud detection: Visual analysis catches inconsistencies (damage doesn't match story) └─ Cost reduction: Fewer adjuster reviews (only high-value or edge cases)

Product support (visual troubleshooting)

Use case: Tech support agente (B2B SaaS)

Current (text-only): ├─ Customer: "Getting error code 0x502" ├─ Agent searches: Text KB for "0x502" ├─ Result: Generic solution (often wrong)

With Marengo 3.0 (vision-enabled): ├─ Customer: "Getting error code 0x502" + screenshot of error ├─ Agent searches image KB: "Error 0x502 with Windows 11" ├─ Result: Finds tutorial VIDEO + similar customer screenshots ├─ Agent: "Você está no Windows 11. Erro 0x502 é compatibilidade. Veja este tutorial: [video link]" ├─ Result: Specific solution (with visual context)

Business impact: ├─ First-contact resolution: +40% (better solutions) ├─ Customer satisfaction: +30% (visual guides help) ├─ Support cost: -20% (fewer follow-ups) └─ Time-to-resolution: 5 min (vs 30 min text back-and-forth)


Roadmap: When vision becomes mandatory (market timeline)

Evolution of agentes (next 12 months)

Market adoption of multimodal agentes:

Sep 2026 (now): ├─ Amazon Bedrock: Marengo 3.0 (GA, production-ready) ├─ Early adopters: Building multimodal agentes ├─ Text-only agentes: Still majority (>80%) ├─ Competitive advantage: Huge (vision = better UX) └─ Your decision: Adopt Marengo 3.0 now or wait?

Dec 2026 (3 months): ├─ Competitors start adding vision ├─ Customers notice: "Your agente can see my photo?" ├─ Market expectation: Agentes should support images ├─ Text-only agentes look cheap (vision becomes baseline) ├─ Your agente: If no vision, losing deals to multimodal competitors └─ Adoption: ~20% of agentes using multimodal

June 2027 (9 months): ├─ Multimodal agentes are standard (not nice-to-have) ├─ Most serious agentes have vision capabilities ├─ Text-only agentes are niche (simple use cases only) ├─ Customers expect: "Can your agente see attachments?" ├─ Your agente: If no vision, major liability └─ Adoption: ~60% of agentes using multimodal

Sep 2027 (12 months): ├─ Multimodal is expected (like auth for SaaS) ├─ Customers assume agentes have vision by default ├─ Text-only agentes are obsolete (customers won't use) ├─ Market has consolidated (multimodal winners, others out) ├─ Your agente: No vision = customer churn └─ Adoption: ~90% of agentes using multimodal

=== IMPLICATION FOR YOU === Now (Sep 2026): Add Marengo 3.0 = 12-month competitive advantage 3 months (Dec 2026): Advantage shrinks as competitors catch up 9 months (June 2027): Advantage gone (everyone has vision) 12 months (Sep 2027): No vision = losing customers

=== DECISION MATRIX === Add Marengo 3.0 NOW: ├─ Cost: R$ 10-40K (integration + knowledge base upload) ├─ Timeline: 4-8 weeks to market ├─ Advantage: 12 months (huge) ├─ ROI: Massive (competitive moat) └─ Recommendation: DO IT NOW

Wait and add later: ├─ Cost: Same R$ 10-40K (same work, later) ├─ Timeline: Same 4-8 weeks (but delayed) ├─ Advantage: 0 months (everyone will have it) ├─ ROI: Zero (no competitive edge) └─ Recommendation: NOT RECOMMENDED (lose market)


Conclusion: Multimodal agentes are competitive standard (vision is mandatory)

The reality (Amazon confirmed):

  • Video + image search is now production-ready (Marengo 3.0, GA)
  • Multimodal agentes answer better (richer context)
  • Text-only agentes are losing (missing visual knowledge)
  • Vision will be table-stakes in 6-12 months (market adopting)
  • Marengo 3.0 is easiest path to multimodal (integrated with Bedrock)

Your choice (2 paths):

Path 1: Stay text-only (no vision)

  • Capability: Text-only (customer photos ignored)
  • Performance: Slower decisions (no visual context)
  • Timeline: 6-12 months until multimodal is expected
  • Competitive position: Behind (will lose to vision-enabled competitors)
  • Recommendation: Not recommended (losing market)

Path 2: Add Marengo 3.0 (go multimodal)

  • Capability: Vision + text (understand photos, videos, documents)
  • Performance: 3-10x faster resolutions (visual context)
  • Timeline: 4-8 weeks to deployment
  • Competitive position: Ahead (12-month advantage)
  • Recommendation: Essential (capture market before saturation)

At OpenClaw, we help SaaS add multimodal vision to agentes:

  • MULTIMODAL AUDIT: Is your agente missing visual knowledge?
  • MARENGO 3.0 INTEGRATION: Connect video/image search to your agente
  • KNOWLEDGE BASE MIGRATION: Upload photos, videos, documents (multimodal indexing)
  • AGENTE ENHANCEMENT: Update prompts + logic to use visual context
  • PERFORMANCE OPTIMIZATION: Fine-tune search accuracy (image + video + text)
  • COMPETITIVE POSITIONING: Win deals with multimodal capability
  • SCALING STRATEGY: Handle high volume of visual queries efficiently
  • ONGOING MONITORING: Track metrics (resolution speed, customer satisfaction)

Result: Your agente sees customer photos/videos. Your agente makes faster decisions (with visual evidence). You're ahead of competitors (multimodal when competitors are still text-only). You win more deals (customers appreciate vision). You retain customers (better UX).

Seu agente é só texto?

Seu agente perde tempo pedindo descrições (quando poderia ver foto)?

Você quer visão 3.0 multimodal antes que seja obrigatório?

Se quer expert guidance (Marengo 3.0 integration, multimodal architecture, knowledge base migration, competitive positioning, scaling strategy):

Agente Multimodal Vision | Marengo 3.0 | Video + Image Search | Bedrock Integration →


Publicado em 11 de setembro de 2026

Leia também