Seu agent tá preso em data caótica? Knowledge Bases solve tudo.
Amazon Bedrock Knowledge Bases: Query unstructured data via natural language. Seu agent agora acessa emails, PDFs, tickets (tudo junto).
Equipe OpenClaw · Time de Engenharia & Produto
A Equipe OpenClaw é formada por engenheiros, designers e especialistas em IA dedicados a construir a melhor plataforma de agentes conversacionais para negócios brasileiros. Combinamos expertise…
Seu agent tá preso em data caótica? Knowledge Bases solve tudo.
Você é founder de SaaS.
Seu SaaS tem agent de IA (WhatsApp, atendimento ao cliente, automação de vendas).
Current data problem:
Your agent's data access today: │ ├─ What your agent can access: │ ├─ Structured data (databases, APIs) │ │ ├─ Customer table (name, email, phone) │ │ ├─ Order table (order ID, date, amount) │ │ ├─ Invoice table (invoice number, amount, status) │ │ └─ CRM records (company info, interaction history) │ │ │ └─ Unstructured data (what's scattered everywhere): │ ├─ NOT accessible: Customer emails (scattered in inbox) │ ├─ NOT accessible: Support tickets (in helpdesk system) │ ├─ NOT accessible: PDFs (repair estimates, quotes, contracts) │ ├─ NOT accessible: Scanned documents (photos of receipts) │ ├─ NOT accessible: Chat transcripts (WhatsApp, Slack) │ ├─ NOT accessible: Notes (adjuster diary, support notes) │ ├─ NOT accessible: Attachments (email attachments) │ └─ Reality: 80% of customer data is unstructured │ ├─ What customer asks: │ ├─ "Was my claim approved?" │ ├─ "What's the status of my refund?" │ ├─ "Do you have my insurance documents?" │ ├─ "What did the technician write in the report?" │ └─ Agent response: "I don't know (data is scattered, I can't find it)" │ ├─ Reality check: │ ├─ Claim approval info: In email (customer sent it) │ ├─ Refund status: In support ticket notes (adjuster wrote it) │ ├─ Insurance documents: In PDF files (customer uploaded them) │ ├─ Technician report: In scanned attachment (email from tech) │ ├─ Agent can see: None of this (all unstructured) │ └─ Customer frustrated: "Your agent is useless" │ └─ The problem: ├─ Your agent: Smart (Claude/GPT quality) ├─ Your data: Rich (detailed, useful information) ├─ Connection: MISSING (agent can't access data) ├─ Result: Agent is blind (can't help customers) ├─ Customer experience: Terrible (agent can't answer basic questions) └─ Your liability: High (agent should know customer info)
Then AWS launched Knowledge Bases.
Agent data access just changed.
The Problem: Agents Are Blind to Unstructured Data
80% of customer data is unstructured. Your agent can't see it. Knowledge Bases fix this.
Why unstructured data matters for agents
STRUCTURED DATA (Database):
Example (Insurance claim): ├─ Customer ID: 12345 ├─ Claim ID: CLM-2024-001 ├─ Status: "Pending" (single field) ├─ Amount: €5,000 (number field) ├─ Date opened: 2024-09-15 (date field) └─ Agent can access: ALL (simple query)
How agent queries it: ├─ Query: "SELECT status FROM claims WHERE customer_id = 12345" ├─ Result: "Pending" ├─ Agent response: "Your claim status is pending" └─ Speed: Instant (milliseconds)
Limitation: ├─ Agent only knows: Status is "Pending" ├─ Agent doesn't know: WHY it's pending (details in notes) ├─ Agent doesn't know: What happened on Sept 20 (event in email) ├─ Agent doesn't know: What tech found (report in PDF) └─ Customer wants: Context (not just status)
UNSTRUCTURED DATA (Everything else):
Example (Insurance claim details): ├─ Claim notes: "Visited site. Damage minimal. Estimate pending review." ├─ Tech report: "Left side door has minor dent. Recommend repair." ├─ Customer email: "The issue started after the storm on Sept 12." ├─ Adjuster diary: "Reviewed estimate. Need second opinion from panel." ├─ PDF attachment: "Repair estimate: €3,200. 5-day turnaround." └─ Agent can access: NONE (all scattered, unstructured)
How agent tries to query it: ├─ Agent: "I need to know why claim is pending" ├─ System: "That info is not in database (it's in notes/emails/PDFs)" ├─ Agent: "Can I search for it?" (no search across all documents) ├─ Agent: "I give up" (returns generic answer) └─ Speed: Impossible (can't query unstructured data)
Reality: ├─ Data exists: YES (in emails, notes, PDFs) ├─ Data is valuable: YES (contains claim details) ├─ Agent can find it: NO (not integrated) ├─ Customer frustrated: YES (agent can't help) └─ Problem: Unstructured data is inaccessible to agent
THE GAP (Before Knowledge Bases):
What customer needs: ├─ "Why is my claim pending?" ├─ Answer requires: Structured data (status) + Unstructured data (notes, emails, reports) └─ Current system: Can't connect both
What agent can do: ├─ Agent: "Your claim status is pending" (from database) ├─ Agent: "I don't know why" (can't access notes/emails/PDFs) ├─ Customer: "That's not helpful" (frustrated) └─ Result: Agent failure (despite having the answer)
Where the data actually is: ├─ Status: Database (easy to query) ├─ Why pending: Email from adjuster (unstructured) ├─ What needs review: PDF attachment (unstructured) ├─ Next steps: Adjuster notes (unstructured) ├─ Timeline: Tech report + email thread (unstructured) └─ Answer requires: Combining structured + unstructured data
WHY UNSTRUCTURED DATA IS HARD FOR AGENTS:
Traditional approach (before Knowledge Bases): ├─ Option 1: Manually digitize everything (months of work) │ ├─ Cost: €10K-100K (depends on volume) │ ├─ Time: 2-6 months (OCR, manual entry) │ ├─ Accuracy: 90-95% (always some errors) │ └─ Scalability: Breaks (can't handle new documents) │ ├─ Option 2: Build custom ETL (engineering effort) │ ├─ Cost: €20K-50K (engineer time) │ ├─ Time: 1-3 months (design, build, test) │ ├─ Complexity: High (data integration nightmare) │ └─ Maintenance: Ongoing (breaks when formats change) │ └─ Option 3: Give up (accept agent limitations) ├─ Cost: €0 (but customer experience suffers) ├─ Time: Immediate (but competitors will beat you) ├─ Agent quality: Poor (can't access 80% of data) └─ Customer churn: Inevitable (when agent can't help)
Reality: ├─ Most companies: Choose Option 3 (give up) ├─ Result: Agents are limited (can't access real data) ├─ Outcome: Agent ROI is poor (limited by data access) └─ Problem: Unstructured data = bottleneck
Knowledge Bases: Connect Agents to Unstructured Data
Amazon Bedrock Knowledge Bases use RAG (Retrieval Augmented Generation) to let agents query any document. Emails, PDFs, notes, transcripts—everything.
How Knowledge Bases work (high level)
RAG (Retrieval Augmented Generation) in 3 steps:
STEP 1: INGEST (Index your unstructured data)
What you upload: ├─ Customer emails (Outlook, Gmail) ├─ Support tickets (Zendesk, Intercom) ├─ PDFs (Contracts, estimates, reports) ├─ Word documents (Specs, procedures) ├─ Scanned images (Receipts, documents) ├─ Notes (Confluence, Notion, Jira) ├─ Chat transcripts (Slack, Teams) └─ Any document format (txt, csv, json, etc)
What Knowledge Base does: ├─ Extract text (from PDFs, images via OCR) ├─ Break into chunks (pages into searchable pieces) ├─ Vectorize chunks (convert text to embeddings) ├─ Store in vector database (searchable index) └─ Result: Documents now searchable
Timeline: 1-5 minutes (depending on document volume)
STEP 2: RETRIEVE (Agent asks a question)
Agent: "Why is claim CLM-2024-001 pending?"
Knowledge Base: ├─ Convert question to embedding (same way as documents) ├─ Search database (find most similar documents) ├─ Return top results (e.g., adjuster notes, email, report) ├─ Score relevance (which documents are most relevant?) └─ Result: Relevant documents retrieved
Example results: ├─ Document 1: Adjuster email ("Pending second opinion") ├─ Document 2: Tech report ("Damage assessed, estimate provided") ├─ Document 3: Claim notes ("Waiting for panel review") └─ Context: Agent now has documents to read
Timeline: < 1 second (search is fast)
STEP 3: GENERATE (Agent grounds response in documents)
Agent now has: ├─ Customer question: "Why is my claim pending?" ├─ Retrieved documents: Adjuster notes, tech report, email └─ LLM capability: Claude/GPT can read documents
Agent reasoning: ├─ Read documents (understand context) ├─ Identify relevant information (claim is waiting for review) ├─ Ground response in documents (not hallucinating) ├─ Generate answer (personalized to customer) └─ Result: Good answer
Agent response: ├─ "Your claim is pending a second opinion from our panel." ├─ "We received the tech report (damage is minimal)." ├─ "We're also waiting for the repair estimate review." ├─ "Expected timeline: 3-5 business days." └─ Context: Full picture, grounded in documents
Timeline: 2-3 seconds (LLM reasoning)
KEY BENEFIT: Grounded Responses
Without Knowledge Base: ├─ Agent: "I don't have that information" ├─ Agent: Guesses/hallucinates (makes up answer) ├─ Result: Wrong answer (or no answer) └─ Customer trust: Broken
With Knowledge Base: ├─ Agent: Retrieves relevant documents ├─ Agent: Reads documents (grounding) ├─ Agent: Generates response based on real data ├─ Result: Accurate answer (sourced from actual documents) └─ Customer trust: Earned (agent knows my data)
Why this matters: ├─ Hallucination prevention (LLM can't make things up) ├─ Accuracy guarantee (response is grounded in documents) ├─ Trust building (customer feels heard + understood) ├─ Liability protection (you have audit trail) └─ Agent credibility (becomes actually useful)
Real-World Use Cases: Where Knowledge Bases Transform Agents
Insurance, support, sales—any industry with unstructured data + agents.
Three industries where Knowledge Bases unlock agent value
USE CASE 1: INSURANCE (Claims management)
Current problem: ├─ Claim data: Scattered (emails, notes, PDFs, attachments) ├─ Adjuster needs: "All open claims > €10,000 from last month" ├─ Policyholder asks: "Was my claim approved?" ├─ Agent access: Limited (structured data only) └─ Result: Slow responses, manual work, customer frustration
With Knowledge Bases: ├─ Agent ingests: All claim documents (emails, reports, notes) ├─ Agent queries: "All open claims > €10,000" (searches documents) ├─ Agent retrieves: Relevant claims with full context ├─ Agent responds: Adjuster gets answer in seconds (not hours) ├─ Policyholder asks: "Was my claim approved?" ├─ Agent responds: Full context (status + reasoning + timeline) └─ Result: Fast, accurate, detailed responses
Benefit: ├─ Adjuster efficiency: +50% (less manual document searching) ├─ Customer satisfaction: +40% (agent has full context) ├─ Response time: 10 minutes → 30 seconds └─ Agent ROI: High (actually useful for complex queries)
USE CASE 2: CUSTOMER SUPPORT (Ticket resolution)
Current problem: ├─ Support tickets: Scattered across helpdesk system ├─ Previous conversations: Buried in email threads ├─ Product docs: In Confluence (not searchable by agent) ├─ Bug reports: In Jira (agent can't access) ├─ Customer needs: Full context (what happened before?) └─ Agent limitations: Can only see current ticket
With Knowledge Bases: ├─ Agent ingests: All tickets, emails, docs, bug reports ├─ Agent queries: "Similar issues in past 6 months" ├─ Agent retrieves: Relevant tickets + solutions ├─ Agent learns: How was this solved before? ├─ Agent responds: Uses past solution (not inventing) └─ Result: Consistent, accurate, faster resolutions
Benefit: ├─ Resolution time: -60% (uses past solutions) ├─ First-contact resolution: +50% (knows answer already) ├─ Support escalation: -40% (fewer escalations needed) ├─ Agent ROI: Very high (handles complex tickets) └─ Customer loyalty: High (consistent good support)
USE CASE 3: SALES (Proposal generation)
Current problem: ├─ Customer requirements: In emails + Slack messages ├─ Previous proposals: Scattered in email archives ├─ Pricing info: In different spreadsheets ├─ Product specs: In various documents ├─ Past deals: Sales notes (hard to search) ├─ Sales rep needs: To manually gather all this └─ Result: Proposal generation takes hours (not minutes)
With Knowledge Bases: ├─ Agent ingests: All customer emails, past proposals, pricing, specs ├─ Agent queries: "Customer requirements + similar past deals" ├─ Agent retrieves: Relevant info (what they need, what worked before) ├─ Agent generates: Customized proposal (in minutes) ├─ Sales rep reviews: Proposal is 80% done (just minor edits) └─ Result: Proposal ready in 30 minutes (vs 3 hours)
Benefit: ├─ Proposal time: 3 hours → 30 minutes (-83%) ├─ Sales velocity: +3x (more proposals, more deals) ├─ Win rate: +15% (better customized proposals) ├─ Agent ROI: Extremely high (accelerates entire sales cycle) └─ Revenue impact: Major (faster deals = more revenue)
Implementation: How to Set Up Knowledge Bases for Your Agent
Three steps to connect your agent to unstructured data.
Knowledge Base setup process
STEP 1: CHOOSE DATA SOURCE (2-4 hours)
What to ingest: ├─ Option A: Email archive (Gmail, Outlook) ├─ Option B: Helpdesk tickets (Zendesk, Intercom, Jira) ├─ Option C: Document repository (SharePoint, Confluence) ├─ Option D: File storage (Google Drive, OneDrive) ├─ Option E: Custom database (your internal system) └─ Recommendation: Start with highest-value source (e.g., customer emails)
How to prepare: ├─ Export documents from source system ├─ Organize into folders (by category/type) ├─ Remove sensitive data (PII, passwords) ├─ Format standardization (consistent naming) └─ Timeline: 1-2 hours (depending on volume)
STEP 2: CREATE KNOWLEDGE BASE (15-30 minutes)
Amazon Bedrock process: ├─ Create Knowledge Base (web console, 3 clicks) ├─ Configure vector storage (AWS Titan embeddings) ├─ Upload documents (drag & drop or API) ├─ Set chunking strategy (how to split documents) ├─ Configure retrieval settings (how many documents to return) └─ Timeline: 15-30 minutes (fully automated)
Configuration options: ├─ Chunk size: 512-2048 tokens (smaller = more granular) ├─ Overlap: 10-20% (for context preservation) ├─ Retrieval count: 5-10 documents (more = slower but better context) ├─ Filtering: By date, metadata, category (optional) └─ Recommendation: Start with defaults (works well for most)
STEP 3: INTEGRATE WITH AGENT (1-3 days)
Connect Knowledge Base to your agent: ├─ Add Knowledge Base as tool to agent (API call) ├─ Teach agent when to use Knowledge Base (prompt engineering) ├─ Test queries (does agent retrieve right documents?) ├─ Validate responses (is agent grounding correctly?) ├─ Deploy to production (monitor performance) └─ Timeline: 1-3 days (depending on complexity)
Example agent prompt:
You are a support agent. When customer asks about their claim/ticket:
- First, search Knowledge Base (use tool: bedrock_knowledge_base)
- Find relevant documents (emails, notes, reports)
- Read documents carefully
- Generate response grounded in documents
- If no relevant documents, tell customer honestly
- Never hallucinate (only use what documents say)
Testing checklist: ├─ Query test 1: "Why is my claim pending?" (should find adjuster notes) ├─ Query test 2: "What was the estimate amount?" (should find PDF) ├─ Query test 3: "When will it be done?" (should find timeline) ├─ Query test 4: "Nonsense query" (should say documents don't answer) └─ All tests pass? Launch to customers
COST & TIMELINE SUMMARY:
Setup cost: ├─ AWS Knowledge Base service: ~€100-500/month (usage-based) ├─ Engineering time: 3-5 days (€3K-7K salary) ├─ Document preparation: 1-2 days (€1K-2K salary) └─ Total: €5K-10K first month (then €500-1000/month ongoing)
Timeline to launch: ├─ Planning: 1 day ├─ Setup: 1 day ├─ Integration: 1-2 days ├─ Testing: 1-2 days ├─ Deployment: 1 day └─ Total: 5-8 days (can launch in 1-2 weeks)
ROI calculation: ├─ Cost: €5-10K setup + €500-1000/month ongoing ├─ Benefit: Reduce agent training time by 50% ├─ Benefit: Reduce support response time by 60% ├─ Benefit: Increase customer satisfaction by 30% ├─ Benefit: Handle 3x more queries (agent is smarter) └─ ROI: Break-even in 1-2 months (very high)
Knowledge Bases Change Everything for Agent-Driven SaaS
Before Knowledge Bases: Agents are limited by structured data. After: Agents can access ALL your data.
The shift in agent capability
BEFORE KNOWLEDGE BASES:
Agent capability: ├─ Can access: Structured data (databases, APIs) ├─ Can't access: Unstructured data (emails, PDFs, notes) ├─ Agent quality: Limited (only knows what's in database) ├─ Customer satisfaction: Moderate (agent lacks context) └─ Agent ROI: Low (many queries agent can't answer)
Example: ├─ Customer: "Why is my claim pending?" ├─ Agent: "Status is pending" (from database) ├─ Agent: "I don't know why" (can't access adjuster notes) ├─ Customer: "That's not helpful" └─ Result: Agent failure (has the answer, can't access it)
AFTER KNOWLEDGE BASES:
Agent capability: ├─ Can access: Structured data + Unstructured data (everything) ├─ Can answer: Complex questions (requires reading multiple documents) ├─ Agent quality: High (full context available) ├─ Customer satisfaction: High (agent is actually helpful) └─ Agent ROI: Very high (can handle most customer needs)
Example: ├─ Customer: "Why is my claim pending?" ├─ Agent: Searches documents (finds adjuster notes, tech report, email) ├─ Agent: Reads documents (understands full context) ├─ Agent: "Your claim is waiting for panel review. The tech assessment is complete (minor damage). Once approved, repair will take 3-5 days." ├─ Customer: "That's exactly what I needed to know" └─ Result: Agent success (helpful, accurate, complete)
COMPETITIVE IMPACT:
You (with Knowledge Bases): ├─ Agent: Can answer ANY question (has access to all data) ├─ Customer support: Fast + accurate (grounded in documents) ├─ Agent ROI: High (handles complex queries) ├─ Customer loyalty: High (feels understood) └─ Competitive advantage: Clear (agent is actually useful)
Competitors (without Knowledge Bases): ├─ Agent: Limited to structured data ├─ Customer support: Generic (lacks context) ├─ Agent ROI: Low (many queries need escalation) ├─ Customer loyalty: Low (agent is not helpful) └─ Competitive disadvantage: Clear (agent is toy)
Market shift: ├─ 2024: Agents with structured data access = competitive ├─ 2025: Agents with Knowledge Bases = table stakes ├─ 2026: Agents without Knowledge Bases = market losers ├─ Your window: NOW (2025 is when this becomes standard) └─ Early adopters: Win customer loyalty (others catching up)
Next Steps: Add Knowledge Bases to Your Agent
At OpenClaw, we help SaaS founders add Knowledge Bases to their agents (connect unstructured data, implement RAG architecture, optimize retrieval), measure agent quality improvement (accuracy metrics, customer satisfaction, support efficiency), and scale knowledge base systems (handle growing data, maintain performance):
- Agent knowledge assessment (what data should your agent access? where is it now?)
- Knowledge Base architecture design (which data sources? chunking strategy? retrieval optimization?)
- Implementation & integration (connect Knowledge Base to your agent, test, deploy)
- Performance optimization (improve retrieval accuracy, reduce latency, lower cost)
- Training & documentation (teach your team how to use Knowledge Bases)
Get a free knowledge base assessment: Schedule 30 minutes with our agent architect. We'll evaluate your current agent (what data can it access?), identify unstructured data you're missing (emails, PDFs, notes?), recommend Knowledge Base strategy (which data source first?), and design your implementation roadmap (how to launch in 2 weeks?).
[Book your free knowledge base assessment] → [Button: Schedule 30-Minute Call]
Your agent is blind to 80% of your customer data. Knowledge Bases fix this. Your agent just became actually useful.
FAQ
Q: Não é perigoso dar acesso a agente em todos os documentos? (Segurança/privacidade)
A: Ótima pergunta. Você controla completamente:
- Acesso: Você escolhe quais documentos entram na Knowledge Base
- Filtragem: Pode remover dados sensíveis (PII) antes de ingerir
- Segregação: Knowledge Bases diferentes para clientes diferentes
- Auditoria: Log completo de quais documentos agent acessou
- LGPD compliance: Você mantém controle total
Smart play: Comece com dados menos sensíveis (ex: claims públicos), teste segurança, depois escale.
Q: Qual é o custo real de rodar Knowledge Base?
A: Três componentes:
- Embedding storage (€50-200/month): Armazenar vetorizações
- Retrieval queries (€100-500/month): Cada busca custa centavos
- LLM calls (varies): Seu agent lê documentos (custa como qualquer LLM)
- Total: €200-1000/month (pra maioria das empresas)
Mais barato que: Contratar pessoa pra buscar documentos manualmente.
Q: Como Knowledge Base se compara com simples vector search?
A: Mesma coisa, mas empacotado melhor:
- Vector search (DIY): You build it (Pinecone, Weaviate)
- Knowledge Base (managed): Amazon builds it (você usa)
- Benefit of Bedrock: Integração nativa com agentes + managed service
- Recommendation: Use Bedrock Knowledge Base (simpler, not DIY)
Publicado em 30 de setembro de 2026