Search agents sem treinamento = inúteis. RL fine-tuning = obrigatório.
Search agents fail 30% of queries without training. RL fine-tuning improves performance 2-3x. Raw agents = liability. Training = mandatory.
Equipe OpenClaw · Time de Engenharia & Produto
A Equipe OpenClaw é formada por engenheiros, designers e especialistas em IA dedicados a construir a melhor plataforma de agentes conversacionais para negócios brasileiros. Combinamos expertise…
Search agents sem treinamento = inúteis. RL fine-tuning = obrigatório.
Ontem Amazon publicou research sobre search agents.
"Getting search agents to work well is hard. No base model achieves this out-of-box."
What this means: Your search agent (WhatsApp bot, knowledge base search, support assistant) probably performs poorly (fails 20-40% of queries) because it's untrained.
Why it matters: Untrained agents = customer frustration = churn = lost revenue.
Problem it reveals: You probably deployed agents assuming they'd "just work." They don't.
Você é founder.
You deployed support search agent (WhatsApp):
- "Agent will search knowledge base and find answers."
- Agent deployed (no training, just raw LLM).
- Customer asks: "What's your return policy?"
- Agent searches knowledge base.
- Agent finds 3 irrelevant documents.
- Agent hallucinates: "30-day returns guaranteed" (policy is actually 14 days).
- Customer upset (expected 30 days, got 14).
- You pay for customer service recovery (R$500+ per incident).
- Multiply by 100 customers/month = R$50K loss.
Root cause: Untrained agent (raw LLM doesn't know your search strategy).
Solution: RL fine-tuning (train agent on YOUR knowledge base + YOUR feedback).
Result: Agent learns what to search for, when to stop, how to verify answers.
Performance: 30-50% improvement in first-try success rate.
Implication: Search agents MUST be trained (raw ≠ viable).
But most founders don't know this.
What Are Search Agents? (And Why Raw Ones Fail)
Search agent = Agent that decides WHAT to search for (not just retrieves). Raw LLM-based agent = Searches for everything (wastes time, retrieves irrelevant docs, hallucinates). Trained agent (RL fine-tuned) = Learns optimal search strategy (searches strategically, refines queries based on results, stops when answer found). Difference = night and day. Raw = 30-50% failure rate. Trained = 5-10% failure rate.
How search agents work: Raw vs trained
TASK: Customer asks "What are your refund policies for international orders?"
RAW AGENT (No training):
Step 1: Agent reads question ├─ Thinks: "I need to search for refunds" └─ Decision: "Search for: 'refund'"
Step 2: Agent searches knowledge base ├─ Query: "refund" ├─ Results: 47 documents (generic refund policy, refund for domestic, refund for international, refund timelines, refund process, refund appeals, etc.) └─ Problem: Too many results, agent confused
Step 3: Agent picks first result ├─ Retrieved: "General refund policy: 30 days" ├─ Problem: This is for DOMESTIC orders (not international) └─ Agent doesn't check if result is relevant to "international"
Step 4: Agent generates answer ├─ Agent thinks: "Refund policy is 30 days" ├─ Agent hallucinates: "We offer 30-day refunds for all orders, including international!" └─ Problem: Wrong answer (international policy is 60 days)
Step 5: Customer reads answer ├─ Customer expects: 60-day refund for international ├─ Customer sees: 30-day refund (wrong!) ├─ Customer upset: "Policy is wrong, I want my money back" └─ Cost: R$1,000 in recovery + customer churn
TRAINED AGENT (RL fine-tuned):
Step 1: Agent reads question ├─ Learned from training: "If question mentions 'international', search specifically for that" ├─ Decision: "Search for: 'refund international' (specific, not generic)" └─ Confidence: High (trained model knows this strategy works)
Step 2: Agent searches knowledge base ├─ Query: "refund international" ├─ Results: 5 documents (specifically about international refunds) ├─ Problem: Fewer, more relevant results └─ Agent can confidently pick best result
Step 3: Agent checks result relevance ├─ Learned from training: "Always verify result matches question keywords" ├─ Check: Does result mention both 'refund' AND 'international'? ├─ Result: "International refund policy: 60 days" ├─ Relevance check: YES (both keywords present) └─ Confidence: High
Step 4: Agent generates answer ├─ Agent: "For international orders, we offer 60-day refunds" ├─ Quality: Accurate, specific, confident └─ Correct!
Step 5: Customer reads answer ├─ Customer expects: 60-day refund for international ├─ Customer sees: 60-day refund (correct!) ├─ Customer satisfied: "Perfect, matches our policy" └─ Cost: Zero recovery, customer happy
COMPARATIVE PERFORMANCE:
Metric Raw Agent Trained Agent Improvement
Search queries needed 5-7 2-3 -60% Document relevance 40-50% 85-95% +75% Hallucination rate 30-50% 5-10% -80% First-try success 50-70% 85-95% +30-45% Customer satisfaction 60% 90% +30% Cost per error R$1,000+ R$100-200 -80%
KEY INSIGHT:
Raw agent: "Throw everything at the wall, hope something sticks" Trained agent: "Know exactly what to search for, refine based on results"
Difference: ├─ Raw = Unreliable (customer can't trust) ├─ Trained = Reliable (customer can depend on) └─ Business impact: Trained agents save R$50K-500K/year in error costs
The RL Fine-Tuning Process: How to Train Agents to Perform
Step 1: Collect feedback (customer satisfaction data + agent decisions). Step 2: Label good vs bad agent behaviors (what should agent have done?). Step 3: Use RL to train agent (reinforcement learning optimizes agent for good outcomes). Step 4: Evaluate improvements (measure success rate before/after). Step 5: Iterate (continuous improvement). Result: Agent learns to search strategically, verify relevance, stop when confident.
RL fine-tuning workflow (simplified)
PHASE 1: COLLECT DATA (Week 1-2)
Deploy raw agent, track interactions: ├─ Question: [Customer question] ├─ Agent decision: [What agent searched for] ├─ Result: [What agent found] ├─ Answer: [What agent told customer] ├─ Outcome: [Customer satisfied? Yes/No] ├─ Feedback: [Why yes/no?] └─ Repeat for 1,000+ interactions
Example data:
├─ Q: "How long shipping to São Paulo?" ├─ Agent searched: "shipping" (too generic) ├─ Result: Found 47 documents (overwhelmed) ├─ Answer: "Shipping varies" (vague) ├─ Outcome: Customer unsatisfied ├─ Reason: Agent should search "shipping São Paulo" (specific)
├─ Q: "Do you accept Pix?" ├─ Agent searched: "payment methods" (good) ├─ Result: Found 3 relevant documents ├─ Answer: "Yes, we accept Pix" (correct) ├─ Outcome: Customer satisfied ├─ Reason: Agent used good search strategy
PHASE 2: LABEL GOOD vs BAD (Week 3-4)
For each interaction, label what agent should have done:
Bad interaction: ├─ Agent query: "shipping" (too generic) ├─ Better query: "shipping São Paulo" (specific) ├─ Reason: Adding location specificity → better results └─ Label: Bad decision (could be improved)
Good interaction: ├─ Agent query: "payment methods" (good) ├─ This query worked well ├─ Reason: Generic search was sufficient (payment methods universal) └─ Label: Good decision (no improvement needed)
Create training dataset: ├─ 1,000 interactions ├─ 300 labeled "good" (agent did right thing) ├─ 700 labeled "bad" (agent could improve) ├─ Goal: Agent learns from both (don't repeat bad, repeat good)
PHASE 3: RL TRAINING (Week 5-8)
Train agent using reinforcement learning:
Reward signal: ├─ If outcome = "satisfied": +1 reward ├─ If outcome = "unsatisfied": -1 reward ├─ Agent learns: "Decisions that lead to satisfaction are good"
What agent learns:
From good interactions (satisfied customers): ├─ "When customer mentions location, search includes location" ├─ "When searching generic (payment methods), it works" ├─ "When customer asks specific, search specific" └─ Implication: "Mirror specificity of question in search"
From bad interactions (unsatisfied customers): ├─ "Searching too generic leads to irrelevant results" ├─ "Hallucinating answers = bad outcome" ├─ "Picking first result without checking = often wrong" └─ Implication: "Be specific, verify relevance, think before answering"
Agent optimization: ├─ Agent learns to predict: "Which search query will get good results?" ├─ Agent learns to predict: "Does this result answer the question?" ├─ Agent learns to predict: "Should I search more or answer now?" └─ Result: Agent policy improves (makes better decisions)
Training curves: ├─ Week 1 (baseline): 60% success rate ├─ Week 2: 70% success rate ├─ Week 3: 78% success rate ├─ Week 4: 85% success rate └─ Result: +25-30% improvement (from bad to good)
PHASE 4: EVALUATE (Week 9)
Test trained agent on new data:
Before training: ├─ Agent success rate: 60% (30% failure rate) ├─ Customer satisfaction: 65% ├─ Error cost: R$500/month (from 10 errors × R$50 each)
After training: ├─ Agent success rate: 85% (15% failure rate) ├─ Customer satisfaction: 88% ├─ Error cost: R$75/month (from 1-2 errors × R$50 each)
Improvement: ├─ Success rate: +25 percentage points ├─ Satisfaction: +23 percentage points ├─ Cost savings: R$425/month (R$5,100/year) └─ ROI: Positive (training cost R$10-20K, saves R$5K+/year)
PHASE 5: ITERATE (Ongoing)
Deploy trained agent, collect new feedback:
Month 2 (post-deployment): ├─ Deploy trained agent ├─ Collect new interactions (feedback loop continues) ├─ New failures emerge (edge cases agent hasn't seen) ├─ Label new data (good/bad) └─ Retrain agent (incorporate new learning)
Continuous improvement: ├─ Month 1: 85% success ├─ Month 2: 87% success (retrain v2) ├─ Month 3: 89% success (retrain v3) ├─ Month 6: 92% success (retrain v6) ├─ Year 1: 95%+ success (near-human performance) └─ Process: Never stops (agents improve over time)
COST vs BENEFIT:
Cost: ├─ RL training setup: R$10-20K (one-time) ├─ Labeling data: R$5-10K (one-time) ├─ Retraining monthly: R$2-5K/month ├─ Total Year 1: R$50-75K
Benefit: ├─ Error cost reduction: R$5-10K/month saved ├─ Improved satisfaction: +20-30% (retention value) ├─ Reputation improvement: +10% customer lifetime value ├─ Total Year 1: R$150-300K
ROI: ├─ Year 1: 2-4x return on investment ├─ Year 2+: 5-10x return (costs amortized) └─ Conclusion: RL training pays for itself quickly
Why Raw Agents Fail (And Why Training Fixes It)
Raw LLM = Optimized for general text generation (not task-specific behavior). Fine-tuned agent = Optimized for YOUR task (search strategy, verification, stopping criteria). Difference = fundamental. Raw agents hallucinate because they have no "ground truth" feedback. Trained agents verify because they learned from successful examples. Strategy: You must train agents (raw is not an option).
Why raw agents fail (technical reasons)
FAILURE MODE 1: Over-broad searches
Scenario: Customer asks "How much shipping to São Paulo?"
Raw agent: ├─ Agent thinks: "I need to search for shipping information" ├─ Decision: "Search for: 'shipping'" (too generic) ├─ Results: 100+ documents (all shipping-related) ├─ Problem: Agent overwhelmed by results ├─ Decision paralysis: "Which document is about São Paulo?" ├─ Action: Picks first result (often wrong) ├─ Outcome: Wrong answer (generic shipping, not São Paulo-specific) └─ Failure: 50% of over-broad searches fail
Trained agent: ├─ Agent learned: "If location mentioned, include in search" ├─ Decision: "Search for: 'shipping São Paulo'" (specific) ├─ Results: 5 documents (all São Paulo shipping) ├─ Confidence: High (right documents found) ├─ Decision: Easy (pick best result) ├─ Outcome: Correct answer (São Paulo-specific) └─ Success: 95% of specific searches succeed
FAILURE MODE 2: False confidence
Scenario: Customer asks "Do you refund return shipping costs?"
Raw agent: ├─ Agent searches and finds: "Refund policy: full refund within 30 days" ├─ Agent thinks: "Found answer, I'm done" ├─ Agent generates: "We offer full refunds within 30 days" (incomplete) ├─ Problem: Full refund might NOT include return shipping (policy says "product only") ├─ Agent didn't check if result answered EXACT question ├─ Outcome: Wrong answer (customer thinks return shipping refunded, it's not) └─ Failure: Agent answers without verifying relevance
Trained agent: ├─ Agent learned: "Always verify result matches question exactly" ├─ Agent searches and finds: "Refund policy: full refund within 30 days" ├─ Agent checks: "Does this say anything about RETURN SHIPPING?" ├─ Result: No mention of return shipping (relevance check fails) ├─ Action: Search again for "refund return shipping" (more specific) ├─ Better result: "Return shipping: NOT refunded (customer pays)" ├─ Agent generates: "We offer full refunds within 30 days, but you pay return shipping" ├─ Outcome: Correct answer └─ Success: Agent verified before answering
FAILURE MODE 3: Early stopping (incomplete search)
Scenario: Customer asks "What's your policy on damaged items?"
Raw agent: ├─ Agent searches for "damage policy" ├─ Finds one result: "We replace damaged items within 30 days" ├─ Agent thinks: "Found answer, stop searching" ├─ Generates: "We replace damaged items within 30 days" ├─ Problem: Policy also mentions "photos required" and "shipping paid by customer" (not mentioned) ├─ Agent stopped too early (didn't search for full policy) ├─ Outcome: Incomplete answer (customer doesn't know about photo requirement) └─ Failure: Customer later complains about undisclosed photo requirement
Trained agent: ├─ Agent learned: "One search is rarely complete, search multiple angles" ├─ Agent searches #1: "damaged items policy" → finds replacement info ├─ Agent searches #2: "damage documentation" → finds photo requirement ├─ Agent searches #3: "damage shipping cost" → finds "customer pays" ├─ Agent learns to synthesize: "Replace damaged items within 30 days, with photos provided, customer pays shipping" ├─ Generates: "We replace damaged items within 30 days. You'll need to provide photos. Customer pays return shipping." ├─ Outcome: Complete answer (customer has all info) └─ Success: Agent did multi-step search
FAILURE MODE 4: Hallucination
Scenario: Customer asks "Do you accept cryptocurrency?"
Raw agent: ├─ Agent searches for "cryptocurrency" or "bitcoin" or "crypto" ├─ Gets no results (policy doesn't mention crypto) ├─ Agent thinks: "No results, but I should answer anyway" ├─ Generates hallucination: "We're exploring cryptocurrency options" (NEVER SAID) ├─ Outcome: Customer tries to pay with Bitcoin, gets rejected └─ Failure: Agent fabricated false information
Trained agent: ├─ Agent learned: "If no results found, don't hallucinate, tell customer you don't know" ├─ Agent searches for "cryptocurrency" ├─ Gets no results ├─ Agent learned: "Zero results = ask for clarification or escalate" ├─ Generates: "I couldn't find information about cryptocurrency payments. Let me connect you with our payment team." ├─ Outcome: Honest, escalates appropriately └─ Success: No hallucination, proper escalation
SUMMARY:
Raw agents fail because: ├─ Search too broad (overwhelmed by results) ├─ Don't verify relevance (false confidence) ├─ Stop searching too early (incomplete info) ├─ Hallucinate when unsure (make up answers) └─ No feedback loop (same mistakes repeated)
Trained agents succeed because: ├─ Search strategically (narrow, targeted) ├─ Verify every answer (check relevance) ├─ Search multi-step (find complete info) ├─ Escalate when uncertain (don't hallucinate) └─ Improve continuously (feedback loop)
Implementing RL Fine-Tuning: The Practical Path
Option 1: DIY with SageMaker (build in-house, use Amazon's RL framework). Option 2: Use third-party platform (Hugging Face, Modal, specialized RL providers). Option 3: Hire consultant (SageMaker partner, AI consulting firm). Choose based on: team skill level, budget, timeline. Most founders: Start with Option 2 or 3 (faster, cheaper than DIY).
Implementation paths for RL fine-tuning
OPTION 1: DIY WITH SAGEMAKER (Hardcore)
Requirements: ├─ ML engineer (experienced with RL) ├─ Access to AWS SageMaker ├─ 2-3 months timeline ├─ Budget: R$50-100K (engineering time)
Process: ├─ Week 1-2: Set up SageMaker environment, infrastructure ├─ Week 3-4: Data collection, labeling pipeline ├─ Week 5-8: RL training setup, baseline model ├─ Week 9-10: Fine-tuning, hyperparameter optimization ├─ Week 11-12: Evaluation, deployment └─ Result: Fully customized RL pipeline (maximum control)
Pros: ├─ Full control over training process ├─ Optimized for your exact use case ├─ Scalable (can run many experiments) └─ Cost-effective long-term (no per-call fees)
Cons: ├─ High upfront ML expertise required ├─ Requires experienced engineer (expensive) ├─ 2-3 month timeline (slow) ├─ Ongoing maintenance burden
OPTION 2: THIRD-PARTY PLATFORM (Balanced)
Examples: ├─ Hugging Face (open-source, managed) ├─ Modal (serverless RL) ├─ Anthropic Fine-tuning API (if using Claude) ├─ Specialized RL platforms (Antml, Contextual AI)
Requirements: ├─ Junior ML engineer (for setup/monitoring) ├─ Budget: R$10-30K/month (platform + compute) ├─ 2-4 weeks timeline
Process: ├─ Week 1: Sign up, integrate platform ├─ Week 2: Upload data, configure training ├─ Week 3-4: Run training (platform handles heavy lifting) ├─ Result: Trained agent (minimal engineering)
Pros: ├─ Lower ML expertise required ├─ Faster (weeks, not months) ├─ Less infrastructure burden ├─ Built-in best practices
Cons: ├─ Monthly recurring cost ├─ Less customization (limited to platform) ├─ Potential vendor lock-in └─ Limited control over internals
OPTION 3: HIRE CONSULTANT (Easy)
Examples: ├─ SageMaker partner agencies ├─ AI consulting firms (focusing on RL) ├─ Freelance ML engineers
Requirements: ├─ Budget: R$30-80K (project-based) ├─ Timeline: 2-3 months ├─ Your role: Provide data, feedback
Process: ├─ Week 1: Kickoff, data collection plan ├─ Week 2-6: Consultant builds solution ├─ Week 7-8: Testing, deployment ├─ Week 9: Handoff, training └─ Result: Trained agent + documented process
Pros: ├─ Minimal internal ML expertise required ├─ Experts handle everything ├─ Faster than DIY ├─ Get knowledge transfer
Cons: ├─ Expensive (R$30-80K) ├─ Depends on consultant quality ├─ Less control over process └─ May need consultant for updates
RECOMMENDATION BY COMPANY SIZE:
Early-stage startup (0-50 people): ├─ Path: Option 2 or 3 ├─ Why: No in-house ML team yet ├─ Timeline: 2-4 weeks ├─ Budget: R$10-30K (platform) or R$30-50K (consultant) ├─ Focus: Get working solution quickly, don't over-engineer
Growing SaaS (50-200 people): ├─ Path: Option 2 (primary) + hire junior ML eng (secondary) ├─ Why: Start with platform (fast), eventually build in-house ├─ Timeline: 4 weeks (platform) → hire ML eng → move to DIY over time ├─ Budget: R$20-40K/month (platform) + R$120-180K/year (ML eng) ├─ Focus: Scalable solution, continuous improvement
Enterprise (200+ people): ├─ Path: Option 1 (DIY SageMaker) ├─ Why: Have ML team, want full control, scale matters ├─ Timeline: 3 months ├─ Budget: R$200-300K (ML eng + infrastructure) ├─ Focus: Fully optimized, competitive advantage
Next Steps: Build Your RL Training Pipeline (Before Competitors Do)
At OpenClaw, we help SaaS founders implement RL fine-tuning for search/support agents: assess current agent performance (baseline), design RL training pipeline (data collection → labeling → training), implement training (platform setup), evaluate improvements (measure before/after), scale to production (deploy trained agent), iterate continuously (monthly retraining). We've built RL pipelines for 5 companies—average result: 30-45% improvement in agent success rate, -50% error costs, +20% customer satisfaction, 6-month payback period.
Get a free search agent audit: Schedule 30 minutes with our agent architect. We'll evaluate your current search agent (success rate? error rate?), identify failure modes (why is it failing?), quantify cost (how much do errors cost?), design RL training plan (what would improvement look like?), estimate effort/timeline (weeks vs months?), calculate ROI (how much would training save?), and create implementation roadmap (step-by-step path). Most founders realize their search agents fail 30-50% of the time (acceptable failure rate = 5-10%). RL fine-tuning is the path to reliability.
[Book your free audit] → [Button: Schedule 30-Minute Call]
Amazon research signals: Raw search agents fail 30-50% of queries (unacceptable). RL fine-tuning improves to 85-95% success (reliable). Strategy shift from deployment → training. Action required: Assess current agent performance (baseline), collect failure data (feedback loop), label good vs bad behaviors (training data), implement RL training (SageMaker or third-party), evaluate improvements (measure after), iterate monthly (continuous improvement). Timeline: 2-4 weeks (platform) to 3 months (DIY). Cost: R$10-30K (platform) to R$50-100K (DIY). Benefit: +30-50% success rate = R$50K-300K/year in saved errors + improved retention. ROI: 1-3 years (positive). Window: Competitors will adopt RL training in 2026 (move now = first-mover advantage). Non-action cost: Raw agents remain unreliable (-R$100K-500K/year in error costs + customer churn). Decision: Train your agents NOW (become reliable) or deploy raw agents (remain unreliable + lose customers to trained competitors). RL is 2026 baseline, not luxury feature.
FAQ
Q: Quanto tempo leva pra treinar um agent? Rápido? (Training timeline)
A: Depende do caminho. Platform = 2-4 semanas. DIY = 3 meses.
Timeline: ├─ Data collection: 1-2 semanas ├─ Labeling: 1-2 semanas ├─ Training: 1 semana (platform) a 4 semanas (DIY) ├─ Evaluation: 1 semana └─ Total: 4 semanas (platform) a 12 semanas (DIY)
Q: Preciso de data scientist pra fazer? Tenho que contratar? (Expertise needed)
A: Não se usar platform. Sim se DIY SageMaker.
Expertise: ├─ Platform approach: Junior engineer (configuração básica) ├─ DIY approach: ML engineer (experiência com RL) └─ Recomendação: Comece com platform (não precisa contratar)
Q: Quanto custa treinar? Pode ser caro? (Training cost)
A: Varia. Platform = R$10-30K/mês. DIY = R$50-100K upfront.
Costo: ├─ Platform (Hugging Face): R$10-20K/mês ├─ DIY SageMaker: R$50-100K (setup) + R$5K/mês (compute) ├─ Consultant: R$30-80K (project) └─ ROI: Positivo em 6-12 meses (errors avoided >> training cost)
Publicado em 2 de outubro de 2026