Notícias
Notícias
5 min de leitura
28 de setembro de 2026

Seu agent tá overfitado. Por isso quebra em casos novos.

TabPFN vence XGBoost (sem treinar). Seu agent tá treinado demais. Zero-shot = flexível. Como estruturar agent pra generalizar?

Equipe OpenClaw

Equipe OpenClaw · Time de Engenharia & Produto

A Equipe OpenClaw é formada por engenheiros, designers e especialistas em IA dedicados a construir a melhor plataforma de agentes conversacionais para negócios brasileiros. Combinamos expertise…


Seu agent tá overfitado. Por isso quebra em casos novos.

Você é founder de SaaS.

Seu SaaS tem agent no WhatsApp (atendimento ao cliente).

You trained the agent:

Training process: ├─ Collected 500 real customer chats (from your company) ├─ Labeled them (correct/incorrect responses) ├─ Fine-tuned model on your data ("learn your style") ├─ Tested on your customers (works great: 95% accuracy) └─ Deployed to production (celebrating success)

But then:

Customer A (same as training): Agent responds perfectly Customer B (slightly different): Agent confused (weird edge case) Customer C (new scenario): Agent hallucinates (completely wrong) Customer D (new industry): Agent breaks (can't answer)

Realization: Agent works only on cases I trained it on. New problem: Every new case = new training loop. Cost: R$ 50K/month on continuous training (to keep up). Frustration: "Why doesn't it generalize?"

You think: "Fine-tuning makes models smarter."

Or: "Custom training = competitive advantage."

Or: "My data is unique. Generic models won't work."

Then you read research (2026):

Headline: "TabPFN and TabICL vs. tuned XGBoost: the model that doesn't train won 14/14" │ What happened: ├─ Experiment: TabPFN (foundation model, ZERO training) vs XGBoost (custom trained) ├─ Contest: 14 different datasets (real-world tabular data) ├─ Result: TabPFN won 14/14 (100% win rate) ├─ The kicker: TabPFN never trained on any of these datasets ├─ Implication: Foundation models > custom training (sometimes) │

The Problem: Custom Training = Overfitting Risk

What Overfitting Means

In plain English:

Overfitting = memorizing examples instead of learning rules

Example: ├─ Training on: "How do I cancel my subscription?" ├─ Agent learns: "If customer says 'cancel', respond: 'Here's how...'" ├─ But: Customer says "unsubscribe" (different word, same intent) ├─ Agent response: "I don't understand" ├─ Problem: Agent memorized specific phrase, not the concept └─ Result: Brittle (breaks on variations)

In your agent:

What you're training: ├─ 500 customer chats from YOUR company ├─ YOUR specific language (how your customers talk) ├─ YOUR specific problems (unique to your industry) ├─ YOUR specific solutions (custom to your business) └─ Result: Model optimized for YOUR data

What the model learns: ├─ "This is how OUR customers communicate" ├─ "This is how WE solve problems" ├─ "Anything different = probably wrong" └─ Result: Confident but fragile

When new customer arrives (slightly different): ├─ Different phrasing ├─ Different context ├─ Different problem ├─ Agent thinks: "Doesn't match training data" ├─ Agent response: Hallucinates or declines └─ Result: Failure

The TabPFN Lesson: Foundation > Fine-Tuning

What TabPFN did:

TabPFN model: ├─ Pre-trained on 1,000+ tabular datasets (diverse) ├─ Learned general patterns (features, distributions) ├─ Never saw any of the test datasets ├─ Zero fine-tuning (no custom training) └─ Still won all 14 benchmarks

XGBoost (the old way): ├─ Custom-trained on each dataset ├─ Optimized for that specific data ├─ Took hours/days to train ├─ Memorized the patterns └─ Still lost to untrained foundation model

Why TabPFN won: ├─ Foundation model learned GENERAL rules ├─ General rules apply to new data too ├─ No overfitting (because no specific training) ├─ Flexible (handles variations automatically) └─ Faster (no training = instant)

What this means for your agent:

Your approach (training-heavy): ├─ Fine-tune on 500 chats (YOUR company) ├─ Works great on YOUR patterns ├─ Breaks on NEW patterns ├─ Requires retraining (expensive, slow) └─ Fragile (overfitted)

Foundation model approach: ├─ Use pre-trained model (trained on 1M+ customer chats from everywhere) ├─ Works ok on YOUR patterns (already generalized) ├─ Still works on NEW patterns (wasn't memorizing yours) ├─ No retraining (stays flexible) └─ Robust (handles variations)

Trade-off: ├─ Lose: Hyper-optimization (won't be 98% accurate on YOUR specific cases) ├─ Gain: Robustness (85% accurate on anything) └─ Net: Better real-world performance (handles novelty)

Real Examples: Overfitting in Action

Example 1: Pricing Questions (E-commerce Agent)

Overfitting approach (custom-trained):

Training data included: ├─ "What's the price of the blue shirt?" ├─ "How much is the red dress?" ├─ "Do you have discounts?" └─ (500 similar questions)

Agent learned: ├─ "If customer asks price → query product database" ├─ "If customer asks discount → apply rule X" └─ Pattern: "PRODUCT KEYWORD + PRICE ACTION"

What breaks: ├─ Customer: "Is the blue shirt still R$ 50?" ├─ Agent: Confused (doesn't match training pattern "what's the price") ├─ Response: Hallucinates (might say wrong price) └─ Issue: Memorized exact phrasing, not the concept

Cost of overfitting: ├─ Wrong answer (customer angry) ├─ Needs retraining (R$ 10K, 1 week) ├─ Downtime (agent broken while retraining) └─ Churn (customers see errors, leave)

Foundation model approach (zero-shot):

No training needed: ├─ Model already trained on 1M+ price/product questions (from everywhere) ├─ Learned: "INTENT: price inquiry, ACTION: look up product, RESPOND: price + context" ├─ General rule (not your specific data) └─ Flexible (applies to any variation)

What works: ├─ Customer: "What's the price of the blue shirt?" ├─ Agent: Responds correctly (learned general rule) ├─ Customer: "Is the blue shirt still R$ 50?" ├─ Agent: Responds correctly (generalizes to variation) ├─ Customer: "Can I get the blue shirt for less?" ├─ Agent: Responds correctly (understands intent despite phrasing) └─ Result: Handles variations (not overfitted)

Benefit of foundation model: ├─ No training (zero cost) ├─ No retraining (always works on new variations) ├─ Handles novelty (generalizes automatically) └─ Happy customers (consistent accuracy)

Example 2: Support Escalation (SaaS Support Agent)

Overfitting approach:

Training data: ├─ 500 support tickets from YOUR customers ├─ YOUR product (specific features) ├─ YOUR issues (specific bugs) ├─ YOUR escalation rules (specific departments) └─ Agent learns: "ONLY these issues matter"

What breaks: ├─ New customer: "How do I integrate with Stripe?" ├─ Agent: "Not in training data, probably can't help" ├─ Response: Escalates unnecessarily (takes human time) │ ├─ Another customer: "Your API is broken" ├─ Agent: "Doesn't match 'bug report' pattern" ├─ Response: Misclassifies (wrong department) └─ Issue: Overfitted to YOUR specific language

Cost: ├─ Unnecessary escalations (human team overwhelmed) ├─ Misrouted tickets (customer waits longer) ├─ Retraining cycles (continuous fine-tuning) └─ Churn (customers frustrated)

Foundation model approach:

No training needed: ├─ Model trained on 1M+ support tickets (from every SaaS) ├─ Learned: "INTENT: integration question, ACTION: provide docs, ESCALATE: if custom needed" ├─ General rules (apply to any SaaS) └─ Flexible (works on variations)

What works: ├─ New customer: "How do I integrate with Stripe?" ├─ Agent: Recognizes integration pattern (learned from general training) ├─ Response: Provides integration guidance (doesn't need YOUR training) │ ├─ Another customer: "Your API is broken" ├─ Agent: Recognizes bug report pattern ├─ Response: Routes to right team (routes correctly) └─ Result: Handles novelty (not overfitted)

Benefit: ├─ No unnecessary escalations (humans free) ├─ Correct routing (customers satisfied) ├─ No retraining (always works) └─ Happy team (predictable volume)

How to Avoid Overfitting: The Right Architecture

Strategy 1: Use Foundation Models, Don't Fine-Tune

Current approach (fine-tuning trap):

Architecture: ├─ Collect 500 chats from YOUR company ├─ Fine-tune base model on your data ├─ Deploy custom model to production └─ Monitor for failures (will happen)

Problem: ├─ Model memorized YOUR data ├─ Breaks on anything different ├─ Requires continuous retraining └─ Cost: High, maintenance: High

Better approach (foundation-first):

Architecture: ├─ Use pre-trained foundation model (GPT-4, Claude, etc) ├─ No fine-tuning (keep it generic) ├─ Add YOUR context via prompting (not training) ├─ Deploy unchanged model to production └─ Monitor for issues (fewer than custom-trained)

Benefit: ├─ Model already trained on 1M+ examples (generalizes) ├─ Handles variations (not overfitted) ├─ No retraining (always works on new data) ├─ Cost: Lower, maintenance: Lower └─ Robust: Works on YOUR data + new variations

Example: ├─ Don't do: Fine-tune GPT-4 on YOUR 500 chats ├─ Do: Use GPT-4 as-is, add YOUR context in system prompt ├─ System prompt: "You are support agent for [company]. Rules: [...]." ├─ Result: Model generalizes because no overfitting └─ Cost: R$ 5K/month (no training, just inference)

Strategy 2: Prompt Engineering > Fine-Tuning

Fine-tuning approach (overfitting):

Cost: R$ 50K/month ├─ Training: R$ 20K ├─ Retraining: R$ 20K (every month, new data) ├─ Monitoring: R$ 10K └─ Overtime: R$ 5K (fixing failures)

Process: ├─ Month 1: Train on 500 chats ├─ Month 2: Model breaks on new variations ├─ Month 2: Collect 100 new chats ├─ Month 3: Retrain on 600 chats (old + new) ├─ Month 3: Model still breaks (infinite loop) └─ Result: Constant firefighting

Prompt engineering approach (no overfitting):

Cost: R$ 5K/month ├─ Model inference: R$ 3K ├─ Monitoring: R$ 1K ├─ Prompt refinement: R$ 1K └─ No training (zero overhead)

Process: ├─ Week 1: Write system prompt ("You are support agent for [company]") ├─ Week 2: Test on customer chats (works ok) ├─ Week 3: Refine prompt (add edge cases, rules) ├─ Week 4: Deploy (model unchanged, prompt improved) ├─ Month 2-12: Iterate on prompt (no retraining) └─ Result: Stable, evolving, no catastrophic failures

Example prompt (much cheaper than fine-tuning):

You are a support agent for [Company].

Your rules:

  • If customer asks about refunds: [company refund policy]
  • If customer asks about integration: [integration docs link]
  • If you're unsure: Escalate to [team name]

Your personality:

  • Be helpful, not pushy
  • Use Portuguese (for Brazil)
  • Be honest about limitations

Cost: R$ 0 training (just text prompt) Flexibility: Update anytime (no retraining cycle) Robustness: Model generalizes (not overfitted)

Strategy 3: Hybrid Approach (If You Must Train)

If you need customization (some fine-tuning is ok):

Limited fine-tuning: ├─ DO fine-tune on: Domain-specific vocabulary (your product names, terms) ├─ DO NOT fine-tune on: General reasoning, conversation patterns ├─ DO use: Foundation model for core logic, custom tokens for YOUR domain └─ Result: Customized but not overfitted

Example: ├─ Foundation model: GPT-4 (generalizes well) ├─ Custom tokens: [PRODUCT_NAME], [COMPANY_FEATURE], [INDUSTRY_TERM] ├─ Fine-tune only: Token embeddings (not core weights) ├─ Result: Knows YOUR terms but still generalizes to new scenarios └─ Cost: R$ 10-20K one-time (much less than R$ 50K/month)

Decision Tree: Should You Fine-Tune?

Do you have 10,000+ labeled examples? ├─ YES → Fine-tuning might help (but usually foundation is better) └─ NO → Use foundation model + prompting (recommended)

Is your domain completely unique? ├─ YES → Light fine-tuning OK (domain vocabulary only) └─ NO → Foundation model sufficient

Do you have budget for continuous retraining? ├─ YES → Fine-tuning sustainable └─ NO → Foundation model only (zero ongoing cost)

Do you need to handle edge cases not in training data? ├─ YES → Foundation model better (generalizes) └─ NO → Fine-tuning OK (narrow use case)

Final decision: └─ Most founders: USE FOUNDATION MODEL + PROMPTING └─ Reason: Better ROI, less maintenance, handles novelty

Cost Comparison: Fine-Tuning vs Foundation Model

12-month cost analysis:

Fine-tuning approach: ├─ Initial training: R$ 20K ├─ Monthly retraining: R$ 20K × 12 = R$ 240K ├─ Monitoring: R$ 5K × 12 = R$ 60K ├─ Fixing failures: R$ 10K × 12 = R$ 120K └─ TOTAL YEAR 1: R$ 440K

Foundation model approach: ├─ Model inference: R$ 3K × 12 = R$ 36K ├─ Monitoring: R$ 1K × 12 = R$ 12K ├─ Prompt refinement: R$ 1K × 12 = R$ 12K └─ TOTAL YEAR 1: R$ 60K

Savings: R$ 380K (86% reduction) Quality: Better (foundation model handles novelty) Speed: Faster (no training cycles) Maintenance: Easier (just prompt updates)

ROI: Immediate (break-even in month 1)

Next Steps: Audit Your Agent for Overfitting

At OpenClaw, we help founders evaluate whether their agent is overfitted:

  • Overfitting risk assessment (is fine-tuning hurting you?)
  • Foundation model vs custom analysis (which is right for you?)
  • Prompt engineering optimization (get better results without training)
  • Migration planning (how to move from custom-trained to foundation models)
  • Cost analysis (how much are you wasting on retraining?)

Get a free agent architecture audit: Schedule 30 minutes with our AI architect. We'll analyze your current agent (fine-tuned or not), measure overfitting risk, compare costs (fine-tuning vs foundation model), and recommend whether to continue custom training or switch to foundation models.

[Book your free agent architecture audit] → [Button: Schedule Now]


FAQ

Q: If I use foundation model without fine-tuning, won't it give generic answers?

A: No. Foundation models (GPT-4, Claude) already trained on 1M+ customer service scenarios. They know how to be helpful, specific, and accurate. You add YOUR context via system prompts ("Our refund policy is..."). Result: Specific answers without overfitting.

Q: Doesn't fine-tuning make agents smarter for my specific use case?

A: It does, for cases in training data. But it breaks on variations (overfitting). You gain 95% accuracy on known cases, lose 60% on unknown cases. Foundation model gives 85% on everything. Real-world wins for foundation model (handles surprise questions).

Q: How do I prevent overfitting if I do fine-tune?

A: Use techniques: early stopping (stop training before memorization), regularization (penalize overfitting), validation set (test on unseen data), and continuous monitoring. But honestly? Use foundation model instead. Prevention is easier than cure.

Q: Can I combine fine-tuning with foundation models?

A: Yes, but carefully. Fine-tune only on domain vocabulary/tokens, not core reasoning. Or use LoRA (Low-Rank Adaptation) to fine-tune 1% of weights, not 100%. This keeps generalization while adding customization. Best of both worlds, but complex.


Publicado em 28 de setembro de 2026

Leia também