Seu agente tá baseado em hype? (arXiv diz sim)
arXiv: 6x AI papers em 2 anos (flooding). Qualidade? Ruim. Seu agente tá built on hype? Como separar innovation de trend vazio.
Equipe OpenClaw · Time de Engenharia & Produto
A Equipe OpenClaw é formada por engenheiros, designers e especialistas em IA dedicados a construir a melhor plataforma de agentes conversacionais para negócios brasileiros. Combinamos expertise…
Seu agente tá baseado em hype? (arXiv diz sim)
Notícia: arXiv (o maior repositório de papers científicos do mundo) CAPPED submissions a 2 por mês (por pessoa). Razão: AI paper flooding. Dados: Submissions dobraram em 2 anos. CS.AI category cresceu 6x. Resultado: Moderadores voluntários overwhelmed. Quality? Ruins (muitos papers são low-quality hype).
Implicação: arXiv had to BLOCK submissions. That's unprecedented. Why? Because researchers (e startups) tão SPAMMING arXiv com AI papers que são bullshit. Resultado: Real innovation is buried under 1000s of low-quality papers. Implicação: Seu agente? Provavelmente baseado em papers que são HYPE, não innovation.
Problema: Você tá construindo agente:
TYPICAL STARTUP LOGIC (WRONG): ├─ Year 1: "Vamos usar AI para [problem]" ├─ Research: Read AI papers (seems promising) ├─ Papers say: "This technique is amazing!" (50% of papers say this) ├─ Decision: "Let's build on this paper!" (excited) ├─ Year 2: Build agente (took 18 months) ├─ Launch: "This is revolutionary AI!" (marketing says) ├─ Reality: Feature barely works (paper was overhyped) ├─ Competition: Already solved this 6 months ago (simpler way) ├─ Result: Your "innovative" agente is outdated before launch └─ Lesson: Papers told you hype, not truth
WHAT ARXIV FLOODING MEANS: ├─ Baseline: 30,000 papers/year (pre-2024) ├─ Now: 60,000 papers/year (2x in 2 years) ├─ CS.AI category: 5,000 → 30,000 papers (6x!!!) ├─ Quality: Dropped dramatically (volume ≠ quality) ├─ Signal-to-noise: Terrible (finding good papers = hard) ├─ Impact: Researchers can't keep up (drowning in papers) ├─ Implication: Most AI papers = hype (not real innovation) └─ Your decision: Built on hyped papers = wrong direction
HYPE PAPERS vs REAL INNOVATION: ├─ Hype papers: │ ├─ Claim: "50% improvement!" (sounds amazing) │ ├─ Reality: Only works in lab on toy dataset │ ├─ Production: Fails when real data used │ ├─ Reason: Evaluated on unrealistic conditions │ ├─ Quote: "We tested on ImageNet" (but real world ≠ ImageNet) │ ├─ Adoption: Months later, technique is forgotten │ └─ Example: "This agente is 90% accurate!" (but only on demo) │ └─ Real innovation papers: ├─ Claim: "2% improvement, but reliable" (realistic) ├─ Reality: Works in production, tested on real data ├─ Adoption: Years later, still used (became standard) ├─ Reason: Solved actual problem (not lab problem) ├─ Quote: "We tested on 100M real customer conversations" ├─ Impact: Simplicity + reliability (not flashy) └─ Example: "Our agente works 95% of the time on real data"
HOW TO SPOT HYPE: ├─ Red flag 1: "Amazing results!" (vague language) ├─ Red flag 2: Tested on toy dataset (not real conditions) ├─ Red flag 3: No reproducibility (code not open-source) ├─ Red flag 4: No comparison to simple baseline ├─ Red flag 5: Published 6 months ago, now forgotten ├─ Red flag 6: No real-world case studies ├─ Red flag 7: Doesn't explain why it works ├─ Red flag 8: Same author has 10 papers this month (flooding) └─ Red flag 9: Claims to "solve" problem (nothing solves anything)
WHAT ARXIV CAP MEANS FOR YOU: ├─ Before cap: Researchers published anything (even bullshit) ├─ Now (Oct 2026): Max 2 papers/month (filters out junk) ├─ Implication: Future papers = higher quality (fewer noise) ├─ But: Your agente is built on OLD papers (pre-cap, lots of hype) ├─ Action: Audit your technical foundation (is it on hype?) ├─ Risk: Your agente might be based on paper that's now "known to be wrong" └─ Opportunity: Switch to simpler, more reliable approach
Entender: Hype vs Innovation
The hype cycle
THE AI HYPE CYCLE (typically 18-24 months): ├─ Month 0: Paper published │ ├─ Claim: "Revolutionary new technique!" │ ├─ Media: "This will change everything!" │ ├─ Startups: "We need to build on this!" │ ├─ Investors: "This is the next big thing!" │ └─ Reality: Unknown │ ├─ Month 3-6: Hype peak │ ├─ Action: 10s of startups building on it │ ├─ Papers: 100s of derivative papers (extend original) │ ├─ Conferences: "All papers use this technique" │ ├─ Assumption: "Everyone is using it, must be good" │ └─ Reality: Still untested at scale │ ├─ Month 6-12: Reality check │ ├─ Problem: Technique doesn't work in production │ ├─ Issue: Doesn't scale, too slow, too expensive │ ├─ Realization: Simpler approach was always better │ ├─ Startups: "Our product isn't working" (shock) │ └─ Papers: Suddenly critical papers published │ ├─ Month 12-18: Decline │ ├─ Adoption: Startups pivot away from technique │ ├─ Papers: Now say "technique has limitations" │ ├─ Media: "The hype is overblown" (opposite of month 0) │ ├─ Reality: Technique is useful, but only for specific cases │ └─ Startups: Lost 18 months + $500K building on hype │ └─ Month 18+: Maturity ├─ Understanding: Now we know what it's actually good for ├─ Usage: Only applied to right problems (not everything) ├─ Papers: Now realistic about limitations ├─ Startups: Those who survived are competitive └─ Lesson: Don't chase hype, chase fundamentals
TIMELINE OF ACTUAL TECHNOLOGIES: ├─ Transformers (2017): Hyped → Now standard (success) ├─ GPT-2 (2019): "Scary!" → Useful (success) ├─ BERT fine-tuning (2018): Hyped → Still used (success) ├─ Few-shot learning (2020): "Will replace all learning!" → Only works sometimes (partial) ├─ Reinforcement learning from human feedback (2022): Hyped → Works for LLMs (success) ├─ Chain of thought prompting (2022): "Will solve reasoning!" → Helps but limited (partial) ├─ Retrieval-augmented generation (2023): "Game changer!" → Actually useful (success) ├─ AI agents (2024): "Will replace humans!" → Still failing 50% (hype peak) ├─ Multimodal models (2023): Hyped → Useful (success) └─ Constitutional AI (2023): "Perfect safety solution!" → Helps but limited (partial)
THE PATTERN: ├─ 50% of new techniques: Oversold initially, useful later ├─ 30% of new techniques: Never solve the problem they promised ├─ 20% of new techniques: Actually revolutionary (rare) ├─ Cost of being wrong: 18 months + $500K-2M (if you bet on hype) ├─ Cost of being late: 6 months behind (but on solid ground) └─ Lesson: Slow down, don't chase hype
Why papers are hyped
WHY RESEARCHERS PUBLISH HYPE: ├─ Incentive 1: Career (need papers for PhD, tenure) │ ├─ Pressure: "Publish or perish" (culture) │ ├─ Strategy: Make results sound amazing (better chance of acceptance) │ ├─ Reality: "Incremental improvement" sounds bad, "breakthrough" sounds good │ ├─ Result: Researchers oversell their work │ └─ Implication: Most papers are overstated │ ├─ Incentive 2: Attention (need citations, media coverage) │ ├─ Strategy: Make paper sound revolutionary │ ├─ Media: "New AI breakthrough!" (clickbait works) │ ├─ Citations: Hyped papers get more citations (everyone cites them) │ ├─ Result: Good for researcher, bad for field │ └─ Implication: Exciting papers ≠ True papers │ ├─ Incentive 3: Funding (need grants for research) │ ├─ Pitch: "This technique will change the world!" │ ├─ Funding: Grants to researchers, not to those who say "meh" │ ├─ Result: Researchers oversell potential │ └─ Implication: Papers are marketing pitches │ ├─ Incentive 4: Startup funding (researchers turn into founders) │ ├─ Pitch: "Our paper + startup = unicorn!" │ ├─ Funding: VCs invest if potential seems huge │ ├─ Result: Paper gets oversold to raise funds │ └─ Implication: Academic credibility ≠ Product potential │ └─ ARXIV FLOODING REASON: ├─ All these incentives + easy publishing = EXPLOSION of papers ├─ Researchers: "More papers = more impact" (false, but believed) ├─ Startups: "Need papers to prove we're AI company" (trend) ├─ Result: 30,000+ papers/year (most are low-quality) └─ arXiv: "We need to stop this (cap submissions)"
TYPES OF HYPE PAPERS: ├─ Type 1: "Works in lab, not in production" │ ├─ Example: "90% accuracy on ImageNet!" (but slow, expensive) │ ├─ Reality: 50% accuracy in production (trade-offs not explained) │ ├─ Lesson: Evaluate on real constraints (latency, cost, etc) │ └─ Your agente: Don't build on lab results, test at scale │ ├─ Type 2: "Improvement on toy dataset" │ ├─ Example: "BERT fine-tuning: 2% improvement on small dataset" │ ├─ Reality: 0.1% improvement on real dataset │ ├─ Lesson: Evaluate on realistic data distribution │ └─ Your agente: Don't trust papers that only show toy results │ ├─ Type 3: "Compares to old baseline" │ ├─ Example: "Our technique is 30% better than 2015 method!" │ ├─ Reality: Compared to simple baseline, not state-of-art │ ├─ Lesson: Compare to current best method │ └─ Your agente: Demand comparison to latest approach │ ├─ Type 4: "No open-source code" │ ├─ Example: "Amazing results!" (but code not available) │ ├─ Reality: Results not reproducible (might be wrong) │ ├─ Lesson: Can't trust results you can't verify │ └─ Your agente: Only trust papers with open-source code │ └─ Type 5: "Forgotten 6 months later" ├─ Example: Paper published, hyped, then... silence ├─ Reality: Technique didn't work, researchers moved on ├─ Lesson: If paper is from 2 years ago and no one uses it, probably hype └─ Your agente: Check if technique is still used (adoption = real)
Como escolher papers fundamentados (não hype)
Strategy 1: Check adoption, not hype
IDEIA: ├─ Ignore: What paper claims (marketing) ├─ Check: Who is actually using it (truth) ├─ Logic: If technique works, people adopt it ├─ If technique is hype: Everyone forgets it in 6 months └─ Lesson: Adoption = reality check
IMPLEMENTATION: ├─ Step 1: Identify candidate technique │ ├─ From: Paper that sounds promising │ └─ Next: Check if people actually use it │ ├─ Step 2: Search for adoption │ ├─ Search: GitHub for open-source implementations │ ├─ Search: Papers citing this paper (2-3 years later) │ ├─ Search: Companies using this technique (Google, Meta, etc) │ ├─ Search: Conferences mentioning this technique (current talks) │ └─ Result: If nothing found = hype (not adopted) │ ├─ Step 3: Evaluate adoption quality │ ├─ Check: GitHub stars (1000+ = real adoption) │ ├─ Check: Papers citing (100+ citations = influence) │ ├─ Check: Companies using (public case studies) │ ├─ Check: Recent papers still using it (vs forgotten) │ └─ Result: High adoption = likely real innovation │ ├─ Step 4: Validate with production data │ ├─ Test: Technique on your real data │ ├─ Compare: To simple baseline (not paper's best case) │ ├─ Measure: Latency, cost, accuracy together (not just accuracy) │ ├─ Real-world: Does it actually help your agente? │ └─ Result: If helps = fundamentally sound, if not = hype │ └─ Timeline: 2 weeks of research + 1 week of testing
EXAMPLE: ├─ Paper: "Revolutionary attention mechanism!" (2023) ├─ Claim: "50% better than transformers!" ├─ Check adoption: │ ├─ GitHub: 10 implementations, 500 stars = some interest │ ├─ Papers: 50 citations = decent influence │ ├─ Companies: Meta using in research (not production) │ ├─ Conferences 2024: Still mentioned (positive sign) │ ├─ Conferences 2025: Rarely mentioned (hype fading?) │ └─ Conclusion: Interesting, but not proven at scale │ ├─ Test yourself: │ ├─ Implement mechanism on your agente │ ├─ Test on toy data: 45% improvement (matches paper) │ ├─ Test on production data: 5% improvement (very different!) │ ├─ Latency cost: 2x slower than transformer │ ├─ Tradeoff: "Not worth 5% improvement for 2x slowness" │ └─ Conclusion: Paper was hype (lab conditions ≠ production) │ └─ Decision: Skip this paper, use standard transformer
OPPOSITE EXAMPLE: ├─ Paper: "Retrieval-augmented generation" (2020) ├─ Claim: "Reduce hallucinations by 30%" ├─ Check adoption: │ ├─ GitHub: 5000+ implementations, 10000+ stars = massive │ ├─ Papers: 1000+ citations = huge influence │ ├─ Companies: Google, Meta, OpenAI using in production │ ├─ Conferences 2024: Still heavily mentioned │ ├─ Conferences 2025: Still in use (not forgotten) │ └─ Conclusion: This is the real deal │ ├─ Test yourself: │ ├─ Implement RAG on your agente │ ├─ Test on toy data: 30% improvement (matches paper) │ ├─ Test on production data: 28% improvement (still good!) │ ├─ Latency cost: 50ms slower (acceptable) │ ├─ Tradeoff: "28% improvement for 50ms cost = worth it" │ └─ Conclusion: Paper was accurate (still works in production) │ └─ Decision: Use RAG in your agente (real innovation)
Strategy 2: Demand reproducibility
IDEIA: ├─ Real innovation: Code available, results reproducible ├─ Hype papers: Results only in paper, can't verify ├─ Logic: If you can't reproduce results, probably wrong └─ Action: Only trust papers with open-source code
IMPLEMENTATION: ├─ Red flag 1: No code available │ ├─ Paper says: "Code available on request" (don't trust) │ ├─ Reality: Authors don't want you checking their work │ ├─ Implication: Results might be wrong │ └─ Decision: Skip this paper │ ├─ Red flag 2: Code is messy/undocumented │ ├─ Code quality: Matters (bad code = bad research) │ ├─ Check: Is code well-written? (or spaghetti?) │ ├─ Check: Are there tests? (or just raw code?) │ ├─ Implication: Bad code = probably bad results │ └─ Decision: Caution (maybe still hype) │ ├─ Red flag 3: Code is hard to run │ ├─ Check: Can you run it in <30 min? (or need PhD?) │ ├─ Implication: Authors don't care about reproduction │ ├─ Reality: If they wanted you to verify, they'd make it easy │ └─ Decision: Skip │ ├─ Green flag 1: Code is on GitHub, easy to run │ ├─ Sign: Authors want you to verify (confidence) │ ├─ Reality: If results are wrong, they'd hide code │ ├─ Implication: Authors confident in results │ └─ Decision: Trust this more │ ├─ Green flag 2: Code has tests, CI/CD │ ├─ Sign: Professional researchers (care about quality) │ ├─ Reality: Tests = they're serious │ ├─ Implication: Results probably real │ └─ Decision: Trust this │ └─ Green flag 3: Reproduce results yourself ├─ Action: Run their code, verify they get same results ├─ If you get: Different results = paper might be wrong ├─ If you get: Same results = paper is probably real ├─ Effort: 1-2 hours └─ Decision: Only trust papers you can reproduce
EXAMPLE: ├─ Paper A: "Amazing agente!" (code not available) │ ├─ Trust: Low (can't verify) │ ├─ Decision: Skip (too risky) │ └─ Lesson: Hype papers hide code │ ├─ Paper B: "Good agente" (code on GitHub, easy to run) │ ├─ Test: Clone repo, run example in 10 minutes │ ├─ Results: Get same numbers as paper (verified!) │ ├─ Trust: High (reproduction successful) │ └─ Decision: Consider for your agente │ └─ Conclusion: Reproducibility = trustworthiness
Strategy 3: Focus on fundamentals, not novelty
IDEIA: ├─ Hype: "New technique!" (trendy) ├─ Fundamentals: "Simple, reliable approach" (boring) ├─ Reality: Fundamentals win in production ├─ Lesson: Choose boring over trendy └─ Timeline: 18 months later, you'll be right
IMPLEMENTATION: ├─ When choosing technique: │ ├─ Evaluate: Is it simple? (or complex?) │ ├─ Evaluate: Is it reliable? (or fragile?) │ ├─ Evaluate: Is it understanding? (or magic?) │ ├─ Evaluate: Is it battle-tested? (or brand new?) │ ├─ Evaluate: Is it slow/expensive? (or efficient?) │ └─ Score: If >3/5, probably good (if <3/5, probably hype) │ ├─ Simple examples: │ ├─ Boring (fundamentals): Regular expressions for parsing │ │ ├─ Pro: Fast, reliable, understand exactly │ │ ├─ Con: Limited to simple cases │ │ └─ When to use: Simple text matching │ │ │ ├─ Trendy (hype): Neural network for parsing │ │ ├─ Pro: Works for complex cases │ │ ├─ Con: Slow, unpredictable, black-box │ │ └─ When to use: When regex fails │ │ │ └─ Best: Start with regex, use neural if needed │ ├─ Lesson: Use simplest tool that works │ ├─ Don't use neural network for simple problem │ └─ Result: Faster, cheaper, more reliable │ ├─ For agente design: │ ├─ Boring: LLM + prompt engineering (simple, works) │ ├─ Trendy: Fine-tuning + RAG + LoRA + agents (complex) │ ├─ Reality: Boring approach works 80% of the time │ ├─ Reality: Trendy approach works 85% (not worth complexity) │ └─ Lesson: Use boring first, add complexity only if needed │ └─ Decision framework: ├─ "Is this new technique necessary?" ├─ If yes: Use it (might be real innovation) ├─ If maybe: Skip it (not worth the risk) ├─ If no: Definitely skip (why add complexity?) └─ Result: Avoid 90% of hype
EXAMPLE: ├─ Decision: Should we use agents (trendy, 2024) or simple LLM (boring)? │ ├─ Evaluate agents: │ ├─ Simple? No (complex framework, many components) │ ├─ Reliable? No (fail 50% of time, unpredictable) │ ├─ Understanding? No (black box, hard to debug) │ ├─ Battle-tested? No (new, only 2 years old) │ ├─ Efficient? No (slow, expensive, multiple LLM calls) │ └─ Score: 0/5 (SKIP) │ ├─ Evaluate simple LLM: │ ├─ Simple? Yes (prompt + API call) │ ├─ Reliable? Yes (works 80% of time, predictable) │ ├─ Understanding? Yes (you control prompt, know behavior) │ ├─ Battle-tested? Yes (3+ years, billions of calls) │ ├─ Efficient? Yes (1 LLM call, fast, cheap) │ └─ Score: 5/5 (USE THIS) │ └─ Decision: Use simple LLM (boring wins, agents are hype)
Conclusão: arXiv cap = reality check
Fatos:
✓ arXiv capped submissions (unprecedented) ✓ Reason: 6x AI papers in 2 years (flooding with hype) ✓ Quality: Terrible (many low-quality papers) ✓ Your agente: Probably built on hyped papers (if modern) ✓ Risk: Agente based on hype = fails in production ✓ Solution 1: Check adoption, not hype (what people actually use?) ✓ Solution 2: Demand reproducibility (open-source code) ✓ Solution 3: Focus on fundamentals (simple > trendy) ✓ Timeline: Hype cycle = 18 months (then everyone knows) ✓ Cost of hype: 18 months + $500K-2M (if you bet wrong) ✓ Cost of boring: 6 months behind (but on solid ground) ✓ Reality: Fundamentals win in production (every time)
ACTION ITEMS (THIS WEEK):
- TODAY: Audit agente technical foundation (based on what papers?)
- TODAY: For each paper: Check GitHub adoption (500+ stars?)
- TODAY: For each paper: Check citations 2024-2025 (still used?)
- WEEK 1: Verify reproducibility (can you run their code?)
- WEEK 1: Test on your real data (does it actually work?)
- WEEK 1: Evaluate complexity vs benefit (worth the hype?)
- WEEK 2: Identify non-essential components (remove hype)
- WEEK 2: Simplify agente (fewer trendy techniques)
- WEEK 3: Re-test (simpler version works better?)
- ONGOING: Only adopt papers with 2+ years of adoption
Problema resolvido quando: └─ Agente foundation: Solid (not on hype) └─ Technologies used: Battle-tested (not brand new) └─ Complexity: Only when necessary (not trendy) └─ Reliability: High (not fragile) └─ Cost: Low (not paying for hype) └─ Production: Works well (not overcomplicated) └─ Timeline: On schedule (not chasing hype └─ Result: Agente is fundamentally sound
→ OpenClaw: Agentes Fundamentados (Não Hype) + Paper Evaluation
arXiv capping submissions: 6x AI papers, qualidade ruim. Seu agente? Pode estar built on hype. Audit agente foundation. Check GitHub adoption (não papers). Demand reproducibility. Teste em real data. Foco em fundamentals (não trends). Simples > Trendy = Production wins. 📊
Publicado em 11 de outubro de 2026