Seu agent content tá sendo scraped. Google paga <0,1%. Você perde.
Google extracts your agent content (0.1% revenue share). AI scraping economy = you create, Google profits. Content value stolen.
Equipe OpenClaw · Time de Engenharia & Produto
A Equipe OpenClaw é formada por engenheiros, designers e especialistas em IA dedicados a construir a melhor plataforma de agentes conversacionais para negócios brasileiros. Combinamos expertise…
Seu agent content tá sendo scraped. Google paga <0,1%. Você perde.
Ontem você descobriu.
Google's AI Contribution Pilot (relatório novo).
What it shows: Quando Google (ou OpenAI, ou Anthropic) extrai conteúdo pra treinar seus modelos (ou pra usar em AI answers), eles pagam publishers <0.1% do ad revenue gerado.
Translation: Você cria conteúdo (suporte, FAQ, documentação). Google extrai sem pedir. Google lucra R$1 milhão. Você recebe R$500 (0.05% de R$1M).
Você criou, Google se apropriou.
Por que importa pra você: Seu agent tá gerando conteúdo (respostas, documentos, guias). Esse conteúdo tá sendo scraped por AI models (Google, OpenAI, Anthropic). Você não recebe um centavo. Eles lucram bilhões.
Você é founder.
Seu agent escreve 1M documentos/month (suporte, produto docs, FAQs).
Esse conteúdo é genuinamente útil (ajuda clientes).
Google Gemini vê esse conteúdo.
Google extrai (treina modelo com seu conteúdo).
Google lucra (Gemini responde perguntas usando seu conteúdo).
Você lucra: R$0.
Extraction economy = you lose.
The Signal: AI Extraction Economy Is Real (and Unfair)
Google's AI Contribution Pilot pays <0.1% ad revenue. This signals: AI companies extracting creator content without fair compensation. Your agent content is being scraped (trained into models, used in AI answers). You're creating value. They're capturing it. The extraction economy is broken.
Google's AI Contribution Pilot: The reality of extraction economics
GOOGLE AI EXTRACTION ECONOMY (How it works):
Step 1: You create content (your agent generates docs) ├─ Your agent: Writes FAQ, support docs, product guides ├─ Effort: Real (your LLM + infrastructure costs) ├─ Value: High (helps your customers) ├─ Monetization: None (you own it, customer benefits) └─ Cost to you: R$0.01 - R$0.10 per document (LLM API cost)
Step 2: Google crawls your site (Googlebot discovers content) ├─ Detection: Automatic (Google's web crawler) ├─ Permission: You didn't explicitly grant (your robots.txt allows it) ├─ Extraction: Google copies your content ├─ Your awareness: Zero (you don't know it happened) └─ Your consent: Not asked
Step 3: Google trains Gemini (your content trains their model) ├─ Training: Your content used to fine-tune Gemini ├─ Value: Your content makes Gemini better ├─ Your compensation: None ├─ Google's gain: Millions of documents improve model quality └─ Result: Gemini gets smarter, you get nothing
Step 4: Google uses your content (Gemini answers queries with your data) ├─ User asks: "How do I reset my password?" ├─ Gemini responds: Uses your FAQ content to answer ├─ User satisfied: Gets answer (from your content) ├─ Google profitable: User stays in Google ecosystem (Ads, products) ├─ You profitable: Nothing (user didn't visit your site) └─ Value transfer: Your content → Google profit
Step 5: Google optionally pays you (AI Contribution Pilot) ├─ Payment terms: <0.1% of ad revenue generated ├─ Calculation: User interaction = R$1 ad revenue ├─ Your share: R$0.0005 (0.05% of R$1) ├─ Your reaction: "That's it?" ├─ Google's position: "We're paying something (be grateful)" └─ Reality: You created value, Google captured 99.95%
EXAMPLE SCENARIO: Your agent generates support docs
Setup: ├─ Your company: SaaS with 10,000 customers ├─ Support docs: 1,000 articles (FAQ, troubleshooting, guides) ├─ Content quality: High (helps customers resolve issues) ├─ Google extraction: All 1,000 articles indexed + used in Gemini training └─ Impact: Gemini now answers your customer questions
Value creation (your side): ├─ Your effort: 1,000 articles × 1 hour each = 1,000 hours ├─ Your cost: 1,000 hours × R$200/hour = R$200,000 ├─ Value created: Customers satisfied (retention benefit) ├─ Monetization: Indirect (customer stays, pays subscription) └─ Total value: R$200,000 + customer retention value
Value extraction (Google's side): ├─ Google's effort: Zero (crawler automated) ├─ Google's cost: Negligible (already crawling for search) ├─ Value extracted: 1,000 articles train Gemini ├─ Gemini improvement: Better answers (attracts users) ├─ Google's monetization: More Gemini users → more ads → more revenue └─ Total value: R$100 million+ (Gemini ads, API sales)
Revenue split: ├─ You created: R$200,000+ in value ├─ You received: R$0 (without AI Contribution Pilot) ├─ You received (with pilot): R$100 - R$500 (0.05% - 0.25% of extraction value) ├─ Google extracted: R$100,000,000+ ├─ Google paid: R$100 - R$500 (0.0001% - 0.0005%) └─ Fairness: Absolutely not
WHY GOOGLE'S PILOT IS INSUFFICIENT:
Google's argument: "We're paying creators through AI Contribution Pilot" ├─ Payment model: <0.1% of ad revenue ├─ Calculation: Transparent (you can see the math) ├─ Participation: Opt-in (you can join the pilot) └─ Google's position: "We're being fair"
Creator reality: "This is essentially theft" ├─ Your content created value: R$200,000 minimum ├─ You received: R$200 maximum (even with pilot) ├─ Google's extraction: R$100,000,000+ value ├─ Creator extraction: 0.0001% (you get crumbs) ├─ Fairness assessment: 1% would be fair, 0.1% is an insult └─ Reality: Google structured this to extract value + appear generous
Your Agent Content: The New Extraction Target
Google's AI Contribution Pilot pays <0.1% of extraction value. Your agent is generating content (docs, responses, guides). That content is being extracted (trained into models, used in AI answers). You receive nothing (or near-nothing). This is the new extraction economy.
How your agent content is being scraped (and you don't even know it)
YOUR AGENT'S CONTENT PIPELINE (The extraction funnel):
Stage 1: Your agent generates content ├─ Agent type: Support bot (WhatsApp, email, website) ├─ Content generated: FAQ answers, troubleshooting guides, product docs ├─ Scale: 100K - 1M documents per month (depends on usage) ├─ Quality: High (AI-generated, but reviewed by your team) ├─ Placement: Your website, knowledge base, chat interface ├─ Visibility: Public (searchable, indexable by Google) └─ Problem: All content is extractable
Stage 2: Google discovers your content ├─ Method: Automated crawl (Googlebot visits your site) ├─ Frequency: Daily to weekly (depends on your traffic) ├─ Scope: Everything public (all agent-generated docs) ├─ Extraction: Content indexed in Google's database ├─ Your permission: Granted via robots.txt (by default) └─ Your awareness: Zero (it happens silently)
Stage 3: Content used for model training ├─ Training target: Gemini, Bard, and future Google models ├─ Purpose: Improve AI quality (your content makes models smarter) ├─ Scale: Millions of documents from creators like you ├─ Your consent: Not asked (Google claims "public content is fair use") ├─ Your compensation: None (yet, unless you join AI Contribution Pilot) └─ Google's value: Massive (billions in model improvement)
Stage 4: Content used in AI answers ├─ Use case: User asks question (similar to your FAQ) ├─ Gemini responds: Uses your content to generate answer ├─ User satisfied: Gets answer from your content (but thinks it's Google's) ├─ User traffic: Stays in Google (doesn't visit your site) ├─ Your benefit: Zero (no customer acquisition) ├─ Your loss: Content freely given to Google └─ Google's gain: User stays in ecosystem (Ads, services)
Stage 5: Optional compensation (if you opt in) ├─ Program: Google AI Contribution Pilot ├─ Requirement: You apply + Google approves ├─ Compensation: <0.1% of ad revenue your content generates ├─ Example: Your content drives R$1,000 in ads → you get R$0.50 - R$1.00 ├─ Payment method: AdSense integration (minimal friction) ├─ Transparency: You see attribution + earnings (appears fair) └─ Reality: Google still extracted 99.9% of value
SCALE OF EXTRACTION (Your agent content × millions of creators):
Individual creator extraction (you): ├─ Documents created: 1M per month (agent-generated) ├─ Google extraction: 100% of public documents ├─ Value created: R$200,000 - R$500,000/month ├─ Compensation (with pilot): R$200 - R$500/month (0.1% of value) ├─ Extraction rate: 99.9% unpaid └─ Your loss: R$199,500/month
Aggregated extraction (all creators × Google): ├─ Total creators: 1,000,000+ websites + content creators ├─ Avg extraction per creator: R$200,000/month ├─ Total value extracted: R$200 billion/month (estimated) ├─ Compensation paid by Google: R$200 million/month (0.1%) ├─ Google's net gain: R$199.8 billion/month ├─ Creator extraction rate: 99.9% unpaid globally └─ Implication: Google's entire AI economics built on extraction
WHY YOU SHOULD CARE (The economic impact on your SaaS):
Impact 1: Content as competitive advantage (eroded) ├─ Your strategy: Build knowledge base (agent-generated docs) ├─ Expected benefit: Organic traffic (SEO, content ranking) ├─ Reality: Google's AI answers your questions (users don't visit you) ├─ SEO impact: Traffic to your site decreases (Gemini steals it) ├─ Economic impact: -50% - 80% organic traffic loss (estimated) └─ Lesson: Content no longer drives traffic (AI answers do)
Impact 2: Margins compressed (extraction kills profitability) ├─ Your model: Content generation costs R$100/article ├─ Value capture: Customers visit, buy = R$1,000 per customer ├─ Reality: Google Gemini answers same question (customer doesn't visit) ├─ Value capture (actual): R$0 (customer leaves) ├─ Margin impact: Content generation cost = negative ROI └─ Lesson: Content generation costs you money (extraction profits Google)
Impact 3: Competitive moat destroyed (everyone has your content) ├─ Your strategy: Build proprietary knowledge base ├─ Moat: Unique documentation (hard for competitors to replicate) ├─ Reality: Google trained Gemini on your docs (published model quality) ├─ Competitor benefit: Copycats use Gemini to match your quality ├─ Moat erosion: Your documentation advantage → everyone's advantage └─ Lesson: Documentation doesn't create competitive advantage anymore
The Real Problem: Extraction Economics Are Fundamentally Broken
Google's AI Contribution Pilot pays <0.1% of extraction value. This signals deeper problem: AI extraction economy assumes creators should give away value (content, training data, labor) while AI companies capture 99%+. This math doesn't work. You can't build a profitable SaaS if your content is extracted without compensation.
How to protect your content (and your margins)
STRATEGY 1: Block extraction (at the source)
Tactic: robots.txt + legal notices ├─ robots.txt: Disallow Googlebot from crawling certain pages ├─ Legal: Add "noindex" meta tag to prevent indexing ├─ Notice: "This content may not be used for AI training" ├─ Effectiveness: Medium (Google crawlers respect robots.txt, but may ignore) ├─ Problem: Google may still crawl + train (they claim fair use) └─ Recommendation: Use, but don't rely on it
Implementation:
robots.txt
User-agent: Googlebot Disallow: /ai-knowledge-base/ # Block AI training extraction Disallow: /support/faq/ # Block FAQ extraction
HTML meta tag (in your agent docs)
Pros: ├─ No cost (just configuration) ├─ Easy to implement ├─ Shows intent (you're protecting content) └─ Some crawlers respect it (Anthropic's Claude respects noai tag)
Cons: ├─ Google may ignore it (they don't recognize "noai" tag) ├─ Doesn't protect existing extraction (already crawled) ├─ May reduce SEO (if you block all crawlers) └─ Compliance unclear (legal status uncertain)
Recommendation: Use + combine with other tactics.
STRATEGY 2: Paywall extraction (charge for access)
Tactic: Put your best content behind paywall ├─ Model: Public summaries (searchable) + paid details (protected) ├─ Benefit: Google crawls summary (SEO), can't access full content ├─ Extraction protection: Paywalled content not easily scraped ├─ Monetization: Selling content access directly └─ Trade-off: Lower visibility (less organic traffic)
Implementation:
Public: FAQ summary (300 words) - searchable, extractable Paywall: Full FAQ guide (5,000 words) - not extractable, monetized
Advantage: Google ranks your summary (brings traffic). Users want full version (convert to customers).
Pros: ├─ Protects your best content (not easily extracted) ├─ Direct monetization (users pay for content) ├─ Creates moat (exclusive content) └─ Sustainable (you control economics)
Cons: ├─ Reduces organic traffic (less content visible) ├─ Conversion friction (users must pay) ├─ Cannibalization (free vs paid confusion) └─ Complexity (manage two-tier content)
Recommendation: Use for high-value content (not all docs).
STRATEGY 3: License your content (controlled extraction)
Tactic: Offer licensing terms for AI training ├─ Model: Creators opt-in to licensing (explicit consent) ├─ Terms: AI companies pay fair rate to use your content ├─ Enforcement: Legal contract (creators have recourse) ├─ Benefit: You control extraction + get compensated └─ Reality: Few companies accept these terms
Implementation:
License terms:
- AI companies can use your content for training
- Requirement: Pay R$0.10 - R$1.00 per 1,000 tokens
- Attribution: Cite your company in model outputs
- Exclusivity: Competitor AI models can't use (competitive advantage)
Pros: ├─ Fair compensation (you set price) ├─ Explicit consent (legal protection) ├─ Competitive advantage (exclusive content) ├─ Sustainable (you control terms) └─ Scalable (license to multiple AI companies)
Cons: ├─ Adoption challenge (companies ignore licenses) ├─ Enforcement hard (legal battles expensive) ├─ Extraction continues (even without license) └─ Reality: Not effective against major AI companies
Recommendation: Use for enterprise negotiations (custom deals).
STRATEGY 4: Diversify (don't rely on organic content)
Tactic: Build value beyond content (extraction-proof moat) ├─ Model: Content attracts users, community keeps them ├─ Moat: Community (network effects) > Content (easily extractable) ├─ Benefit: Even if content extracted, community stays ├─ Reality: Extraction-proof strategy └─ Example: Discord community + private documentation
Implementation:
Public content: Basic FAQs, tutorials (extractable, fine) Private content: Advanced guides, community docs (members only)
Advantage: Gemini answers basic questions (good). Users want advanced knowledge (join community). Community is extraction-proof.
Pros: ├─ Extraction-resistant (private content protected) ├─ Community moat (network effects) ├─ Differentiated (community unique to you) ├─ Sustainable (community harder to replicate) └─ Customer loyalty (community members stay longer)
Cons: ├─ High effort (manage community) ├─ Slower growth (less viral) ├─ Lower visibility (some content private) └─ Complexity (two-tier strategy)
Recommendation: Best long-term strategy (combines everything).
STRATEGY 5: Demand fair compensation (advocacy)
Tactic: Join creator coalitions demanding fair AI revenue share ├─ Model: Collective action (creators negotiate together) ├─ Goal: Force AI companies to pay 5-10% extraction value (not 0.1%) ├─ Mechanism: Legal action, industry standards, regulation ├─ Benefit: Market correction (extraction becomes expensive for AI companies) └─ Reality: Slow change, but necessary
Implementation:
Join: Industry groups advocating for creator rights Examples: Authors Guild, Journalism Organizations, Creator Unions
Action: Sign letters, support lawsuits against AI companies Goal: Establish precedent (extraction requires compensation)
Timeline: 2-5 years (legal processes slow)
Pros: ├─ Long-term fix (systemic change) ├─ Fair outcome (creators compensated fairly) ├─ Legal protection (established precedent) ├─ Industry-wide benefit (all creators benefit) └─ Moral correctness (extraction becomes costly)
Cons: ├─ Slow (years, not months) ├─ Uncertain outcome (legal battles unpredictable) ├─ Doesn't protect your content today ├─ Requires collective action (coordination hard) └─ AI companies lobby against (well-funded)
Recommendation: Support while pursuing other tactics.
Next Steps: Protect Your Agent Content (and Your Economics)
At OpenClaw, we help SaaS founders protect their agent-generated content (prevent extraction without compensation), evaluate AI training exposure (what's being extracted?), design extraction-resistant content strategies (paywalls, community, licensing), and build sustainable content economics (content + community + exclusivity = moat):
- Content extraction audit (what of your content is being extracted by AI companies?)
- Extraction risk assessment (how much revenue are you losing to extraction?)
- Protection strategy design (robots.txt + licensing + community)
- Licensing negotiation (how to get paid for AI training data)
- Content economics modeling (sustainable margins with extraction happening)
Get a free agent content protection assessment: Schedule 30 minutes with our content economics strategist. We'll audit your public content (what's extractable?), quantify extraction impact (how much traffic/revenue lost to AI answers?), design multi-tier content strategy (public + paywall + community), evaluate licensing opportunities (AI companies willing to pay?), and create 90-day content protection roadmap (how to reduce extraction while maintaining SEO).
[Book your free assessment] → [Button: Schedule 30-Minute Call]
Google's AI Contribution Pilot pays <0.1% of extraction value. Your agent content is being extracted (trained into models, used in AI answers). You receive nothing (or near-nothing). This extraction economy is broken. You can't build profitable SaaS if your content is free labor for Google's AI. Protect your content now: block extraction (robots.txt), paywall best content, build community (extraction-proof moat), demand fair compensation (join advocacy groups). Don't let AI companies capture your content's value without paying. Your margins depend on it.
FAQ
Q: Mas se eu bloquear Google, meu site perde SEO rankings, certo? (Trade-off)
A: Não exatamente. Depende do que você bloqueia.
Opções:
- Bloquear apenas AI training crawlers (Googlebot-Extended): Site ainda aparece em search normalmente
- Bloquear all Googlebots: Você sai do search (bad idea)
- Bloquear specific paths (ex: /faq/): Only FAQ blocked, rest indexed normally
- Meta tag noai: Signals "don't use for training" (Google may respect)
Recommendação: Use selective blocking (protect key content, allow search indexing). Don't block all Google (you need search traffic).
Q: Quanto vale meu conteúdo realmente? Como calcular? (Valuation)
A: Três métodos:
-
Cost method: Quanto custou gerar o conteúdo?
- 1M docs × R$0.10 per doc = R$100,000 cost
- Fair compensation: 2-5x cost = R$200,000 - R$500,000 value
-
Revenue method: Quanto vale em terms of customer acquisition?
- 1M docs × 0.1% convert to customer = 1,000 customers
- 1,000 customers × R$1,000 lifetime value = R$1M value
- Google extraction: 50% of traffic loss = R$500,000 lost value
-
Market method: What are AI companies paying for training data?
- Market rate: R$0.001 - R$0.10 per 1,000 tokens
- Your content: 1B tokens = R$1,000 - R$100,000 fair value
Recommendação: Use revenue method (most accurate for your business). Google paying <0.1% = you're undercompensated by 1000x.
Q: Quais AI companies estão respeitando licenças de criador? (Compliance)
A: Honestidade: Poucas.
Companies respeitando (alguns):
- Anthropic: Respeita meta tag "noai" (alguns crawlers obedecem)
- OpenAI: Menos compliance (mais agressiva em extraction)
- Google: Minimal compliance (respects robots.txt only)
- Others: Varying (depends on company policy)
Reality: Self-regulation not working. Creators need legal protection (licenses with teeth, enforcement mechanism).
Recommendação: Assume non-compliance. Protect content defensively (block + paywall + community). Don't rely on AI company goodwill.
Publicado em 1 de outubro de 2026