Notícias
Notícias
5 min de leitura
11 de outubro de 2026

Seu agente tá sendo bloqueado como bot? (Yandex insight)

Yandex: Massive bot traffic (Cloudflare data). Seu agente? Pode estar sendo bloqueado. Como passar em detecção anti-bot legitimamente.

Equipe OpenClaw

Equipe OpenClaw · Time de Engenharia & Produto

A Equipe OpenClaw é formada por engenheiros, designers e especialistas em IA dedicados a construir a melhor plataforma de agentes conversacionais para negócios brasileiros. Combinamos expertise…


Seu agente tá sendo bloqueado como bot? (Yandex insight)

Notícia: Cloudflare publicou dados: Yandex MASSIVAMENTE bloqueado como bot traffic. Dados: https://radar.cloudflare.com/bots/as13238 (458 pts trending, 173 comments).

Significado: Cloudflare (major CDN/security provider) está FLAGGING Yandex como "malicious bot". Implicação: Se Yandex (search engine) tá sendo bloqueado, QUALQUER agente de IA também pode estar.

Problema: Seu agente tá tentando acessar APIs, websites, ou scrapar dados. Website vê traffic vindo do seu agente (IP ou user-agent). Sistema anti-bot: "This looks like bot, BLOCK IT." Seu agente: Morre. Customer: "Why didn't agente fetch data?" Você: "System blocked us."

Implicação: Você TEM que entender: Como agentes contornam detecção anti-bot (legitimamente). Como fazer seu agente parecer "humano" (sem enganar). Como respeitar robots.txt e policies (while still functioning).

Problema: Agentes são bloqueados constantemente:

WHY AGENTS GET BLOCKED: ├─ Reason 1: User-Agent header identifies as bot │ ├─ Your agent sends: "User-Agent: OpenAI-Bot/1.0" │ ├─ Website sees: "This is a bot, block it" │ ├─ Result: 403 Forbidden (blocked) │ ├─ Why: Bot header triggers anti-bot rules │ └─ Solution: Spoof user-agent? (sketchy, avoid) │ ├─ Reason 2: Request pattern looks automated │ ├─ Your agent sends: 1,000 requests/minute │ ├─ Normal human: 5 requests/minute │ ├─ Website detects: "This is bot behavior, block IP" │ ├─ Result: IP blocked (all your agents blocked) │ ├─ Why: Rate pattern is inhuman │ └─ Solution: Slow down requests (rate limit yourself) │ ├─ Reason 3: No cookies/session management │ ├─ Your agent sends: Each request is stateless │ ├─ Normal human: Browser maintains session (cookies) │ ├─ Website detects: "No session = bot" │ ├─ Result: Blocked (treated as malicious) │ ├─ Why: Bot behavior = no session │ └─ Solution: Implement cookie jar, session handling │ ├─ Reason 4: Browser behavior missing │ ├─ Your agent sends: Plain HTTP requests │ ├─ Normal human: JavaScript executed, CSS loaded, images shown │ ├─ Website detects: "No JavaScript = bot" │ ├─ Result: Blocked (treated as bot) │ ├─ Why: Bots often skip JavaScript │ └─ Solution: Use headless browser (Selenium, Puppeteer) │ ├─ Reason 5: Cloudflare/WAF rules │ ├─ Your agent hits: Cloudflare protected site │ ├─ Cloudflare sees: "Bot-like behavior" │ ├─ Cloudflare shows: CAPTCHA (or blocks) │ ├─ Your agent: Can't solve CAPTCHA (fails) │ ├─ Result: Blocked from site │ └─ Solution: Use residential proxy, or accept CAPTCHA │ └─ Reason 6: IP reputation ├─ Your agent IP: Marked as "datacenter/cloud" ├─ Website sees: "This IP runs bots, block it" ├─ Result: Blocked (IP blacklist) ├─ Why: Datacenter IPs are common for bots └─ Solution: Use residential proxy (looks like real user)

YANDEX CASE (from Cloudflare data): ├─ Yandex traffic: Heavily flagged as bot ├─ Why: Yandex search crawler (legitimate bot, but...) ├─ But: Sites might block Yandex (Russian search, geopolitical) ├─ Or: Yandex bot pattern looks "too automated" (rate, user-agent, etc) ├─ Implication: Even legitimate search engines get blocked ├─ Lesson: If Yandex gets blocked, so can your agent └─ Takeaway: You MUST implement proper bot behavior (or get blocked)

YOUR AGENT SCENARIO: ├─ Your agent tries to fetch: Customer data from 3rd party API ├─ API site: Protected by Cloudflare ├─ Your agent sends: Request with bot-like headers ├─ Cloudflare detects: "This looks like bot (Yandex-like pattern)" ├─ Cloudflare shows: CAPTCHA (or blocks) ├─ Your agent: Fails (can't solve CAPTCHA, blocked) ├─ Customer: "Agent didn't fetch data?" ├─ You: "Site is blocking us as bot" ├─ Customer: "Why is my agente broken?" ├─ You: "Need to use residential proxy or implement proper bot behavior" └─ Lesson: Bot detection blocks legitimate agents too


Entender: Como detecção anti-bot funciona

The bot detection stack

WHAT IS BOT DETECTION? ├─ Definition: Systems that identify automated traffic (bots) ├─ Goal: Block malicious bots (hackers, scrapers, DDoS) ├─ Problem: Also blocks legitimate bots (search engines, agents) ├─ Tools: Cloudflare, Akamai, AWS WAF, custom rules └─ Sophistication: From simple (user-agent check) to complex (ML)

THE DETECTION PYRAMID: ├─ Level 1: User-Agent header │ ├─ Check: Does header say "bot"? │ ├─ If yes: Block immediately │ ├─ False positives: None (explicit bot declaration) │ └─ Difficulty to bypass: Trivial (spoof header) │ ├─ Level 2: Rate limiting │ ├─ Check: Requests per second/minute │ ├─ If high: Likely bot (slow down = human-like) │ ├─ Threshold: 10-100 req/sec (varies by site) │ ├─ False positives: Yes (legitimate traffic spikes) │ └─ Difficulty to bypass: Easy (throttle requests) │ ├─ Level 3: Session/Cookie behavior │ ├─ Check: Does client maintain cookies? │ ├─ If no: Likely bot (humans keep sessions) │ ├─ False positives: Some (privacy-focused browsers) │ └─ Difficulty to bypass: Medium (implement cookie jar) │ ├─ Level 4: JavaScript execution │ ├─ Check: Does client execute JS? │ ├─ If no: Likely bot (humans have JS engine) │ ├─ False positives: Some (JS disabled in browser) │ └─ Difficulty to bypass: Medium-Hard (use headless browser) │ ├─ Level 5: Browser fingerprinting │ ├─ Check: Is client a real browser? │ ├─ Tests: Canvas, WebGL, fonts, plugins, timing │ ├─ If fails: Likely bot (bots don't have real browser) │ ├─ False positives: Rare (very accurate) │ └─ Difficulty to bypass: Hard (need real browser) │ ├─ Level 6: Behavioral analysis (ML) │ ├─ Check: Does behavior match human pattern? │ ├─ Tests: Mouse movement, scroll patterns, click timing │ ├─ If fails: Likely bot (behavior too perfect/robotic) │ ├─ False positives: Very rare (ML trained) │ └─ Difficulty to bypass: Very hard (need realistic behavior) │ └─ Level 7: IP reputation + geolocation ├─ Check: Is IP from datacenter or residential? ├─ If datacenter: Likely bot (bots run on cloud) ├─ False positives: Some (VPN users, corporate networks) └─ Difficulty to bypass: Hard (need residential proxy)

CLOUDFLARE'S APPROACH: ├─ They use: All levels above + proprietary ML ├─ They track: Bot patterns (Yandex, Googlebot, etc) ├─ They allow: Known good bots (Google, Bing search) ├─ They block: Unknown/suspicious bots ├─ They challenge: CAPTCHA for suspicious requests ├─ Yandex case: Flagged as bot (not on whitelist? Or pattern suspicious?) └─ Your agent: Will likely be challenged (unknown user-agent)

HOW YOUR AGENT GETS BLOCKED: ├─ Step 1: Agent sends HTTP request │ ├─ Headers: No User-Agent or bot-like header │ ├─ Cookies: None (stateless request) │ ├─ JavaScript: Not executed │ └─ IP: Datacenter IP (AWS, GCP, Azure) │ ├─ Step 2: Website/CDN analyzes request │ ├─ Check 1: "No cookies = bot?" │ ├─ Check 2: "Missing headers = bot?" │ ├─ Check 3: "Datacenter IP = bot?" │ ├─ Check 4: "No JavaScript = bot?" │ └─ Result: Multiple signals = "Probably bot" │ ├─ Step 3: Decision │ ├─ If high confidence bot: Block (403) │ ├─ If medium confidence bot: Challenge (CAPTCHA) │ ├─ If low confidence bot: Allow (but monitor) │ └─ Your agent: Blocked or CAPTCHA'd │ └─ Step 4: Your agent fails ├─ If blocked: 403 error (can't proceed) ├─ If CAPTCHA: Can't solve (bots can't do CAPTCHA) ├─ Fallback: Retry? (gets rate limited) └─ Result: Agent dies

Why this matters for your agent

SCENARIO: Your agent needs to fetch external data ├─ Task: Agent should fetch customer billing from SaaS provider ├─ Provider: Uses Cloudflare (common for SaaS) ├─ Your agent: Tries to call API ├─ Cloudflare: "This looks like bot, CAPTCHA" ├─ Your agent: "I can't solve CAPTCHA" ├─ Result: Agent fails to fetch data ├─ Customer: "Agent didn't work, it's broken" ├─ You: "Site is blocking us as bot" └─ Impact: Feature is broken (no workaround)

SCALE: ├─ 1 API call fails: Minor (try again) ├─ 10 API calls fail: Problems (some features don't work) ├─ 100s API calls fail: Major (most features broken) ├─ 1000s API calls fail: Disaster (product is unusable) └─ Risk: Even small % of blocked calls = noticeable failures

TIMING: ├─ 2024: Bot detection was simple (user-agent check) ├─ 2025: Started getting sophisticated (level 3-4) ├─ 2026: Now sophisticated (level 5-7, ML-based) ├─ 2027: Will be very sophisticated (near-perfect) └─ Implication: Your agent MUST be bot-detection-aware

LESSON: ├─ Yandex getting blocked: Proof bot detection is aggressive ├─ If search engine blocked: Your agent will be too ├─ Solution: Implement proper "bot behavior" ├─ Or: Use residential proxy (expensive but works) ├─ Or: Use official API (if available, best option) └─ Timeline: Fix NOW (before it becomes a problem)


Como contornar (legitimamente) detecção anti-bot

Strategy 1: Use official APIs (best option)

IDEIA: ├─ Don't hit websites with agent ├─ Use official API (if available) ├─ APIs are bot-friendly (expecting bots) ├─ Result: No bot detection (APIs allow robots) └─ Why: This is the intended method

IMPLEMENTATION: ├─ Step 1: Check if provider has API │ ├─ Example: Need customer data from Stripe │ ├─ Don't: Scrape stripe.com (will be blocked) │ ├─ Do: Use Stripe API (official, bot-friendly) │ └─ Result: Works perfectly │ ├─ Step 2: Get API credentials │ ├─ Auth: API key, OAuth, mTLS │ ├─ Rate limits: Usually generous (1000s req/sec) │ ├─ Reliability: Guaranteed (SLA) │ └─ Result: Scalable, reliable │ ├─ Step 3: Implement API client │ ├─ Library: Use official SDK (if available) │ ├─ Headers: Set proper User-Agent (describe your app) │ ├─ Auth: Include credentials (API key in header) │ ├─ Error handling: Respect rate limits, backoff │ └─ Result: Proper integration │ └─ Success: ├─ No bot detection: APIs expect automated requests ├─ Reliability: 99.9%+ uptime ├─ Scalability: Can handle 1000s requests/sec ├─ Cost: Transparent (per request or fixed) └─ Lesson: Always prefer official API

EXAMPLE: ├─ Provider: Shopify store ├─ Need: Get customer order history ├─ DON'T: Scrape shopify.com (Cloudflare blocks) ├─ DO: Use Shopify GraphQL API │ ├─ No bot detection (API expects bots) │ ├─ Reliable (99.9% SLA) │ ├─ Scalable (1000s req/sec) │ └─ Proper (OAuth authentication) │ └─ Result: Works perfectly, no detection issues

Strategy 2: Implement proper bot behavior

IDEIA: ├─ If no API available: Act like real browser ├─ Implement human-like behavior ├─ Respect robots.txt and rate limits ├─ Identify yourself properly ├─ Result: Not flagged as malicious bot └─ Why: Legitimate bots look like this

IMPLEMENTATION:

Level 1: User-Agent (basic, required) ├─ Set: User-Agent header (describe your bot) ├─ Example: "User-Agent: MyAgent/1.0 (+http://mysite.com/bot)" ├─ Why: Identifies as bot (but not suspicious) ├─ Include: Website URL (shows legitimacy) ├─ Include: Contact info (shows you're real) └─ Code: headers = {"User-Agent": "MyAgent/1.0 (+http://mysite.com/bot)"}

Level 2: Rate limiting (critical, required) ├─ Throttle: Requests to 1-5 req/sec (human pace) ├─ Why: Humans don't make 1000 req/sec ├─ Implement: Delay between requests (random 0.2-1sec) ├─ Respect: robots.txt (check Crawl-delay) ├─ Backoff: If 429 response, double wait time └─ Code: time.sleep(random.uniform(0.5, 1.5))

Level 3: Session/Cookies (medium, recommended) ├─ Maintain: Cookie jar (like browser) ├─ Why: Humans keep cookies across requests ├─ Implement: Store & resend Set-Cookie headers ├─ Library: requests.Session() (Python) handles this ├─ Benefit: Passes cookie check (not a simple bot) └─ Code: session = requests.Session()

Level 4: Headers (medium, recommended) ├─ Include: Common browser headers │ ├─ Referer: Set correctly (where you came from) │ ├─ Accept: "text/html,application/xhtml+xml,..." │ ├─ Accept-Language: "en-US,en;q=0.9" │ ├─ Cache-Control: "no-cache" │ └─ (Many more, like real browser) ├─ Why: Real browsers send these ├─ Tool: Use library that mimics browser (easier) └─ Benefit: Looks more human

Level 5: JavaScript (hard, optional) ├─ If site requires JS: Use headless browser ├─ Tools: Puppeteer (Node), Selenium (Python), Playwright ├─ Why: Real browsers execute JavaScript ├─ Impact: Can handle dynamic sites (loads faster) ├─ Cost: Slower (headless browser overhead) ├─ Benefit: Passes JS check (not a simple HTTP client) └─ Code: browser = await puppeteer.launch()

Level 6: Residential proxy (hard, optional) ├─ If IP blocked: Use residential proxy ├─ What: Real home internet IPs (not datacenter) ├─ Why: Datacenter IPs are blacklisted ├─ Cost: Expensive (R$ 100-1000/mês) ├─ Benefit: Bypasses IP reputation checks ├─ Note: Use only if legitimate bot └─ Caution: Expensive, use as last resort

Level 7: robots.txt respect (critical, required) ├─ Check: robots.txt first (what's allowed?) ├─ Example: GET /robots.txt ├─ Respect: User-agent rules ├─ Why: Ethical (shows you're legitimate bot) ├─ Legal: Following robots.txt = safe legally └─ Code:

 import requests
 from urllib.robotparser import RobotFileParser
 rp = RobotFileParser()
 rp.set_url("https://example.com/robots.txt")
 rp.read()
 allowed = rp.can_fetch("MyAgent/1.0", "/api/data")
 

Success checklist: ├─ User-Agent: Identified as bot (not malicious) ├─ Rate: 1-5 req/sec (human pace) ├─ Cookies: Maintained (like browser) ├─ Headers: Complete (like real browser) ├─ robots.txt: Respected (ethical) ├─ Behavior: Realistic (not too perfect) └─ Result: Passes bot detection (not flagged)

EXAMPLE CODE (Python): python import requests import time from urllib.robotparser import RobotFileParser

Setup

session = requests.Session() # Maintain cookies session.headers.update({ "User-Agent": "MyAgent/1.0 (+http://mysite.com/bot)", "Accept": "text/html,application/xhtml+xml", "Accept-Language": "en-US,en;q=0.9", })

Check robots.txt

rp = RobotFileParser() rp.set_url("https://example.com/robots.txt") rp.read()

Make request

if rp.can_fetch("MyAgent/1.0", "/api/data"): response = session.get("https://example.com/api/data") time.sleep(1) # Respect rate limits print(response.json()) else: print("Not allowed by robots.txt")

Strategy 3: Use proxy service

IDEIA: ├─ Can't modify request behavior ├─ Or: Site is too aggressive with blocking ├─ Use: Proxy service (handles bot detection) ├─ They: Deal with CAPTCHA, bot detection ├─ You: Get clean data back └─ Cost: Moderate (per request or monthly)

IMPLEMENTATION: ├─ Service: Use proxy provider │ ├─ Options: Bright Data, Oxylabs, Apify │ ├─ They offer: Browser automation + proxy │ ├─ Feature: Auto-solve CAPTCHA │ ├─ Feature: Residential IPs │ ├─ Feature: Rate limiting management │ └─ Cost: R$ 500-5000/mês (depends on usage) │ ├─ How it works: │ ├─ You: Send request to proxy service │ ├─ Proxy: Handles bot detection │ ├─ Proxy: Solves CAPTCHA (if needed) │ ├─ Proxy: Returns clean data │ └─ You: Get reliable results │ └─ When to use: ├─ If: API not available ├─ If: Site is very aggressive (Cloudflare, etc) ├─ If: Bot detection too sophisticated ├─ Benefit: Reliable (99.9% success) ├─ Tradeoff: Expensive + slower └─ Timeline: Last resort (when native fails)


Conclusão: Bot detection is a real problem

Fatos:

✓ Bot detection: Getting sophisticated (Level 5-7) ✓ Yandex blocked: Proves even search engines flagged ✓ Your agent: WILL be detected as bot (if not careful) ✓ Impact: Agent fails, features don't work ✓ Solution 1: Use official APIs (best) ✓ Solution 2: Implement proper bot behavior (good) ✓ Solution 3: Use proxy service (fallback) ✓ Timing: Fix NOW (before it becomes bigger problem) ✓ robots.txt: MUST respect (legal + ethical) ✓ Rate limiting: MUST implement (human-like pace) ✓ User-Agent: MUST identify (show you're legitimate) ✓ Cookies: Should maintain (like real browser) ✓ Test: Before deploying (verify bot detection) ✓ Monitor: Track blocked requests (know when it happens)

ACTION ITEMS (ASAP):

  1. TODAY: Audit agent requests (which APIs, which sites?)
  2. TODAY: Check if official APIs available (prefer these)
  3. WEEK 1: Implement proper User-Agent (describe bot)
  4. WEEK 1: Add rate limiting (1-5 req/sec, not 1000)
  5. WEEK 1: Maintain cookies (use session, not stateless)
  6. WEEK 2: Test with Cloudflare site (verify not blocked)
  7. WEEK 2: Check robots.txt compliance (ethical + legal)
  8. ONGOING: Monitor blocked requests (alert if increases)
  9. ONGOING: Switch to official APIs (when available)
  10. FALLBACK: Use proxy service (if bot detection too hard)

Problema resolvido quando: └─ Agent: Never blocked (passes bot detection) └─ Requests: Succeed 99%+ (not rate limited) └─ Behavior: Looks human-like (proper headers, cookies) └─ robots.txt: Respected (ethical + legal) └─ Monitoring: Alerts if blocked (know ASAP) └─ APIs: Used where available (best reliability) └─ Fallback: Proxy service ready (if needed) └─ Result: Agent is bot-detection-aware, works reliably

→ OpenClaw: Agentes com Bot Detection Awareness + Proxy Integration

Yandex massively bloqueado (Cloudflare data). Seu agente? Será bloqueado também. Use official APIs (best). Ou: Implemente proper bot behavior (User-Agent, rate limit, cookies). Respeite robots.txt (ético + legal). Teste antes de deploy. Proxy service como fallback (caro mas funciona). 🤖


Publicado em 11 de outubro de 2026

Leia também