Claude's new rules: Seu agente IA precisa de compliance
Anthropic bane fake sources em respostas Claude. Implicação: Seu agente WhatsApp/SaaS precisa governance agora. Era de prompt rápido acabou.
Equipe OpenClaw · Time de Engenharia & Produto
A Equipe OpenClaw é formada por engenheiros, designers e especialistas em IA dedicados a construir a melhor plataforma de agentes conversacionais para negócios brasileiros. Combinamos expertise…
Claude's new rules: Seu agente IA precisa de compliance
Notícia: Anthropic (criadora de Claude) atualizou sua usage policy: proíbe espeificamente seeding de fake sources em respostas Claude. Empresas podem NÃO mais:
- ✗ Inserir links falsos em respostas (pra influenciar Claude)
- ✗ Criar fontes fictícias ("estudos" que não existem)
- ✗ Treinar Claude com dados intencionalmente enganosos
- ✗ Automatizar publicação de conteúdo enganoso (pra depois Claude usar)
Implicação: Se você tem agente de IA (WhatsApp, suporte, vendas, integrado em produto SaaS), você precisa de governance + compliance agora. A era de "prompt engineering rápido, depois vê o que sai" acabou.
**"Você é CEO de fintech com agente WhatsApp que dá recomendações de investimento.
Cenário antigo (sem governance): ├─ Agente treinado com "dados de mercado" (alguns reais, alguns inventados) ├─ Agente recomenda: "Ação XYZ vai subir 50% (baseado em análise)" ├─ Cliente investe R$ 100K baseado em recomendação ├─ Ação sobe... 2% (não 50%) ├─ Cliente descobre: "Fonte que você citou não existe" ├─ Cliente processa você (fake sources, misleading) ├─ Custo: R$ 500K (lawsuit) + R$ 200K (regulação BACEN) ├─ Reputação: Destruída └─ Agente: Desligado
Cenário novo (com governance): ├─ Agente só usa FONTES REAIS (CVM, B3, Bloomberg) ├─ Agente recomenda: "Ação XYZ pode subir 2-5% (baseado em dados públicos)" ├─ Cliente investe R$ 100K (expectativas corretas) ├─ Ação sobe 3% ├─ Cliente satisfeito (expectativa vs reality = match) ├─ Zero problemas legal (tudo documentado, rastreável) ├─ Reputação: Protegida └─ Agente: Escalável com confiança "**
O problema: Fake sources em agentes IA
Por que isso é risco
Cenário real (como agentes caem em fake sources):
Caso 1: Agente de atendimento (e-commerce) ├─ Agente treinado com: "Nosso produto é o melhor (segundo pesquisa interna)" ├─ Pesquisa interna: Não existe (foi feita por um estagiário em Excel) ├─ Agente responde cliente: "97% de satisfação (fonte: pesquisa nossa)" ├─ Cliente depois descobre: "Essa pesquisa não é pública, como vocês sabem?" ├─ Resultado: "Vocês tão mentindo pra vender" └─ Dano: Reputação online (1 star reviews)
Caso 2: Agente de suporte (SaaS) ├─ Agente treinado com: "Feature X não é possível (segundo limitações do sistema)" ├─ Limitações: Outdated (sistema evoluiu, mas training data não) ├─ Agente responde: "Não conseguimos fazer integração com SAP" ├─ Cliente depois descobre: "Vocês conseguem, eu vi no roadmap" ├─ Resultado: "Seu agente me mentiu, não confio" └─ Dano: Churn customer
Caso 3: Agente de vendas (fintech) ├─ Agente treinado com: "Nossa taxa é 0.5% (segundo comparativo de 2023)" ├─ Comparativo: Outdated (concorrência lançou taxa mais baixa) ├─ Agente vende: "Nós somos os mais baratos" ├─ Cliente depois descobre: "Competitor tem 0.3%" ├─ Resultado: "Vocês enganaram a gente" ├─ Cliente pede chargeback (false advertising) └─ Dano: Custo legal + perda de customer
Por que Anthropic bane isso (3 razões):
-
LEGAL RISK ├─ Fake sources = misleading advertising ├─ Regulação (FTC, CVM, BACEN) vai investigar ├─ Liabilities: empresa de IA responsável por output └─ Anthropic quer evitar: "Claude foi usado pra enganar"
-
ACCURACY DEGRADATION ├─ Agente treinado com fake data = output degradado ├─ Fake source contém erro → Claude reproduz erro ├─ Erro amplificado (memes na internet, influença) └─ Claude's reputation: afetada
-
BUSINESS RISK (pra Anthropic) ├─ Enterprise clientes: "Podemos confiar em Claude?" ├─ Se Claude foi enganado: Não ├─ Anthropic perde confiança → perda de revenue └─ Policy update é defesa preemptiva
O impacto pra SU
Se seu agente usa fake sources:
├─ CURTO PRAZO (1-3 meses) │ ├─ Regulador avisa (warning letter) │ ├─ Platform (Claude, OpenAI) pode desligar sua conta │ ├─ Media picks up ("Startup enganava clientes com IA fake") │ └─ Dano: Reputação, brand trust │ ├─ MÉDIO PRAZO (3-6 meses) │ ├─ Lawsuit de cliente enganado │ ├─ Settlement: Pode chegar a R$ 1M+ │ ├─ Insurance: Pode negar cobertura (fraud não cobre) │ └─ Dano: Cash, legal distraction │ └─ LONGO PRAZO (6+ meses) ├─ Licensing: Pode ser banido de usar IA APIs ├─ Future funding: VC não investe em "IA fraud" company ├─ Enterprise sales: "Clientes não confiam em vocês" └─ Dano: Business existência
Solução: Governance framework pra agentes IA
Como implementar compliance
Princípio geral:
Sem governance (arriscado): ├─ Agente: "Diga o que quer ouvir" ├─ Fonte: Qualquer uma (real ou fake) ├─ Verificação: Nenhuma └─ Risco: Alto
Com governance (seguro): ├─ Agente: "Diga apenas o que sabemos ser verdade" ├─ Fonte: Apenas oficiais (public data) ├─ Verificação: Automática (antes de enviar) └─ Risco: Baixo
Framework em 4 pilares:
╔═══════════════════════════════════════════════════════════╗ ║ AI Agent Governance Framework (4 Pilares) ║ ╚═══════════════════════════════════════════════════════════╝
PILAR 1: SOURCE VERIFICATION ├─ Lista branca de fontes (apenas público, verificável) ├─ Agente: Citará APENAS fontes verificadas ├─ Exemplo (bom): │ └─ "Segundo CVM" (fonte: cvm.gov.br) │ └─ "Segundo B3" (fonte: b3.com.br) │ └─ "Segundo BACEN" (fonte: bacen.gov.br) ├─ Exemplo (ruim): │ └─ "Segundo pesquisa interna" (não verificável) │ └─ "Segundo análise" (qual análise?) │ └─ "Segundo estudos mostram" (quais estudos?) └─ Implementation: Before agent responds, check: is source whitelisted?
PILAR 2: DATA FRESHNESS ├─ Agente usa apenas dados RECENT (max 90 dias old) ├─ Outdated data = source de fake claims ├─ Exemplo: │ ├─ "Nossa taxa é 0.5%" (from 2023) │ ├─ Reality: 0.5% em 2023, mas 0.3% em 2024 │ └─ Agente quer falar informação desatualizada │ └─ Sistema bloqueia: "Atualizar dataset" └─ Implementation: Tag toda source com timestamp, expire automatic
PILAR 3: CONTEXT BOUNDARIES ├─ Agente sabe "o que EU NÃO SEI" ├─ Se perguntado sobre algo fora de training data │ └─ Responde: "Não tenho informação atualizada" │ └─ NÃO inventa resposta ├─ Exemplo: │ ├─ Pergunta: "Qual é o melhor investimento agora?" │ ├─ Agente (sem governance): "Ação XYZ (eu acho)" │ └─ Agente (com governance): "Não posso recomendar (fora de minha expertise)" └─ Implementation: Hard-coded boundaries, agente rejeita perguntas fora de scope
PILAR 4: AUDIT TRAIL ├─ Cada resposta do agente é LOGGADA ├─ Log inclui: pergunta, resposta, fontes usadas, timestamp ├─ Se cliente reclama: "Você mentiu" │ └─ Company: "Aqui tá o log, com fonte e timestamp" │ └─ Prova que agente usou fonte verificada ├─ Legal protection: Se agente foi enganado (source fake), culpa é de quem plantou fake, não de você └─ Implementation: Database de todas as respostas, queryable
Implementação prática
Step 1: Define whitelist de fontes
yaml
sources_whitelist.yaml
sources: financial: - name: "CVM" url: "www.cvm.gov.br" refresh_frequency: "daily" authority: "official" - name: "B3" url: "www.b3.com.br" refresh_frequency: "realtime" authority: "official" - name: "BACEN" url: "www.bacen.gov.br" refresh_frequency: "daily" authority: "official" - name: "Bloomberg (Enterprise API)" url: "api.bloomberg.com" refresh_frequency: "realtime" authority: "third-party" requires_verification: true
product: - name: "Our Public Roadmap" url: "product.company.com/roadmap" refresh_frequency: "weekly" authority: "internal" requires_approval: true - name: "Our Documentation" url: "docs.company.com" refresh_frequency: "realtime" authority: "internal" requires_approval: true - name: "Customer Reviews (Verified)" url: "api.verified-reviews.com" refresh_frequency: "daily" authority: "third-party" requires_verification: true
disallowed_sources:
- "internal research (unverifiable)"
- "team opinions"
- "marketing claims (unverified)"
- "competitor intelligence (potentially false)"
Step 2: Agent validation logic
python
agent_governance.py
class GovernedAgent: """ Agent que respeita compliance rules """
def __init__(self, sources_whitelist_file: str):
self.whitelist = load_whitelist(sources_whitelist_file)
self.audit_log = []
def respond_to_user(self, question: str) -> dict:
"""
Main method: user asks question
Agent responds with governance checks
"""
# STEP 1: Check if question is within bounds
if not self._is_question_in_scope(question):
response = "I don't have reliable information about that topic."
self._log_response(question, response, "out_of_scope")
return {"response": response, "sources": [], "confidence": 0}
# STEP 2: Generate candidate response (with Claude)
candidate = self._generate_response(question)
# candidate = {
# "text": "Our product has 99% uptime",
# "sources": ["Internal monitoring system"]
# }
# STEP 3: Validate sources
validated_sources = []
for source in candidate["sources"]:
if self._is_source_whitelisted(source):
validated_sources.append(source)
else:
# Source not whitelisted: remove from response
self._log_violation(question, source, "unapproved_source")
if not validated_sources:
# No valid sources: can't answer confidently
response = "I don't have reliable information from official sources about this."
self._log_response(question, response, "no_valid_sources")
return {"response": response, "sources": [], "confidence": 0}
# STEP 4: Check data freshness
for source in validated_sources:
if not self._is_data_fresh(source):
self._log_violation(question, source, "outdated_data")
response = "The information I have is outdated. Please check our official sources."
self._log_response(question, response, "data_too_old")
return {"response": response, "sources": [], "confidence": 0}
# STEP 5: Log successful response
self._log_response(
question=question,
response=candidate["text"],
sources=validated_sources,
status="approved"
)
return {
"response": candidate["text"],
"sources": validated_sources,
"confidence": self._calculate_confidence(validated_sources),
"timestamp": datetime.now().isoformat(),
"audit_id": self.audit_log[-1]["id"]
}
def _is_question_in_scope(self, question: str) -> bool:
"""
Is agent supposed to answer this?
"""
out_of_scope = [
"investment recommendation", # only educational
"legal advice", # not a lawyer
"medical advice", # not a doctor
"guaranteed returns", # impossible to guarantee
]
for keyword in out_of_scope:
if keyword.lower() in question.lower():
return False
return True
def _is_source_whitelisted(self, source: str) -> bool:
"""
Is this source on the approved list?
"""
for whitelisted in self.whitelist["sources"]:
if whitelisted["name"].lower() in source.lower():
return True
return False
def _is_data_fresh(self, source: str) -> bool:
"""
Is data from this source recent enough?
"""
for whitelisted in self.whitelist["sources"]:
if whitelisted["name"].lower() in source.lower():
last_update = self._get_last_update(source)
max_age = whitelisted["max_age_days"]
if (datetime.now() - last_update).days > max_age:
return False
return True
def _log_response(self, question: str, response: str, sources: list, status: str):
"""
Log every response (audit trail)
"""
log_entry = {
"id": str(uuid.uuid4()),
"timestamp": datetime.now().isoformat(),
"question": question,
"response": response,
"sources": sources,
"status": status, # "approved", "rejected", "out_of_scope", etc
}
self.audit_log.append(log_entry)
# Also save to database (perm persistence)
save_to_audit_db(log_entry)
def _log_violation(self, question: str, source: str, violation_type: str):
"""
Log when agent tries to cite non-whitelisted source
"""
violation = {
"timestamp": datetime.now().isoformat(),
"question": question,
"attempted_source": source,
"violation_type": violation_type, # unapproved_source, outdated_data, etc
}
log_violation(violation)
# Alert ops if suspicious patterns
if self._is_suspicious_pattern(violation):
alert_ops(f"Potential data injection attempt: {violation}")
def _calculate_confidence(self, sources: list) -> float:
"""
Higher confidence = more official sources
"""
official_sources = 0
for source in sources:
if self._is_official(source):
official_sources += 1
return official_sources / len(sources) if sources else 0
def _generate_response(self, question: str) -> dict:
"""
Use Claude to generate candidate response
(will be validated in respond_to_user)
"""
# Call Claude API
response = claude.messages.create(
model="claude-3-5-sonnet",
max_tokens=500,
system="You are a helpful assistant. Always cite your sources.",
messages=[{"role": "user", "content": question}]
)
return parse_response_with_sources(response.content[0].text)
Usage
agent = GovernedAgent("sources_whitelist.yaml")
User asks question
result = agent.respond_to_user("Qual é a melhor ação pra comprar agora?")
System response (with governance)
print(result)
Output:
{
"response": "Não posso recomendar ações específicas. Você pode verificar análises oficiais da B3 em www.b3.com.br",
"sources": [],
"confidence": 0,
"audit_id": "abc-123"
}
Step 3: Audit dashboard (monitorar violations)
python
audit_dashboard.py
class AuditDashboard: """ Monitor agent behavior, detect violations """
def get_violation_summary(self, days: int = 7) -> dict:
"""
Show violations in last N days
"""
violations = get_violations_from_db(days)
return {
"total_violations": len(violations),
"by_type": Counter([v["violation_type"] for v in violations]),
"timeline": self._group_by_hour(violations),
"alert_level": self._assess_risk(violations),
}
def get_coverage_report(self) -> dict:
"""
What % of agent responses are properly sourced?
"""
all_responses = get_all_responses_from_db()
approved = [r for r in all_responses if r["status"] == "approved"]
return {
"total_responses": len(all_responses),
"approved_responses": len(approved),
"coverage": len(approved) / len(all_responses) if all_responses else 0,
}
def alert_if_suspicious(self):
"""
Alert ops if something looks wrong
"""
violations = get_violations_from_db(days=1)
if len(violations) > 10: # threshold
alert_ops(f"ALERT: {len(violations)} violations in last 24h")
unapproved_sources = [v for v in violations if v["violation_type"] == "unapproved_source"]
if len(unapproved_sources) > 5:
alert_ops("ALERT: Agent trying to cite unapproved sources")
Impacto legal (por que isso importa)
Framework regulatório (Brasil):
Lei Geral de Proteção de Dados (LGPD) ├─ Agente deve ser "transparente" (dizer o que tá fazendo) ├─ Agente não pode usar dados pessoais pra enganar └─ Fine: Até R$ 50M
Código de Defesa do Consumidor ├─ Publicidade enganosa = proibida (se agente faz) ├─ Responsabilidade: Empresa é responsável pelo agente └─ Fine: Até 10% do faturamento
Laws específicas (por setor) ├─ Fintech (BACEN): Recomendações precisam de disclaimer ├─ Saúde (ANVISA): Conselhos médicos = proibidos ├─ Educação: Diplomas falsos = crime └─ Todas: Fake sources = fraud
Caso recente (EUA):
Casos que já aconteceram: ├─ Lawyer usou ChatGPT pra pesquisa legal │ └─ ChatGPT inventou cases que não existem │ └─ Lawyer apresentou "fake cases" pra court │ └─ Judge não achava graça │ └─ Lawyer: Processo + multa + desligamento │ └─ SaaS company usou Claude em agente de vendas └─ Agente citava "estudos" que não existiam └─ Customer descobriu (Googled e não achou) └─ Lawsuit: Misrepresentation └─ Settlement: R$ 500K
"Fake sources" = responsabilidade YOUR:
AnthropicPolicy says: └─ "Customers cannot use Claude to seed fake sources"
But that doesn't mean: └─ "You're not responsible"
It means: ├─ You broke Anthropic's terms of service ├─ Your account: Banned ├─ Your customer: Still damaged ├─ Lawsuits: Still coming └─ Your liability: 100%
Conclusão: Compliance agora, ou desastre depois
Before (sem governance):
- Agente rápido pra deploy (1 semana)
- Risco alto (fake sources possible)
- Lawsuit quando descobrem
- Custo legal: R$ 500K-5M
After (com governance):
- Agente mais lento pra deploy (2 semanas)
- Risco baixo (só fontes verificadas)
- Zero lawsuits (tudo documentado)
- Custo: R$ 50K (infra compliance)
- Benefit: Peace of mind, legal protection, customer trust
Ação prática (faça HOJE):
- Audit seu agente (Ele cita fake sources? Como você sabe?)
- Define whitelist (Quais fontes são OK?)
- Implement validation (Código que bloqueia bad sources)
- Setup audit trail (Log tudo)
- Monitor violations (Dashboard de problemas)
Anthropic's update é canário na mina de carvão. Reguladores virão depois. Prepare-se agora.
→ OpenClaw: AI Agent Governance Framework
Seu agente está seguro? Ou é uma bomba esperando explodir? 🎯⚖️
Publicado em 9 de outubro de 2026