Observability em agentes IA (você sabe o que ele faz?)
Seu agente IA é uma black box. Não sabe o que decide, por quê, onde erra. Observability = visibilidade total. Guia prático com Datadog insights.
Equipe OpenClaw · Time de Engenharia & Produto
A Equipe OpenClaw é formada por engenheiros, designers e especialistas em IA dedicados a construir a melhor plataforma de agentes conversacionais para negócios brasileiros. Combinamos expertise…
Observability em agentes IA (você sabe o que ele faz?)
Notícia: Datadog (plataforma de monitoramento) está expandindo suporte para observability de agentes IA não-determinísticos. Yanbing Li (Chief Product Officer, Datadog) explica: agentes de IA são diferentes de software tradicional — você não consegue prever outputs, não consegue debugar com logs simples, não consegue rastrear custos facilmente.
Implicação: Se você tem agente em produção (WhatsApp, suporte, vendas), você provavelmente não sabe o que ele tá fazendo.
**"Você é CEO de SaaS com agente WhatsApp de vendas.
Cenário: Agente fecha 100 deals/dia.
Perguntas que você NÃO CONSEGUE RESPONDER: ├─ "Por que agente aprovou esse cliente (crédito score 200)?" ├─ "Quanto custou rodar esse agente hoje? (tokens gastos?)" ├─ "Onde agente está errando? (quais clientes reclamando?)" ├─ "Agente tá evoluindo ou degradando? (taxa conversão está subindo ou caindo?)" ├─ "Qual foi a decision path? (10 passos? 100 passos?)" ├─ "Agente tá hallucinating? (criando informações falsas?)" └─ "Qual é o custo por deal? (quanto LLM API tá comendo?)"
Sem observability: ├─ Agente é CAIXA PRETA ├─ Você não consegue debugar ├─ Você não consegue otimizar ├─ Você não consegue confiar └─ Você tá rodando na fé "**
O problema: Agentes IA são black box
Por que observability é diferente de software tradicional
Software tradicional (você consegue debugar):
Code tradicional: ├─ Input: User clica em "comprar" ├─ Logic: if user.balance > price: ├─ Output: Order created ├─ Determinístico: Sempre mesmo output pra mesmo input ├─ Debug: Put breakpoint, rastreia variable └─ Observability: Simples (logs, traces, alerts)
Agente IA (você não consegue debugar):
Agente IA: ├─ Input: "Quero pedir R$ 10K de crédito" ├─ Logic: Agente chama 5 ferramentas diferentes │ ├─ Tool 1: Consulta score de crédito │ ├─ Tool 2: Consulta histórico de pagamento │ ├─ Tool 3: Consulta limite anterior │ ├─ Tool 4: Consulta market conditions │ └─ Tool 5: Faz reasoning (isso leva 50 tokens) ├─ Output: "Sua aprovação: R$ 8K (não R$ 10K)" ├─ NÃO-determinístico: Cada vez executa diferente │ (token randomness, temperature, etc) ├─ Black box: POR QUÊ agente limitou pra R$ 8K? │ (Você não consegue saber exatamente) └─ Observability: Complexa (quais tokens? qual modelo? qual reasoning?)
Diferenças críticas:
┌─────────────────┬──────────────────────┬──────────────────────┐ │ Aspecto │ Software Tradicional │ Agente IA │ ├─────────────────┼──────────────────────┼──────────────────────┤ │ Determinismo │ Determinístico │ Não-determinístico │ │ Debug │ Breakpoints fáceis │ Impossível breakpoint│ │ Rastreamento │ Call stack claro │ Token path opaco │ │ Custo │ Previsível │ Varia muito │ │ Erros │ Exceções claras │ Hallucinations? │ │ Performance │ Tempo fixo │ Tempo muito variável │ │ Observability │ Logs + Traces │ Precisa de LLM logs │ └─────────────────┴──────────────────────┴──────────────────────┘
Cenários onde observability salva seu negócio
Cenário 1: Agente tá hallucinating (criando info falsa)
Sem observability: ├─ Cliente reclama: "Agente disse que tenho R$ 50K de crédito" ├─ Você checa: "Seu crédito é R$ 10K" ├─ Problema: Agente mentiu (hallucinated) ├─ Solução: Desligar agente (capotaço) └─ Custo: Clientes perdidos, reputação, lawsuit
Com observability: ├─ Dashboard mostra: "Agente citou fonte que não existe" ├─ Você vê exatamente onde hallucinated ├─ Solução: Ajustar prompt (30 min) └─ Custo: Mínimo, problema resolvido
Cenário 2: Custo tá explodindo (tokens caros)
Sem observability: ├─ Invoice da OpenAI: R$ 50K/mês (era R$ 5K) ├─ Você não sabe por quê ├─ Agente tá usando prompts muito longos? ├─ Agente tá looping (chamando ferramentas 100x)? ├─ Agente tá usando modelo caro (GPT-4 vs GPT-3.5)? └─ Solução: Especulação, pode errar
Com observability: ├─ Dashboard mostra: "Token usage por agent action" ├─ Você vê exatamente: "Tool X usa 500 tokens, poderia usar 50" ├─ Você vê: "Agente tá looping 10x (deveria ser 2x)" ├─ Você vê: "Modelo XYZ tá usando 80% do custo" └─ Solução: Otimização cirúrgica (economiza R$ 40K/mês)
Cenário 3: Taxa de conversão caiu (mas agente tá "funcionando")
Sem observability: ├─ Métrica: Conversão 25% → 15% (12 dias depois) ├─ Problema: Agente? LLM mudou? Dataset desatualizado? ├─ Você não sabe ├─ Solução: Testar coisas aleatoriamente └─ Custo: 1-2 semanas pra descobrir
Com observability: ├─ Dashboard mostra timeline de todas as mudanças ├─ Você vê: "Conversão caiu dia X (quando OpenAI fez update em Claude)" ├─ Você vê: "Agora usa modelo V2 (que é diferente)" ├─ Você vê: "Accuracy de predictions caiu de 92% → 78%" └─ Solução: Reverter modelo, recuperar em 1 hora
Cenário 4: Compliance/auditoria ("Prove que agente não fez coisa ruim")
Sem observability: ├─ Regulador: "Cliente reclamou que agente aprovou crédito falso" ├─ Você: "Não tenho logs, não consigo provar nada" ├─ Regulador: "Fine: R$ 100K" └─ Realidade: Você tá cobrindo o vazio (você não sabe)
Com observability: ├─ Dashboard mostra: "Exatamente que dados agente usou" ├─ Você vê: "Agente usou Score oficial (CVM), não inventou" ├─ Você vê: "Agente seguiu regra corretamente" ├─ Regulador: "Tudo está auditável, ok" └─ Fine: Zero (você provou compliance)
Solução: Observability framework pra agentes IA
4 pilares de observability (segundo Datadog)
╔════════════════════════════════════════════════════════════╗ ║ AI Agent Observability (4 Pilares) ║ ╚════════════════════════════════════════════════════════════╝
PILAR 1: TRACING (Rastrear cada decision) ├─ Cada chamada de agente = trace completo ├─ Trace inclui: │ ├─ Input que agente recebeu │ ├─ Tools que agente chamou (ordem) │ ├─ Outputs de cada tool │ ├─ Reasoning do agente (prompts intermediários) │ ├─ Tokens gastos em cada step │ ├─ Tempo de execução por step │ ├─ Erros ou warnings que aconteceram │ └─ Output final ├─ Exemplo trace: │ └─ Agent.start() → Tool.creditScore() → Tool.paymentHistory() │ → Tool.marketAnalysis() → Agent.decide() → Agent.respond() ├─ Benefit: Debug exato (viu cada passo) └─ Cost: Storage (logs podem ser grandes)
PILAR 2: METRICS (Monitorar indicadores) ├─ Métricas por agente: │ ├─ Calls per hour (volume) │ ├─ Success rate (% de sucesso) │ ├─ Average latency (quanto tempo demora) │ ├─ Token usage (custo) │ ├─ Error rate (% de erro) │ ├─ Hallucination rate (% que inventa info) │ └─ Cost per request (quanto custa cada call) ├─ Exemplo dashboard: │ ├─ Sales Agent: 100 calls/hour, 92% success, 1.2s latency, 0.05 $/call │ ├─ Support Agent: 500 calls/hour, 85% success, 0.8s latency, 0.02 $/call │ └─ Onboarding Agent: 50 calls/hour, 98% success, 2.5s latency, 0.10 $/call ├─ Benefit: Visão total (rápido de entender) └─ Easy to alert (setup alerta se success <80%)
PILAR 3: LOGS (Entender detalhes) ├─ Structured logs pra cada agent action: │ ├─ {"agent": "sales", "action": "approve_credit", "amount": 10000, "score": 750} │ ├─ {"agent": "sales", "tool": "credit_check", "duration_ms": 150, "status": "success"} │ └─ {"agent": "support", "hallucinated": true, "claim": "feature X exists", "reality": "feature X planned Q4"} ├─ Queryable (SQL-like queries): │ └─ SELECT * FROM agent_logs WHERE hallucinated=true AND timestamp > now()-1h ├─ Benefit: Drill-down (entender causa raiz) └─ Cost: Storage + processing
PILAR 4: ALERTS (Notificar quando algo vai mal) ├─ Smart alerts: │ ├─ IF success_rate < 80% → Alert "Agent degrading" │ ├─ IF avg_latency > 5s → Alert "Agent slow" │ ├─ IF cost_per_request > $0.50 → Alert "Tokens exploding" │ ├─ IF hallucination_rate > 5% → Alert "Agent confusing" │ └─ IF error_rate > 10% → Alert "Agent broken" ├─ Escalation: │ ├─ Level 1: Slack notification │ ├─ Level 2: PagerDuty (if critical) │ └─ Level 3: Auto-disable agent (if catastrophic) ├─ Benefit: Detectar problemas ANTES de customer reclamar └─ Cost: Low (monitoring é cheap)
Implementação prática
Step 1: Setup basic tracing
python
agent_observability.py
import json import time from datetime import datetime from datadog import initialize, api
class ObservableAgent: """ Agent com observability built-in """
def __init__(self, agent_name: str, dd_api_key: str = None):
self.agent_name = agent_name
self.dd_enabled = dd_api_key is not None
self.traces = []
if self.dd_enabled:
initialize(api_key=dd_api_key)
def execute(self, user_input: str) -> dict:
"""
Execute agent, trace every step
"""
trace = {
"trace_id": self._generate_trace_id(),
"agent_name": self.agent_name,
"timestamp": datetime.now().isoformat(),
"user_input": user_input,
"steps": [],
"output": None,
"total_tokens": 0,
"total_cost": 0,
"total_duration_ms": 0,
}
start_time = time.time()
try:
# Step 1: Understand input
trace["steps"].append(self._trace_step(
name="parse_input",
action=lambda: self._parse_input(user_input)
))
parsed = trace["steps"][-1]["output"]
# Step 2: Retrieve context (call tools)
trace["steps"].append(self._trace_step(
name="retrieve_context",
action=lambda: self._call_tools(parsed)
))
context = trace["steps"][-1]["output"]
# Step 3: Reasoning (call LLM)
trace["steps"].append(self._trace_step(
name="reasoning",
action=lambda: self._call_llm(user_input, context)
))
decision = trace["steps"][-1]["output"]
# Step 4: Format response
trace["steps"].append(self._trace_step(
name="format_response",
action=lambda: self._format_response(decision)
))
response = trace["steps"][-1]["output"]
# Calculate totals
trace["output"] = response
trace["total_tokens"] = sum([step["tokens_used"] for step in trace["steps"]])
trace["total_cost"] = sum([step["cost"] for step in trace["steps"]])
trace["total_duration_ms"] = int((time.time() - start_time) * 1000)
trace["status"] = "success"
except Exception as e:
trace["status"] = "error"
trace["error"] = str(e)
trace["total_duration_ms"] = int((time.time() - start_time) * 1000)
# Log to Datadog
self._send_to_datadog(trace)
# Also store locally
self.traces.append(trace)
return trace
def _trace_step(self, name: str, action: callable) -> dict:
"""
Execute step and measure metrics
"""
step = {
"name": name,
"start_time": datetime.now().isoformat(),
"duration_ms": 0,
"tokens_used": 0,
"cost": 0,
"output": None,
"error": None,
}
start = time.time()
try:
output = action()
step["output"] = output
# If tool call, extract token info
if hasattr(output, 'get'):
step["tokens_used"] = output.get("tokens_used", 0)
step["cost"] = output.get("cost", 0)
except Exception as e:
step["error"] = str(e)
finally:
step["duration_ms"] = int((time.time() - start) * 1000)
return step
def _call_tools(self, parsed_input: dict) -> dict:
"""
Call external tools (search, APIs, etc)
"""
# Simulate tool calls
results = {
"tokens_used": 500,
"cost": 0.02,
"data": {"credit_score": 750, "payment_history": "excellent"}
}
return results
def _call_llm(self, user_input: str, context: dict) -> dict:
"""
Call LLM (Claude, GPT, etc)
"""
# Simulate LLM call
response = {
"tokens_used": 200,
"cost": 0.01,
"reasoning": "Based on score 750 and payment history, approve R$ 10K",
"decision": "APPROVE",
"amount": 10000
}
return response
def _format_response(self, decision: dict) -> str:
"""
Format final response
"""
return f"Your approval: R$ {decision['amount']}"
def _send_to_datadog(self, trace: dict):
"""
Send trace to Datadog
"""
if not self.dd_enabled:
return
# Send metrics
api.Metric.send(
metric=f"agent.{self.agent_name}.duration",
points=trace["total_duration_ms"],
tags=[f"agent:{self.agent_name}", f"status:{trace['status']}"]
)
api.Metric.send(
metric=f"agent.{self.agent_name}.tokens",
points=trace["total_tokens"],
tags=[f"agent:{self.agent_name}"]
)
api.Metric.send(
metric=f"agent.{self.agent_name}.cost",
points=trace["total_cost"],
tags=[f"agent:{self.agent_name}"]
)
def _parse_input(self, user_input: str) -> dict:
"""Parse and understand input"""
return {"tokens_used": 100, "cost": 0.001, "intent": "request_credit"}
def _generate_trace_id(self) -> str:
import uuid
return str(uuid.uuid4())
Usage
agent = ObservableAgent("sales_agent", dd_api_key="YOUR_KEY")
Execute with full tracing
result = agent.execute("I want to request R$ 10K credit")
print(json.dumps(result, indent=2))
Output:
{
"trace_id": "abc-123",
"agent_name": "sales_agent",
"timestamp": "2026-10-09T...",
"user_input": "I want to request R$ 10K credit",
"steps": [
{"name": "parse_input", "duration_ms": 50, "tokens_used": 100, ...},
{"name": "retrieve_context", "duration_ms": 150, "tokens_used": 500, ...},
{"name": "reasoning", "duration_ms": 200, "tokens_used": 200, ...},
{"name": "format_response", "duration_ms": 30, "tokens_used": 50, ...}
],
"output": "Your approval: R$ 10K",
"total_tokens": 850,
"total_cost": 0.035,
"total_duration_ms": 430,
"status": "success"
}
Step 2: Setup alerts
python
agent_alerts.py
class AgentAlertManager: """ Monitor agent health, trigger alerts """
def __init__(self, agent_name: str):
self.agent_name = agent_name
self.metrics_window = [] # Last 100 traces
def check_health(self, trace: dict):
"""
After each trace, check if something's wrong
"""
self.metrics_window.append(trace)
if len(self.metrics_window) > 100:
self.metrics_window.pop(0)
# Calculate rolling metrics
success_rate = self._calculate_success_rate()
avg_latency = self._calculate_avg_latency()
avg_cost = self._calculate_avg_cost()
hallucination_rate = self._calculate_hallucination_rate()
# Check thresholds
if success_rate < 0.80:
self._alert("CRITICAL", f"Success rate dropped to {success_rate:.0%}")
if avg_latency > 5000: # 5 seconds
self._alert("WARNING", f"Average latency is {avg_latency:.0f}ms (high)")
if avg_cost > 0.50: # $0.50 per request
self._alert("WARNING", f"Cost per request is ${avg_cost:.2f} (expensive)")
if hallucination_rate > 0.05: # 5%
self._alert("CRITICAL", f"Hallucination rate {hallucination_rate:.0%} (agent confusing)")
def _calculate_success_rate(self) -> float:
successes = len([t for t in self.metrics_window if t["status"] == "success"])
return successes / len(self.metrics_window) if self.metrics_window else 1.0
def _calculate_avg_latency(self) -> float:
if not self.metrics_window:
return 0
total = sum([t["total_duration_ms"] for t in self.metrics_window])
return total / len(self.metrics_window)
def _calculate_avg_cost(self) -> float:
if not self.metrics_window:
return 0
total = sum([t["total_cost"] for t in self.metrics_window])
return total / len(self.metrics_window)
def _calculate_hallucination_rate(self) -> float:
# Detect hallucinations (harder, requires semantic analysis)
# For now, use heuristic: response doesn't match context
hallucinations = 0
for trace in self.metrics_window:
if self._seems_hallucinated(trace):
hallucinations += 1
return hallucinations / len(self.metrics_window) if self.metrics_window else 0
def _seems_hallucinated(self, trace: dict) -> bool:
"""
Heuristic: detect if agent response seems fabricated
(In practice: use semantic matching, fact-checking, etc)
"""
# Example: if output contradicts input context
return False # placeholder
def _alert(self, severity: str, message: str):
"""
Send alert (Slack, PagerDuty, etc)
"""
print(f"[{severity}] Agent {self.agent_name}: {message}")
# In production: send to Slack/PagerDuty
Usage
alert_manager = AgentAlertManager("sales_agent")
After each agent.execute():
alert_manager.check_health(trace)
Dashboard (o que você vai ver)
Exemplo real de observability dashboard:
╔════════════════════════════════════════════════════════════════╗ ║ Agent Observability Dashboard - Sales Agent ║ ║ (Last 24 hours) ║ ╚════════════════════════════════════════════════════════════════╝
KEY METRICS: ├─ Total Calls: 2,400 ├─ Success Rate: 92% ✓ (target: >80%) ├─ Avg Latency: 1.2s ✓ (target: <5s) ├─ Total Cost: R$ 840 ⚠ (was R$ 480 yesterday = +75%) └─ Hallucination Rate: 2% ✓ (target: <5%)
PERFORMANCE TIMELINE: ├─ 00:00-04:00: Success 95%, Cost R$ 120 ├─ 04:00-08:00: Success 88%, Cost R$ 180 ← dip ├─ 08:00-12:00: Success 93%, Cost R$ 200 ├─ 12:00-16:00: Success 91%, Cost R$ 210 ← higher cost ├─ 16:00-20:00: Success 89%, Cost R$ 180 ← another dip └─ 20:00-24:00: Success 94%, Cost R$ 150
TOP ERRORS: ├─ 1. Tool "credit_check" timeout: 45 errors (1.9%) ├─ 2. LLM rate limit: 12 errors (0.5%) └─ 3. Invalid input: 8 errors (0.3%)
COST BREAKDOWN (by step): ├─ Parse input: R$ 30 (4%) ├─ Retrieve context: R$ 350 (42%) ← Most expensive ├─ Reasoning (LLM): R$ 420 (50%) └─ Format response: R$ 40 (4%)
HALLUCINATION EXAMPLES: ├─ 1. "Customer has credit limit R$ 50K" (reality: R$ 20K) - 2026-10-09 14:23 ├─ 2. "We offer 0% interest" (reality: 2-5% depending on credit) - 2026-10-09 11:05 └─ 3. "Instant approval" (actually: 24h review) - 2026-10-09 08:45
ACTIONS TO TAKE: ├─ ⚠ Cost +75%: Investigate "Retrieve context" step │ └─ Action: Check if tool is being called more (or API prices went up) ├─ ⚠ 2% hallucination: Review prompt, add guardrails │ └─ Action: Add fact-checking step └─ ✓ 92% success: Keep current configuration
Conclusão: Observability = Control
Without observability:
- Agente é black box
- Você roda na fé
- Problemas vêm tarde (cliente reclama)
- Debug é especulação
- Custos explodem sem você saber
With observability:
- Agente é transparent
- Você tem controle
- Problemas detectados cedo (alert antes de falhar)
- Debug é cirúrgico (sabe exatamente onde falhou)
- Custos são otimizados (sabe onde gasta mais)
Ação prática (faça HOJE):
- Setup tracing (rastreia cada step do agente)
- Define metrics (success rate, latency, cost, hallucination)
- Configure alerts (notifica se algo vai mal)
- Monitor dashboard (vê saúde do agente em tempo real)
- Iterate fast (ajusta prompt/tools baseado em data)
Datadog + observability é investimento de R$ 5K-10K/mês que salva R$ 100K+ em problemas.
→ OpenClaw: AI Agent Observability Framework
Seu agente é caixa preta ou está transparent? 🎯📊
Publicado em 9 de outubro de 2026