Notícias
Notícias
5 min de leitura
9 de outubro de 2026

Observability em agentes IA (você sabe o que ele faz?)

Seu agente IA é uma black box. Não sabe o que decide, por quê, onde erra. Observability = visibilidade total. Guia prático com Datadog insights.

Equipe OpenClaw

Equipe OpenClaw · Time de Engenharia & Produto

A Equipe OpenClaw é formada por engenheiros, designers e especialistas em IA dedicados a construir a melhor plataforma de agentes conversacionais para negócios brasileiros. Combinamos expertise…


Observability em agentes IA (você sabe o que ele faz?)

Notícia: Datadog (plataforma de monitoramento) está expandindo suporte para observability de agentes IA não-determinísticos. Yanbing Li (Chief Product Officer, Datadog) explica: agentes de IA são diferentes de software tradicional — você não consegue prever outputs, não consegue debugar com logs simples, não consegue rastrear custos facilmente.

Implicação: Se você tem agente em produção (WhatsApp, suporte, vendas), você provavelmente não sabe o que ele tá fazendo.

**"Você é CEO de SaaS com agente WhatsApp de vendas.

Cenário: Agente fecha 100 deals/dia.

Perguntas que você NÃO CONSEGUE RESPONDER: ├─ "Por que agente aprovou esse cliente (crédito score 200)?" ├─ "Quanto custou rodar esse agente hoje? (tokens gastos?)" ├─ "Onde agente está errando? (quais clientes reclamando?)" ├─ "Agente tá evoluindo ou degradando? (taxa conversão está subindo ou caindo?)" ├─ "Qual foi a decision path? (10 passos? 100 passos?)" ├─ "Agente tá hallucinating? (criando informações falsas?)" └─ "Qual é o custo por deal? (quanto LLM API tá comendo?)"

Sem observability: ├─ Agente é CAIXA PRETA ├─ Você não consegue debugar ├─ Você não consegue otimizar ├─ Você não consegue confiar └─ Você tá rodando na fé "**


O problema: Agentes IA são black box

Por que observability é diferente de software tradicional

Software tradicional (você consegue debugar):

Code tradicional: ├─ Input: User clica em "comprar" ├─ Logic: if user.balance > price: ├─ Output: Order created ├─ Determinístico: Sempre mesmo output pra mesmo input ├─ Debug: Put breakpoint, rastreia variable └─ Observability: Simples (logs, traces, alerts)

Agente IA (você não consegue debugar):

Agente IA: ├─ Input: "Quero pedir R$ 10K de crédito" ├─ Logic: Agente chama 5 ferramentas diferentes │ ├─ Tool 1: Consulta score de crédito │ ├─ Tool 2: Consulta histórico de pagamento │ ├─ Tool 3: Consulta limite anterior │ ├─ Tool 4: Consulta market conditions │ └─ Tool 5: Faz reasoning (isso leva 50 tokens) ├─ Output: "Sua aprovação: R$ 8K (não R$ 10K)" ├─ NÃO-determinístico: Cada vez executa diferente │ (token randomness, temperature, etc) ├─ Black box: POR QUÊ agente limitou pra R$ 8K? │ (Você não consegue saber exatamente) └─ Observability: Complexa (quais tokens? qual modelo? qual reasoning?)

Diferenças críticas:

┌─────────────────┬──────────────────────┬──────────────────────┐ │ Aspecto │ Software Tradicional │ Agente IA │ ├─────────────────┼──────────────────────┼──────────────────────┤ │ Determinismo │ Determinístico │ Não-determinístico │ │ Debug │ Breakpoints fáceis │ Impossível breakpoint│ │ Rastreamento │ Call stack claro │ Token path opaco │ │ Custo │ Previsível │ Varia muito │ │ Erros │ Exceções claras │ Hallucinations? │ │ Performance │ Tempo fixo │ Tempo muito variável │ │ Observability │ Logs + Traces │ Precisa de LLM logs │ └─────────────────┴──────────────────────┴──────────────────────┘

Cenários onde observability salva seu negócio

Cenário 1: Agente tá hallucinating (criando info falsa)

Sem observability: ├─ Cliente reclama: "Agente disse que tenho R$ 50K de crédito" ├─ Você checa: "Seu crédito é R$ 10K" ├─ Problema: Agente mentiu (hallucinated) ├─ Solução: Desligar agente (capotaço) └─ Custo: Clientes perdidos, reputação, lawsuit

Com observability: ├─ Dashboard mostra: "Agente citou fonte que não existe" ├─ Você vê exatamente onde hallucinated ├─ Solução: Ajustar prompt (30 min) └─ Custo: Mínimo, problema resolvido

Cenário 2: Custo tá explodindo (tokens caros)

Sem observability: ├─ Invoice da OpenAI: R$ 50K/mês (era R$ 5K) ├─ Você não sabe por quê ├─ Agente tá usando prompts muito longos? ├─ Agente tá looping (chamando ferramentas 100x)? ├─ Agente tá usando modelo caro (GPT-4 vs GPT-3.5)? └─ Solução: Especulação, pode errar

Com observability: ├─ Dashboard mostra: "Token usage por agent action" ├─ Você vê exatamente: "Tool X usa 500 tokens, poderia usar 50" ├─ Você vê: "Agente tá looping 10x (deveria ser 2x)" ├─ Você vê: "Modelo XYZ tá usando 80% do custo" └─ Solução: Otimização cirúrgica (economiza R$ 40K/mês)

Cenário 3: Taxa de conversão caiu (mas agente tá "funcionando")

Sem observability: ├─ Métrica: Conversão 25% → 15% (12 dias depois) ├─ Problema: Agente? LLM mudou? Dataset desatualizado? ├─ Você não sabe ├─ Solução: Testar coisas aleatoriamente └─ Custo: 1-2 semanas pra descobrir

Com observability: ├─ Dashboard mostra timeline de todas as mudanças ├─ Você vê: "Conversão caiu dia X (quando OpenAI fez update em Claude)" ├─ Você vê: "Agora usa modelo V2 (que é diferente)" ├─ Você vê: "Accuracy de predictions caiu de 92% → 78%" └─ Solução: Reverter modelo, recuperar em 1 hora

Cenário 4: Compliance/auditoria ("Prove que agente não fez coisa ruim")

Sem observability: ├─ Regulador: "Cliente reclamou que agente aprovou crédito falso" ├─ Você: "Não tenho logs, não consigo provar nada" ├─ Regulador: "Fine: R$ 100K" └─ Realidade: Você tá cobrindo o vazio (você não sabe)

Com observability: ├─ Dashboard mostra: "Exatamente que dados agente usou" ├─ Você vê: "Agente usou Score oficial (CVM), não inventou" ├─ Você vê: "Agente seguiu regra corretamente" ├─ Regulador: "Tudo está auditável, ok" └─ Fine: Zero (você provou compliance)


Solução: Observability framework pra agentes IA

4 pilares de observability (segundo Datadog)

╔════════════════════════════════════════════════════════════╗ ║ AI Agent Observability (4 Pilares) ║ ╚════════════════════════════════════════════════════════════╝

PILAR 1: TRACING (Rastrear cada decision) ├─ Cada chamada de agente = trace completo ├─ Trace inclui: │ ├─ Input que agente recebeu │ ├─ Tools que agente chamou (ordem) │ ├─ Outputs de cada tool │ ├─ Reasoning do agente (prompts intermediários) │ ├─ Tokens gastos em cada step │ ├─ Tempo de execução por step │ ├─ Erros ou warnings que aconteceram │ └─ Output final ├─ Exemplo trace: │ └─ Agent.start() → Tool.creditScore() → Tool.paymentHistory() │ → Tool.marketAnalysis() → Agent.decide() → Agent.respond() ├─ Benefit: Debug exato (viu cada passo) └─ Cost: Storage (logs podem ser grandes)

PILAR 2: METRICS (Monitorar indicadores) ├─ Métricas por agente: │ ├─ Calls per hour (volume) │ ├─ Success rate (% de sucesso) │ ├─ Average latency (quanto tempo demora) │ ├─ Token usage (custo) │ ├─ Error rate (% de erro) │ ├─ Hallucination rate (% que inventa info) │ └─ Cost per request (quanto custa cada call) ├─ Exemplo dashboard: │ ├─ Sales Agent: 100 calls/hour, 92% success, 1.2s latency, 0.05 $/call │ ├─ Support Agent: 500 calls/hour, 85% success, 0.8s latency, 0.02 $/call │ └─ Onboarding Agent: 50 calls/hour, 98% success, 2.5s latency, 0.10 $/call ├─ Benefit: Visão total (rápido de entender) └─ Easy to alert (setup alerta se success <80%)

PILAR 3: LOGS (Entender detalhes) ├─ Structured logs pra cada agent action: │ ├─ {"agent": "sales", "action": "approve_credit", "amount": 10000, "score": 750} │ ├─ {"agent": "sales", "tool": "credit_check", "duration_ms": 150, "status": "success"} │ └─ {"agent": "support", "hallucinated": true, "claim": "feature X exists", "reality": "feature X planned Q4"} ├─ Queryable (SQL-like queries): │ └─ SELECT * FROM agent_logs WHERE hallucinated=true AND timestamp > now()-1h ├─ Benefit: Drill-down (entender causa raiz) └─ Cost: Storage + processing

PILAR 4: ALERTS (Notificar quando algo vai mal) ├─ Smart alerts: │ ├─ IF success_rate < 80% → Alert "Agent degrading" │ ├─ IF avg_latency > 5s → Alert "Agent slow" │ ├─ IF cost_per_request > $0.50 → Alert "Tokens exploding" │ ├─ IF hallucination_rate > 5% → Alert "Agent confusing" │ └─ IF error_rate > 10% → Alert "Agent broken" ├─ Escalation: │ ├─ Level 1: Slack notification │ ├─ Level 2: PagerDuty (if critical) │ └─ Level 3: Auto-disable agent (if catastrophic) ├─ Benefit: Detectar problemas ANTES de customer reclamar └─ Cost: Low (monitoring é cheap)

Implementação prática

Step 1: Setup basic tracing

python

agent_observability.py

import json import time from datetime import datetime from datadog import initialize, api

class ObservableAgent: """ Agent com observability built-in """

def __init__(self, agent_name: str, dd_api_key: str = None):
    self.agent_name = agent_name
    self.dd_enabled = dd_api_key is not None
    self.traces = []
    
    if self.dd_enabled:
        initialize(api_key=dd_api_key)

def execute(self, user_input: str) -> dict:
    """
    Execute agent, trace every step
    """
    trace = {
        "trace_id": self._generate_trace_id(),
        "agent_name": self.agent_name,
        "timestamp": datetime.now().isoformat(),
        "user_input": user_input,
        "steps": [],
        "output": None,
        "total_tokens": 0,
        "total_cost": 0,
        "total_duration_ms": 0,
    }
    
    start_time = time.time()
    
    try:
        # Step 1: Understand input
        trace["steps"].append(self._trace_step(
            name="parse_input",
            action=lambda: self._parse_input(user_input)
        ))
        
        parsed = trace["steps"][-1]["output"]
        
        # Step 2: Retrieve context (call tools)
        trace["steps"].append(self._trace_step(
            name="retrieve_context",
            action=lambda: self._call_tools(parsed)
        ))
        
        context = trace["steps"][-1]["output"]
        
        # Step 3: Reasoning (call LLM)
        trace["steps"].append(self._trace_step(
            name="reasoning",
            action=lambda: self._call_llm(user_input, context)
        ))
        
        decision = trace["steps"][-1]["output"]
        
        # Step 4: Format response
        trace["steps"].append(self._trace_step(
            name="format_response",
            action=lambda: self._format_response(decision)
        ))
        
        response = trace["steps"][-1]["output"]
        
        # Calculate totals
        trace["output"] = response
        trace["total_tokens"] = sum([step["tokens_used"] for step in trace["steps"]])
        trace["total_cost"] = sum([step["cost"] for step in trace["steps"]])
        trace["total_duration_ms"] = int((time.time() - start_time) * 1000)
        trace["status"] = "success"
        
    except Exception as e:
        trace["status"] = "error"
        trace["error"] = str(e)
        trace["total_duration_ms"] = int((time.time() - start_time) * 1000)
    
    # Log to Datadog
    self._send_to_datadog(trace)
    
    # Also store locally
    self.traces.append(trace)
    
    return trace

def _trace_step(self, name: str, action: callable) -> dict:
    """
    Execute step and measure metrics
    """
    step = {
        "name": name,
        "start_time": datetime.now().isoformat(),
        "duration_ms": 0,
        "tokens_used": 0,
        "cost": 0,
        "output": None,
        "error": None,
    }
    
    start = time.time()
    
    try:
        output = action()
        step["output"] = output
        
        # If tool call, extract token info
        if hasattr(output, 'get'):
            step["tokens_used"] = output.get("tokens_used", 0)
            step["cost"] = output.get("cost", 0)
    
    except Exception as e:
        step["error"] = str(e)
    
    finally:
        step["duration_ms"] = int((time.time() - start) * 1000)
    
    return step

def _call_tools(self, parsed_input: dict) -> dict:
    """
    Call external tools (search, APIs, etc)
    """
    # Simulate tool calls
    results = {
        "tokens_used": 500,
        "cost": 0.02,
        "data": {"credit_score": 750, "payment_history": "excellent"}
    }
    return results

def _call_llm(self, user_input: str, context: dict) -> dict:
    """
    Call LLM (Claude, GPT, etc)
    """
    # Simulate LLM call
    response = {
        "tokens_used": 200,
        "cost": 0.01,
        "reasoning": "Based on score 750 and payment history, approve R$ 10K",
        "decision": "APPROVE",
        "amount": 10000
    }
    return response

def _format_response(self, decision: dict) -> str:
    """
    Format final response
    """
    return f"Your approval: R$ {decision['amount']}"

def _send_to_datadog(self, trace: dict):
    """
    Send trace to Datadog
    """
    if not self.dd_enabled:
        return
    
    # Send metrics
    api.Metric.send(
        metric=f"agent.{self.agent_name}.duration",
        points=trace["total_duration_ms"],
        tags=[f"agent:{self.agent_name}", f"status:{trace['status']}"]
    )
    
    api.Metric.send(
        metric=f"agent.{self.agent_name}.tokens",
        points=trace["total_tokens"],
        tags=[f"agent:{self.agent_name}"]
    )
    
    api.Metric.send(
        metric=f"agent.{self.agent_name}.cost",
        points=trace["total_cost"],
        tags=[f"agent:{self.agent_name}"]
    )

def _parse_input(self, user_input: str) -> dict:
    """Parse and understand input"""
    return {"tokens_used": 100, "cost": 0.001, "intent": "request_credit"}

def _generate_trace_id(self) -> str:
    import uuid
    return str(uuid.uuid4())

Usage

agent = ObservableAgent("sales_agent", dd_api_key="YOUR_KEY")

Execute with full tracing

result = agent.execute("I want to request R$ 10K credit")

print(json.dumps(result, indent=2))

Output:

{

"trace_id": "abc-123",

"agent_name": "sales_agent",

"timestamp": "2026-10-09T...",

"user_input": "I want to request R$ 10K credit",

"steps": [

{"name": "parse_input", "duration_ms": 50, "tokens_used": 100, ...},

{"name": "retrieve_context", "duration_ms": 150, "tokens_used": 500, ...},

{"name": "reasoning", "duration_ms": 200, "tokens_used": 200, ...},

{"name": "format_response", "duration_ms": 30, "tokens_used": 50, ...}

],

"output": "Your approval: R$ 10K",

"total_tokens": 850,

"total_cost": 0.035,

"total_duration_ms": 430,

"status": "success"

}

Step 2: Setup alerts

python

agent_alerts.py

class AgentAlertManager: """ Monitor agent health, trigger alerts """

def __init__(self, agent_name: str):
    self.agent_name = agent_name
    self.metrics_window = []  # Last 100 traces

def check_health(self, trace: dict):
    """
    After each trace, check if something's wrong
    """
    self.metrics_window.append(trace)
    if len(self.metrics_window) > 100:
        self.metrics_window.pop(0)
    
    # Calculate rolling metrics
    success_rate = self._calculate_success_rate()
    avg_latency = self._calculate_avg_latency()
    avg_cost = self._calculate_avg_cost()
    hallucination_rate = self._calculate_hallucination_rate()
    
    # Check thresholds
    if success_rate < 0.80:
        self._alert("CRITICAL", f"Success rate dropped to {success_rate:.0%}")
    
    if avg_latency > 5000:  # 5 seconds
        self._alert("WARNING", f"Average latency is {avg_latency:.0f}ms (high)")
    
    if avg_cost > 0.50:  # $0.50 per request
        self._alert("WARNING", f"Cost per request is ${avg_cost:.2f} (expensive)")
    
    if hallucination_rate > 0.05:  # 5%
        self._alert("CRITICAL", f"Hallucination rate {hallucination_rate:.0%} (agent confusing)")

def _calculate_success_rate(self) -> float:
    successes = len([t for t in self.metrics_window if t["status"] == "success"])
    return successes / len(self.metrics_window) if self.metrics_window else 1.0

def _calculate_avg_latency(self) -> float:
    if not self.metrics_window:
        return 0
    total = sum([t["total_duration_ms"] for t in self.metrics_window])
    return total / len(self.metrics_window)

def _calculate_avg_cost(self) -> float:
    if not self.metrics_window:
        return 0
    total = sum([t["total_cost"] for t in self.metrics_window])
    return total / len(self.metrics_window)

def _calculate_hallucination_rate(self) -> float:
    # Detect hallucinations (harder, requires semantic analysis)
    # For now, use heuristic: response doesn't match context
    hallucinations = 0
    for trace in self.metrics_window:
        if self._seems_hallucinated(trace):
            hallucinations += 1
    return hallucinations / len(self.metrics_window) if self.metrics_window else 0

def _seems_hallucinated(self, trace: dict) -> bool:
    """
    Heuristic: detect if agent response seems fabricated
    (In practice: use semantic matching, fact-checking, etc)
    """
    # Example: if output contradicts input context
    return False  # placeholder

def _alert(self, severity: str, message: str):
    """
    Send alert (Slack, PagerDuty, etc)
    """
    print(f"[{severity}] Agent {self.agent_name}: {message}")
    # In production: send to Slack/PagerDuty

Usage

alert_manager = AgentAlertManager("sales_agent")

After each agent.execute():

alert_manager.check_health(trace)


Dashboard (o que você vai ver)

Exemplo real de observability dashboard:

╔════════════════════════════════════════════════════════════════╗ ║ Agent Observability Dashboard - Sales Agent ║ ║ (Last 24 hours) ║ ╚════════════════════════════════════════════════════════════════╝

KEY METRICS: ├─ Total Calls: 2,400 ├─ Success Rate: 92% ✓ (target: >80%) ├─ Avg Latency: 1.2s ✓ (target: <5s) ├─ Total Cost: R$ 840 ⚠ (was R$ 480 yesterday = +75%) └─ Hallucination Rate: 2% ✓ (target: <5%)

PERFORMANCE TIMELINE: ├─ 00:00-04:00: Success 95%, Cost R$ 120 ├─ 04:00-08:00: Success 88%, Cost R$ 180 ← dip ├─ 08:00-12:00: Success 93%, Cost R$ 200 ├─ 12:00-16:00: Success 91%, Cost R$ 210 ← higher cost ├─ 16:00-20:00: Success 89%, Cost R$ 180 ← another dip └─ 20:00-24:00: Success 94%, Cost R$ 150

TOP ERRORS: ├─ 1. Tool "credit_check" timeout: 45 errors (1.9%) ├─ 2. LLM rate limit: 12 errors (0.5%) └─ 3. Invalid input: 8 errors (0.3%)

COST BREAKDOWN (by step): ├─ Parse input: R$ 30 (4%) ├─ Retrieve context: R$ 350 (42%) ← Most expensive ├─ Reasoning (LLM): R$ 420 (50%) └─ Format response: R$ 40 (4%)

HALLUCINATION EXAMPLES: ├─ 1. "Customer has credit limit R$ 50K" (reality: R$ 20K) - 2026-10-09 14:23 ├─ 2. "We offer 0% interest" (reality: 2-5% depending on credit) - 2026-10-09 11:05 └─ 3. "Instant approval" (actually: 24h review) - 2026-10-09 08:45

ACTIONS TO TAKE: ├─ ⚠ Cost +75%: Investigate "Retrieve context" step │ └─ Action: Check if tool is being called more (or API prices went up) ├─ ⚠ 2% hallucination: Review prompt, add guardrails │ └─ Action: Add fact-checking step └─ ✓ 92% success: Keep current configuration


Conclusão: Observability = Control

Without observability:

  • Agente é black box
  • Você roda na fé
  • Problemas vêm tarde (cliente reclama)
  • Debug é especulação
  • Custos explodem sem você saber

With observability:

  • Agente é transparent
  • Você tem controle
  • Problemas detectados cedo (alert antes de falhar)
  • Debug é cirúrgico (sabe exatamente onde falhou)
  • Custos são otimizados (sabe onde gasta mais)

Ação prática (faça HOJE):

  1. Setup tracing (rastreia cada step do agente)
  2. Define metrics (success rate, latency, cost, hallucination)
  3. Configure alerts (notifica se algo vai mal)
  4. Monitor dashboard (vê saúde do agente em tempo real)
  5. Iterate fast (ajusta prompt/tools baseado em data)

Datadog + observability é investimento de R$ 5K-10K/mês que salva R$ 100K+ em problemas.

→ OpenClaw: AI Agent Observability Framework

Seu agente é caixa preta ou está transparent? 🎯📊


Publicado em 9 de outubro de 2026

Leia também