OpenAI Decisions API: agente sem alucinações (10x rápido)
OpenAI Decisions API: LLM gera respostas tipadas (não texto). 10x mais rápido. Seu agente WhatsApp finalmente determinístico (sem hallucinations).
Equipe OpenClaw · Time de Engenharia & Produto
A Equipe OpenClaw é formada por engenheiros, designers e especialistas em IA dedicados a construir a melhor plataforma de agentes conversacionais para negócios brasileiros. Combinamos expertise…
OpenAI Decisions API: agente sem alucinações (10x rápido)
Notícia: OpenAI lançou Decisions API em public beta: LLM que retorna respostas tipadas (JSON, enums, booleanos) ao invés de texto puro. Resultado: 10x mais rápido + zero parsing errors + determinístico.
Implicação: Seu agente WhatsApp finally pode confiar em outputs do LLM. Não precisa mais de regex crazy pra parsear "sim" vs "yes" vs "yep". API retorna {"decision": true} direto.
**"Você é founder de SaaS de suporte com agente WhatsApp.
Problema antigo (text-based LLM): ├─ Você pede: "Customer quer refund? Responda com 'yes' ou 'no'" ├─ LLM responde: │ ├─ Response 1: "yes" │ ├─ Response 2: "No, but..." │ ├─ Response 3: "I think so, but need to check" │ ├─ Response 4: "Sim" (em português!) │ └─ Response 5: "Yes." (com ponto) ├─ Seu código tenta parsear: │ ├─ if "yes" in response.lower() → SIM │ ├─ if "no" in response.lower() → NÃO │ ├─ else → ERROR (pode ser qual?) │ └─ Resultado: 20% erro rate (unparseable responses) ├─ Latência: 2s (LLM) + 0.5s (parsing) = 2.5s └─ Frustração: Agente não é confiável
Problema novo (Decisions API): ❌ RESOLVIDO ├─ Você pede: "Customer quer refund? Responda com structured decision" ├─ API retorna: {"decision": "yes", "confidence": 0.95} ├─ Seu código: │ ├─ if response.decision == "yes" → SIM │ ├─ Parsing: ZERO errors (tipo é garantido) │ └─ Latência: 2s (só LLM, parsing é instant) └─ Resultado: 100% confiável "**
O problema que ninguém fala: LLM outputs são caóticos
Padrão antigo (text-based)
Você pede ao LLM: "Classifique esse email como: 'urgent', 'normal', 'spam' Responda APENAS com uma palavra."
LLM responde: ├─ "urgent" ✅ ├─ "This is urgent" ❌ (não seguiu instrução) ├─ "Urgent - customer angry" ❌ (adicionou contexto) ├─ "URGENT" ❌ (maiúscula, código espera minúscula) ├─ "urgent." ❌ (com ponto) ├─ "urgente" ❌ (respondeu em português) └─ "i think it's urgent but not super" ❌ (hedge language)
Resultado: ├─ Success rate: ~60% (nem sempre está perfeito) ├─ Você precisa: regex, fuzzy matching, error handling ├─ Code complexity: +500 linhas pra parsear respostas └─ Latência: +500ms (parsing cada resposta)
Padrão novo (Decisions API)
Você pede ao Decisions API: schema = { "decision": {"enum": ["urgent", "normal", "spam"]}, "confidence": {"type": "number", "min": 0, "max": 1} }
API retorna: { "decision": "urgent", "confidence": 0.92 }
Resultado: ├─ Success rate: 100% (output sempre valida schema) ├─ Você precisa: Zero parsing (é JSON válido) ├─ Code complexity: 5 linhas (conforme o schema) └─ Latência: -50% (sem parsing overhead)
Como Decisions API funciona
Architecture
Antes (Completions API): ┌─────────────────────────────────────┐ │ You │ │ "Classify email: urgent/normal/spam"│ └────────────────┬────────────────────┘ │ ▼ ┌─────────────────────────────────────┐ │ OpenAI Completions API │ │ Generates text token by token │ └────────────────┬────────────────────┘ │ ▼ (text) ┌─────────────────────────────────────┐ │ Your code │ │ Parse regex, error handling │ └────────────────┬────────────────────┘ │ ▼ (JSON) ┌─────────────────────────────────────┐ │ Your application logic │ │ if category == "urgent": ... │ └─────────────────────────────────────┘
Depois (Decisions API): ┌─────────────────────────────────────┐ │ You │ │ "Classify email" │ │ + schema: {decision: enum[...]} │ └────────────────┬────────────────────┘ │ ▼ ┌─────────────────────────────────────┐ │ OpenAI Decisions API │ │ Generates structured output │ │ respecting schema │ └────────────────┬────────────────────┘ │ ▼ (JSON, schema-valid) ┌─────────────────────────────────────┐ │ Your application logic │ │ if decision == "urgent": ... │ └─────────────────────────────────────┘
Diferença: ├─ Parsing layer: Eliminada (API retorna JSON) ├─ Latência: 10x mais rápido (sem parsing) ├─ Reliability: 100% (schema validation garantido) └─ Code: 10x simpler
Exemplo prático: Email classification
Antes (Completions API):
python import openai import re
def classify_email_old(email_text: str) -> str: """ Classify email using Completions API (text-based) """ response = openai.ChatCompletion.create( model="gpt-4", messages=[ { "role": "system", "content": "Classify as: urgent, normal, spam. Answer with ONE word only." }, {"role": "user", "content": email_text} ] )
text = response.choices[0].message.content.strip().lower()
# Parse (lots of error handling)
if "urgent" in text:
return "urgent"
elif "spam" in text:
return "spam"
elif "normal" in text:
return "normal"
else:
# Unparseable! What to do?
return "unknown" # Fallback (loss of information)
Problem: 20% of responses unparseable
Latency: 2s (LLM) + parsing = 2.5s total
Reliability: 80%
Depois (Decisions API):
python import openai
def classify_email_new(email_text: str) -> str: """ Classify email using Decisions API (structured) """ response = openai.beta.decisions.create( model="gpt-4", schema={ "type": "object", "properties": { "decision": { "enum": ["urgent", "normal", "spam"] }, "confidence": { "type": "number", "minimum": 0, "maximum": 1 } }, "required": ["decision", "confidence"] }, messages=[ { "role": "system", "content": "Classify email into one of: urgent, normal, spam." }, {"role": "user", "content": email_text} ] )
# No parsing needed! JSON is already valid
decision = response.parsed.decision # Direct access
confidence = response.parsed.confidence
return decision
Benefit: 100% parseable (schema enforced)
Latency: 2s (LLM only, no parsing overhead) = 10x faster
Reliability: 100%
Real-world use cases para agentes
Use case 1: Lead qualification (sales)
python
Your sales agent receives: Customer inquiry
Decision needed: Is this a qualified lead?
schema = { "type": "object", "properties": { "qualified": {"type": "boolean"}, "segment": { "enum": ["enterprise", "mid-market", "startup"] }, "priority": { "enum": ["hot", "warm", "cold"] }, "confidence": {"type": "number"}, "next_action": { "enum": [ "call_immediately", "send_demo", "nurture_sequence", "reject" ] } }, "required": ["qualified", "segment", "priority", "next_action"] }
Before (text parsing):
Agent output: "This looks like a warm mid-market lead, maybe enterprise?"
Your code: Confused (is it qualified? is it mid-market?)
Result: Manual review needed
After (Decisions API):
Agent output: {"qualified": true, "segment": "mid-market", "priority": "warm", "next_action": "send_demo"}
Your code: Automatic routing
Result: Lead goes to sales, demo scheduled
Use case 2: Customer support triage
python schema = { "type": "object", "properties": { "issue_category": { "enum": [ "billing", "technical", "feature_request", "complaint", "general" ] }, "urgency": { "enum": ["critical", "high", "medium", "low"] }, "should_escalate": {"type": "boolean"}, "suggested_team": { "enum": [ "billing_team", "engineering_team", "product_team", "executive_team" ] }, "estimated_resolution_time": { "type": "string", "enum": ["<1 hour", "<4 hours", "<1 day", ">1 day"] } } }
Before: Agent outputs "URGENT - NEEDS ESCALATION!!!"
Your code: Guess what means escalate (is ??? urgent enough?)
After: {"urgency": "critical", "should_escalate": true, "suggested_team": "engineering_team"}
Your code: Auto-escalate to engineering immediately
Use case 3: Content moderation
python schema = { "type": "object", "properties": { "is_appropriate": {"type": "boolean"}, "violation_type": { "enum": [ "none", "harassment", "spam", "misinformation", "nsfw", "other" ] }, "action": { "enum": [ "approve", "shadow_ban", "remove", "flag_for_review" ] }, "confidence": {"type": "number"} } }
Before: Agent outputs: "This might be spam but I'm not sure"
Your code: Guess (remove? flag? approve?)
After: {"is_appropriate": false, "violation_type": "spam", "action": "remove", "confidence": 0.98}
Your code: Remove immediately (98% confident)
Performance: Decisions API vs Completions API
Latência
Completions API (text-based): ├─ LLM inference: 2.0s ├─ Token generation: 0.5s (generate 50 tokens) ├─ Network: 0.2s ├─ Client parsing: 0.5s (regex, error handling) └─ Total: 3.2s
Decisions API (structured): ├─ LLM inference: 2.0s ├─ Schema generation: 0.1s (generate JSON that validates schema) ├─ Network: 0.2s ├─ Client parsing: ~0s (no parsing needed) └─ Total: 2.3s
Improvement: 3.2s → 2.3s = 28% faster But OpenAI claims "10x faster" → probably including: ├─ Faster LLM (Decision-optimized model) ├─ Less token generation (shorter JSON vs prose) └─ Native schema validation (no post-processing)
Realistic improvement: 2-5x faster depending on complexity
Reliability
Completions API: ├─ Success rate: 60-80% (parsing works) ├─ Failure rate: 20-40% (unparseable output) ├─ Human review needed: YES (when parsing fails) └─ SLA: 95% for truly automated decisions
Decisions API: ├─ Success rate: 100% (schema enforced) ├─ Failure rate: 0% (API validates output) ├─ Human review needed: NO (automated) └─ SLA: 99.9% (fully automated)
Cost
Completions API: ├─ Tokens used: 50-100 (longer responses) ├─ Cost: ~$0.001 per request ├─ Infrastructure: Parsing servers, error handling ├─ Human overhead: 20% (manual review of failures) └─ Total cost per decision: $0.001 + labor
Decisions API: ├─ Tokens used: 10-20 (shorter, structured responses) ├─ Cost: ~$0.0002 per request (5x cheaper tokens) ├─ Infrastructure: None (native parsing) ├─ Human overhead: 0% (fully automated) └─ Total cost per decision: $0.0002
Savings: 5-10x cheaper when including human overhead
Quando usar Decisions API
Use Decisions API se:
✅ Output é "categorical" (enum, boolean, number) └─ Ex: Is this urgent? (yes/no), Category? (a/b/c), Score? (0-100)
✅ Você precisa de parsing confiável └─ Ex: Customer support routing (must be 100% accurate)
✅ Latência é crítica (<2 segundos) └─ Ex: Real-time decisions, WhatsApp agent response
✅ Volume é alto (1M+ decisions/mês) └─ Razão: 5x cheaper = saves money at scale
✅ Você quer fully automated (sem human review) └─ Ex: Content moderation, lead routing
NÃO use Decisions API se:
❌ Output é "generative" (texto longo, narrativa) └─ Ex: "Write a customer email", "Summarize this doc" └─ Razão: Decisionsapi é pra structured output
❌ Output é "fuzzy" (não cabe em schema) └─ Ex: "Explain why this bug happened" └─ Razão: Schema não consegue capturar complexidade
❌ Você precisa de justificativa (explain your decision) └─ Ex: "Why you approved this loan?" └─ Razão: Decisions API retorna decisão, não raciocínio └─ Solution: Use Decisions API + Completions API together
Implementação: Decisioning agent com WhatsApp
Arquitetura
WhatsApp Message │ ▼ ┌─────────────────────┐ │ Your Agent (Claude) │ │ - Receives message │ │ - Understands context └──────────┬──────────┘ │ ▼ ┌─────────────────────┐ │ Decisions API │ │ - Classifies intent │ │ - Extracts decision │ │ - Returns {decision}│ └──────────┬──────────┘ │ ▼ ┌─────────────────────┐ │ Your Code (Logic) │ │ if decision == X: │ │ → Route/Execute │ └──────────┬──────────┘ │ ▼ Action (Send response, Create ticket, Approve order)
Code
python from openai import OpenAI import json
client = OpenAI()
def whatsapp_agent(customer_message: str) -> str: """ WhatsApp support agent that decides what to do """
# Step 1: Decisions API classifies intent
decision = client.beta.decisions.create(
model="gpt-4",
schema={
"type": "object",
"properties": {
"intent": {
"enum": [
"ask_status",
"request_refund",
"report_issue",
"feature_request",
"general_question"
]
},
"priority": {
"enum": ["high", "medium", "low"]
},
"needs_human": {"type": "boolean"},
"suggested_action": {
"enum": [
"auto_respond",
"create_ticket",
"escalate",
"send_info"
]
}
},
"required": ["intent", "priority", "needs_human", "suggested_action"]
},
messages=[
{
"role": "system",
"content": "You are a WhatsApp support agent. Analyze customer intent and decide what to do."
},
{"role": "user", "content": customer_message}
]
)
# Step 2: Extract decision (100% valid JSON, no parsing errors)
intent = decision.parsed.intent
priority = decision.parsed.priority
needs_human = decision.parsed.needs_human
action = decision.parsed.suggested_action
# Step 3: Route based on decision
if action == "auto_respond":
return generate_auto_response(intent)
elif action == "create_ticket":
ticket_id = create_support_ticket(customer_message, priority)
return f"Ticket #{ticket_id} created. We'll help you soon!"
elif action == "escalate":
escalate_to_human(customer_message, priority)
return "Escalating to specialist..."
else: # send_info
return send_relevant_info(intent)
def generate_auto_response(intent: str) -> str: """Generate automatic response based on intent""" responses = { "ask_status": "Your order is being shipped. Tracking: [link]", "request_refund": "You can request refund in Account Settings.", "report_issue": "We're sorry! Please describe the issue...", "feature_request": "Thanks for the suggestion! We'll consider it.", "general_question": "How can we help?" } return responses.get(intent, "How can we help?")
Usage
message = "I want to return my order and get a refund" response = whatsapp_agent(message) print(response) # → "You can request refund in Account Settings."
Comparação: APIs OpenAI pra agentes
┌─────────────────────┬──────────────┬─────────────┬─────────────┐ │ API │ Output type │ Speed │ Reliability │ ├─────────────────────┼──────────────┼─────────────┼─────────────┤ │ Completions │ Text │ 1x (base) │ 60% (parse) │ │ Structured Output │ JSON schema │ 3x faster │ 95% (valid) │ │ Decisions API │ JSON enum │ 10x faster │ 100% (enum) │ │ Vision API │ Multi-modal │ 2x slower │ Good (image)│ │ Reasoning API │ Step-by-step │ 5x slower │ Excellent │ └─────────────────────┴──────────────┴─────────────┴─────────────┘
Quando usar cada: ├─ Text output? → Completions API ├─ JSON schema? → Structured Output ├─ Simple enum decision? → Decisions API ⭐ (fastest, most reliable) ├─ Image understanding? → Vision API └─ Complex reasoning? → Reasoning API
Conclusão: Determinismo em agentes
O problema antigo:
LLM outputs são probabilísticos (não determinísticos) ├─ "Classifique em A, B, C" ├─ LLM pode responder: "A and B", "Maybe A?", "AB", "a" └─ Seu código: Tenta parsear (20% fail)
Resultado: ├─ Agentes não são confiáveis ├─ Precisam de fallback/human review └─ Enterprise não permite (compliance)
A solução nova:
Decisions API força LLM a respeitar schema ├─ "Classifique em A, B, C" ├─ API retorna: {"decision": "A"} (garantido) └─ Seu código: Confia cegamente (100% valid)
Resultado: ├─ Agentes são determinísticos ├─ Sem fallback (nunca falha) └─ Enterprise-ready (fully automated)
Bottom line:
Antes: Agente + parsing = fragile Depois: Agente + Decisions API = solid
Decisions API = "compilar" outputs LLM └─ De: "pode ser qualquer coisa" └─ Para: "é sempre válido"
Para agentes em produção: └─ Decisions API é não-negociável (reliability)
→ OpenClaw: Agentes com Decisions API (pronto)
Seu agente WhatsApp tá falhando em parsear? Decisions API resolve. 10x rápido, 100% confiável. 🚀
Publicado em 9 de outubro de 2026