Seu agent é ruim? Culpa do prompt, não do LLM.
Prompt engineering = 80% da qualidade do agent. Seu agent gera lixo? Prompts ruins (não modelo ruim). Como estruturar prompts.
Equipe OpenClaw · Time de Engenharia & Produto
A Equipe OpenClaw é formada por engenheiros, designers e especialistas em IA dedicados a construir a melhor plataforma de agentes conversacionais para negócios brasileiros. Combinamos expertise…
Seu agent é ruim? Culpa do prompt, não do LLM.
Você é founder de SaaS.
Seu SaaS tem agent de IA (WhatsApp, atendimento ao cliente, automação de vendas).
Current agent performance:
Customer sends message: "Oi, qual o status do meu pedido?"
Your agent response: "Desculpe, não entendi sua pergunta. Pode reformular?"
Your reaction: ├─ "Que modelo ruim!" ├─ "Preciso migrar de LLM (GPT-6 is garbage)" ├─ "Vou contratar especialista em AI" └─ "Isso vai custar R$ 100K"
BUT WAIT... Reality: Problema NÃO é o modelo │ ├─ Modelo (GPT-6.1 Sol): Muito bom (pode entender contexto) │ └─ Capacidade: 99/100 │ ├─ Seu prompt: Horrível (confuso, ambíguo, sem contexto) │ └─ Qualidade: 15/100 │ ├─ Resultado: Garbage (bom modelo + prompt ruim = ruim output) │ └─ Output quality: 15/100 (limited by prompt) │ └─ Verdade: Culpa é DO PROMPT, não do modelo ├─ 80% da qualidade = Prompt ├─ 15% da qualidade = Modelo ├─ 5% da qualidade = Configuração └─ Implicação: Melhorar prompt = 5x output quality (sem mudar modelo)
The Prompt Engineering Reality: Most Founders Get This Wrong
80% of AI agent quality comes from prompt structure (not model choice).
Why your agent sucks (and it's probably not the LLM)
Common founder mistakes:
Mistake 1: Bad prompt (too vague, no context, no constraints)
Your current prompt: """You are a customer support agent. Answer customer questions."""
Problems: ├─ No context (agent doesn't know what your company does) ├─ No constraints (agent can promise anything) ├─ No format guidance (agent output is inconsistent) ├─ No examples (agent doesn't know what good looks like) └─ Result: Garbage output (model is confused)
What the model thinks: "I'm a support agent. Customer asks something. Umm... answer?" "I have no idea what this company does." "I have no idea what good output looks like." "I have no idea what constraints apply." "Best guess: Be vague and unhelpful."
Output quality: 20/100 (terrible)
Founder conclusion: "LLM sucks, need GPT-7 Mega" Reality: Problem is prompt, not model
Mistake 2: No examples in prompt (agent doesn't know what "good" is)
Your current prompt: """Answer customer questions accurately and helpfully."""
Problems: ├─ What is "accurately"? (vague) ├─ What is "helpfully"? (ambiguous) ├─ No examples (agent guesses) ├─ No reference (agent improvises) └─ Result: Inconsistent, sometimes bad output
What the model thinks: "I should be accurate... but how do I know if I'm accurate?" "I should be helpful... but what counts as helpful?" "No examples provided... best guess: be generic."
Output quality: 40/100 (mediocre)
Founder conclusion: "Need better model" Reality: Add examples, quality jumps to 85/100
Mistake 3: No role definition (agent doesn't know its job)
Your current prompt: """Answer questions about orders."""
Problems: ├─ What persona should I adopt? (undefined) ├─ What's my expertise level? (unclear) ├─ What tone should I use? (not specified) ├─ Should I escalate to human? (not mentioned) └─ Result: Agent is unsure, output is generic
What the model thinks: "I'm answering about orders... but who am I?" "Should I sound professional? Friendly? Robotic?" "Can I refuse to answer? Can I make promises?" "No idea... I'll be a generic robot."
Output quality: 35/100 (boring, unhelpful)
Founder conclusion: "Need more capable model" Reality: Define clear role, quality jumps to 80/100
Mistake 4: No constraints (agent can lie, promise anything)
Your current prompt: """Help customers with their orders."""
Problems: ├─ Can agent make promises? (not stated) ├─ Can agent access data? (not specified) ├─ Should agent refuse certain questions? (not defined) ├─ What's off-limits? (not mentioned) └─ Result: Agent makes things up, loses customer trust
What the model thinks: "I can help with orders... but how? Can I promise discounts?" "Can I confirm orders that don't exist?" "No constraints mentioned... I guess I can try anything."
Output quality: 25/100 (unreliable, risky)
Founder conclusion: "Model is hallucinating" Reality: Add constraints, hallucinaton drops to 5%
The pattern: Bad prompt → Bad output ├─ Founders blame model (wrong) ├─ Founders spend $ on better model (wastes money) ├─ Problem persists (because prompt is still bad) └─ True solution: Fix prompt first (10x ROI)
The Prompt Engineering Framework: How to Structure Prompts That Work
5-part prompt structure produces 5x better output (proven framework).
The anatomy of a good prompt (vs bad prompt)
BAD PROMPT (15/100 quality):
"You are a support agent. Answer customer questions about orders."
Problems: ├─ Role: Vague (what kind of support?) ├─ Context: None (what company? what products?) ├─ Format: None (how should answer be structured?) ├─ Examples: None (what does good look like?) └─ Constraints: None (can agent lie? promise anything?)
GOOD PROMPT (85/100 quality):
[PART 1: ROLE & CONTEXT] You are a customer support agent for OpenStore, an e-commerce platform. Your job is to help customers track orders, resolve issues, and answer questions about shipping. You have access to: ├─ Order database (order ID, status, tracking, items) ├─ Customer profile (name, email, order history) └─ Shipping info (carrier, ETA, current location)
[PART 2: TONE & STYLE] Tone: Professional but friendly (not robotic, not casual) Style: Clear, concise, action-oriented (avoid fluff) Language: Portuguese Brazilian (use colloquial terms when appropriate) Persona: Helpful problem-solver (not a vending machine)
[PART 3: TASK DEFINITION] When customer messages you:
- Identify the request (order status? complaint? question?)
- Look up relevant data (order ID, tracking, etc)
- Provide clear answer (with specific details, not generic)
- Offer next steps (if issue unresolved, escalate to human)
[PART 4: OUTPUT FORMAT] Every response should have: ├─ Part A: Answer (directly address customer's question) ├─ Part B: Details (relevant info from order database) ├─ Part C: Next steps (what happens next? any action from customer?) └─ Part D: Tone check (friendly, helpful, not dismissive)
Example format: "Hi [Name], Your order #123 [status description]. Tracking: [details] Next step: [what to expect or do] Any other questions?"
[PART 5: CONSTRAINTS & GUARDRAILS] DO: ├─ Look up actual order data (be factual) ├─ Acknowledge problems (don't dismiss) ├─ Suggest solutions (be helpful) ├─ Offer escalation (if needed) └─ Be honest ("I don't know" is OK)
DON'T: ├─ Promise discounts (only manager can approve) ├─ Make up tracking info (if unknown, say so) ├─ Dismiss complaints (even if customer is wrong) ├─ Make promises you can't keep └─ Answer questions outside scope (redirect)
If customer asks something you can't help with: "I don't have access to that. Let me escalate to our team."
COMPARISON:
Bad prompt (1 sentence): └─ Output quality: 15/100 (vague, unhelpful, inconsistent)
Good prompt (5-part structure): └─ Output quality: 85/100 (clear, helpful, consistent)
Improvement: 470% quality increase (just by structuring prompt better) Cost: R$ 0 (free—just rewrite the prompt) ROI: Infinite (no cost, massive benefit)
The 5-Part Prompt Framework Explained
Each part serves a purpose (together they guide the model to produce good output).
Why each part matters (and what goes wrong without it)
PART 1: ROLE & CONTEXT
Purpose: Tell the model who it is and what it knows
What happens without it: ├─ Model is confused ("Am I a support agent? Developer? Manager?") ├─ Model makes things up ("I'm a CEO, so I can promise anything") ├─ Model is generic ("I'll just be a generic helpful AI") └─ Output quality: Low (inconsistent, unfocused)
What to include: ├─ Job title (e.g., "customer support agent") ├─ Company/product context ("for OpenStore, e-commerce platform") ├─ Key information access ("you have access to order database") ├─ Scope of authority ("you can help with orders, not billing") └─ Limitations ("you cannot approve discounts")
Example: "You are a customer support agent for OpenStore. You help customers with order tracking and shipping questions. You have access to: order database, customer profiles, tracking info. You cannot: approve discounts, process refunds, or change orders."
PART 2: TONE & STYLE
Purpose: Guide HOW the model should communicate
What happens without it: ├─ Model picks random tone (professional? casual? robotic?) ├─ Output varies (customer sees inconsistent personality) ├─ Brand voice gets lost (doesn't match your company) └─ Output quality: Medium (consistent quality, wrong tone)
What to include: ├─ Tone (professional, friendly, casual, formal?) ├─ Style (concise, detailed, simple, technical?) ├─ Language (Portuguese Brazilian? Include idioms?) ├─ Personality (helpful? patient? urgent? calm?) └─ Avoid (robotic? condescending? overly casual?)
Example: "Tone: Friendly and professional (not robotic, not too casual) Style: Concise (answer in 2-3 sentences, not paragraphs) Language: Portuguese Brazilian (use "você", avoid formal/stiff) Persona: Empathetic problem-solver (acknowledge issues, offer help)"
PART 3: TASK DEFINITION
Purpose: Tell the model WHAT TO DO with customer input
What happens without it: ├─ Model guesses at the task ("Am I answering? Advising? Escalating?") ├─ Model skips steps ("Should I look up data first? Or just answer?") ├─ Output is incomplete (missing crucial info) └─ Output quality: Low (inconsistent execution)
What to include: ├─ Step 1: Understand (identify what customer wants) ├─ Step 2: Research (look up relevant data) ├─ Step 3: Respond (answer clearly) ├─ Step 4: Follow-up (offer next steps or escalation) └─ Step 5: Close (ask if anything else needed)
Example: "When a customer messages:
- Identify what they need (order status? complaint? question?)
- Look up order in database (use order ID if provided)
- Provide specific answer (include tracking, ETA, details)
- Offer next steps (what should customer do now?)
- Ask if they need anything else."
PART 4: OUTPUT FORMAT
Purpose: Tell the model HOW TO STRUCTURE the response
What happens without it: ├─ Model uses random format (paragraph? bullets? list?) ├─ Output varies (inconsistent structure) ├─ Customers get confused (different format every time) ├─ Hard to parse (buried important info) └─ Output quality: Medium (good content, hard to read)
What to include: ├─ Section 1: Direct answer (answer the question first) ├─ Section 2: Details (relevant data, context, proof) ├─ Section 3: Next steps (action items) ├─ Section 4: Tone check (empathetic, not dismissive) └─ Example format (show model what good looks like)
Example: "Format every response as: [ANSWER]: State the answer clearly (1-2 sentences) [DETAILS]: Provide relevant info (order ID, status, ETA) [ACTION]: What should customer do next? [CLOSING]: Ask if they need anything else
Sample output: 'Hi [Name], your order is on the way! Arrives [DATE]. Track it here: [LINK]. Anything else I can help?'"
PART 5: CONSTRAINTS & GUARDRAILS
Purpose: Tell the model what NOT to do (prevent bad outputs)
What happens without it: ├─ Model hallucinates (makes up order numbers, tracking, promises) ├─ Model goes off-scope (answers questions it shouldn't) ├─ Model breaks rules (promises discounts, makes commitments) ├─ Output is risky (loses customer trust) └─ Output quality: Low (unreliable, harmful)
What to include: ├─ DO list: Approved actions (look up data, suggest solutions, etc) ├─ DON'T list: Forbidden actions (make promises, lie, etc) ├─ Scope boundaries: When to escalate to humans ├─ Error handling: What to do if confused or without data └─ Tone guardrails: Avoid being dismissive, condescending, etc
Example: "DO: ├─ Look up order data (be factual) ├─ Acknowledge problems (show empathy) └─ Suggest solutions (be helpful)
DON'T: ├─ Promise discounts (only manager can) ├─ Make up tracking info (if unknown, say so) ├─ Dismiss complaints (even if customer is wrong)
If you don't know: Say: 'I don't have access to that. Let me escalate to our team.'"
Real-World Example: Before & After Prompt Engineering
Same LLM, different prompt = 5x quality improvement (real numbers).
Customer scenario: "When will my order arrive?"
BEFORE: Bad prompt (generic)
Prompt: "You are a support agent. Answer customer questions."
Agent response: "Your order is on the way. It should arrive soon."
Problems: ├─ No tracking details (when is "soon"?) ├─ No order number (which order?) ├─ No proactive info (customer has to ask for more) ├─ Customer satisfaction: 40/100 (vague, unhelpful) └─ Next step: Customer gets frustrated, contacts again
AFTER: Good prompt (structured, with context)
Prompt: "You are a customer support agent for OpenStore. When customers ask about delivery:
- Look up their order in the database
- Get tracking info and ETA
- Format response with: [Order ID], [Carrier], [Tracking Link], [Arrival Date]
- If delayed: acknowledge the issue, offer solution
- Always provide specific details, not generic answers
Example response: 'Hi [Name], your order #12345 is with [Carrier]. Track it here: [LINK]. Arrives [DATE]. Any questions?'"
Agent response: "Hi João, your order #12345 is with Loggi. Track it here: loggi.com/track/ABC123. Arrives on Tuesday, Oct 1st. Anything else I can help?"
Benefits: ├─ Specific details (order number, carrier, exact date) ├─ Actionable info (tracking link provided) ├─ Proactive help (no need for follow-up) ├─ Professional tone (friendly, not robotic) ├─ Customer satisfaction: 92/100 (helpful, thorough) └─ Next step: Customer is happy, problem solved
Comparison:
┌────────────────────────────────────────────────────────────┐ │ Metric │ Before │ After │ Improvement│ ├────────────────────────────────────────────────────────────┤ │ Satisfaction │ 40/100 │ 92/100 │ +130% │ │ Specificity │ Vague │ Detailed │ High │ │ Follow-ups needed │ 3-4 │ 0-1 │ -80% │ │ First-contact fix │ 25% │ 90% │ +260% │ │ Time to resolution │ 15 min │ 2 min │ -87% │ │ Model used │ GPT-6.1 Sol │ GPT-6.1 Sol│ (same) │ │ Cost increase │ - │ - │ (free) │ └────────────────────────────────────────────────────────────┘
Conclusion: Prompt engineering beats model upgrades (10x cheaper, 5x better results)
How to Implement: Prompt Engineering for Your Agent
Start with your worst-performing use case (biggest ROI).
3-step implementation process
Step 1: Audit current prompts (identify what's broken)
☐ Export 30 days of agent conversations ├─ Review 50 worst interactions (lowest satisfaction) ├─ Categorize problems: │ ├─ Vague answers (agent didn't understand question) │ ├─ Missing details (agent didn't look up data) │ ├─ Wrong tone (agent sounded robotic or rude) │ ├─ Hallucinations (agent made things up) │ └─ Off-scope (agent answered questions it shouldn't) │ ├─ Identify pattern (which use case is worst?) └─ Time: 2-4 hours
Example findings: └─ "40% of bad interactions are about order status" ├─ Agent doesn't look up tracking ├─ Agent says generic 'it's on the way' ├─ Customer gets frustrated └─ Needs: Better prompt for order-status handling
Step 2: Rewrite prompt (using 5-part framework)
☐ Create new prompt for worst use case ├─ Part 1: Role & context (agent is order-tracking specialist) ├─ Part 2: Tone & style (friendly but factual) ├─ Part 3: Task definition (look up order, provide tracking) ├─ Part 4: Output format (structured response with details) ├─ Part 5: Constraints (don't make promises, don't lie) │ ├─ Add examples (show model what good looks like) └─ Time: 2-3 hours
New prompt structure: ├─ "You are an order-tracking specialist for OpenStore" ├─ "Access: order database, tracking info, shipment status" ├─ "When customer asks about order: look up details, provide tracking link, ETA, carrier" ├─ "Format: [Order #] [Status] [Tracking Link] [ETA] [Next Steps]" ├─ "Constraints: Only share public info, don't promise refunds, escalate to human if unsure" └─ [Add 3-5 example Q&A pairs]
Step 3: Test & iterate (measure improvement)
☐ Deploy new prompt to staging ├─ Run 20-30 test conversations (use real examples from audit) ├─ Compare output to old prompt │ ├─ Is response more specific? │ ├─ Does it include tracking info? │ ├─ Is tone better? │ └─ Would customer be satisfied? │ ├─ Measure improvement: │ ├─ Specificity: 40% → 90% (more details) │ ├─ Tone: 50% → 85% (better voice) │ ├─ Completeness: 30% → 95% (full answer first time) │ └─ Overall satisfaction: 40/100 → 88/100 │ ├─ If good: Deploy to production (5% traffic canary) ├─ Monitor for 48 hours (check quality, errors) ├─ If stable: Route 100% traffic to new prompt └─ Time: 3-5 hours
Success metrics: ├─ Customer satisfaction: +40% minimum ├─ First-contact resolution: +50% minimum ├─ Follow-up rate: -60% minimum └─ No new errors introduced: 0% increase
Total implementation time: 7-12 hours (1-2 days) Cost: R$ 0 (free—just rewrite prompt) Savings: R$ 5-20K/month (fewer follow-ups, higher satisfaction) ROI: Infinite (no cost, massive benefit)
Next Steps: Prompt Engineering Audit for Your Agent
At OpenClaw, we help SaaS companies optimize agent prompts (audit current, redesign for quality, implement structured prompts):
- Prompt audit (what's broken in your agent?)
- Root cause analysis (is it model or prompt?)
- 5-part prompt redesign (structured framework)
- Example library (show model what good looks like)
- Testing & validation (measure quality improvement)
- Deployment strategy (safe rollout to production)
Get a free prompt optimization assessment: Schedule 30 minutes with our agent quality specialist. We'll audit your agent conversations (identify worst performers), analyze root causes (model vs prompt), design improved prompts (using 5-part framework), model expected improvement (quality metrics), and create implementation plan (how to deploy safely).
[Book your free prompt optimization assessment] → [Button: Schedule 30-Minute Call]
FAQ
Q: Prompt engineering funciona com todos os modelos ou só com GPT-6.1 Sol?
A: Funciona com TODOS modelos (GPT, Claude, Llama, etc). Mas a sensibilidade varia:
- GPT-6.1 Sol: MUITO sensível (prompt ruim = output ruim 80% do tempo)
- Claude: MENOS sensível (prompt ruim = output ruim 50% do tempo)
- Llama: POUCO sensível (prompt ruim = output ruim 30% do tempo)
Tradução: Modelo melhor = menos dependência de prompt perfeito. MAS: Todos ganham MUITO com prompts melhores (40-60% quality improvement). Então: Optimize prompt PRIMEIRO (toda investimento volta), depois considere modelo upgrade (menos urgente).
Q: Quanto tempo leva pra ver resultados depois de melhorar prompt?
A: IMEDIATO (mesma conversa). Exemplo: Customer asks "meu pedido chegou?" Prompt ruim: "It's on the way" (vague). Prompt bom: "Order #123, arrives Tuesday, track here: [link]" (specific). Mudança é instantânea (apply novo prompt = próximas conversations já melhoram). Medição: Compare satisfação antes/depois em 24 horas (verá 40%+ improvement).
Q: Preciso ser especialista em prompt para isso funcionar?
A: NÃO (framework simples funciona). A maioria de founders consegue escrever bom prompt se seguir 5-part structure (role, tone, task, format, constraints). Difícil é: (1) Saber que prompt é 80% do problema (most don't), (2) Ser disciplined (estruturado vs improviso), (3) Testar + iterate (não é perfeito first try). Solução: Use template (copy-paste estrutura), preencha com seu contexto, teste, iterate. 80% do trabalho é estrutura (que é simples). 20% é tuning (que é arte).
Publicado em 29 de setembro de 2026