Notícias
Notícias
5 min de leitura
7 de outubro de 2026

Decisões IA sem treinamento: Laya engine (zero-shot pronto)

Laya = engine open-source (421M params) que decide sem fine-tuning (zero-shot). Deploy decisions hoje (não meses). Custo 80% menos, 10x mais rápido.

Equipe OpenClaw

Equipe OpenClaw · Time de Engenharia & Produto

A Equipe OpenClaw é formada por engenheiros, designers e especialistas em IA dedicados a construir a melhor plataforma de agentes conversacionais para negócios brasileiros. Combinamos expertise…


Decisões IA sem treinamento: Laya engine (zero-shot pronto)

Notícia: Convai Innovations lançou Laya: engine open-source (421 milhões de parâmetros) que toma decisões complexas sem fine-tuning (zero-shot). Tornou-se um dos repositórios ML mais starred de setembro 2026.

Implicação: Você pode fazer seu agente IA DECIDIR hoje (não meses). Sem treinar modelo. Sem esperar dataset. Deploy production-ready em horas.

"Seu agente IA roda WhatsApp. Cliente pergunta: 'Posso devolver?' Agente antigo: Precisa treinar modelo (3-6 meses) pra aprender refund policy. Agente novo (Laya): Já sabe decidir (zero-shot), aprova/rejeita em <2s. Deploy hoje, não daqui a 6 meses."

What this means: Zero-shot = mudou o jogo (não precisa treinar mais).

Why it matters: Tempo = dinheiro. Você economiza 6 meses (time + GPUs). Concorrente que usa Laya = já está em produção (você ainda treinando).

Problem it reveals: Founders acreditam "decisions = precise training (6-12 meses)". Laya provou "decisions = zero-shot ready (1 dia)". Treinamento = unnecessary (na maioria dos casos).

Você quer agentes que DECIDAM rápido?

Laya mudou quanto tempo leva.


O que é Laya (exatamente)

Definition: zero-shot decision engine (não chatbot)

Laya vs. ChatGPT:

ChatGPT (texto generativo):

  • Lê pergunta → Gera texto → Retorna resposta
  • Exemplo: "Por que devo comprar seu produto?"
  • GPT: "Porque oferece 10 features, é rápido, é barato..."
  • Modelo: Autoregressive (token by token)
  • Latência: 5-10 segundos
  • Custo: R$ 0.0003 por chamada (adds up)

Laya (decisões estruturadas):

  • Lê contexto → Analisa opções → Retorna probabilidade
  • Exemplo: "Aprovar ou rejeitar refund?"
  • Laya: "APPROVE (92% confidence), REJECT (8%)"
  • Modelo: Non-autoregressive (single forward pass)
  • Latência: <500ms (150x mais rápido)
  • Custo: R$ 0 (open-source, roda local)

Diferença fundamental:

  • ChatGPT = "Generate text" (criativo, lento, caro)
  • Laya = "Choose option" (preciso, rápido, grátis)

How Laya works (simplified)

Architecture:

Input:

  • Context: "Customer bought R$ 299 item, 5 days ago, return reason: 'Changed mind'"
  • Options: ["APPROVE", "REJECT", "ESCALATE"]
  • Constraints: Policy allows returns within 30 days

Laya processes (single forward pass):

  • Reads context (encoder)
  • Analyzes options
  • Computes probability for each option

Output:

  • APPROVE: 87%
  • REJECT: 8%
  • ESCALATE: 5%

Result:

  • Decision: APPROVE (highest probability)
  • Confidence: 87%
  • Time: <500ms
  • Cost: R$ 0.00 (local, open-source)

Why zero-shot matters (game-changer)

Traditional approach (6-12 months):

Month 1: Collect training data (1000+ examples of decisions) Month 2-3: Label data (is this decision right? yes/no) Month 4-6: Train model (fine-tune LLM or build classifier) Month 7: Test + validation (does it work?) Month 8-9: Deploy + monitor (is it accurate in production?) Month 10-12: Improve + iterate (feedback loop)

Result: 1 year before agente can decide Cost: R$ 100K+ (data, infrastructure, people)

Laya approach (1 day):

Hour 1: Understand your decision (what are the options?) Hour 2: Define rules (when to approve/reject/escalate) Hour 3: Write prompt (context + options) Hour 4: Test with Laya (does it decide correctly?) Hour 5-8: Adjust + deploy (put in production)

Result: Same day deployment Cost: R$ 0 (open-source)

Speed comparison:

Traditional: 12 months deployment time Laya: 1 day deployment time Difference: 360x faster (12 months = 365 days)

Competitive advantage: Your agente is deciding. Competitor's is still training.


Why Laya is revolutionary (for SaaS)

Problem 1: Training data is hard

Old approach:

You want AI to decide: "Approve refund or not?"

Step 1: Collect 1000+ past refund decisions

  • Problem: Takes time (months)
  • Problem: Privacy (customer data)
  • Problem: Bias (historical decisions might be biased)

Step 2: Label each decision (correct or incorrect)

  • Problem: Labor-intensive (hire person to label)
  • Problem: Expensive (R$ 50-100 per decision)
  • Problem: Subjective (what's "correct" is opinion)

Result: You need R$ 50K-100K + 3-6 months just to get training data ready

Laya approach:

You want AI to decide: "Approve refund or not?"

Step 1: No data collection needed

  • Laya is pre-trained (learns from 421M parameters)
  • Already knows refund patterns (from general web data)
  • Can make decent decisions immediately

Step 2: No labeling needed

  • Laya is zero-shot (works without examples)
  • Can decide without seeing your specific examples
  • Accuracy: 80-90% out of the box

Result: Deploy today (zero data collection, zero labeling)

Problem 2: Model training is expensive + slow

Old approach:

Train refund classifier (LLM fine-tuning):

  • Cost: R$ 10K-50K (GPU time)
  • Time: 2-4 weeks (training)
  • Hardware: Need GPU (A100 = R$ 5K/month rental)
  • Expertise: Need ML engineer (R$ 200K/year)

Result: Expensive + slow + requires specialist

Laya approach:

Use Laya (zero-shot):

  • Cost: R$ 0 (open-source)
  • Time: 1 hour (setup)
  • Hardware: CPU fine (runs on laptop)
  • Expertise: Any developer (simple API)

Result: Free + fast + anyone can do it

Problem 3: Accuracy vs. cost tradeoff

ChatGPT approach (high cost, decent accuracy):

1000 refund decisions/day:

  • Cost: 1000 × R$ 0.0003 = R$ 0.30/day = R$ 9K/year
  • Accuracy: ~85% (good, but not perfect)
  • Latency: 5-10 seconds (customer waits)
  • Dependency: OpenAI API (if down, you're stuck)

Problems:

  • Cost adds up (R$ 100K+ for big volume)
  • Latency is annoying (5-10s is slow)
  • API dependency (outages affect you)
  • Privacy (data goes to California)

Laya approach (low cost, good accuracy):

1000 refund decisions/day:

  • Cost: R$ 0 (open-source, local)
  • Accuracy: ~82-88% (comparable, sometimes better)
  • Latency: <500ms (instant)
  • Dependency: Your servers (under your control)

Benefits:

  • Cost zero (scale without worrying about price)
  • Latency fast (customer happy)
  • No API dependency (yours to control)
  • Privacy 100% (data never leaves your servers)

Which is better?

ChatGPT: 85% accuracy, R$ 100K/year, 5-10s latency Laya: 85% accuracy, R$ 0/year, <500ms latency

Winner: Laya (same accuracy, 1000x cheaper, 10x faster)


How to use Laya (practical examples)

Use case 1: Refund approval (e-commerce)

Setup: python from laya import ZeroShotDecisionEngine

Initialize Laya

model = ZeroShotDecisionEngine.load("convai/laya-421m")

Define decision task

refund_decision = model.define_task( name="refund_approval", options=["APPROVE", "REJECT", "ESCALATE"], constraints=[ "Approve if within 30 days AND item unused", "Reject if outside return window OR used item", "Escalate if ambiguous OR high-value (>R$ 1K)" ] )

Make decision (zero-shot, no training!)

context = """ Customer: bought R$ 299 T-shirt 5 days ago Reason: Changed mind, found similar elsewhere cheaper Item condition: Unused, with tags Customer history: 10 purchases, 0 returns (trusted) """

decision = model.decide( task=refund_decision, context=context )

print(decision)

Output:

APPROVE: 89%

ESCALATE: 8%

REJECT: 3%

Result: Decision made in <500ms, no training required.

Use case 2: Lead qualification (sales)

Setup: python

Define lead scoring

lead_qualification = model.define_task( name="lead_scoring", options=["HOT (close deal now)", "WARM (nurture)", "COLD (skip)"], constraints=[ "HOT if budget confirmed AND timeline <30 days", "WARM if interested but timeline unclear", "COLD if no budget OR no urgency" ] )

Score a lead (zero-shot)

context = """ Company: Tech startup (50 employees) Product interest: CRM automation Budget: "R$ 50K-100K available this quarter" Timeline: "Need implementation by end of Q4" Decision maker: VP Sales (confirmed meeting scheduled) """

lead_score = model.decide( task=lead_qualification, context=context )

print(lead_score)

Output:

HOT: 78%

WARM: 18%

COLD: 4%

Result: Sales team knows HOT lead (close today) without manual scoring.

Use case 3: Support escalation (customer service)

Setup: python

Define escalation logic

escalation_rule = model.define_task( name="support_escalation", options=["RESOLVE (bot only)", "ESCALATE (human)"], constraints=[ "RESOLVE if simple question (FAQ-able)", "ESCALATE if complex OR emotional OR angry" ] )

Check ticket (zero-shot)

context = """ Ticket: "Why is my order late??? This is unacceptable! I paid extra for fast shipping!" Customer sentiment: Angry (caps, multiple ?) Issue type: Order status Customer history: First purchase, no prior issues """

routing = model.decide( task=escalation_rule, context=context )

print(routing)

Output:

ESCALATE: 92% (human should handle angry customer)

RESOLVE: 8%

Result: Angry customers go to human (keep them happy). Boring tickets stay with bot.

Use case 4: Content moderation (platforms)

Setup: python

Define moderation rules

moderation = model.define_task( name="content_moderation", options=["ALLOW", "REMOVE", "FLAG_FOR_REVIEW"], constraints=[ "ALLOW if normal, safe content", "REMOVE if clearly violates policy", "FLAG if borderline (let human decide)" ] )

Check post (zero-shot)

context = """ Post text: "I hate this new feature, it's terrible and whoever designed it should feel bad." Post type: Product feedback/criticism Author history: Active user, 100+ posts, no prior violations Report count: 0 (no one flagged it) """

moderation_decision = model.decide( task=moderation, context=context )

print(moderation_decision)

Output:

ALLOW: 85% (criticism is allowed, not harassment)

FLAG: 12%

REMOVE: 3%

Result: Legitimate criticism stays. Clear violations removed. Edge cases flagged.


Implementation guide (step-by-step)

Step 1: Download + setup (15 minutes)

Install Laya: bash

Clone from HuggingFace

git clone https://huggingface.co/convai/laya-421m cd laya-421m

Or use pip

pip install laya-engine

Requirements:

RAM: 2GB (tiny!) GPU: Optional (CPU works fine, 500ms latency) Storage: 500MB (model weights) Python: 3.8+

Step 2: Define your decision task (30 minutes)

Specify what AI decides: python from laya import ZeroShotDecisionEngine

model = ZeroShotDecisionEngine.load("convai/laya-421m")

Define task

your_decision = model.define_task( name="your_decision_name", options=["Option A", "Option B", "Option C"], # The choices constraints=[ # The rules "Choose A if [condition 1]", "Choose B if [condition 2]", "Choose C if [edge case]" ] )

Step 3: Test with examples (1 hour)

Validate decisions: python

Test cases

test_cases = [ { "context": "Customer bought R$ 50 item, 2 days ago, unused", "expected": "APPROVE" }, { "context": "Customer bought R$ 50 item, 60 days ago, used", "expected": "REJECT" }, { "context": "Customer bought R$ 5000 item, 5 days ago, damaged in shipping", "expected": "ESCALATE" } ]

Run tests

for test in test_cases: decision = model.decide( task=your_decision, context=test["context"] ) top_choice = max(decision, key=decision.get) accuracy = top_choice == test["expected"] print(f"{test['context'][:30]}... → {top_choice} (expected {test['expected']}) ✓" if accuracy else "✗")

Step 4: Deploy (1-2 hours)

Create API endpoint: python from fastapi import FastAPI import uvicorn

app = FastAPI()

@app.post("/decide/refund") def decide_refund(scenario: dict): decision = model.decide( task=your_decision, context=scenario["context"] ) return {"decision": decision}

if name == "main": uvicorn.run(app, host="0.0.0.0", port=8000)

Run: bash python app.py

Server running at http://localhost:8000

Step 5: Integrate with your agente IA (1-2 hours)

Call from WhatsApp agent: python import requests

def handle_refund_request(customer_id, order_id, reason): # Get context customer = db.get_customer(customer_id) order = db.get_order(order_id)

# Prepare context for Laya
context = f"""
Customer: {customer['name']}
Purchase: R$ {order['amount']} ({order['days_ago']} days ago)
Item condition: {order['condition']}
Return reason: {reason}
Customer history: {customer['return_count']} returns (out of {customer['total_orders']} orders)
"""

# Get Laya decision
response = requests.post(
    "http://localhost:8000/decide/refund",
    json={"context": context}
)

decision = response.json()["decision"]
top_choice = max(decision, key=decision.get)
confidence = decision[top_choice]

# Send to customer
if top_choice == "APPROVE":
    send_message(f"Sua devolução foi aprovada! {confidence:.0%} confiança.")
    process_refund(order_id)
elif top_choice == "REJECT":
    send_message("Desculpa, sua devolução não foi aprovada (fora da janela de 30 dias).")
else:  # ESCALATE
    send_message("Vou encaminhar seu caso para um agente. Aguarde...")
    escalate_to_human(customer_id, order_id)

Step 6: Monitor + iterate (ongoing)

Track performance: python

Log every decision

log_decision({ "task": task_name, "context": context, "laya_decision": decision, "actual_outcome": actual_outcome, # Did we make the right call? "timestamp": datetime.now() })

Weekly audit

accuracy = count_correct_decisions() / count_all_decisions() print(f"Laya accuracy: {accuracy:.1%}")

If accuracy drops, adjust constraints

if accuracy < 0.85: print("Accuracy dropped! Adjust decision constraints.")


Cost + Speed comparison

Scenario: SaaS making 10.000 decisions/day

Traditional approach (fine-tuned LLM):

Setup cost:

  • Training data: R$ 50K (collection + labeling)
  • Model training: R$ 20K (GPU time, 2 weeks)
  • Deployment: R$ 10K
  • Total setup: R$ 80K
  • Time to production: 12 months

Operating cost:

  • Fine-tuned model hosting: R$ 500/month
  • Maintenance: R$ 2K/month
  • Total operating: R$ 2.5K/month = R$ 30K/year

Total 1-year cost: R$ 80K + R$ 30K = R$ 110K Decisions you can make after 1 year: 3.65M (10K/day × 365 days) Cost per decision: R$ 0.03

Laya approach (zero-shot):

Setup cost:

  • Download model: R$ 0 (open-source)
  • Integration: R$ 5K (1 developer, 1 week)
  • Testing: R$ 5K (QA)
  • Total setup: R$ 10K
  • Time to production: 1 week

Operating cost:

  • Model hosting: R$ 200/month (tiny infra)
  • Maintenance: R$ 500/month
  • Total operating: R$ 700/month = R$ 8.4K/year

Total 1-year cost: R$ 10K + R$ 8.4K = R$ 18.4K Decisions you can make in 1 year: 3.65M (same) Cost per decision: R$ 0.005 (6x cheaper!)

Comparison:

Traditional: R$ 110K (12 months to deploy) Laya: R$ 18.4K (1 week to deploy)

Savings: R$ 91.6K (83% cheaper) Speed: 52x faster deployment Accuracy: Same (~85%)


Checklist: Is Laya right for your use case?

Answer these questions:

✓ Do you need to make decisions quickly (decisions/second)? → YES ✓ Are decisions "classification" (choose from options)? → YES (not text generation) ✓ Do you need low latency (<1 second)? → YES ✓ Do you care about cost? → YES (want cheap) ✓ Do you want to avoid training data hassle? → YES (zero-shot) ✓ Can you define decision rules clearly? → YES ✓ Do you have structured data (context + options)? → YES ✓ Do you need privacy (local, not cloud)? → YES

If YES to all: Laya is perfect. If YES to 6+: Laya is good fit. If YES to 4-5: Consider Laya (with ChatGPT fallback). If YES to <4: Maybe ChatGPT is better.


Conclusão: Zero-shot decisions are production-ready NOW

For your SaaS:

If you want agentes IA que DECIDAM (refund, lead score, escalate, moderate):

  1. Try Laya (this week)

    • Download model (R$ 0)
    • Define decision task (30 min)
    • Test on 10 examples (1 hour)
    • Deploy (2 hours)
    • Total effort: 4 hours (not 12 months)
  2. Measure accuracy (first month)

    • Log all decisions
    • Track correct vs. incorrect
    • Target: 85%+ accuracy
    • If <85%: Tweak decision rules
  3. Compare with alternatives (optional)

    • Run ChatGPT in parallel (A/B test)
    • Compare accuracy (same?)
    • Compare cost (Laya wins)
    • Compare latency (Laya wins)
    • Compare privacy (Laya wins)
  4. Scale (3-6 months)

    • Deploy to 100% of decisions
    • Monitor accuracy (maintain 85%+)
    • Improve rules (based on feedback)
    • Save R$ 100K+ vs. traditional approach

Expected outcome: Your agentes decide rápido (no delays), barato (R$ 0 vs. R$ 100K), privado (local, not cloud), e sem treinar nada (zero-shot magic).


Agentes IA com decisões zero-shot (framework pronto)

Se você quer implementar agentes IA que DECIDAM rápido (sem treinar meses), você precisa de framework que:

  • Usa Laya (421M params, zero-shot, open-source)
  • Define decision tasks (options + constraints)
  • Implements fast decisions (<500ms latency)
  • Logs all decisions (audit trail)
  • Monitors accuracy (85%+ target)
  • Provides fallback logic (escalate if uncertain)
  • Integrates with your agente (API simple)
  • Tracks cost savings (vs. ChatGPT baseline)
  • Compares alternatives (ChatGPT, frontier models)
  • Provides guardrails (prevent bad decisions)

OpenClaw Zero-Shot Decision Framework:

  • Pre-built Laya setup (download + deploy em 1 hora)
  • Decision templates (refund, fraud, lead scoring, escalation, moderation)
  • Accuracy monitoring (track 85%+ target)
  • A/B testing (compare Laya vs. ChatGPT)
  • Guardrails (prevent edge cases)
  • Cost dashboard (track savings vs. frontier models)
  • Scaling playbook (1 decision/day → 1M decisions/day)
  • Integration examples (WhatsApp, Slack, HTTP API)

Use case: "Implemented Laya with OpenClaw Framework. Deployed refund decisions in 4 hours (vs. 12 months traditional). Accuracy 87% (vs. 85% ChatGPT). Cost R$ 18K/year (vs. R$ 110K traditional). Paid for itself in 2 weeks. Now making 10K decisions/day, zero delays."

Decisões zero-shot + production-ready → OpenClaw Zero-Shot Framework

Não espere 12 meses. Laya é open-source, grátis, pronto. Deploy hoje. Economize 83% em custo de decisões. Acelere 52x. 🚀


Publicado em 7 de outubro de 2026

Leia também