Notícias
Notícias
5 min de leitura
9 de outubro de 2026

Compartilhe GPU entre times (SageMaker HyperPod)

SageMaker HyperPod: GPU cluster compartilhado entre times (agentes, ML, pesquisa). Isolation + fairness. Custo -60%, GPUs eficientes. Multi-tenant infrastructure.

Equipe OpenClaw

Equipe OpenClaw · Time de Engenharia & Produto

A Equipe OpenClaw é formada por engenheiros, designers e especialistas em IA dedicados a construir a melhor plataforma de agentes conversacionais para negócios brasileiros. Combinamos expertise…


Compartilhe GPU entre times (SageMaker HyperPod)

Notícia: AWS lançou SageMaker HyperPod: plataforma pra compartilhar GPU clusters entre múltiplos times (agentes IA, ML teams, research) com isolation + fairness automática. Cada time consegue rodar seus workloads (agente WhatsApp, LLM training, computer vision, etc.) no MESMO cluster sem interferir um no outro.

Implicação: Suas GPUs caras finalmente trabalham 24h/dia (não ficam ociosas). Custo cai 60%, porque infra é shared.

**"Você é CTO de fintech com múltiplos times de IA:

  • Time A (agentes): Precisa de 2x GPU pra rodar agente WhatsApp (8h/dia)
  • Time B (ML): Precisa de 2x GPU pra treinar modelos (16h/dia)
  • Time C (pesquisa): Precisa de 2x GPU pra experimentar (8h/dia)

Cenário antigo (3 clusters separados): ├─ Custo total: 6 GPUs = R$ 300K/mês ├─ Utilização: 30% (muita ociosidade) ├─ Problema: Time A tem ociosidade, Time B precisa mais ├─ Resultado: Dinheiro jogado fora └─ Desperdício: R$ 210K/mês

Cenário novo (1 cluster compartilhado com HyperPod): ├─ Custo total: 4 GPUs = R$ 200K/mês ├─ Utilização: 90% (fair share automático) ├─ Problema: Resolvido (HyperPod aloca dinamicamente) ├─ Resultado: Mesmo output, 1/3 do custo └─ Economia: R$ 100K/mês

Annual impact: R$ 1.2M economizado."**


O problema: GPU clusters são caros e ficam ociosos

Por que você está desperdiçando dinheiro em GPUs

Economias de GPU (realidade atual):

Custo de GPU hoje (outubro 2026): ├─ 1x H100 (Nvidia): R$ 50K/mês (rental) ├─ 1x A100: R$ 35K/mês ├─ 1x RTX 4090: R$ 10K/mês ├─ 1x L4: R$ 5K/mês └─ Total cluster (8x GPUs mixed): R$ 200K+/mês

Problema real (multi-team): ├─ Time A (agentes): Quer 2x GPUs │ ├─ Usa: 8h/dia (agente WhatsApp online) │ ├─ Ociosa: 16h/dia (ninguém usando) │ ├─ Custo: R$ 60K/mês × 67% ociosidade = R$ 40K desperdício │ └─ Problema: Você paga 24h por 8h de uso │ ├─ Time B (ML): Quer 2x GPUs │ ├─ Usa: 16h/dia (training de modelos) │ ├─ Ociosa: 8h/dia (modelo pronto) │ ├─ Custo: R$ 60K/mês × 33% ociosidade = R$ 20K desperdício │ └─ Problema: Precisa de mais GPU, mas não tem budget │ ├─ Time C (pesquisa): Quer 2x GPUs │ ├─ Usa: 8h/dia (experimentos) │ ├─ Ociosa: 16h/dia (aguardando resultados) │ ├─ Custo: R$ 60K/mês × 67% ociosidade = R$ 40K desperdício │ └─ Problema: Pior utilização da empresa │ └─ Total desperdício: R$ 100K/mês (50% de R$ 200K)

Por que isso acontece? ├─ GPUs são "resource contended" (todos querem) ├─ Mas cycles não são sempre: Time A quer 8h, depois ociosa ├─ Você aloca 2x GPU pra Time A (paga 24h) ├─ Time B não consegue usar Time A's GPU (isolamento/segurança) ├─ Resultado: Desperdício inevitable (sem multi-tenancy) └─ Solução: Compartilhar dinamicamente (HyperPod)

Casos reais de desperdício (seu infra provavelmente está assim)

Caso 1: Startup de agentes IA ├─ Agente WhatsApp: Precisa H100 (inference) │ ├─ Pico: 12h (horário comercial) │ ├─ Ocioso: 12h (noite/madrugada) │ ├─ Custo: R$ 50K/mês │ └─ Desperdiço: R$ 25K/mês │ ├─ Agente WhatsApp roda 24h, mas precisa ser ONLINE │ ├─ Não pode "desligar" (cliente pode chamar 3am) │ ├─ Então paga full price │ └─ Desperdício: Inevitável (without HyperPod) │ └─ Solução (HyperPod): Compartilhe com team de ML ├─ Team ML treina durante o dia (quando agente está calminho) ├─ Agente fica em background (low-resource mode) ├─ Noite: Team pesquisa usa GPU (quando agente de novo low-resource) ├─ Resultado: 1 GPU faz trabalho de 3 └─ Economia: R$ 100K/mês

Caso 2: SaaS de código-generation (como você) ├─ API roda inference (model serving): Precisa 2x A100 │ ├─ Pico: 8h (office hours, devs usando) │ ├─ Ocioso: 16h (night, weekend) │ ├─ Custo: R$ 70K/mês │ └─ Desperdício: R$ 47K/mês │ ├─ Fine-tuning team treina modelos: Precisa 2x A100 │ ├─ Usa: 24h (training é 24h job) │ ├─ Mas: Caro demais, então treina 1x semana │ ├─ Custo: R$ 70K/mês │ └─ Problema: Team quer treinar mais, não tem budget │ └─ Solução (HyperPod): Compartilhe 4 A100 entre os 2 workloads ├─ API inference: 2x quando precisa (peak demand) ├─ Training: Usa 2x quando API tá baixo (night) ├─ Fair scheduler: Automático (não precisa de operador) ├─ Resultado: 4 GPUs fazem trabalho de 4 (vs. 5 antes) └─ Economia: R$ 35K/mês

Caso 3: Empresa com muitos times pequenos ├─ Team 1 (agentes): 1x GPU (8h/dia) ├─ Team 2 (ML): 1x GPU (16h/dia) ├─ Team 3 (pesquisa): 1x GPU (irregular) ├─ Team 4 (fine-tune): 1x GPU (episódico) │ ├─ Status quo: 4 GPUs separadas = R$ 140K/mês │ └─ Utilização média: 40% (60% desperdício) │ └─ Com HyperPod: 2 GPUs compartilhadas = R$ 70K/mês ├─ Utilização média: 85% (15% margem) ├─ Fair allocation automático └─ Economia: R$ 70K/mês


Solução: SageMaker HyperPod (GPU compartilhada com isolation)

O que é HyperPod (e como funciona)

Arquitetura:

╔══════════════════════════════════════════════════════════════╗ ║ Amazon SageMaker HyperPod ║ ║ ─────────────────────────────────────────────────────────── ║ ║ Multi-tenant GPU Cluster com Isolation + Fair Scheduling ║ ╚══════════════════════════════════════════════════════════════╝

  1. Physical GPU Cluster (shared) ├─ 8x H100 / 4x A100 / 2x L4 (whatever you want) └─ ~R$ 200K/mês

  2. HyperPod Overlay (virtual partition) ├─ Partition A: Team Agentes (gets 40% of GPUs) │ ├─ 3.2x H100 equivalent │ ├─ Guarantees minimum (non-preemptible) │ ├─ Can burst to more if available │ └─ Workload: Agente WhatsApp inference │ ├─ Partition B: Team ML (gets 50% of GPUs) │ ├─ 4x H100 equivalent │ ├─ Guarantees minimum │ ├─ Can burst if Team A idle │ └─ Workload: LLM training │ └─ Partition C: Team Research (gets 10% of GPUs) ├─ 0.8x H100 equivalent ├─ Best-effort (preemptible) ├─ Uses "spare capacity" └─ Workload: Experiments

  3. Fair Scheduler (automatic) ├─ Monitors each partition's usage ├─ Reallocates dynamically │ ├─ If Team A idle: Team B can use more │ ├─ If Team B idle: Team A can burst │ ├─ If both idle: Team C gets access │ └─ All automatic (no ops needed) │ └─ Result: GPUs ALWAYS working (95%+ utilization)

  4. Isolation (security/compliance) ├─ Team A can't see Team B's data ├─ Team B can't steal Team A's GPU cycles ├─ Each team has "resource quota" (enforced) ├─ Multi-tenant security (like AWS VPCs) └─ Compliance: HIPAA/SOC2 ready

Key features:

✓ Fair Allocation └─ Each team gets guaranteed minimum └─ Can burst if others are idle

✓ Isolation Boundary └─ Teams can't interfere with each other └─ Even if sharing same GPU

✓ Automatic Scheduling └─ No ops team needed └─ AWS manages the complexity

✓ Cost Tracking └─ See exactly how much each team costs └─ Chargeback model (if you want)

✓ Preemption Policy └─ Define what's interruptible └─ Critical workloads (agentes) never preempted

Implementação prática (como setup)

Step 1: Create HyperPod cluster (AWS CLI)

bash

Install AWS CLI

pip install awscli aws configure # Add your credentials

Create cluster

aws sagemaker create-cluster
--cluster-name "shared-gpu-cluster"
--vpc-config SubnetIds=subnet-xxx,SecurityGroupIds=sg-xxx
--instance-groups '[ { "InstanceGroupName": "gpu-nodes", "InstanceType": "ml.p5.48xlarge", # H100 cluster "InstanceCount": 4, "VolumeSizeInGB": 1000 } ]'

Check status

aws sagemaker describe-cluster
--cluster-name "shared-gpu-cluster"

Step 2: Define team partitions (resource quotas)

yaml

hyperpod-config.yaml

partitions:

  • name: "team-agents" resource_request: gpu_percentage: 40 memory_gb: 320 # 40% of 800GB guaranteed: true # Min guarantee workload_type: "inference" users:

    • "agente-team"
    • "agente-deployment"
  • name: "team-ml" resource_request: gpu_percentage: 50 memory_gb: 400 guaranteed: true workload_type: "training" users:

    • "ml-team"
    • "training-pipeline"
  • name: "team-research" resource_request: gpu_percentage: 10 memory_gb: 80 guaranteed: false # Best-effort (preemptible) workload_type: "experiment" users:

    • "research-team"

fair_scheduling: burst_allowed: true # Teams can use idle capacity preemption_policy: "graceful" # Shutdown jobs cleanly monitoring_interval: "30s" # Check utilization every 30s

Isolation: network_isolation: true data_isolation: true resource_limits: true # Enforce quotas

Step 3: Deploy workloads to partitions

python

Python: Deploy agente to "team-agents" partition

import boto3

sagemaker = boto3.client('sagemaker')

Agente inference job

agent_job = sagemaker.create_training_job( TrainingJobName='agente-whatsapp-inference', RoleArn='arn:aws:iam::ACCOUNT:role/SageMakerRole', AlgorithmSpecification={ 'TrainingImage': 'ACCOUNT.dkr.ecr.us-east-1.amazonaws.com/agente:latest', 'TrainingInputMode': 'File' }, InputDataConfig=[ { 'ChannelName': 'training', 'DataSource': { 'S3DataSource': { 'S3Uri': 's3://bucket/agente-data/', 'S3DataType': 'S3Prefix', 'S3DataDistributionType': 'FullyReplicated' } } } ], OutputDataConfig={ 'S3OutputPath': 's3://bucket/agente-output/' }, ResourceConfig={ 'InstanceType': 'ml.p5.48xlarge', # H100 'InstanceCount': 1, 'VolumeSizeInGB': 100 }, StoppingCondition={'MaxRuntimeInSeconds': 3600}, # KEY: Specify partition ClusterConfig={ 'PartitionName': 'team-agents', 'MaxRunDurationInSeconds': 3600 } )

print(f"Agente job submitted to partition: team-agents") print(f"GPU allocation: 40% of cluster (guaranteed)")

Step 4: Monitor utilization (real-time dashboard)

python

Monitor GPU utilization by partition

import boto3 import pandas as pd

cloudwatch = boto3.client('cloudwatch')

Get metrics for each partition

partitions = ['team-agents', 'team-ml', 'team-research']

for partition in partitions: metrics = cloudwatch.get_metric_statistics( Namespace='AWS/SageMaker', MetricName='GPUUtilizationPercentage', Dimensions=[ {'Name': 'ClusterName', 'Value': 'shared-gpu-cluster'}, {'Name': 'PartitionName', 'Value': partition} ], StartTime=datetime.now() - timedelta(hours=1), EndTime=datetime.now(), Period=60 )

# Calculate average utilization
datapoints = metrics['Datapoints']
avg_utilization = sum([d['Average'] for d in datapoints]) / len(datapoints)

print(f"{partition}: {avg_utilization:.1f}% utilized")

Example output:

team-agents: 92.5% utilized (agente running)

team-ml: 87.3% utilized (training job running)

team-research: 45.2% utilized (experiment running, can burst if others idle)

Step 5: Cost tracking (chargeback)

python

Chargeback by partition (bill each team)

ce = boto3.client('ce') # Cost Explorer

costs = ce.get_cost_and_usage( TimePeriod={ 'Start': '2026-10-01', 'End': '2026-10-31' }, Granularity='DAILY', Filter={ 'Dimensions': { 'Key': 'RESOURCE_ID', 'Values': ['shared-gpu-cluster'] } }, GroupBy=[ {'Type': 'DIMENSION', 'Key': 'TAG'}, # Partition tag ], Metrics=['UnblendedCost'] )

Extract costs

for result in costs['ResultsByTime']: for group in result['Groups']: partition = group['Keys'][0] cost = group['Metrics']['UnblendedCost']['Amount'] print(f"{partition}: R$ {float(cost):,.2f}")

Example output:

team-agents: R$ 8,000 (40% of R$ 200K)

team-ml: R$ 10,000 (50% of R$ 200K)

team-research: R$ 2,000 (10% of R$ 200K)

Total: R$ 20,000 (monthly cost)


Use Cases: Onde HyperPod muda a economia

Use Case 1: Multi-team IA na startup

Startup de fintech com 3 times:

Before (3 clusters separados): ├─ Team Agentes: 2x H100 = R$ 100K/mês ├─ Team ML: 2x A100 = R$ 70K/mês ├─ Team Research: 1x RTX 4090 = R$ 10K/mês ├─ Total: R$ 180K/mês ├─ Utilização média: 35% └─ Desperdício: R$ 117K/mês

After (1 cluster com HyperPod): ├─ Cluster compartilhado: 4x H100 + 1x A100 = R$ 170K/mês ├─ Fair allocation automática ├─ Utilização média: 85% ├─ Total: R$ 170K/mês └─ Economia: R$ 10K/mês (pequena) + mais capacity

Real benefit: ├─ Antes: Team ML tava 50% idle (precisava mais GPU) ├─ Depois: Team ML consegue usar idle time de Team A ├─ Resultado: Team ML treina 2x mais modelos (mesmo cost!) └─ Value: +R$ 500K em productivity

Use Case 2: SaaS com inference + training

SaaS de code generation:

Before: ├─ Inference (API): 2x A100 24h/day = R$ 70K/mês │ ├─ Peak: 80% utilization │ ├─ Off-peak: 20% utilization │ └─ Average: 40% ├─ Training: 1x A100 2h/day = R$ 35K/mês │ ├─ Used: 2h/day (budget limited) │ └─ Idle: 22h/day (wish we could train more) ├─ Total: R$ 105K/mês └─ Problem: Can't do more fine-tuning (cost-prohibitive)

After (HyperPod with smart scheduling): ├─ Cluster: 3x A100 = R$ 105K/mês (SAME COST!) ├─ Smart scheduler: │ ├─ Peak hours: Inference gets 2.5x A100 │ ├─ Off-peak: Training gets 2x A100 │ ├─ Fair allocation: Both workloads thrive │ └─ Automatic (no ops needed) ├─ Result: │ ├─ Inference: Same performance (actually better) │ ├─ Training: Can now run 10h/day (5x more!) │ └─ Impact: Models improve 5x faster └─ Benefit: SAME COST, WAY MORE VALUE

Use Case 3: Enterprise with many small teams

Enterprise (500+ people, 10+ teams using AI):

Before (each team had own GPU): ├─ Team 1: 1x H100 = R$ 50K/mês ├─ Team 2: 1x A100 = R$ 35K/mês ├─ Team 3: 1x A100 = R$ 35K/mês ├─ Team 4: 0.5x GPU = R$ 25K/mês ├─ ... (6 more teams) ├─ Total: R$ 500K/mês ├─ Utilization: 30% (70% waste) └─ Waste: R$ 350K/mês

After (centralized with HyperPod): ├─ Central pool: 8x mixed GPUs = R$ 300K/mês ├─ HyperPod partitions: │ ├─ 10 teams share 8 GPUs │ ├─ Fair scheduler allocates dynamically │ ├─ Utilization: 85% │ └─ No team fights over GPU ├─ Total: R$ 300K/mês └─ Savings: R$ 200K/mês (40%!)

Extra benefit: ├─ Reduces "GPU hoarding" (teams allocate what they need, not max) ├─ Enables chargeback (see exactly who costs what) ├─ Improves ops (centralized monitoring) └─ Strategic: Reserve capacity for new initiatives


Conclusão: Era de GPU compartilhada começou

Before HyperPod:

  • Cada time tem seu próprio cluster (isolated)
  • Resultado: Desperdício de 50-70% (ociosa)
  • Custo explosivo (R$ 500K+/mês pra empresa grande)
  • Operacional nightmare (cada cluster precisa de ops)

After HyperPod:

  • Times compartilham GPU (mas com isolation/fairness)
  • Resultado: 85%+ utilization (quase zero desperdício)
  • Custo otimizado (40-60% economia)
  • Operacional simples (AWS gerencia tudo)

Impacto econômico (empresa média):

Investimento: ├─ HyperPod setup: R$ 20K (one-time) ├─ Migration: R$ 50K (one-time) └─ Total: R$ 70K

Benefício (anual): ├─ Reduced GPU cost: R$ 240K/ano (40% of R$ 600K baseline) ├─ Improved productivity: +30% (teams train 3x more models) └─ Reduced ops burden: R$ 100K/ano (less manual work) └─ Total: R$ 340K/ano

ROI: 485% (Year 1) Payback period: 8 weeks

Próximo passo (faça HOJE):

  1. Audit seu infra (quantas GPUs você tem? Qual utilização?)
  2. Calculate desperdício (GPUs × idle% × cost)
  3. Pilot HyperPod (setup 1 cluster, 2-3 teams)
  4. Measure gains (cost + productivity)
  5. Roll out (consolidate all GPUs to shared pool)

GPU compartilhada é o futuro. Seu dinheiro está sendo desperdiçado agora. → OpenClaw: HyperPod Setup & Multi-team Orchestration

Era de GPU isolada (e desperdício) acabou. 🎯🚀


Publicado em 9 de outubro de 2026

Leia também