Get ready to compound your AI. August 11.

Save your spot
Oumi AI

Start free. Scale when you're ready

Try it free. Ship with your team. Scale with ours.

No credit card required

Free

$0

Up to $50 in credits for 3 months

Explore and build your first custom model, no credit card required

  • Full automation with Oumi agent (Pat. Pend.)
  • Data synthesis with open/closed models
  • Data analysis and curation
  • Open/closed model evaluation with failure modes
  • SFT and PEFT (LoRA, QLoRA) training

For teams and builders

Pro

$25/month

pay-as-you-go after

The power and flexibility to ship real workloads for professional model development

  • Everything in Free
  • Training with on-policy distillation
  • Multiple concurrent jobs for higher efficiency
  • Download model weights: 1 / month
  • Production inference with autoscaling
  • $25/month credits platform-wide

For Production at Scale

Enterprise

Custom

For enterprise scale, with enterprise-grade security & controls – full support from Oumi

  • Everything in Pro
  • BYOC - VPC & on-prem
  • Dedicated capacity and autoscaling to 1000s of GPUs
  • Guaranteed throughput and SLAs
  • Advanced training methods (RL, custom pipelines)
  • Dedicated experts embedded with your team
  • Bespoke engagements scoped to your goals

Hosted Platform – detailed pricing

Detailed breakdown of agent, tools, storage, training, and inference pricing.

Oumi Agent

TokensPrice / 1M
Input
$6.25
Output
$31.25
Cache write
$7.80
Cache read
$0.65

Tools & Storage

Evaluation
1,000 judgments / $1
Data Synthesis
1,000 rows / $1
Storage
4 GB/month / $1

Supervised Fine-Tuning

Priced per 1M training tokens — calculated as the number of tokens in your training dataset multiplied by the number of epochs.

Model SizePrice
Up to 16B
$0.49
16.1–32B
$2.00
32.1–80B
$3.00
80.1–300B
$6.00

On-Policy Distillation

Priced per GPU-hour on dedicated GPUs. Training currently runs on 8 GPUs; GPU used is subject to availability.

GPUPrice / GPU-hr
A100-80GB
$2.90
H100-80GB
$4.00

Inference

ModelInput / 1MOutput / 1M
DeepSeek-V4-Pro
$1.91
$3.83
Gemma 3 4B
$0.06
$0.11
Gemma 4 31B
$0.45
$1.10
GLM-5
$1.10
$3.50
GLM-5.1
$1.55
$4.85
GLM-5.2
$1.55
$4.85
gpt-oss-20b
$0.03
$0.15
gpt-oss-120b
$0.15
$0.60
Kimi K2.5
$0.66
$3.30
Kimi K2.6
$1.05
$4.40
Kimi K3
$3.30
$16.50
Llama 3.1 8B Instruct
$0.17
$0.32
Llama 3.2 1B Instruct
$0.03
$0.22
Llama 3.2 3B Instruct
$0.06
$0.37
Llama 3.3 70B
$1.00
$1.00
Llama 4 Scout 17B 16E Instruct
$0.11
$0.33
Mistral Large 3 2512
$0.55
$1.65
Muse Glimmer 30B
$0.40
$1.65
Nemotron 3.5 Lightning
$0.11
$0.28
Qwen2.5 7B Instruct
$0.33
$0.33
Qwen3 8B
$0.13
$0.50
Qwen3 32B
$0.11
$0.46
Qwen3 235B A22B Instruct 2507
$0.25
$0.90
Qwen3.5 9B
$0.11
$0.17
Qwen3.5 27B
$0.21
$1.72
Qwen3.5 35B A3B
$0.18
$1.43
Qwen3.5 397B A17B
$0.70
$4.00
Qwen3.6 27B
$0.49
$2.97
Qwen3.6 35B A3B
$0.11
$1.04

Inference is only charged when you utilize models hosted by Oumi to power an action on the platform — evaluation, data synthesis, or a Model Playground comparison.

OpenAI, Anthropic and Google models are not listed above. Add your own key for those providers and Oumi charges you nothing for inference — you pay your provider directly. Without a key, Oumi falls back to its own and passes through the provider's published price at cost, with no margin added.

Production Inference

Deploy fully fine-tuned or LoRA models on dedicated GPUs, priced per GPU-hour and billed for uptime. LoRA and full fine-tunes cost the same because each runs on its own dedicated GPU.

GPUPrice / GPU-hr
A100-80GB
$3.00
H100-80GB
$7.00
H200-141GB
$7.00
B200-180GB
$10.00

Frequently asked questions

Sign up today with a corporate email for $50 in credits, or a personal email for $25.