Oumi Summer ’26 is here.
See what we launchedStart free. Scale when you're ready
Try it free. Ship with your team. Scale with ours.
No credit card required
Free
$0
Up to $50 in credits for 3 months
Explore and build your first custom model, no credit card required
- Full automation with Oumi agent (Pat. Pend.)
- Data synthesis with open/closed models
- Data analysis and curation
- Open/closed model evaluation with failure modes
- SFT and PEFT (LoRA, QLoRA) training
For teams and builders
Pro
$25/month
pay-as-you-go after
The power and flexibility to ship real workloads for professional model development
- Everything in Free
- Training with on-policy distillation
- Multiple concurrent jobs for higher efficiency
- Download model weights: 1 / month
- Production inference with autoscaling
- $25/month credits platform-wide
For Production at Scale
Enterprise
Custom
For enterprise scale, with enterprise-grade security & controls – full support from Oumi
- Everything in Pro
- BYOC - VPC & on-prem
- Dedicated capacity and autoscaling to 1000s of GPUs
- Guaranteed throughput and SLAs
- Advanced training methods (RL, custom pipelines)
- Dedicated experts embedded with your team
- Bespoke engagements scoped to your goals
Hosted Platform – detailed pricing
Detailed breakdown of agent, tools, storage, training, and inference pricing.
Oumi Agent
| Tokens | Price / 1M |
|---|---|
Input | $6.25 |
Output | $31.25 |
Cache write | $7.80 |
Cache read | $0.65 |
Tools & Storage
Evaluation | 1,000 judgments / $1 |
Data Synthesis | 1,000 rows / $1 |
Storage | 4 GB/month / $1 |
Supervised Fine-Tuning
Priced per 1M training tokens — calculated as the number of tokens in your training dataset multiplied by the number of epochs.
| Model Size | Price |
|---|---|
Up to 16B | $0.49 |
16.1–32B | $2.00 |
32.1–80B | $3.00 |
80.1–300B | $6.00 |
On-Policy Distillation
Priced per GPU-hour on dedicated GPUs. Training currently runs on 8 GPUs; GPU used is subject to availability.
| GPU | Price / GPU-hr |
|---|---|
A100-80GB | $2.90 |
H100-80GB | $4.00 |
Inference
| Model | Input / 1M | Output / 1M |
|---|---|---|
DeepSeek-V4-Pro | $1.91 | $3.83 |
Gemma 3 4B | $0.06 | $0.11 |
Gemma 4 31B | $0.45 | $1.10 |
GLM-5 | $1.10 | $3.50 |
GLM-5.1 | $1.55 | $4.85 |
GLM-5.2 | $1.55 | $4.85 |
gpt-oss-20b | $0.03 | $0.15 |
gpt-oss-120b | $0.15 | $0.60 |
Kimi K2.5 | $0.66 | $3.30 |
Kimi K2.6 | $1.05 | $4.40 |
Kimi K3 | $3.30 | $16.50 |
Llama 3.1 8B Instruct | $0.17 | $0.32 |
Llama 3.2 1B Instruct | $0.03 | $0.22 |
Llama 3.2 3B Instruct | $0.06 | $0.37 |
Llama 3.3 70B | $1.00 | $1.00 |
Llama 4 Scout 17B 16E Instruct | $0.11 | $0.33 |
Mistral Large 3 2512 | $0.55 | $1.65 |
Muse Glimmer 30B | $0.40 | $1.65 |
Nemotron 3.5 Lightning | $0.11 | $0.28 |
Qwen2.5 7B Instruct | $0.33 | $0.33 |
Qwen3 8B | $0.13 | $0.50 |
Qwen3 32B | $0.11 | $0.46 |
Qwen3 235B A22B Instruct 2507 | $0.25 | $0.90 |
Qwen3.5 9B | $0.11 | $0.17 |
Qwen3.5 27B | $0.21 | $1.72 |
Qwen3.5 35B A3B | $0.18 | $1.43 |
Qwen3.5 397B A17B | $0.70 | $4.00 |
Qwen3.6 27B | $0.49 | $2.97 |
Qwen3.6 35B A3B | $0.11 | $1.04 |
Inference is only charged when you utilize models hosted by Oumi to power an action on the platform — evaluation, data synthesis, or a Model Playground comparison.
OpenAI, Anthropic and Google models are not listed above. Add your own key for those providers and Oumi charges you nothing for inference — you pay your provider directly. Without a key, Oumi falls back to its own and passes through the provider's published price at cost, with no margin added.
Production Inference
Deploy fully fine-tuned or LoRA models on dedicated GPUs, priced per GPU-hour and billed for uptime. LoRA and full fine-tunes cost the same because each runs on its own dedicated GPU.
| GPU | Price / GPU-hr |
|---|---|
A100-80GB | $3.00 |
H100-80GB | $7.00 |
H200-141GB | $7.00 |
B200-180GB | $10.00 |