Get ready to compound your AI. August 11.
Save your spotStart free. Scale when you're ready
Try it free. Ship with your team. Scale with ours.
No credit card required
Free
$0
Up to $50 in credits for 3 months
Explore and build your first custom model, no credit card required
- Full automation with Oumi agent (Pat. Pend.)
- Data synthesis with open/closed models
- Data analysis and curation
- Open/closed model evaluation with failure modes
- SFT and PEFT (LoRA, QLoRA) training
For teams and builders
Pro
$25/month
pay-as-you-go after
The power and flexibility to ship real workloads for professional model development
- Everything in Free
- Training with on-policy distillation
- Multiple concurrent jobs for higher efficiency
- Download model weights: 1 / month
- Production inference with autoscaling
- $25/month credits platform-wide
For Production at Scale
Enterprise
Custom
For enterprise scale, with enterprise-grade security & controls – full support from Oumi
- Everything in Pro
- BYOC - VPC & on-prem
- Dedicated capacity and autoscaling to 1000s of GPUs
- Guaranteed throughput and SLAs
- Advanced training methods (RL, custom pipelines)
- Dedicated experts embedded with your team
- Bespoke engagements scoped to your goals
Hosted Platform – detailed pricing
Detailed breakdown of agent, tools, storage, training, and inference pricing.
Oumi Agent
| Tokens | Price / 1M |
|---|---|
Input | $6.25 |
Output | $31.25 |
Cache write | $7.80 |
Cache read | $0.65 |
Tools & Storage
Evaluation | 1,000 judgments / $1 |
Data Synthesis | 1,000 rows / $1 |
Storage | 4 GB/month / $1 |
Supervised Fine-Tuning
Priced per 1M training tokens — calculated as the number of tokens in your training dataset multiplied by the number of epochs.
| Model Size | Price |
|---|---|
Up to 16B | $0.49 |
16.1–32B | $2.00 |
32.1–80B | $3.00 |
80.1–300B | $6.00 |
On-Policy Distillation
Priced per GPU-hour on dedicated GPUs. Training currently runs on 8 GPUs; GPU used is subject to availability.
| GPU | Price / GPU-hr |
|---|---|
A100-80GB | $2.90 |
H100-80GB | $4.00 |
Inference
| Model | Input / 1M | Output / 1M |
|---|---|---|
DeepSeek-V4-Pro | $1.91 | $3.83 |
Gemma 3 4B | $0.06 | $0.11 |
Gemma 4 31B | $0.45 | $1.10 |
GLM-5 | $1.10 | $3.50 |
GLM-5.1 | $1.55 | $4.85 |
GLM-5.2 | $1.55 | $4.85 |
gpt-oss-20b | $0.03 | $0.15 |
gpt-oss-120b | $0.15 | $0.60 |
Kimi K2.5 | $0.66 | $3.30 |
Kimi K2.6 | $1.05 | $4.40 |
Kimi K3 | $3.30 | $16.50 |
Llama 3.1 8B Instruct | $0.17 | $0.32 |
Llama 3.2 1B Instruct | $0.03 | $0.22 |
Llama 3.2 3B Instruct | $0.06 | $0.37 |
Llama 3.3 70B | $1.00 | $1.00 |
Llama 4 Scout 17B 16E Instruct | $0.11 | $0.33 |
Mistral Large 3 2512 | $0.55 | $1.65 |
Muse Glimmer 30B | $0.40 | $1.65 |
Nemotron 3.5 Lightning | $0.11 | $0.28 |
Qwen2.5 7B Instruct | $0.33 | $0.33 |
Qwen3 8B | $0.13 | $0.50 |
Qwen3 32B | $0.11 | $0.46 |
Qwen3 235B A22B Instruct 2507 | $0.25 | $0.90 |
Qwen3.5 9B | $0.11 | $0.17 |
Qwen3.5 27B | $0.21 | $1.72 |
Qwen3.5 35B A3B | $0.18 | $1.43 |
Qwen3.5 397B A17B | $0.70 | $4.00 |
Qwen3.6 27B | $0.49 | $2.97 |
Qwen3.6 35B A3B | $0.11 | $1.04 |
Inference is only charged when you utilize models hosted by Oumi to power an action on the platform — evaluation, data synthesis, or a Model Playground comparison.
OpenAI, Anthropic and Google models are not listed above. Add your own key for those providers and Oumi charges you nothing for inference — you pay your provider directly. Without a key, Oumi falls back to its own and passes through the provider's published price at cost, with no margin added.
Production Inference
Deploy fully fine-tuned or LoRA models on dedicated GPUs, priced per GPU-hour and billed for uptime. LoRA and full fine-tunes cost the same because each runs on its own dedicated GPU.
| GPU | Price / GPU-hr |
|---|---|
A100-80GB | $3.00 |
H100-80GB | $7.00 |
H200-141GB | $7.00 |
B200-180GB | $10.00 |