models & pricing

pricing is usage-based, in USD, billed per 1M tokens. we offer the following models for training and serverless inference.

  • training is metered on the tokens your training run processes: rollout generation plus weight updates.
  • inference pricing is the same for base model inference and serving trained checkpoints. other models are included for use within our synthetic generation pipeline and judge rewards.

training

per 1M training tokens
model size context training cost
Qwen3-VL 4B Instruct Qwen/Qwen3-VL-4B-Instruct
small 4.4b
256k
$0.10 launch price
Qwen3.5 4B Qwen/Qwen3.5-4B
small 4.7b
256k
$0.10 launch price
Qwen3.5 35B-A3B Qwen/Qwen3.5-35B-A3B
medium 36b
256k
$0.20 launch price
Gemma 4 26B-A4B google/gemma-4-26B-A4B-it
medium 26b
256k
$0.20 launch price
Step 3.7 Flash stepfun-ai/Step-3.7-Flash
large 198b
256k contact us
Hy3 tencent/Hy3
large 295b
256k contact us

need a different model? get in touch with us.

inference

per 1M input / output tokens
model input output
GPT-5.6 Sol gpt-5.6-sol
$5.00 $30.00
GPT-5.6 Luna gpt-5.6-luna
$1.00 $6.00
GPT-5.6 Terra gpt-5.6-terra
$2.50 $15.00
GPT-5.4 gpt-5.4
$2.50 $15.00
GPT-5.4 Mini gpt-5.4-mini
$0.75 $4.50
GPT-5.4 Nano gpt-5.4-nano
$0.20 $1.25
Grok 4.1 Fast Non-Reasoning grok-4-1-fast-non-reasoning
$0.20 $0.50
Grok 4.3 grok-4.3
$1.25 $2.50
Qwen3.5 4B qwen3.5-4b
$0.03 $0.15
Qwen3.5 35B-A3B qwen3.5-35b-a3b
$0.25 $2.00
Gemma 4 26B-A4B gemma-4-26b-a4b-it
$0.06 $0.33

embeddings

model input
Text Embedding 3 Large text-embedding-3-large
$0.143

get in touch with our team

join our slack ask questions directly to our team and see what others are building schedule a call get 30 minutes with us for a live demo and help with scoping your project