gpt-oss-120B API pricing
Together AI lists gpt-oss-120B at $0.15 per million input tokens and $0.6 per million output tokens. The cheapest input rate need not give the cheapest total bill.
| Provider / tier / region | Input / 1M | Output / 1M | Cached input / 1M | Updated |
|---|---|---|---|---|
| Together AI | $0.15 | $0.6 | Not published | 2026-09-27 |
| Fireworks AI · Standard | $0.15 | $0.6 | $0.015 | 2026-09-27 |
| Nebius Token Factory · FP4 · eu-north1 | $0.15 | $0.6 | Not published | 2026-09-27 |
| Fireworks AI · Priority | $0.18 | $0.72 | $0.018 | 2026-09-27 |
gpt-oss-120B usage costs
These are arithmetic examples, not measured workloads. Change the volumes in the calculator to match your use.
| Example | Provider / tier | Input tokens | Output tokens | Total |
|---|---|---|---|---|
| Short replies | Together AI | 1,000,000 | 100,000 | $0.21 |
| Short replies | Fireworks AI · Standard | 1,000,000 | 100,000 | $0.21 |
| Short replies | Nebius Token Factory · FP4 · eu-north1 | 1,000,000 | 100,000 | $0.21 |
| Short replies | Fireworks AI · Priority | 1,000,000 | 100,000 | $0.25 |
| Equal input and output | Together AI | 1,000,000 | 1,000,000 | $0.75 |
| Equal input and output | Fireworks AI · Standard | 1,000,000 | 1,000,000 | $0.75 |
| Equal input and output | Nebius Token Factory · FP4 · eu-north1 | 1,000,000 | 1,000,000 | $0.75 |
| Equal input and output | Fireworks AI · Priority | 1,000,000 | 1,000,000 | $0.90 |
| Document-heavy workload | Together AI | 10,000,000 | 1,000,000 | $2.10 |
| Document-heavy workload | Fireworks AI · Standard | 10,000,000 | 1,000,000 | $2.10 |
| Document-heavy workload | Nebius Token Factory · FP4 · eu-north1 | 10,000,000 | 1,000,000 | $2.10 |
| Document-heavy workload | Fireworks AI · Priority | 10,000,000 | 1,000,000 | $2.52 |
Cost = (input tokens × input rate + output tokens × output rate) ÷ 1,000,000. Each total keeps one provider and tier's rates together. These examples exclude cached input, batch pricing, long-context changes and separately billed reasoning tokens.
Calculate API versus GPU rental costCompare other API rates
These have the nearest listed input prices in this table. Price proximity does not mean equal quality, speed or context support.
Check memory estimates for self-hosted models before assuming an API model can be replaced by a rented GPU.