$GPU Rental Prices.com

gpt-oss-120B API pricing

Together AI lists gpt-oss-120B at $0.15 per million input tokens and $0.6 per million output tokens. The cheapest input rate need not give the cheapest total bill.

Provider / tier / regionInput / 1MOutput / 1MCached input / 1MUpdated
Together AI$0.15$0.6Not published2026-09-27
Fireworks AI · Standard$0.15$0.6$0.0152026-09-27
Nebius Token Factory · FP4 · eu-north1$0.15$0.6Not published2026-09-27
Fireworks AI · Priority$0.18$0.72$0.0182026-09-27

gpt-oss-120B usage costs

These are arithmetic examples, not measured workloads. Change the volumes in the calculator to match your use.

ExampleProvider / tierInput tokensOutput tokensTotal
Short repliesTogether AI1,000,000100,000$0.21
Short repliesFireworks AI · Standard1,000,000100,000$0.21
Short repliesNebius Token Factory · FP4 · eu-north11,000,000100,000$0.21
Short repliesFireworks AI · Priority1,000,000100,000$0.25
Equal input and outputTogether AI1,000,0001,000,000$0.75
Equal input and outputFireworks AI · Standard1,000,0001,000,000$0.75
Equal input and outputNebius Token Factory · FP4 · eu-north11,000,0001,000,000$0.75
Equal input and outputFireworks AI · Priority1,000,0001,000,000$0.90
Document-heavy workloadTogether AI10,000,0001,000,000$2.10
Document-heavy workloadFireworks AI · Standard10,000,0001,000,000$2.10
Document-heavy workloadNebius Token Factory · FP4 · eu-north110,000,0001,000,000$2.10
Document-heavy workloadFireworks AI · Priority10,000,0001,000,000$2.52

Cost = (input tokens × input rate + output tokens × output rate) ÷ 1,000,000. Each total keeps one provider and tier's rates together. These examples exclude cached input, batch pricing, long-context changes and separately billed reasoning tokens.

Calculate API versus GPU rental cost

Compare other API rates

These have the nearest listed input prices in this table. Price proximity does not mean equal quality, speed or context support.

GLM-5.3-FlashRnj-1 InstructDeepSeek V4 Flash 0731

Check memory estimates for self-hosted models before assuming an API model can be replaced by a rented GPU.