Bertrande

Catalogue / Groq / GPT OSS 120B (on Groq)

GPT OSS 120B (on Groq)

Prices and limits, each value carrying the date it was observed and the source it came from.

As of 18 September 2026, GPT OSS 120B costs $0.150 per million input tokens and $0.600 per million output tokens on Groq. Observed on Groq's own published price list on 18 September 2026. This is Groq's resale price, not the model vendor's list price.

Input $0.150 per 1M
Output $0.600 per 1M
Context 131 072 tokens

What does GPT OSS 120B cost?

ParameterValueUnit ObservedSourceProof
Input price $0.150 USD per 1M tokens 18 September 2026
checked 18 September 2026
official API L3 5cae5bcec35a
Output price $0.600 USD per 1M tokens 18 September 2026
checked 18 September 2026
official API L3 5cae5bcec35a
Context window 131 072 tokens 18 September 2026
checked 18 September 2026
official API L3 5cae5bcec35a
Maximum output 65 536 tokens 18 September 2026
checked 18 September 2026
official API L3 5cae5bcec35a
Cache read price $0.075 USD per 1M tokens 18 September 2026
checked 18 September 2026
official API L3 5cae5bcec35a

Does the price change with the hour or the size of the request?

Yes. This provider publishes conditional rates, so the figure above is not the only price you can pay. Each rate below applies under a stated condition, taken verbatim from the provider's own catalogue.

ParameterRate AppliesObserved
Rate limit, requests per minute 1 000.0 developer plan 18 September 2026
Rate limit, tokens per minute 250 000.0 developer plan 18 September 2026

We record each conditional rate as its own dated observation, so a discount window that disappears shows up as a change like any other.

How fast is GPT OSS 120B here?

Speed is the thing price tables leave out, and it is the thing developers complain about. These figures are measured by the Hugging Face router, not by us, on the traffic it sends to this host. They are not what you would measure calling the host yourself: neither the network path nor the load is the same.

MeasureMedian of our readings ReadingsRange seen
Time to first token 274 ms 17 131 to 462, 2 September 2026 to 18 September 2026
Output throughput 433.4 17 402 to 444, 2 September 2026 to 18 September 2026

One reading a day, and we keep every one. A single measurement of latency covers a handful of calls and moves a great deal, so we never publish the latest figure on its own: the median above is computed from our own dated readings, and the range shows how far they spread. The computation is ours; the measurements are not.

How has GPT OSS 120B moved?

Output price, 3 observations between 2 September 2026 and 18 September 2026. Range $0.600 to $0.600. Drawn as a step: a value holds until the next observation, we do not interpolate between them.

Has GPT OSS 120B changed price?

No change has been recorded since we started observing this service on 2 September 2026. We record a change only when two successive observations differ, so an empty table here means the values have held.

What does GPT OSS 120B cost elsewhere?

We track this model at 9 sellers. The same model rarely costs the same everywhere, and the difference is not visible from any single seller's catalogue.

Compare GPT OSS 120B across 9 sellers, side by side, or see every model tracked at more than one seller.

Other services from Groq

All Groq services

Where these figures come from

Source: Groq. Every figure on this page was read from Groq's own published price list, and every row above links straight back to it. We publish factual values with attribution, we do not reproduce Groq's content, and we send readers to them. If you publish this source and want us to stop tracking it, write to contact@bertrande.com and we will act on it without argument.

How is this page verified?

Every figure above comes from Groq's own published price list, which we rank as a level 3 source. We re-check it every day. Each observation stores the exact source URL and a SHA-256 fingerprint of the source content at that moment, which is what lets us prove what the source said on a given day.

The Proof column shows the first twelve characters of the SHA-256 fingerprint of the source content at the moment we read it. Hover it for the full value. It is what lets us demonstrate later what the source said on a given day, after that page has changed or vanished. The full fingerprints are in data.json.

No value on this page was produced, estimated or completed by a language model. Where we do not have a verified figure, the row is absent rather than filled in. Read the full methodology, or report an error.

Take the data, or watch this page

This page is available as data.json (including the full observation history of every parameter) and data.csv, under CC-BY-4.0, free, with attribution. See all downloads.

Want to know when this changes? Point a feed reader or a script at the Atom feed, which carries every confirmed change and is rebuilt hourly, or poll this page's data.json, where every figure carries its observation date and the fingerprint of its source. Both work today and need no address.

Email alerts are open, and nothing is sent until you confirm from a link we email you. Subscribe, or see what an alert looks like.