Catalogue / Novita AI / meta-llama/llama-4-maverick-17b-128e-instruct-fp8 (on Novita AI)
meta-llama/llama-4-maverick-17b-128e-instruct-fp8 (on Novita AI)
Prices and limits, each value carrying the date it was observed and the source it came from.
As of 17 September 2026, meta-llama/llama-4-maverick-17b-128e-instruct-fp8 costs $0.270 per million input tokens and $0.850 per million output tokens on Novita AI. Observed on Novita AI's own published price list on 17 September 2026. This is Novita AI's resale price, not the model vendor's list price.
What does meta-llama/llama-4-maverick-17b-128e-instruct-fp8 cost?
| Parameter | Value | Unit | Observed | Source | Proof |
|---|---|---|---|---|---|
| Input price | $0.270 | USD per 1M tokens | 17 September 2026 checked 18 September 2026 |
official API L1 | c50e1f80682b |
| Output price | $0.850 | USD per 1M tokens | 17 September 2026 checked 18 September 2026 |
official API L1 | c50e1f80682b |
| Context length | 1 048 576 | tokens | 17 September 2026 checked 18 September 2026 |
official API L1 | c50e1f80682b |
| Maximum output | 8 192 | tokens | 17 September 2026 checked 18 September 2026 |
official API L1 | c50e1f80682b |
| Input modalities |
| list | 17 September 2026 checked 18 September 2026 |
official API L1 | c50e1f80682b |
| Output modalities |
| list | 17 September 2026 checked 18 September 2026 |
official API L1 | c50e1f80682b |
How fast is meta-llama/llama-4-maverick-17b-128e-instruct-fp8 here?
Speed is the thing price tables leave out, and it is the thing developers complain about. These figures are measured by the Hugging Face router, not by us, on the traffic it sends to this host. They are not what you would measure calling the host yourself: neither the network path nor the load is the same.
| Measure | Median of our readings | Readings | Range seen |
|---|---|---|---|
| Time to first token | 400 ms | 17 | 343 to 524, 2 September 2026 to 18 September 2026 |
| Output throughput | 80.4 | 17 | 53 to 103, 2 September 2026 to 18 September 2026 |
One reading a day, and we keep every one. A single measurement of latency covers a handful of calls and moves a great deal, so we never publish the latest figure on its own: the median above is computed from our own dated readings, and the range shows how far they spread. The computation is ours; the measurements are not.
How has meta-llama/llama-4-maverick-17b-128e-instruct-fp8 moved?
Has meta-llama/llama-4-maverick-17b-128e-instruct-fp8 changed price?
No change has been recorded since we started observing this service on 2 September 2026. We record a change only when two successive observations differ, so an empty table here means the values have held.
What does meta-llama/llama-4-maverick-17b-128e-instruct-fp8 cost elsewhere?
We track this model at 2 sellers. The same model rarely costs the same everywhere, and the difference is not visible from any single seller's catalogue.
Compare meta-llama/llama-4-maverick-17b-128e-instruct-fp8 across 2 sellers, side by side, or see every model tracked at more than one seller.
Other services from Novita AI
- meta-llama/llama-3.2-1b-instruct
- meta-llama/llama-3.2-3b-instruct
- meta-llama/llama-3.3-70b-instruct
- meta-llama/llama-4-scout-17b-16e-instruct
- microsoft/wizardlm-2-8x22b
- mindai/macaron-v1-tall
Where these figures come from
Source: Novita AI. Every figure on this page was read from
Novita AI's own published price list, and every row above links straight back to it.
We publish factual values with attribution, we do not reproduce Novita AI's content, and
we send readers to them. If you publish this source and want us to stop tracking it, write to
contact@bertrande.com and we will act on it without argument.
How is this page verified?
Every figure above comes from Novita AI's own published price list, which we rank as a level 1 source. We re-check it every 12 hours. Each observation stores the exact source URL and a SHA-256 fingerprint of the source content at that moment, which is what lets us prove what the source said on a given day.
The Proof column shows the first twelve characters of the SHA-256 fingerprint of the source content at the moment we read it. Hover it for the full value. It is what lets us demonstrate later what the source said on a given day, after that page has changed or vanished. The full fingerprints are in data.json.
No value on this page was produced, estimated or completed by a language model. Where we do not have a verified figure, the row is absent rather than filled in. Read the full methodology, or report an error.
Take the data, or watch this page
This page is available as data.json (including the full observation history of every parameter) and data.csv, under CC-BY-4.0, free, with attribution. See all downloads.
Want to know when this changes? Point a feed reader or a script at the Atom feed, which carries every confirmed change and is rebuilt hourly, or poll this page's data.json, where every figure carries its observation date and the fingerprint of its source. Both work today and need no address.
Email alerts are open, and nothing is sent until you confirm from a link we email you. Subscribe, or see what an alert looks like.