Which model for what you are building
Say what you are building. We show the offers that fit, ranked on a criterion we name, with every figure carrying the day we observed it.
A chatbot people talk to. Answers have to start fast, so time to first token matters more than raw throughput, and a cached system prompt is what makes it affordable. Ranked by lowest time to first token, then lowest output price, across 1 045 offers for which we hold a published output price.
| # | Offer | Output | Cached input | First token | Entity | Observed |
|---|---|---|---|---|---|---|
| 1 | Muse Glimmer 30BFireworks AI | $1.50 | $0.040 | 102Â ms | not established | 16 September 2026 |
| 2 | openai/gpt-oss-120bCerebras | $0.750 | not measured | 188 ms | 🇺🇸 | 16 September 2026 |
| 3 | Qwen/Qwen3.8-27BCerebras | $1.49 | not measured | 192 ms | 🇺🇸 | 17 September 2026 |
| 4 | NousResearch/Hermes-3-Llama-3.1-70BDeepInfra | $0.700 | not measured | 212 ms | 🇺🇸 | 15 September 2026 |
| 5 | Qwen/Qwen3-14BDeepInfra | $0.240 | not measured | 220 ms | 🇺🇸 | 15 September 2026 |
| 6 | ibm-granite/granite-4.2-3bDeepInfra | $0.120 | not measured | 220 ms | 🇺🇸 | 15 September 2026 |
| 7 | Qwen/Qwen3-30B-A3BDeepInfra | $0.500 | not measured | 228 ms | 🇺🇸 | 15 September 2026 |
| 8 | Qwen/Qwen3.6-27BDeepInfra | $3.20 | not measured | 242 ms | 🇺🇸 | 15 September 2026 |
| 9 | meta-models/Muse-Glimmer-30BDeepInfra | $1.20 | not measured | 246 ms | 🇺🇸 | 15 September 2026 |
| 10 | GPT OSS 20BGroq | $0.300 | $0.037 | 247 ms | 🇬🇧🇺🇸 | 18 September 2026 |
| 11 | InklingBaseten | $4.05 | $0.170 | 249 ms | 🇺🇸 | 18 September 2026 |
| 12 | ibm-granite/granite-4.2-8bDeepInfra | $0.250 | not measured | 257 ms | 🇺🇸 | 15 September 2026 |
What this does, and what it refuses to do
It computes no score. It applies the filters named above and sorts on the criterion printed with the result. A single score would mix dollars with milliseconds and weight them according to an opinion, and the opinion would be ours rather than yours. You can see the criterion, so you can disagree with it.
It says what it does not know. An offer with no measured latency is not slow, it is unmeasured, and it stays in the list saying so rather than being pushed down. Speed comes from the Hugging Face router and availability from OpenRouter's router, neither is our own measurement, and both are the median of our own daily readings rather than the latest figure alone.
It ignores what it cannot check. Nothing here says a model is good at your task. Capability scores are relayed from Artificial Analysis and measure the model, not the seller. Whether a model suits you is a judgement we do not make.
Every row links to the offer page, where each figure carries its date, its source URL and a fingerprint of that source. The complete set of 1 045 offers is in the catalogue and in prices.json.