Back to Benchmarks Hub
Research Report 02 • September 2026
Token Generation Velocity (t/s)
Measuring continuous stream decoding throughput across 50,000 live completion events. Learn how model quantization, architecture parameters, and serving clusters impact token velocity.
Peak Decoding Rate
142.0 t/s
Gemini 1.5 Flash
Fastest Frontier Reasoning
84.1 t/s
GPT-4o (OpenAI)
Fastest Open Model
96.5 t/s
Llama 3.3 70B
Generation Velocity Telemetry
| Model | Mean Throughput (t/s) | Observed Range | Tier |
|---|---|---|---|
| Gemini 1.5 Flash | 142 t/s | 110.5 - 175.2 t/s | High Velocity |
| Llama 3.3 70B | 96.5 t/s | 78 - 115.4 t/s | Open Weights |
| GPT-4o | 84.1 t/s | 65.2 - 104 t/s | Frontier Flagship |
| Claude 3.5 Sonnet | 72.4 t/s | 55 - 88.6 t/s | Frontier Code |
| Gemini 1.5 Pro | 64.2 t/s | 48.5 - 80.1 t/s | Deep Context |
| DeepSeek R1 | 54.8 t/s | 41 - 68.9 t/s | Reasoning CoT |

