API100 LLM Benchmarks
Direct telemetry measuring Time to First Token (TTFT), streaming token throughput, and pricing efficiency across Claude 3.5 Sonnet, GPT-4o, Gemini 1.5 Pro, DeepSeek R1, and Llama 3.3.
Empirical Model Telemetry Matrix
Measured across Mumbai, US-East, Frankfurt, and Singapore ingress edge nodes.
| Model & Provider | Median TTFT (P50) | Tail TTFT (P95) | Velocity (t/s) | Input / 1M | Output / 1M | Availability |
|---|---|---|---|---|---|---|
| Claude 3.5 SonnetAnthropic • Frontier Reasoning & Code | 290 ms | 480 ms | 72.4 t/s | $3.00 | $15.00 | 99.98% |
| GPT-4oOpenAI • Multimodal & Vision | 275 ms | 440 ms | 84.1 t/s | $2.50 | $10.00 | 99.96% |
| DeepSeek R1DeepSeek • Chain-of-Thought Math | 340 ms | 590 ms | 54.8 t/s | $0.55 | $2.19 | 99.94% |
| Gemini 1.5 ProGoogle • 2M Token Long Context | 310 ms | 520 ms | 64.2 t/s | $1.25 | $5.00 | 99.97% |
| Gemini 1.5 FlashGoogle • High-Velocity Ingestion | 195 ms | 310 ms | 142 t/s | $0.07 | $0.30 | 99.99% |
| Llama 3.3 70BMeta • Open Weights Frontier | 230 ms | 390 ms | 96.5 t/s | $0.40 | $0.40 | 99.97% |
Dedicated Research Reports
Frontier LLM Latency Breakdown (TTFT & P95)
An exhaustive evaluation of Time to First Token across 4 geographic regions. Learn how network serialization and model weight sharding impact initial latency.
Token Throughput & Generation Velocity (t/s)
Comparing decoding throughput from 140+ tokens/second on Flash models down to 50+ tokens/second on deep reasoning models.
Token Economics & Cost-Per-Intelligence Frontier
Benchmarking real cost per 1M input and output tokens across providers. Discover how DeepSeek and Gemini Flash reduce inference spend by up to 90%.
GPT-4o vs Claude 3.5 Sonnet vs Gemini 1.5 Pro
Head-to-head empirical comparison across reasoning accuracy, tool use, streaming performance, and production gateway availability.

