Real-world throughput, latency and reliability numbers measured against the BLACKBOX AI inference router under load.
gpt-5.6-sol
Capacity-discovery stress test of openai/gpt-5.6-sol served through the BLACKBOX AI router. Load was ramped from 100 → 10,000 RPM over 10 minutes and held at 10,000 RPM, with ~50k input tokens per request over a streaming connection. The run drove the router to its saturation point while measuring time-to-first-token (TTFT), end-to-end latency and error rate the whole way up.
11.9K RPM
Peak completed throughput
592M TPM
Peak token throughput
100.5K
Total requests
99.36%
Success rate
2.45s
TTFT p50 @ load
Throughput under load
The router tracked the scheduled ramp almost perfectly up to the 10k RPM target and sustained it through the hold, peaking at 11,856 completed RPM in the best 10-second bucket before reaching saturation at the 15-minute mark.
Achieved throughput
Completed requests/min · ramp 100 → 10,000 RPM, then hold
Higher is better
Achieved (completed RPM)Scheduled RPM
Peak completed 11,856 RPM · peak dispatched 10,908 RPM · both exceed the 10k target because the 10s peak-bucket captures bursty completions. Chart ends at the 15m saturation point.
Time-to-first-token distribution
Even at full load, half of all requests began streaming in under 2.45s, and 95% within 6.51s — with 50k tokens of input per request.
TTFT percentiles
Seconds to first streamed token · 99,858 successful requests
Lower is better
min
min TTFT · 1.14s1.14s
p50
p50 TTFT · 2.45s2.45s
mean
mean TTFT · 3.32s3.32s
p95
p95 TTFT · 6.51s6.51s
p99
p99 TTFT · 15.82s15.82s
min 1.14s · p50 2.45s · mean 3.32s · p95 6.51s · p99 15.82s. End-to-end latency tracked TTFT closely (p50 2.52s · p95 6.60s) since responses were short.
Reliability
Across all 100,500 requests driven at up to 10k RPM, only 613 returned a real error — a 0.61% error rate, dominated by HTTP 400s (601) with a handful of timeouts.
Ramp 100 → 10,000 RPM over 10 min, then hold 10,000 RPM
Input size
~50,000 tokens/request
Transport
Streaming, no output token cap
Rolling window
30s
Total requests
100,500
Run date
2026-07-24
Peak throughput of 11,856 completed RPM was measured in the best 10-second bucket while holding at the 10,000 RPM target. Numbers reflect a single capacity-discovery run and represent the router’s saturation ceiling for this model and payload size — real-world sustained throughput depends on payload, region and account limits.