Each model is probed with a synthetic request every 20 minutes (Cloudflare cron) · monitoring staging.freeinference.org . Click a model to zoom in on its latency and throughput history.
bge-m3
Status UP
Latency 394 ms
TTFT —
Throughput —
Uptime 92%
Checked 2026-09-12T14:00:58
Latency trend (92) 394 ms · 255–893 ms
click to zoom ↗
deepseek-v4-flash
Status UP
Latency 2104 ms
TTFT 566 ms
Throughput 434.98 tok/s
Uptime 90%
Checked 2026-09-12T14:00:56
TTFT trend (71) 566 ms · 422–50654 ms
click to zoom ↗
diffusiongemma
Status UP
Latency 4640 ms
TTFT 3652 ms
Throughput 205.39 tok/s
Uptime 92%
Checked 2026-09-12T14:00:56
TTFT trend (73) 3652 ms · 3111–20485 ms
click to zoom ↗
glm-5.2
Status UP
Latency 16352 ms
TTFT 3251 ms
Throughput 78.09 tok/s
Uptime 84.9%
Checked 2026-09-12T14:00:33
TTFT trend (62) 3251 ms · 2153–22133 ms
click to zoom ↗
glm-5.3
Status UP
Latency 16210 ms
TTFT 3131 ms
Throughput 78.22 tok/s
Uptime 84.9%
Checked 2026-09-12T14:00:33
TTFT trend (62) 3131 ms · 2165–21974 ms
click to zoom ↗
glm-5.3-flash
Status UP
Latency 21033 ms
TTFT 4521 ms
Throughput 61.95 tok/s
Uptime 81%
Checked 2026-09-12T14:00:33
TTFT trend (62) 4521 ms · 2263–18250 ms
click to zoom ↗
kimi-k2.7-code
Status UP
Latency 9898 ms
TTFT 1535 ms
Throughput 44.96 tok/s
Uptime 32.9%
Checked 2026-09-12T14:00:50
TTFT trend (24) 1535 ms · 1515–8042 ms
click to zoom ↗
kimi-k3
Status DOWN
Latency 1105 ms
TTFT —
Throughput —
Uptime 6.9%
Checked 2026-09-12T14:00:54
TTFT trend (5) 11025 ms · 1603–11025 ms
click to zoom ↗
stream error: You've reached your concurrent request limit. Please wait for your ongoing requests to finish and try again. (request_id: req_d198822e21db4ccb81c5eb2e6debaa74)
minimax-m3
Status UP
Latency 4936 ms
TTFT 1235 ms
Throughput 276.41 tok/s
Uptime 100%
Checked 2026-09-12T14:00:50
TTFT trend (73) 1235 ms · 589–4325 ms
click to zoom ↗
qwen3.6-35b
Status UP
Latency 1213 ms
TTFT 350 ms
Throughput 281.58 tok/s
Uptime 92%
Checked 2026-09-12T14:00:55
TTFT trend (73) 350 ms · 330–3115 ms
click to zoom ↗