Each model is probed with a synthetic request every 20 minutes (Cloudflare cron) · monitoring staging.freeinference.org . Click a model to zoom in on its latency and throughput history.
bge-m3
Status DOWN
Latency 30159 ms
TTFT —
Throughput —
Uptime 96%
Checked 2026-08-04T06:21:00
Latency trend (96) 247 ms · 118–480 ms
click to zoom ↗
HTTP 500
deepseek-v4-flash
Status UP
Latency 3522 ms
TTFT 338 ms
Throughput 227.7 tok/s
Uptime 98%
Checked 2026-08-04T06:20:58
TTFT trend (98) 338 ms · 179–34805 ms
click to zoom ↗
diffusiongemma
Status UP
Latency 3829 ms
TTFT 3829 ms
Throughput 229.83 tok/s
Uptime 42%
Checked 2026-08-04T06:20:56
TTFT trend (42) 3829 ms · 2938–8613 ms
click to zoom ↗
glm-5.1
Status UP
Latency 11098 ms
TTFT 4857 ms
Throughput 130.91 tok/s
Uptime 100%
Checked 2026-08-04T06:20:44
TTFT trend (100) 4857 ms · 2244–44141 ms
click to zoom ↗
glm-5.2
Status UP
Latency 12239 ms
TTFT 5362 ms
Throughput 111.68 tok/s
Uptime 100%
Checked 2026-08-04T06:20:44
TTFT trend (100) 5362 ms · 1977–43953 ms
click to zoom ↗
kimi-k2.7-code
Status UP
Latency 31534 ms
TTFT 8165 ms
Throughput 41.81 tok/s
Uptime 99%
Checked 2026-08-04T06:20:55
TTFT trend (99) 8165 ms · 1282–21260 ms
click to zoom ↗
minimax-m3
Status UP
Latency 11947 ms
TTFT 1458 ms
Throughput 74.08 tok/s
Uptime 95%
Checked 2026-08-04T06:20:44
TTFT trend (95) 1458 ms · 759–9714 ms
click to zoom ↗
qwen3.6-35b
Status UP
Latency 1677 ms
TTFT 233 ms
Throughput 256.23 tok/s
Uptime 99%
Checked 2026-08-04T06:20:56
TTFT trend (99) 233 ms · 113–2042 ms
click to zoom ↗