Not what a model costs — whether it's actually fast and reliable enough to sit inside an agent loop. Time to first token, output tokens/sec, and whether structured JSON output comes back valid, every week, measured against real running calls rather than claimed benchmarks.
Methodology: We don't hold direct API keys for OpenAI, Anthropic, or xAI, so every provider here — including Google and DeepSeek, for a consistent comparison — is measured through OpenRouter's API, which proxies to the real underlying provider. That almost certainly adds some fixed proxy overhead on top of each provider's own number, so read these as relative comparisons between providers, not what you'd measure calling a provider directly. Each row uses that provider's fast/cheap-tier model on OpenRouter — the one you'd actually reach for in an agent loop, picked fresh each run from OpenRouter's live model list, never hardcoded. Structured-output success is checked in code against a fixed JSON schema, not self-reported by an LLM.
Speed data couldn't be loaded right now. Try again shortly.