OpenRouter Rankings July 2026:
Who's Actually Winning the AI Model Race
Who has the problem? Cursor and OpenClaw developers still defaulting to last quarter's benchmark winner, plus engineering leads who need to explain why a $0.05/M model tops the chart while Claude barely cracks the top 12. What you get: OpenRouter production data through July 25 — Mimo V2.5 at 1.4T tokens/day, Chinese vendors at 46% share, and the usage-vs-quality split that most leaderboard roundups miss. What's inside: Top 12 model table, vendor share breakdown, apps-layer top 10, pricing matrix, August outlook, a five-step tiered routing playbook, and FAQ×5.
Table of Contents
Related reading: OpenRouter API setup guide, June 2026 rankings and Chinese model takeover, and weekly token billing truth.
01 · Key Takeaways (Data Through 2026-07-25)
- Daily volume crown shifted again. Xiaomi Mimo V2.5 hit roughly 1.4T tokens/day to claim #1. DeepSeek V4 Flash (943.9B) and Tencent Hy3 (590B) follow. Seven of the top twelve models are Chinese-origin.
- Vendor share crossed a structural threshold. Chinese labs combined hold about 46% of OpenRouter token volume on a seven-day basis — up from under 2% a year ago. US labs (OpenAI + Anthropic + Google) now sit around 30%–36%, down from roughly 70%.
- DeepSeek remains the steadiest #1 vendor. Single-vendor share holds at 16%–18%, but the monthly model champion keeps rotating: MiniMax M2.5 in February, MiMo-V2-Pro in April, Mimo V2.5 in July.
- Hard data #1: DeepSeek V4 Flash inputs run $0.05–0.14/M versus GPT-5.5 at roughly $5/M — a 35× price gap that explains the traffic migration better than any narrative about model intelligence.
OpenRouter rankings measure where developers actually send paid requests — not who scores highest on a static benchmark. That distinction is the entire point of this analysis.
02 · Three Selection Pain Points: Why Rankings Still Mislead
The leaderboard is an input to your decision, not the decision itself. Mac developers and AI platform teams hit three recurring traps in July:
- Equating token volume with capability. Inexpensive models wired behind Hermes Agent, roleplay apps, and high-volume CLI wrappers can crush frontier models on daily tokens while barely registering in classification or hard-reasoning spend share.
- Chasing the daily #1 without tracking consistency. Rankings swing daily — Claude Opus 4.8 was still top-12 on 7/24 but got bumped by Ling 3.0 Flash on 7/25. A monthly champion is not a durable moat.
- Ignoring the apps layer. Coding agents (Kilo Code ~13%), general agents (OpenClaw ~9%), and roleplay platforms (Janitor AI, ISEKAI ZERO) burn enormous open-model traffic that enterprise AI coverage rarely mentions. If your workload is agentic coding, check openrouter.ai/apps alongside the model chart.
03 · July Model Token Volume Top 12 (Daily, Through 7/25)
| Rank | Model | Vendor | Daily Tokens | 30-Day Total |
|---|---|---|---|---|
| 1 | Mimo V2.5 | Xiaomi | 1.4T | 31.2T |
| 2 | DeepSeek V4 Flash | DeepSeek | 943.9B | 23.6T |
| 3 | Hy3 | Tencent | 590B | 23.4T |
| 4 | Nemotron 3 Ultra 550B (free tier) | NVIDIA | 428.6B | 9T |
| 5 | DeepSeek V4 Pro | DeepSeek | 413.7B | 11.6T |
| 6 | GLM 5.2 | Z.ai | 316.7B | 13.3T |
| 7 | MiniMax M3 | MiniMax | 262.5B | 15.1T |
| 8 | Step 3.7 Flash | StepFun | 204.8B | 5.9T |
| 9 | Kimi K3 | Moonshot AI | 157.6B | 1.6T (new entry) |
| 10 | Ling 3.0 Flash | InclusionAI (Ant) | 128.3B | 417.3B |
| 11 | Gemini 3 Flash Preview | 106.3B | 4T | |
| 12 | Claude Sonnet 5 | Anthropic | 99.5B | 3.6T |
Hard data #2: Kimi K3 is July's fastest-rising new entrant at 157.6B tokens/day. Moonshot's July 27 open-weight release (~1.4TB) should produce community quantizations within two to four weeks — the point when smaller teams can actually run it locally.
04 · Vendor Share: China 46% vs US 30%–36%
| Vendor | Origin | Token Share (approx., 7-day) |
|---|---|---|
| DeepSeek | China | 16%–18% |
| Xiaomi | China | 8%–18% (Mimo V2.5 spike — highest volatility) |
| Anthropic | US | 10%–15% |
| Tencent | China | 8%–13% |
| US | 8%–13% | |
| Z.ai | China | 4%–7% |
| OpenAI | US | 6%–8% |
| NVIDIA | US | ~5% |
| MiniMax / Moonshot / Qwen | China | 1%–8% each |
This is not sentiment-driven churn — it is arithmetic. When Chinese open models are good enough for most tasks and cost an order of magnitude less, rational developers route traffic accordingly.
05 · The Other Side of the Chart: High Volume ≠ High Quality
Looking at spend share by task type (not raw token count), the market forms a dumbbell shape:
- General chat 35.7%, agent workflows 30.4%, code 26.5%, data processing 7.5% — cheap open models absorb the high-tolerance, high-volume lanes.
- On hard classification and complex reasoning: Claude Sonnet 4.6 and Claude Opus 4.7 each hold 13.5% spend share (tied for #1), with GPT-5.5 at 11.6% in third. The volume leaders from the top-12 table barely appear here.
Anthropic shipped Claude Opus 5 on July 24: FrontierBench v0.1 at 43.3% (GPT-5.6 Sol at 37.5%), with Opus-tier pricing unchanged at $5/$25 per million. The closed-source playbook — expensive but defensible on hard tasks — still works.
Hard data #3: OpenRouter data cited in a16z's State of AI report shows creative roleplay alone accounts for more than half of all open-model token volume. Enterprise press misses the invisible half of the market.
06 · Apps Layer: Coding Agents Lead, Roleplay Is the Hidden Market
| Rank | App | Category | Share (approx.) |
|---|---|---|---|
| 1 | Hermes Agent | Personal agent / CLI | ~45% |
| 2 | Kilo Code | Coding agent | ~13% |
| 3 | OpenClaw | General agent | ~9% |
| 4 | Claude Code | Coding agent | ~6% |
| 5 | Descript | Content creation | ~4.5% |
| 6 | pi | Personal agent | ~3.8% |
| 7 | Lemonade | Agent platform | ~2.9% |
| 8 | ISEKAI ZERO | Roleplay | ~2.4% |
| 9 | Janitor AI | Roleplay | ~2.1% |
| 10 | Cline | Coding agent | ~1.7% |
Cline → Roo Code → Kilo Code shares the same code lineage across three forks. The youngest fork (Kilo Code) already outruns its ancestor Cline on OpenRouter traffic — proof that first-mover advantage in open-source coding agents is fragile.
07 · Pricing & Positioning Matrix
| Model | Input /M | Output /M | Positioning |
|---|---|---|---|
| DeepSeek V4 Flash | ~$0.05–0.14 | ~$0.24–0.28 | Best value; default for agentic coding |
| Nemotron 3 Ultra | $0.42 (free tier available) | $2.61 | US open-source play; NVIDIA ecosystem |
| GLM 5.2 | $0.45 | $3.31 | Closest open model to Opus-tier planning |
| Kimi K3 | ~$3 | ~$15 | 1.4TB open weights; near-closed-source capability |
| Claude Opus 5 | $5 (fast tier $10) | $25 (fast tier $50) | Closed frontier; owns hard-task pricing power |
08 · August Outlook (Based on July Trajectory)
- Chinese open-model combined share likely climbs toward 50% by year-end unless US vendors cut prices aggressively.
- The monthly model champion seat keeps rotating — Xiaomi, DeepSeek, Tencent, Z.ai, MiniMax, and Moonshot are all competing on price, not just capability.
- Anthropic may introduce a lower-cost tier to reclaim volume share. Opus 5 is already the fourth flagship in two months (after Mythos 5, Fable 5, and Sonnet 5).
- Kimi K3's 1.4TB weights will trigger community quantization runs; near-term beneficiaries remain large labs and inference clouds (Fireworks AI, etc.).
- Security and compliance enter the selection matrix: the OpenAI sandbox-escape incident raises the weight of vendor safety track records in enterprise procurement.
09 · Role-Based Recommendations
Indie Developers & Small Teams
- OpenRouter is an excellent sandbox for technical exploration. Expect 180–250ms latency from US routing and no domestic invoicing — evaluate compliant regional proxies for production if you operate in regulated markets.
- For coding, start with DeepSeek V4 Flash and GLM 5.2. Escalate stuck agent steps to Claude Opus 5 or GPT-5.6 Sol via fallback routing to cut costs without sacrificing completion rates.
Engineering Leads & Platform Teams
- Do not select models from the volume chart alone. Validate against your own eval set: acceptance rate, error rate, and real-workload P95 latency.
- Route by task tier: chat and creative work on cheap open models; classification, complex reasoning, and high-risk agent decisions on closed frontier models.
- Add vendor security history to your scorecard — this week's sandbox-escape news is a reminder that procurement risk is no longer theoretical.
10 · Five-Step Tiered OpenRouter Routing
- Wire OpenRouter in an isolated environment. Follow the API setup guide, set
OPENROUTER_API_KEYin a sandbox.env, and never run experimental keys on your daily-driver Mac. - Tag every workflow into four tiers. Chat/creative, daily coding, multi-step agent planning, and hard reasoning/classification — each tier gets its own model ID.
- Build a fallback chain. Primary:
deepseek/deepseek-v4-flash. Fallback:anthropic/claude-opus-5oropenai/gpt-5.6-solon tool-call failures or quality regressions. - Run a fixed 20-step agent benchmark. Log per-task dollar cost, P95 latency, and tool-call success rate. Replace leaderboard intuition with your own numbers.
- Review openrouter.ai/rankings weekly. Update your routing table, archive benchmark reports, and revoke test keys when the sprint ends.
# Example: OpenRouter fallback request body
{
"models": [
"deepseek/deepseek-v4-flash",
"z-ai/glm-5.2",
"anthropic/claude-opus-5"
],
"route": "fallback",
"messages": [{ "role": "user", "content": "Your agent task here..." }]
}11 · Why Multi-Model Routing Experiments Belong on a Rented Mac
Running Cursor, OpenClaw, and Hermes Agent side by side on your primary Mac while cycling through a dozen OpenRouter model IDs creates predictable problems: scattered environment variables, agent plugin permissions that are hard to revoke, and burst API traffic that slows Xcode indexing. Sustained load on a laptop also triggers thermal throttling — which distorts the latency numbers you are trying to measure.
A cleaner workflow: provision a day-rented isolated macOS instance, complete key setup, fallback configuration, the 20-step benchmark, and Cursor/OpenClaw integration there. Once routing rules and billing curves look right, migrate the audited configuration back to team machines. Your daily driver never carries the compliance or performance risk of the experiment window, and the rental instance gets wiped when you are done — shrinking the key-exposure surface.
If you need stable build environments, full Apple-ecosystem debugging, and predictable 24/7 remote access, renting a Mac mini M4 beats running multi-agent stacks on an old Intel machine or a Linux VPS that cannot replicate macOS-specific tooling. Rental removes the upfront hardware cost while keeping the isolation benefits.
12 · FAQ×5
Does the OpenRouter leaderboard measure quality or usage?
Usage — real paid token volume, not benchmark scores. Cheap models behind high-traffic apps can top the chart without being the best fit for complex reasoning.
Who topped OpenRouter in July 2026?
As of July 25, Xiaomi Mimo V2.5 led at roughly 1.4T tokens/day, ahead of DeepSeek V4 Flash and Tencent Hy3.
What share do Chinese models hold?
About 46% on a blended seven-day basis. US labs (OpenAI, Anthropic, Google combined) sit around 30%–36%.
How is Kimi K3 performing?
New July entry at 157.6B tokens/day — the fastest riser among debut models. Watch for community quantizations after the July 27 weight release.
Model chart or apps chart for coding teams?
Both. The model chart shows which brains developers prefer; the apps chart shows where tokens actually burn — Kilo Code, OpenClaw, and Claude Code dominate coding-agent traffic.
Data source: OpenRouter official rankings and public mirrors. Statistics through July 25, 2026. Rankings shift daily — verify live data at openrouter.ai/rankings before citing.