OpenRouter Rankings Chinese Models 46% 2026-07-27 · Data through 7/25

OpenRouter Rankings July 2026:
Who's Actually Winning the AI Model Race

Who has the problem? Cursor and OpenClaw developers still defaulting to last quarter's benchmark winner, plus engineering leads who need to explain why a $0.05/M model tops the chart while Claude barely cracks the top 12. What you get: OpenRouter production data through July 25 — Mimo V2.5 at 1.4T tokens/day, Chinese vendors at 46% share, and the usage-vs-quality split that most leaderboard roundups miss. What's inside: Top 12 model table, vendor share breakdown, apps-layer top 10, pricing matrix, August outlook, a five-step tiered routing playbook, and FAQ×5.

July 2026 OpenRouter AI model rankings chart showing vendor token share and top models by daily volume

Related reading: OpenRouter API setup guide, June 2026 rankings and Chinese model takeover, and weekly token billing truth.

01 · Key Takeaways (Data Through 2026-07-25)

  • Daily volume crown shifted again. Xiaomi Mimo V2.5 hit roughly 1.4T tokens/day to claim #1. DeepSeek V4 Flash (943.9B) and Tencent Hy3 (590B) follow. Seven of the top twelve models are Chinese-origin.
  • Vendor share crossed a structural threshold. Chinese labs combined hold about 46% of OpenRouter token volume on a seven-day basis — up from under 2% a year ago. US labs (OpenAI + Anthropic + Google) now sit around 30%–36%, down from roughly 70%.
  • DeepSeek remains the steadiest #1 vendor. Single-vendor share holds at 16%–18%, but the monthly model champion keeps rotating: MiniMax M2.5 in February, MiMo-V2-Pro in April, Mimo V2.5 in July.
  • Hard data #1: DeepSeek V4 Flash inputs run $0.05–0.14/M versus GPT-5.5 at roughly $5/M — a 35× price gap that explains the traffic migration better than any narrative about model intelligence.
OpenRouter rankings measure where developers actually send paid requests — not who scores highest on a static benchmark. That distinction is the entire point of this analysis.

02 · Three Selection Pain Points: Why Rankings Still Mislead

The leaderboard is an input to your decision, not the decision itself. Mac developers and AI platform teams hit three recurring traps in July:

  1. Equating token volume with capability. Inexpensive models wired behind Hermes Agent, roleplay apps, and high-volume CLI wrappers can crush frontier models on daily tokens while barely registering in classification or hard-reasoning spend share.
  2. Chasing the daily #1 without tracking consistency. Rankings swing daily — Claude Opus 4.8 was still top-12 on 7/24 but got bumped by Ling 3.0 Flash on 7/25. A monthly champion is not a durable moat.
  3. Ignoring the apps layer. Coding agents (Kilo Code ~13%), general agents (OpenClaw ~9%), and roleplay platforms (Janitor AI, ISEKAI ZERO) burn enormous open-model traffic that enterprise AI coverage rarely mentions. If your workload is agentic coding, check openrouter.ai/apps alongside the model chart.

03 · July Model Token Volume Top 12 (Daily, Through 7/25)

RankModelVendorDaily Tokens30-Day Total
1Mimo V2.5Xiaomi1.4T31.2T
2DeepSeek V4 FlashDeepSeek943.9B23.6T
3Hy3Tencent590B23.4T
4Nemotron 3 Ultra 550B (free tier)NVIDIA428.6B9T
5DeepSeek V4 ProDeepSeek413.7B11.6T
6GLM 5.2Z.ai316.7B13.3T
7MiniMax M3MiniMax262.5B15.1T
8Step 3.7 FlashStepFun204.8B5.9T
9Kimi K3Moonshot AI157.6B1.6T (new entry)
10Ling 3.0 FlashInclusionAI (Ant)128.3B417.3B
11Gemini 3 Flash PreviewGoogle106.3B4T
12Claude Sonnet 5Anthropic99.5B3.6T

Hard data #2: Kimi K3 is July's fastest-rising new entrant at 157.6B tokens/day. Moonshot's July 27 open-weight release (~1.4TB) should produce community quantizations within two to four weeks — the point when smaller teams can actually run it locally.

04 · Vendor Share: China 46% vs US 30%–36%

VendorOriginToken Share (approx., 7-day)
DeepSeekChina16%–18%
XiaomiChina8%–18% (Mimo V2.5 spike — highest volatility)
AnthropicUS10%–15%
TencentChina8%–13%
GoogleUS8%–13%
Z.aiChina4%–7%
OpenAIUS6%–8%
NVIDIAUS~5%
MiniMax / Moonshot / QwenChina1%–8% each

This is not sentiment-driven churn — it is arithmetic. When Chinese open models are good enough for most tasks and cost an order of magnitude less, rational developers route traffic accordingly.

05 · The Other Side of the Chart: High Volume ≠ High Quality

Looking at spend share by task type (not raw token count), the market forms a dumbbell shape:

  • General chat 35.7%, agent workflows 30.4%, code 26.5%, data processing 7.5% — cheap open models absorb the high-tolerance, high-volume lanes.
  • On hard classification and complex reasoning: Claude Sonnet 4.6 and Claude Opus 4.7 each hold 13.5% spend share (tied for #1), with GPT-5.5 at 11.6% in third. The volume leaders from the top-12 table barely appear here.

Anthropic shipped Claude Opus 5 on July 24: FrontierBench v0.1 at 43.3% (GPT-5.6 Sol at 37.5%), with Opus-tier pricing unchanged at $5/$25 per million. The closed-source playbook — expensive but defensible on hard tasks — still works.

Hard data #3: OpenRouter data cited in a16z's State of AI report shows creative roleplay alone accounts for more than half of all open-model token volume. Enterprise press misses the invisible half of the market.

06 · Apps Layer: Coding Agents Lead, Roleplay Is the Hidden Market

RankAppCategoryShare (approx.)
1Hermes AgentPersonal agent / CLI~45%
2Kilo CodeCoding agent~13%
3OpenClawGeneral agent~9%
4Claude CodeCoding agent~6%
5DescriptContent creation~4.5%
6piPersonal agent~3.8%
7LemonadeAgent platform~2.9%
8ISEKAI ZERORoleplay~2.4%
9Janitor AIRoleplay~2.1%
10ClineCoding agent~1.7%

Cline → Roo Code → Kilo Code shares the same code lineage across three forks. The youngest fork (Kilo Code) already outruns its ancestor Cline on OpenRouter traffic — proof that first-mover advantage in open-source coding agents is fragile.

07 · Pricing & Positioning Matrix

ModelInput /MOutput /MPositioning
DeepSeek V4 Flash~$0.05–0.14~$0.24–0.28Best value; default for agentic coding
Nemotron 3 Ultra$0.42 (free tier available)$2.61US open-source play; NVIDIA ecosystem
GLM 5.2$0.45$3.31Closest open model to Opus-tier planning
Kimi K3~$3~$151.4TB open weights; near-closed-source capability
Claude Opus 5$5 (fast tier $10)$25 (fast tier $50)Closed frontier; owns hard-task pricing power

08 · August Outlook (Based on July Trajectory)

  1. Chinese open-model combined share likely climbs toward 50% by year-end unless US vendors cut prices aggressively.
  2. The monthly model champion seat keeps rotating — Xiaomi, DeepSeek, Tencent, Z.ai, MiniMax, and Moonshot are all competing on price, not just capability.
  3. Anthropic may introduce a lower-cost tier to reclaim volume share. Opus 5 is already the fourth flagship in two months (after Mythos 5, Fable 5, and Sonnet 5).
  4. Kimi K3's 1.4TB weights will trigger community quantization runs; near-term beneficiaries remain large labs and inference clouds (Fireworks AI, etc.).
  5. Security and compliance enter the selection matrix: the OpenAI sandbox-escape incident raises the weight of vendor safety track records in enterprise procurement.

09 · Role-Based Recommendations

Indie Developers & Small Teams

  • OpenRouter is an excellent sandbox for technical exploration. Expect 180–250ms latency from US routing and no domestic invoicing — evaluate compliant regional proxies for production if you operate in regulated markets.
  • For coding, start with DeepSeek V4 Flash and GLM 5.2. Escalate stuck agent steps to Claude Opus 5 or GPT-5.6 Sol via fallback routing to cut costs without sacrificing completion rates.

Engineering Leads & Platform Teams

  • Do not select models from the volume chart alone. Validate against your own eval set: acceptance rate, error rate, and real-workload P95 latency.
  • Route by task tier: chat and creative work on cheap open models; classification, complex reasoning, and high-risk agent decisions on closed frontier models.
  • Add vendor security history to your scorecard — this week's sandbox-escape news is a reminder that procurement risk is no longer theoretical.

10 · Five-Step Tiered OpenRouter Routing

  1. Wire OpenRouter in an isolated environment. Follow the API setup guide, set OPENROUTER_API_KEY in a sandbox .env, and never run experimental keys on your daily-driver Mac.
  2. Tag every workflow into four tiers. Chat/creative, daily coding, multi-step agent planning, and hard reasoning/classification — each tier gets its own model ID.
  3. Build a fallback chain. Primary: deepseek/deepseek-v4-flash. Fallback: anthropic/claude-opus-5 or openai/gpt-5.6-sol on tool-call failures or quality regressions.
  4. Run a fixed 20-step agent benchmark. Log per-task dollar cost, P95 latency, and tool-call success rate. Replace leaderboard intuition with your own numbers.
  5. Review openrouter.ai/rankings weekly. Update your routing table, archive benchmark reports, and revoke test keys when the sprint ends.
# Example: OpenRouter fallback request body { "models": [ "deepseek/deepseek-v4-flash", "z-ai/glm-5.2", "anthropic/claude-opus-5" ], "route": "fallback", "messages": [{ "role": "user", "content": "Your agent task here..." }] }

11 · Why Multi-Model Routing Experiments Belong on a Rented Mac

Running Cursor, OpenClaw, and Hermes Agent side by side on your primary Mac while cycling through a dozen OpenRouter model IDs creates predictable problems: scattered environment variables, agent plugin permissions that are hard to revoke, and burst API traffic that slows Xcode indexing. Sustained load on a laptop also triggers thermal throttling — which distorts the latency numbers you are trying to measure.

A cleaner workflow: provision a day-rented isolated macOS instance, complete key setup, fallback configuration, the 20-step benchmark, and Cursor/OpenClaw integration there. Once routing rules and billing curves look right, migrate the audited configuration back to team machines. Your daily driver never carries the compliance or performance risk of the experiment window, and the rental instance gets wiped when you are done — shrinking the key-exposure surface.

If you need stable build environments, full Apple-ecosystem debugging, and predictable 24/7 remote access, renting a Mac mini M4 beats running multi-agent stacks on an old Intel machine or a Linux VPS that cannot replicate macOS-specific tooling. Rental removes the upfront hardware cost while keeping the isolation benefits.

12 · FAQ×5

Does the OpenRouter leaderboard measure quality or usage?

Usage — real paid token volume, not benchmark scores. Cheap models behind high-traffic apps can top the chart without being the best fit for complex reasoning.

Who topped OpenRouter in July 2026?

As of July 25, Xiaomi Mimo V2.5 led at roughly 1.4T tokens/day, ahead of DeepSeek V4 Flash and Tencent Hy3.

What share do Chinese models hold?

About 46% on a blended seven-day basis. US labs (OpenAI, Anthropic, Google combined) sit around 30%–36%.

How is Kimi K3 performing?

New July entry at 157.6B tokens/day — the fastest riser among debut models. Watch for community quantizations after the July 27 weight release.

Model chart or apps chart for coding teams?

Both. The model chart shows which brains developers prefer; the apps chart shows where tokens actually burn — Kilo Code, OpenClaw, and Claude Code dominate coding-agent traffic.

Data source: OpenRouter official rankings and public mirrors. Statistics through July 25, 2026. Rankings shift daily — verify live data at openrouter.ai/rankings before citing.