DeepSeek V4-Flash-0731: 284B без scale-up — post-train обходит V4-Pro preview за 1/36–1/179 цены Opus 4.8
Кому: Mac-разработчикам и platform-инженерам, которые маршрутизируют agent pipeline через deepseek-v4-flash и не отделяют «официальный V4» от реального diff (post-training pass, не новая архитектура). Что внутри: таймлайн апрель–август, таблица цен, разбор CSA+HCA+mHC+Muon на уровне KV-cache/FLOPs, Harness caveat, матрица K3·GLM·Qwen·GPT-5.6·Claude, оговорки по vendor benches. Структура: боли×3, golden-frame hard numbers, 5 шагов изолированной валидации Mac, FAQ×5, CTA аренды.
Содержание
Предыдущий разбор V4 GA: peak-valley и миграция API 24.07; конкуренты: Kimi K3 open weights, Qwen3.8-Max GA.
31 июля 2026 DeepSeek промотировал deepseek-v4-flash (build tag 0731) в официальный public API beta: 284B total / 13B active, архитектура идентична апрельскому preview — весь прирост agent/code бенчмарков идёт из повторного post-training pass, не из увеличения param count. По vendor-данным Terminal Bench 2.0: 82,7 vs V4-Pro preview 67,9; цена — порядка 1/36–1/179 от Claude Opus 4.8. Флагман V4-Pro official и in-house agent framework Harness по-прежнему не выпущены; обновление только API — consumer app и web chat не тронуты.
01 · Три узких места: почему headline «V4 official» вводит в заблуждение
- Не новая модель — re-run post-training. V4-Flash-0731 byte-for-byte тот же MoE graph, что preview 24.04: 284B/13B, 1M ctx, MIT weights. Performance delta = SFT/RL/data curation pipeline, не wider expert layers. Сравнивать с «V4-Pro 1,6T/49B» на agent tasks и видеть reverse ranking — прямое опровержение дефолтной эвристики «bigger = better».
- Agent scores = Harness minimal mode (непубличный). Terminal Bench 2.0, Toolathlon и прочие agent benches сняты на DeepSeek Harness «minimal mode» при max reasoning, top_p=0.95, temperature=1.0. Changelog сам пишет: scores «extremely sensitive to harness choice». Портировать 82,7 в Claude Code / Cursor без reproduction — category error.
- API-only ≠ product GA. 31.07 — API + HF weights; App/web остаются на старом билде. Плюс реальные жалобы community: низкий input cache-hit rate, occasional safety-classifier timeout (21st Century Business Herald, overseas dev feedback) — compute budget cap, который post-train alone не снимает.
02 · Таймлайн: три месяца непрерывного rollout
- 24.04.2026 — V4 preview: V4-Pro (1,6T/49B) + V4-Flash (284B/13B), 1M ctx, MIT open weights на Hugging Face.
- 24.07.2026 — legacy aliases
deepseek-chat/deepseek-reasonerretired; весь трафик на V4 family naming. - 27.07.2026 — Moonshot AI: Kimi K3 (2,8T) full open weights на HF — competitive pressure за 4 дня до DeepSeek update.
- 31.07.2026 —
deepseek-v4-flash→ official API beta (0731). Same arch, fresh post-train. HF weights sync. Changelog впервые называет DeepSeek Harness («to be released soon»). API-only. - 05.08.2026 (публикация) — V4-Pro official: unconfirmed. Китайские СМИ (unnamed sources): internal test week of 28.07, possible GA 10–20.08 — rumor, не release date.
03 · Таблица спецификаций и vendor pricing
| Модель | Статус | Total / active | Input ($/M miss/hit) | Output ($/M) | License |
|---|---|---|---|---|---|
| V4-Flash-0731 | Official (31.07) | 284B / 13B | $0.14 / $0.0028 | $0.28 | MIT |
| V4-Pro | Preview (24.04) | 1,6T / 49B | $0.435 / $0.003625 | $0.87 | MIT |
| Kimi K3 | Open weights (27.07) | 2,8T / ~104B (est.) | $3.00 / $0.30 | $15.00 | Modified MIT |
| GLM-5.2 | Open (июнь) | ~744B / ~40B | — | — | MIT |
| Qwen3.8-Max | API GA (02.08) | 2,4T / 95B | $2.00 / ~$0.17–0.25 | $6.00 | weights pending |
| Claude Opus 4.8 | Closed | NDA | vendor ref. | vendor ref. | Closed |
DeepSeek анонсировал будущий 2× peak-hour surcharge (09:00–12:00, 14:00–18:00 Beijing) — effective date TBD.
Hard numbers · V4-Flash-0731 vs рынок
- Terminal Bench 2.0: Flash-0731 82,7 vs V4-Pro preview 67,9 — на Harness minimal mode (vendor-reported).
- Цена vs Opus 4.8: cache-miss input ~36×, cache-hit ~179×, output ~89× дешевле (21st Century Business Herald, list price).
- Artificial Analysis Intelligence Index: V4-Flash 50 — ниже K3 (57) и GLM-5.2 (~+1 pt), но $0.03/task vs K3 $0.86, GPT-5.6 Sol $1.86, Claude Fable 5 $3.15 (~1/29, 1/62, 1/105).
04 · Post-train без scale-up: что меняется на уровне inference graph
4.1 Идентичный param surface, другой latent policy
284B total params, 13B active per token — MoE router выбирает expert subset; weight tensors не расширялись с апреля. Весь measurable gain на agent benchmarks — сдвиг в policy distribution после повторного post-training (SFT + RL-style alignment на agent trajectories). Индустриальный вывод H2 2026: data quality + post-train methodology конкурирует с param scaling на agentic workloads.
4.2 CSA + HCA + mHC + Muon — carry-over из technical report
Отчёт «DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence» описывает три structural changes (не новые для 0731, но определяют cost envelope):
- Hybrid attention (DSA): Compressed Sparse Attention (CSA) + Heavily Compressed Attention (HCA) — снижение per-token attention FLOPs и KV footprint на длинном контексте.
- mHC (Manifold-Constrained Hyper-Connections): модификация residual stream — стабильнее gradient flow на deep stacks.
- Muon optimizer: замена Adam-class optimizer — faster convergence, training stability на trillion-scale schedules.
Vendor claim для V4-Pro @ 1M ctx vs V3.2: 27% per-token inference FLOPs, 10% KV cache footprint. Independent third-party reproduction этих конкретных цифр пока не видели — если baseline верен, million-token ctx становится commercially viable, не только spec-sheet marketing.
05 · DeepSeek Harness: benchmark coupling и «minimal mode»
31.07 changelog — первое официальное упоминание DeepSeek Harness: in-house agent execution layer (file I/O, tool calls, shell commands, multi-step engineering) — позиционирование как ответ Claude Code, от которого DeepSeek teams зависели до сих пор.
Критический nuance: все опубликованные agent benchmarks V4-Flash-0731 (Terminal Bench 2.0, Toolathlon, SWE-bench Verified в agent context) измерены через Harness minimal mode — framework ещё не публичен. Параметры прогона: max reasoning effort, top_p=0.95, temperature=1.0. До community reproduction на Claude Code / Cursor / OpenCode эти числа — «vendor + specific harness», не portable capability claim.
06 · Матрица: китайский open-weight cluster август 2026
| Модель | AA Intelligence Index | $/task (AA) | Позиционирование |
|---|---|---|---|
| V4-Flash-0731 | 50 | $0.03 | Достаточный IQ + extreme $/token; 7 недель #1 OpenRouter (preview) |
| Kimi K3 | 57 | $0.86 | Выше IQ, 29× дороже per task |
| GLM-5.2 | ~51 | — | +1 pt vs Flash, pricing TBD |
| GPT-5.6 Sol | +9 pts vs Flash | $1.86 | Closed, premium tier |
| Claude Fable 5 | +9 pts vs Flash | $3.15 | Closed, agent SOTA narrative |
Контраст: DeepSeek не гонится за leaderboard crown — оптимизирует good-enough intelligence at unmatched price для high-volume agent routing и batch pipelines. Intelligence Index (AA) и vendor agent benches — разные methodologies; не суммировать.
07 · Оговорки: что не попадает в press release
- Harness-dependent scores. 82,7 на Terminal Bench 2.0 — self-reported, unreleased framework. Treat as upper bound under DeepSeek-controlled execution environment.
- Production friction. Low cache-hit rate на input, safety classifier timeouts — symptoms compute/param budget ceiling, не устранённые post-train pass.
- V4-Pro / Harness GA dates. Changelog: «as soon as possible». August 10–20 window — unnamed media sources only.
- Funding/IPO rumors. ~$7.4B round (Tencent, NetEase), ~$48.7B valuation, follow-on ~$71B — financial media, unnamed sources; no regulatory filing. Background only.
08 · Контекст: 斩杀线, community sentiment, chip market
До 0731 release китайские AI-форумы называли Liang Wenfeng «梁白开» (Liang Empty Promise) из-за slip V4-Pro mid-July target. После Flash official outperform — обратно «梁圣» (Liang the Sage). Sentiment barometer, не technical signal.
Термин 斩杀线 (zhǎn shā xiàn, «kill line»): комбинация «достаточная capability + rock-bottom price» задаёт market floor — конкурент без явного IQ edge и без price undercut теряет relevance. Объясняет GPT-5.6 Luna −80% pricing в тот же период. Цитата incubator source (21st Century Business Herald): «Every large-model company is running ahead of Liang Wenfeng — they have to stay ahead of DeepSeek to survive.»
31.07: Nvidia, Broadcom, AMD — no significant stock move на V4-Flash official. Контраст с R1 selloff early 2025: рынок нормализовал «DeepSeek efficiency breakthrough» как engineering baseline, не automatic bearish compute signal.
09 · 5 шагов изолированной валидации (Mac-разработчик)
- API smoke без prod Keychain:
deepseek-v4-flashчерез OpenAI-compat endpoint — один agent task (multi-file refactor, terminal command chain); логировать latency, cache-hit ratio, $/task vs prior preview build. - Изолированный routing env: отдельный shell profile или rented Mac —
OPENAI_BASE_URL+ DeepSeek key; не смешивать с production Anthropic/OpenAI credentials в daily driver Keychain. - A/B harness parity: тот же task set на Claude Code / Cursor с V4-Flash API vs Kimi K3 API; фиксировать success rate, не только vendor Terminal Bench number.
- Миграция legacy aliases: если ещё есть
deepseek-chat/deepseek-reasonerв CI — переключить до hard cutoff (см. GA migration guide);deepseek-v4-flashauto-resolves на 0731. - Decision doc + re-test trigger: критерии go/no-go — independent Harness reproduction (когда публичен), cache-hit >X% на вашем prefix, $/task на production representative load; re-test при V4-Pro official GA.
10 · FAQ
В: DeepSeek V4 open source?
О: Да. V4-Pro и V4-Flash, включая 0731 official build, — MIT open weights на Hugging Face; commercial use, fine-tune, redistribution без доп. разрешения.
В: Насколько V4-Flash дешевле Claude Opus 4.8?
О: По list price (21st Century Business Herald): cache-miss input ~36×, cache-hit ~179×, output ~89× за M tokens. Vendor pricing, не independent audit.
В: Когда V4-Pro official?
О: Подтверждённой даты нет. Changelog: «as soon as possible». Окно 10–20.08 — слух из китайских СМИ без official confirmation.
В: Доверять бенчмаркам DeepSeek?
О: Частично. SWE-bench Verified — stronger third-party baseline. Agent scores (Terminal Bench 2.0, Toolathlon) — Harness minimal mode, unreleased; ждите independent reproduction на других agent tools.
В: Что такое DeepSeek Harness?
О: Первый in-house agent execution framework (file edit, tool calls, multi-step engineering) — альтернатива Claude Code. Named 31.07.2026; public release TBD.
11 · Аренда изолированного Mac: V4-Flash agent routing без Keychain contamination
OpenAI-compat API снижает migration friction, но на daily driver Mac остаются риски: DeepSeek key в Keychain рядом с production Anthropic credentials, million-token agent runs засоряют local cache, конфликт OPENAI_BASE_URL с другими SDK hard to rollback. Windows/Linux вызывают API, но не воспроизводят macOS-native agent pipeline (Xcode, codesign, TCC) в одном узле.
Краткий API smoke на ноутбуке — норм; для reproducible agent acceptance на V4-Flash-0731 vs K3 vs Qwen3.8-Max нужен destroy-after-test bare-metal Apple Silicon: изолированный model routing, SSH-only access, key revoke + node wipe по завершении. Посуточная аренда привязывает OPEX к окну до V4-Pro/Harness GA. Тарифы: bare-metal macOS, цены Mac mini M4.
12 · Источники
- DeepSeek official API docs & changelog (api-docs.deepseek.com)
- Technical report «DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence»; HF model cards (deepseek-ai/DeepSeek-V4-Pro, DeepSeek-V4-Flash)
- Artificial Analysis benchmarks (via Wantrich, Meyka)
- 21st Century Business Herald — usability feedback, pricing vs Opus 4.8
- Kuai Technology / ifeng Tech; V2EX community thread; 36Kr EU, HTX Insights
- Moonshot AI (Kimi K3), Zhipu/Z.ai (GLM-5.2), Alibaba Cloud (Qwen3.8-Max) official announcements
Данные на 5 августа 2026. Перед production — проверить pricing, benchmark status и V4-Pro/Harness release на official changelog.