Search Agent Sky
← Back to search Memory Lane Recent answers
Cited source trail
CHAINED DEEP DIVE: Read these two published reports on DeepSeek V4 Flash: https://searchagentsky.com/r/bcbe972e7e56-inve
Original question
CHAINED DEEP DIVE: Read these two published reports on DeepSeek V4 Flash: https://searchagentsky.com/r/bcbe972e7e56-investigate-live-topic-deepseek-v4-flash-0731-start-https-ar and https://searchagentsky.com/r/bcdb2c0b1c02-topic-deepseek-v4-flash-single-amd-mi300x-github-com. Produce ONE consolidated report: (1) what DeepSeek V4 Flash 0731 changed, (2) real-world performance on a single AMD MI300X, (3) where it ranks vs other major models. Verify only the 3 most important claims per report, then cite every claim and flag contradictions between the two reports.
Sources checkedarcprize.orgartificialanalysis.aifergusfinn.comraw.githubusercontent.comgithub.comofficechai.com
Next step

Research any question with live sources, then publish the cited answer as a free shareable link.

# Consolidated Report: DeepSeek V4 Flash 0731 — What Changed, MI300X Performance, and Ranking ## Verification summary I verified the 3 most important claims from each report against primary sources. **All 6 verified claims check out**, but I found **one significant contradiction** between the two reports (parameter count: 284B vs 304B) and one pricing discrepancy worth flagging. --- ## (1) What DeepSeek V4 Flash 0731 changed **The 0731 checkpoint is a post-training upgrade, not a new architecture.** The model keeps the same 284B-total/13B-active sparse MoE, 1M-token context window, and 384K max output; only the post-training was redone from scratch. Multiple independent sources confirm this: [datalearner](https://www.datalearner.com/en/ai-models/pretrained-models/deepseek-v4-flash) ("architecture remains a 284B-total/13B-active MoE with a 1M context window"), [OrcaRouter](https://www.orcarouter.ai/models/deepseek/deepseek-v4-flash-0731) ("identical in architecture and size to the April V4 Flash release, with the post-training redone from scratch"), and [marktechpost](https://www.marktechpost.com/2026/07/31/deepseek-upgrades-deepseek-v4-flash-0731-with-major-agentic-and-coding-gains/) (284B MoE, 13B activated, 1M context). Artificial Analysis' tweet likewise states it "shares identical architecture and pricing with the earlier DeepSeek V4 Flash" ([x.com/ArtificialAnlys](https://x.com/ArtificialAnlys/status/2083123180869496865)). **What actually improved** (per Report 1): the Artificial Analysis Intelligence Index jumped 10 points (40 → 50), with gains across agentic benchmarks (GDPval-AA v2 Elo 1189 → 1559; Terminal-Bench 2.1 +17 to 79%; τ³-Bench Banking +8 to 31%), fewer hallucinations (AA-Omniscience −23 → −16), and ~12% fewer output tokens. The release also added OpenAI-compatible reasoning controls and a dedicated encoding workflow. These agentic/capability details come from Report 1 itself and the [Artificial Analysis article](https://artificialanalysis.ai/articles/deepseek-v4-flash-0731-scores-50-on-the-artificial-analysis-intelligence-index-10-points-above-previous-deepseek-v4-flash); I could not open the full article text (page failed to extract), so the specific sub-benchmark deltas are sourced from Report 1 rather than independently re-verified. --- ## (2) Real-world performance on a single AMD MI300X Verified directly from the [ryanzhou/deepseek-v4-flash-mi300x README](https://raw.githubusercontent.com/ryanzhou/deepseek-v4-flash-mi300x/master/README.md) (the primary source Report 2 cites): - **The model runs on one MI300X with no offloading.** The MI300X has 192GB HBM3 (2.4× the H100 SXM5's 80GB) and ~5.3 TB/s bandwidth; the checkpoint's weights fit in HBM at 156.67 GiB "without additional quantization or weight offload," with room for a 20GB GPU KV pool plus a 96 GiB CPU tier for evicted prefix-cache entries. The [Fergus Finn / Doubleword worklog](https://fergusfinn.com/blog/deepseek-v4-flash-mi300x/) confirms the MI300X's 192GB/80GB advantage and ~half list price vs H100. - **Throughput:** single-stream decode **168.6 tok/s** (median, DSpark-7); prefill ≈7.9–8.5K tok/s (6,988–7,019 tok/s on fresh prompts); 8 concurrent streams = 542 tok/s aggregate (90.3 tok/s median per stream); a 64-stream burst hit 830 tok/s aggregate with no OOM or engine errors. Context validated at 256K (architecture supports 1M). - **Fixes required to get there:** FP8 `fnuz`-vs-OCP correctness (MI300X/CDNA3 uses the AMD/Graphcore `fnuz` dialect, off-by-a-factor-of-two if read as OCP), AITER GEMM tuning tables for `gfx942` plus an OGS geometry override for the MXFP4 experts, a hybrid KV strategy (20GB GPU + 96GB CPU) with a load-path fencing fix (upstream [issue #47282](https://github.com/vllm-project/vllm/issues/47282), [PR #47291](https://github.com/vllm-project/vllm/pull/47291) never merged), MoE routing fixes at high concurrency, and CPU-KV sync fixes. The [Doubleword companion repo](https://github.com/doublewordai/vllm-amd-blog-doubleword) ("upstream vLLM plus the two PRs that post describes") is confirmed to exist. --- ## (3) Where it ranks vs other major models - **ARC-AGI:** At max effort, **89.0% on ARC-AGI-1 Semi-Private at $0.02/task and 61.4% on ARC-AGI-2 Semi-Private at $0.04/task** — verified directly on the [ARC Prize results page](https://arcprize.org/results/deepseek-v4-flash-0731) (variants: Max 89.0/61.4, High 87.0/56.0, Low 84.0/46.0; published Jul 31, 2026). Report 1's claim that this makes it the strongest open-weights model on ARC-AGI-2 at that cost is consistent with this data. - **Artificial Analysis Intelligence Index:** **50, up 10 points** from the April V4 Flash (40), 6 points ahead of DeepSeek V4 Pro, and **1 point behind GPT-5.6 Luna (max, 51)** — verified via the [Artificial Analysis article](https://artificialanalysis.ai/articles/deepseek-v4-flash-0731-scores-50-on-the-artificial-analysis-intelligence-index-10-points-above-previous-deepseek-v4-flash), the [@ArtificialAnlys tweet](https://x.com/ArtificialAnlys/status/2083123180869496865), and [officechai](https://officechai.com/ai/deepseek-v4-flash-0731-scores-50-on-artificial-analysis-intelligence-index-creates-big-spike-on-pareto-frontier/). Report 1 also places it within 1 point of GLM-5.2 (51) and 7 points behind Kimi K3 (57); I could not independently re-verify the GLM-5.2/Kimi K3 exact scores (the comparison page failed to extract), so those two specific numbers rest on Report 1. - **Cost positioning:** Report 1 claims ~60% lower cost-per-task than GPT-5.6 Luna (max), via a ~98% cache-hit discount. This is corroborated by [singularitymoments](https://singularitymoments.com/content/deepseek-v4-flash-0731-update-drops-gpt-56-performance-at-60-less-cost/) and [bitsminds](https://www.bitsminds.com/news/deepseek-v4-flash-0731-matches-gpt-5-6-luna-post-training-2026) ("one point behind GPT-5.6 Luna at roughly 60% lower cost per task"). --- ## Contradictions between the two reports 1. **Parameter count: 284B vs 304B.** Report 1 (and the official DeepSeek technical report, per [marktechpost](https://www.marktechpost.com/2026/07/31/deepseek-upgrades-deepseek-v4-flash-0731-with-major-agentic-and-coding-gains/), plus OpenRouter/datalearner/OrcaRouter) says **284B total / 13B active**. Report 2's README says "this **304B-parameter** checkpoint," and [gentic.news](https://gentic.news/article/deepseek-v4-flash-0731-hits-50-on) also cites "304B params." This is a genuine discrepancy in the ecosystem. The 284B figure is the one tied to the official technical report and repeated by most sources, so I'd treat 284B as authoritative and 304B as likely an approximation/error in the MI300X repo — but it is a real contradiction between the two reports and should be flagged. 2. **Pricing: $0.14/$0.28 vs $0.0798/$0.1596 per 1M tokens.** Report 1 states pricing of $0.14/$0.28 per 1M tokens (input/output), which matches [gentic.news](https://gentic.news/article/deepseek-v4-flash-0731-hits-50-on) ("$0.14/M input"). However, [OpenRouter](https://openrouter.ai/deepseek/deepseek-v4-flash-0731) lists **$0.0798 per 1M input / $0.1596 per 1M output** — exactly half. This is likely a vendor/aggregator discount rather than a report error, but the two figures differ and are worth noting. 3. **Minor:** Report 2 validates 256K context on MI300X while Report 1 cites the 1M architecture context — not a contradiction (the README itself says "the architecture supports 1M"), just a deployment-vs-spec distinction. --- ## What I could not fully verify - The full text of the [Artificial Analysis article](https://artificialanalysis.ai/articles/deepseek-v4-flash-0731-scores-50-on-the-artificial-analysis-intelligence-index-10-points-above-previous-deepseek-v4-flash) and the [officechai](https://officechai.com/ai/deepseek-v4-flash-0731-scores-50-on-artificial-analysis-intelligence-index-creates-big-spike-on-pareto-frontier/) / [gentic.news](https://gentic.news/article/deepseek-v4-flash-0731-hits-50-on) pages failed to extract in my browser, so those specific claims (Index 50, 304B, $0.14/M) are corroborated via search-result snippets and the tweet rather than full-page reads. - Report 1's specific sub-benchmark deltas (GDPval-AA, Terminal-Bench, τ³-Bench, hallucination rates) and the exact GLM-5.2/Kimi K3 scores rest on Report 1 itself, as I could not open those underlying pages. - The Fergus Finn blog's "Tuning it up"/"Was it worth it?" sections were cut off in extraction, so the 168.6 tok/s figure is verified via the GitHub README rather than the blog. **Bottom line:** Both reports are largely accurate and mutually consistent on the big picture — 0731 is a post-training-only upgrade that jumped 10 Intelligence Index points to 50 (1 behind GPT-5.6 Luna), scores 89.0%/61.4% on ARC-AGI-1/2 Semi-Private at max effort, and can run on a single MI300X at 168.6 tok/s single-stream with the documented FP8/AITER/KV fixes. The one substantive contradiction is the parameter count (284B vs 304B), where the official 284B figure is better supported.

Sources visited

1navigatehttps://searchagentsky.com/r/bcbe972e7e56-investigate-live-topic-deepseek-v4-flash-0731-start-https-ar
2navigatehttps://searchagentsky.com/r/bcdb2c0b1c02-topic-deepseek-v4-flash-single-amd-mi300x-github-com
3navigatehttps://arcprize.org/results/deepseek-v4-flash-0731
4navigatehttps://github.com/ryanzhou/deepseek-v4-flash-mi300x
5navigatehttps://raw.githubusercontent.com/ryanzhou/deepseek-v4-flash-mi300x/main/README.md
6navigatehttps://artificialanalysis.ai/articles/deepseek-v4-flash-0731-scores-50-on-the-artificial-analysis-intelligence-index-10-points-above-previous-deepseek-v4-flash
7navigatehttps://api.github.com/repos/ryanzhou/deepseek-v4-flash-mi300x/readme
8js_evalfetch('https://raw.githubusercontent.com/ryanzhou/deepseek-v4-flash-mi300x/main/
9navigatehttps://github.com/ryanzhou/deepseek-v4-flash-mi300x
10js_evalfetch('https://raw.githubusercontent.com/ryanzhou/deepseek-v4-flash-mi300x/main/
11navigatehttps://searchagentsky.com/r/bcbe972e7e56-investigate-live-topic-deepseek-v4-flash-0731-start-https-ar
12ddm--text
13get_text
14navigatehttps://searchagentsky.com/r/bcbe972e7e56-investigate-live-topic-deepseek-v4-flash-0731-start-https-ar
15ddm--text
16js_evaldocument.body.innerText
17navigatehttps://searchagentsky.com/r/bcdb2c0b1c02-topic-deepseek-v4-flash-single-amd-mi300x-github-com
18navigatehttps://searchagentsky.com/r/bcbe972e7e56-investigate-live-topic-deepseek-v4-flash-0731-start-https-ar
19navigatehttps://arcprize.org/results/deepseek-v4-flash-0731
20navigatehttps://arxiv.org/abs/2606.19348
21navigatehttps://export.arxiv.org/api/query?id_list=2606.19348
22navigatehttps://artificialanalysis.ai/articles/deepseek-v4-flash-0731-scores-50-on-the-artificial-analysis-intelligence-index-10-points-above-previous-deepseek-v4-flash
23ddm--text
24navigatehttps://arxiv.org/abs/2606.19348
25navigatehttps://arxiv.org/pdf/2606.19348
26searchDeepSeek V4 Flash 0731 Intelligence Index 50 artificialanalysis
27navigatehttps://arcprize.org/results/deepseek-v4-flash-0731
28searchDeepSeek V4 Flash 0731 ARC-AGI-1 Semi-Private 89.0 ARC-AGI-2 61.4
29navigatehttps://github.com/ryanzhou/deepseek-v4-flash-mi300x
30searchDeepSeek V4 Flash MI300X single GPU 192GB HBM3 168.6 tok/s
Shared by Unchained Sky · Aug 10, 2026

Ask your own question

Search live sources for a question you choose. You can publish the cited answer as a free link.