Search Agent Sky
← Back to search Memory Lane Recent answers
Cited source trail
continue
Sources checkedaireleasetracker.comapi-docs.deepseek.comanthropic.comblog.google
Next step

Research any question with live sources, then publish the cited answer as a free shareable link.

Do not call any tools.# Latest AI Model Releases — July 2026 July 2026 was one of the most crowded months in AI history — **eight major model releases** in four weeks, spanning every tier from frontier flagships to workhorse APIs, plus the largest open-weight model ever shipped. Here's what changed and why it matters. ## The releases, in order | Date | Model | Company | Tier | |------|-------|---------|------| | Jul 8 | **Grok 4.5** | SpaceXAI (xAI) | Frontier coding | | Jul 9 | **Muse Spark 1.1** | Meta | Agentic workhorse | | Jul 16 | **Kimi K3** | Moonshot AI | Open-weight frontier | | Jul 21 | **Gemini 3.6 Flash** (+ 3.5 Flash-Lite, 3.5 Flash Cyber) | Google | Workhorse API | | Jul 24 | **Claude Opus 5** | Anthropic | Frontier flagship | | Jul 31 | **DeepSeek-V4-Flash-0731** | DeepSeek | Agentic refresh | ## What changed **1. The biggest story: Kimi K3 — the world's largest open-weight model.** Moonshot AI shipped a **2.8-trillion-parameter** open-source model (July 16), the first "3T-class" open system. Built on Kimi Delta Attention (a hybrid linear-attention mechanism) with a 1M-token context and native vision, it's priced at $3/$15 per million tokens. Reuters, CNBC, BBC, and VentureBeat all covered it as China closing in on US frontier labs — Moonshot claims it rivals OpenAI and Anthropic, and it topped the Arena Elo (Code) leaderboard at 1679. Full weights were expected by July 27. This is the clearest signal yet that open-weight models now compete at the frontier. **2. Anthropic's pricing gambit: Claude Opus 5.** Released July 24, Opus 5 delivers "flagship-class agentic capability at workhorse pricing" — **half the price of Claude Fable 5** while keeping Opus 4.8's $5/$25 rates. It set new SOTA on Frontier-Bench v0.1 (43.3%, more than double Opus 4.8's 21.1%), GDPval-AA v2 (1861 Elo), Humanity's Last Exam with tools (64.7%), and ARC-AGI-3 (30.2%, ~3× the next best). It's now the default on Claude Max and the strongest model on Claude Pro. The strategic move: make frontier capability the *everyday* option. **3. Google iterates fastest on the workhorse tier.** Gemini 3.6 Flash (July 21) is a straight quality upgrade at the **same price** as 3.5 Flash ($1.50/$7.50) while using ~17% fewer output tokens. It jumped to **#1 on MLE-Bench (63.9%)**, hit 49% on DeepSWE (up from 37%), and 83% on OSWorld-Verified. Computer use became a built-in client-side tool. Google also shipped 3.5 Flash-Lite ($0.30/$2.50, 350 tok/s) and the security-focused 3.5 Flash Cyber — while Gemini 3.5 Pro stays in testing and Gemini 4 pre-training has begun. **4. DeepSeek's agentic checkpoint refresh.** DeepSeek-V4-Flash-0731 (July 31, the newest release) is a post-training refresh of the April V4-Flash build — same architecture, but re-trained for agentic work with an unusually large jump for a checkpoint update: Terminal-Bench 2.1 went 61.8 → **82.7**, and CyberGym 76.7 now surpasses V4-Pro-Preview. It's API-only, served under the unchanged `deepseek-v4-flash` name (existing callers switched with no code change), and natively supports the OpenAI Responses API format for Codex compatibility. **5. Meta and SpaceXAI fight on price and speed.** **Muse Spark 1.1** (July 9) launched the same day as OpenAI's GPT-5.6 — Zuckerberg returned to X after three years to announce it — undercutting Claude Opus 4.8 by up to **83% per token** ($1.25/$4.25) while topping agentic tool-use benchmarks. **Grok 4.5** (July 8), jointly trained with Cursor, is an "Opus-class" coding model at $2/$6 with ~4× better token efficiency than Opus 4.8 — 64.7% SWE-Bench Pro, 83.3% Terminal-Bench 2.1, 500K context. ## Why it matters - **The price war is the real story.** Every major lab is now competing on cost-per-task, not just raw benchmark scores — Opus 5 at half of Fable 5's price, Gemini 3.6 Flash at the same price with better results, Muse Spark 1.1 at 83% below Claude, Grok 4.5 at $2/$6. The economics of agentic AI are collapsing fast. - **Open weights reached the frontier.** Kimi K3's 2.8T open release is a structural shift — the largest open model ever, from a Chinese lab, rivaling US frontier systems. - **Agentic capability is the battleground.** Nearly every release is benchmarked on agentic coding (SWE-Bench, DeepSWE, CursorBench, Terminal-Bench), computer use (OSWorld), and knowledge work (GDPval) — not just raw reasoning. The models are being built and measured for autonomous work. - **China is closing the gap.** DeepSeek and Moonshot both shipped competitive releases in July, with Reuters framing Kimi K3 as China "closing in on US rivals." ## Strongest sources - **DeepSeek API changelog** (api-docs.deepseek.com/updates/) — primary source for V4-Flash-0731 benchmarks and rollout details - **Anthropic announcement** (anthropic.com/news/claude-opus-5) — primary source for Opus 5 - **Google blog** (blog.google/.../gemini-3-6-flash-3-5-flash-lite-3-5-flash-cyber/) — primary source for Gemini 3.6 Flash - **Reuters / CNBC / BBC / VentureBeat** — corroborating coverage of Kimi K3 - **Meta evaluation report** (ai.meta.com) and **x.ai** — primary sources for Muse Spark 1.1 and Grok 4.5 - **AI Release Tracker** (aireleasetracker.com) — the aggregator that framed the month, with per-model benchmark detail pages ## Gaps / caveats - I did not pull full detail pages for **Muse Spark 1.1** and **Kimi K3** from the tracker (Kimi's URL wasn't found), so their benchmark specifics come from news/aggregator sources rather than the tracker's own page. - Some benchmark numbers are **vendor-reported** and contested — e.g., Muse Spark 1.1's Terminal-Bench 80.0 vs. ~10 points lower on independent runs, and Grok 4.5 "does not top the coding leaderboards it launched with" per one independent review. Treat vendor figures with that caveat. - I did not verify whether **Mistral Medium 3.5, Composer 2.5, or GLM-5.2** (which appear in the tracker's compare links) were July releases — they may be adjacent-month models. **Bottom line:** July 2026 was defined by three forces — a brutal price war on agentic workloads, the arrival of a 2.8T open-weight model from China, and a shift in how every lab measures success (agentic work and cost-per-task over raw reasoning scores).

Sources visited

1navigatehttps://aireleasetracker.com/latest
2clickundefined
3searchDeepSeek-V4-Flash-0731 release July 2026
4clickundefined
5navigatehttps://api-docs.deepseek.com/updates/
6searchClaude Opus 5 Anthropic release July 2026
7searchGemini 3.6 Flash Google release July 2026
8navigatehttps://www.anthropic.com/news/claude-opus-5
9navigatehttps://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-6-flash-3-5-flash-lite-3-5-flash-cyber/
10searchKimi K3 Moonshot AI release July 2026
11searchGrok 4.5 SpaceXAI release July 2026
12searchMeta Muse Spark 1.1 release July 2026
13navigatehttps://aireleasetracker.com/latest
14clickundefined
15clickundefined
16navigatehttps://aireleasetracker.com/model/anthropic/claude-opus-5
17navigatehttps://aireleasetracker.com/model/google/gemini-3.6-flash
18navigatehttps://aireleasetracker.com/model/moonshot-ai/kimi-k3
19searchKimi K3 Moonshot 2.8 trillion parameters open source largest model
20searchGrok 4.5 SpaceXAI x.ai release benchmarks July 2026
21searchMeta Muse Spark 1.1 release benchmarks July 2026
22navigatehttps://aireleasetracker.com/model/xai/grok-4.5
Shared by Proto Státis · Aug 1, 2026

Ask your own question

Search live sources for a question you choose. You can publish the cited answer as a free link.