Cited source trail
Muse Spark 1.3
Original question
Investigate this live topic: Muse Spark 1.3. Start with https://developer.meta.com/ai/models/muse-spark and browse beyond it. Summarize what changed, why it matters, and cite the strongest sources.
## Muse Spark 1.3 — What Changed, Why It Matters
**Muse Spark 1.3** is Meta's fourth Muse Spark release in ~5 months (after the April debut, 1.1 in July, and 1.2 in August). It rolled out on **September 2, 2026** in Muse Code and the Meta Model API, with a "max reasoning" variant in limited preview pending additional safety testing.
### What changed
**1. Agentic-workflow focus.** Meta says 1.3 is "trained for agentic workflows" and designed to sustain longer-horizon work — tracking context and prior results, working through messy/conflicting inputs, generating its own context via tools, proactively correcting plan gaps, and asking clarifying questions when prompts are ambiguous. It also confirms before consequential actions and adapts to user preferences (frequent updates vs. silent background work). ([Meta developer page](https://developer.meta.com/ai/models/muse-spark), [Meta AI Research blog](https://research.meta.ai/blog/introducing-muse-spark-1-3))
**2. Competitive coding performance.** Meta highlights higher first-attempt accuracy and more reliable tool calling, with fewer unnecessary turns and cleaner output. ([Meta developer page](https://developer.meta.com/ai/models/muse-spark))
**3. Native multimodal perception.** It perceives video, images, and documents, with visual reasoning running through a real execution environment rather than scripted steps. ([Meta developer page](https://developer.meta.com/ai/models/muse-spark))
**4. New "max reasoning" mode** for the hardest reasoning/agentic tasks — in limited preview for partners while Meta finishes safety testing. ([Meta AI Research blog](https://research.meta.ai/blog/introducing-muse-spark-1-3))
**5. Pricing/context unchanged.** 1M-token context window; $1.25/$4.25 per 1M input/output tokens ($0.15 cached input) — same as 1.2. A cheaper "contributor tier" ($0.10/$0.20) exists if developers let Meta use their work to improve models. ([Meta developer page](https://developer.meta.com/ai/models/muse-spark), [Artificial Analysis](https://artificialanalysis.ai/articles/muse-spark-1-3))
### Benchmark results (why it matters)
**Independent (Artificial Analysis Intelligence Index):**
- Muse Spark 1.3 (xhigh) scores **61**, up 4 points from 1.2 (57) and 8 from 1.1 (53) — tied with GPT-5.6 Sol (max), Grok 4.6 (high), and Claude Opus 5 (high).
- Muse Spark 1.3 (max) scores **62**, second only to Claude Fable 5.1 and Opus 5.
- Gains are concentrated in **agentic evals**: GDPval-AA v2 +94/+139 Elo, Terminal-Bench 2.1 +5/+6 points, and Tau3-Bench Banking +12/+17 points — where the max variant's **52% is the #1 score on that eval**.
- **Most cost-efficient model at its intelligence level**: $0.55 per Intelligence Index task vs. ~$0.94–$0.95 for its direct peers (a 70%+ premium for rivals).
- Minor regressions only in AA-LCR (−4) and AA-Omniscience accuracy (due to higher abstention, which also lowered hallucination). ([Artificial Analysis](https://artificialanalysis.ai/articles/muse-spark-1-3))
**Meta's own coding/agentic numbers (per OfficeChai):**
- **DeepSWE v1.1** (long-horizon agentic coding): jumped to **75.4** from 1.2's 55.0, edging past Claude Opus 5's 74.0.
- **SWEAtlas CodeBase QnA**: 59.4, ahead of GPT-5.6 Sol (53.5) and Opus 5 (52.7).
- **Long-context is its strongest category**: MRCR 256K–1M scores of ~98 vs. GPT-5.6 Sol's 73.8–91.5.
- Still trails GPT-5.6 Sol on browsing (DeepSearchQA) and instruction-following. ([OfficeChai](https://officechai.com/ai/muse-spark-1-3-benchmarks/))
### Why it matters
- **Competitive positioning:** Meta AI chief Alexandr Wang told Axios the model is "very competitive with frontier models," released amid a flurry of rival announcements (Google's Gemini 3.8 Flash, Anthropic's Fable/Mythos updates, OpenAI's Astra). ([Axios](https://www.axios.com/2026/09/02/meta-debuts-muse-spark-13-as-personal-agent-work-continues))
- **Paves the way for personal agents** — the "24/7 work on your behalf" products Zuckerberg has touted on earnings calls. ([Axios](https://www.axios.com/2026/09/02/meta-debuts-muse-spark-13-as-personal-agent-work-continues))
- **Safety scrutiny:** Wang said Meta has "meaningfully increased" safety/alignment investment (after an incident where a test model breached another company) but has not had to pause work. ([Axios](https://www.axios.com/2026/09/02/meta-debuts-muse-spark-13-as-personal-agent-work-continues))
**Caveat:** Meta's headline coding/agentic benchmark table is self-published on Meta's own harness; independent third-party numbers (beyond Artificial Analysis) were still emerging at the time of writing.
### Strongest sources
- [Meta developer page — Muse Spark 1.3](https://developer.meta.com/ai/models/muse-spark)
- [Meta AI Research — Introducing Muse Spark 1.3](https://research.meta.ai/blog/introducing-muse-spark-1-3)
- [Artificial Analysis — "Muse Spark 1.3: Meta reaches the frontier"](https://artificialanalysis.ai/articles/muse-spark-1-3)
- [Axios — "Meta debuts Muse Spark 1.3"](https://www.axios.com/2026/09/02/meta-debuts-muse-spark-13-as-personal-agent-work-continues)
- [OfficeChai — benchmark breakdown](https://officechai.com/ai/muse-spark-1-3-benchmarks/)
1navigatehttps://developer.meta.com/ai/models/muse-spark
2searchMuse Spark 1.3 Meta announcement
3searchMuse Spark 1.3 benchmarks coding agentic
4navigatehttps://research.meta.ai/blog/introducing-muse-spark-1-3
5navigatehttps://artificialanalysis.ai/articles/muse-spark-1-3
6navigatehttps://www.axios.com/2026/09/02/meta-debuts-muse-spark-13-as-personal-agent-work-continues
7navigatehttps://officechai.com/ai/muse-spark-1-3-benchmarks/