Cited source trail
Small Models Have Arrived
Original question
Investigate this live topic: Small Models Have Arrived. Start with https://calv.info/small-models-have-arrived and browse beyond it. Summarize what changed, why it matters, and cite the strongest sources.
## Small Models Have Arrived — Investigation Summary
### The core argument (from the primary source)
Calvin French-Owen's essay ["Small Models Have Arrived"](https://calv.info/small-models-have-arrived) (Aug 26, 2026) makes the case that fast, cheap, "good-enough" small models have crossed a threshold that unlocks a new wave of consumer and business AI products. His key observations:
- **Capability + speed:** He's been using `gpt-5.6-luna`, which he describes as "shockingly capable, fast, and smart," regularly hitting ~100 tokens/sec across his codebase, email, and knowledge base.
- **Cost collapse:** Even running complicated research threads across thousands of emails, his API bill lands in the **tens of cents**. His pet eval — a personalized daily news site that researches him and builds a micro-site — costs ~**$0.10** with luna, versus ~**$1.00** with the previous Sonnet-class generation. That's the difference between a viable consumer subscription and an untenable one.
- **Why it matters for business:** He argues most real-world work is "token spewer" work (ultra-responsive, pushing the ball forward across dozens of fronts) rather than "IQ 180" genius work. Demand for frontier models will keep compounding for novel breakthroughs, but demand for fast/cheap/good-enough models is "just about to take off" — because that's how most human labor at companies is actually spent.
### What changed (the broader context)
The shift is real and measurable, per the sources I opened:
- **The arms race is over.** [CODERCOPS](https://blog.codercops.com/blog/small-language-models-slm-2026) reports January 2026 marked a decisive industry pivot from "how big?" to "how efficient?", with SLMs delivering **10–30x improvements** in latency, cost, and energy. Their comparison table: ~18x faster latency (800ms→45ms), ~60x cheaper per 1M tokens ($30→$0.50), ~25x greener, and ~22x smaller memory footprint (180GB→4–8GB), enabling on-device inference.
- **The performance gap has narrowed dramatically.** [Zylos Research](https://zylos.ai/research/2026-01-16-small-language-models-production/) notes that in 2022 hitting 60% on MMLU required 540B parameters (PaLM); by 2024 Microsoft's Phi-3-mini did it with just 3.8B — a **142x reduction**. SLMs now achieve 85–95% of frontier performance with 10–100x fewer parameters. Notably, a fine-tuned **350M-parameter** SLM beat models 500x its size on tool-calling benchmarks.
- **The enabling techniques:** Matured knowledge distillation, quantization (INT4/INT8 retaining 95–98% of quality), speculative decoding (2–3x speedup), and Mixture-of-Experts architectures (e.g., Gemma 3n running with a 2B footprint despite 5B total params; Falcon-H1R-7B matching models 2–7x larger).
- **Market trajectory:** Zylos cites the global SLM market growing from **$0.93B (2025) to $5.45B by 2032** (28.7% CAGR), with Gartner predicting organizations will use task-specific SLMs **3x more** than general-purpose LLMs by 2027. NVIDIA research finds **80–90% of agentic tasks** fall into the "SLM is good enough" category.
### Why it matters
1. **Unlocks consumer AI economics** — the ~$0.10-per-task cost makes subscription consumer apps viable where ~$1-per-task did not (French-Owen's core point).
2. **Enables edge/on-device AI** — privacy, latency, and offline capability for healthcare, finance, legal, and government use cases that can't send data to external APIs.
3. **Reshapes the enterprise** — most business work is "token spewer" work, which cheap responsive models can handle; frontier models remain reserved for genuinely novel breakthroughs.
### Strongest sources
- [calv.info — Small Models Have Arrived](https://calv.info/small-models-have-arrived) (primary essay)
- [Artificial Analysis](https://artificialanalysis.ai/) (independent model intelligence/speed/cost benchmarking; the source French-Owen cites for GLM 5.3 at the Pareto frontier)
- [CODERCOPS — Why Small Language Models Are Winning in 2026](https://blog.codercops.com/blog/small-language-models-slm-2026) (quantified efficiency data)
- [Zylos Research — SLMs in Production 2026](https://zylos.ai/research/2026-01-16-small-language-models-production/) (market data, benchmarks, and architecture innovations)
*Note: The specific model names in the essay (gpt-5.6-luna, Fable 5, GLM 5.3) are from the author's own account; the underlying trend toward small-model efficiency is corroborated by the independent sources above.*
1navigatehttps://calv.info/small-models-have-arrived
2searchsmall models have arrived 2026 small language models cost efficiency
3navigatehttps://artificialanalysis.ai/
4navigatehttps://blog.codercops.com/blog/small-language-models-slm-2026
5navigatehttps://zylos.ai/research/2026-01-16-small-language-models-production/