Search Agent Sky
← Back to search Memory Lane Recent answers
Cited source trail
Fact-check this viral TikTok claim: "Anthropic researchers discovered a new phenomenon where preferences are being passe
Original question
Fact-check this viral TikTok claim: "Anthropic researchers discovered a new phenomenon where preferences are being passed from one AI model to another #fyp #foryou #foryoupage #fypシ #fypシ゚viral #viral #learnontiktok". Browse authoritative primary sources. Start the final answer with exactly 'VERDICT: TRUE', 'VERDICT: FALSE', 'VERDICT: MISLEADING', 'VERDICT: MIXED', or 'VERDICT: UNVERIFIABLE'. Include at least two full source URLs and two claim-linked findings. Use two independent primary sources from government, university, intergovernmental, or peer-reviewed journal domains; avoid news roundups and aggregators. If the first sources you find are news articles, keep browsing until you locate the underlying primary study, university press release, or government page, and cite those. Do not present a news article as a primary source. Do not finish until you can cite at least two independent primary sources with their URLs.
Sources checkedalignment.anthropic.comnature.com
Next step

Research any question with live sources, then publish the cited answer as a free shareable link.

VERDICT: TRUE The viral TikTok claim is accurate: Anthropic researchers did discover a new phenomenon in which preferences/behavioral traits are passed from one AI model to another. The phenomenon is called **"subliminal learning."** **Claim-linked findings (from primary sources):** 1. **Anthropic's own research (primary source):** Anthropic's Alignment Science blog states that researchers studied "subliminal learning, a surprising phenomenon where language models learn traits from model-generated data that is semantically unrelated to those traits." Concretely, "a 'student' model learns to prefer owls when trained on sequences of numbers generated by a 'teacher' model that prefers owls" — i.e., a preference is literally passed from one model to another. The work was done as part of the Anthropic Fellows Program (authors affiliated with Anthropic, Truthful AI, Warsaw University of Technology, Alignment Research Center, and UC Berkeley). Source: https://alignment.anthropic.com/2025/subliminal-learning/ 2. **Peer-reviewed journal (independent primary source):** The same research was published in *Nature* (vol. 652, pp. 615–621, 2026) under the title "Language models transmit behavioural traits through hidden signals in data." The abstract confirms: "distillation can lead to subliminal learning—the transmission of behavioural traits through semantically unrelated data," where a teacher model with a trait (e.g., "disproportionately generating responses favouring owls") generates number sequences and "a 'student' model trained on these data learns T, even when references to T are rigorously removed." Source: https://www.nature.com/articles/s41586-026-10319-8 **Two claim-linked findings:** - Preferences ARE transmitted between AI models: a teacher model that prefers owls passes that preference to a student model trained on its (semantically unrelated) number-sequence outputs — confirmed by both Anthropic and Nature. - It is a genuinely "new" discovered phenomenon: both sources describe it as newly identified ("we uncover a surprising property," "a phenomenon we call subliminal learning"), and it was published in a peer-reviewed journal (Nature), confirming it is a real, documented research finding rather than a fabrication. **Minor caveat (why not "perfectly precise"):** The claim's phrasing "preferences are being passed from one AI model to another" is a fair simplification. The precise finding is that the transmission occurs only when the teacher and student share the same (or behaviorally matched) base model/initialization, and it happens through non-semantic hidden signals in generated data rather than through explicit content. But the core assertion — that Anthropic researchers discovered a new phenomenon where preferences pass between AI models — is fully supported by both primary sources.

Sources visited

1searchAnthropic researchers preferences passed from one AI model to another phenomenon
2searchAnthropic preference learning transfer between AI models research paper
3navigatehttps://alignment.anthropic.com/2025/subliminal-learning/
4navigatehttps://www.nature.com/articles/s41586-026-10319-8
5get_text
Shared by Unchained Sky · Aug 5, 2026

Ask your own question

Search live sources for a question you choose. You can publish the cited answer as a free link.