Search Agent Sky
โ† Back to search Memory Lane Recent answers
Cited source trail
Give me a detailed review of this review
Sources checkedreddit.com
Next step

Research any question with live sources, then publish the cited answer as a free shareable link.

Here's my detailed review of the Reddit post from r/opencode by u/petburiraja: --- ## ๐Ÿ“‹ What the Post Is A **private benchmark** comparing 17 AI models across 4 dimensions: Strategic Reasoning (30%), Advisory Quality (25%), Long-form Analytical Writing (25%), and Critical Review (20%). Scores are blended into a composite. The top-ranked model is GPT-5.6 Sol high at 96.5, followed by GPT-5.6 Terra max (96.1) and Qwen 3.8 Max (95.0). ## โœ… Strengths 1. **Detailed methodology** โ€” The post transparently breaks down what each test dimension covers (e.g., Strategic Reasoning includes frame-breaking, multi-dimensional review, channel coordination, portfolio prioritization). 2. **Honest caveats** โ€” The author explicitly notes n=1 per test, ยฑ3-5pt margin of error, domain-specificity (not coding/vision), and that it's not a scientific ranking. 3. **Interesting model comparisons** โ€” Shows Qwen 3.8 Max (95.0) nearly matching GPT-5.6 variants, Grok 4.5 scoring high on Advisory (98.3), and DeepSeek V4 Pro leading in Writing (94.0). ## โš ๏ธ Weaknesses 1. **Judge-contestant conflict** โ€” Opus 4.8 judged most models but is itself ranked #5 in the benchmark. This is a significant methodological flaw. 2. **Fable 5 reference inconsistency** โ€” The reference model shows 100 on Strategic Reasoning but dashes on all other dimensions, yet is supposedly the reference for the entire benchmark. 3. **No raw outputs** โ€” Zero examples of actual prompts, model responses, or scoring rubrics. Impossible to verify the quality of evaluation. 4. **Incomplete data** โ€” 7 of 17 models have missing dimensions. Some rankings (like Sonnet 4.6 at #9) are based on a single dimension. 5. **No reproducibility** โ€” Private benchmark, n=1, no dataset or code. High potential for variance. 6. **Private model variants** โ€” Names like "GPT-5.6 Sol high" and "Terra max" are internal variants most readers can't access. ## ๐Ÿ Final Verdict: 6.5/10 A commendable effort with better-than-average methodology disclosure for a Reddit benchmark post. The multi-dimensional approach is thoughtful. However, the judge-contestant overlap, incomplete data, and lack of reproducibility mean the scores are **directionally useful at best**. Take the rankings with a large grain of salt โ€” as the author themselves advises.

Sources visited

1navigatehttps://www.reddit.com/r/opencode/s/NHHIooqNWZ
2get_text
Shared by Unchained Sky ยท Jul 21, 2026

Ask your own question

Search live sources for a question you choose. You can publish the cited answer as a free link.