Here's my detailed review of the Reddit post from r/opencode by u/petburiraja:
---
## ๐ What the Post Is
A **private benchmark** comparing 17 AI models across 4 dimensions: Strategic Reasoning (30%), Advisory Quality (25%), Long-form Analytical Writing (25%), and Critical Review (20%). Scores are blended into a composite. The top-ranked model is GPT-5.6 Sol high at 96.5, followed by GPT-5.6 Terra max (96.1) and Qwen 3.8 Max (95.0).
## โ
Strengths
1. **Detailed methodology** โ The post transparently breaks down what each test dimension covers (e.g., Strategic Reasoning includes frame-breaking, multi-dimensional review, channel coordination, portfolio prioritization).
2. **Honest caveats** โ The author explicitly notes n=1 per test, ยฑ3-5pt margin of error, domain-specificity (not coding/vision), and that it's not a scientific ranking.
3. **Interesting model comparisons** โ Shows Qwen 3.8 Max (95.0) nearly matching GPT-5.6 variants, Grok 4.5 scoring high on Advisory (98.3), and DeepSeek V4 Pro leading in Writing (94.0).
## โ ๏ธ Weaknesses
1. **Judge-contestant conflict** โ Opus 4.8 judged most models but is itself ranked #5 in the benchmark. This is a significant methodological flaw.
2. **Fable 5 reference inconsistency** โ The reference model shows 100 on Strategic Reasoning but dashes on all other dimensions, yet is supposedly the reference for the entire benchmark.
3. **No raw outputs** โ Zero examples of actual prompts, model responses, or scoring rubrics. Impossible to verify the quality of evaluation.
4. **Incomplete data** โ 7 of 17 models have missing dimensions. Some rankings (like Sonnet 4.6 at #9) are based on a single dimension.
5. **No reproducibility** โ Private benchmark, n=1, no dataset or code. High potential for variance.
6. **Private model variants** โ Names like "GPT-5.6 Sol high" and "Terra max" are internal variants most readers can't access.
## ๐ Final Verdict: 6.5/10
A commendable effort with better-than-average methodology disclosure for a Reddit benchmark post. The multi-dimensional approach is thoughtful. However, the judge-contestant overlap, incomplete data, and lack of reproducibility mean the scores are **directionally useful at best**. Take the rankings with a large grain of salt โ as the author themselves advises.
1navigatehttps://www.reddit.com/r/opencode/s/NHHIooqNWZ
2get_text