Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Sol 5.6 Ultra doesn't seem worth it because its long-horizon task scaling is not that good.

Remove the one thing that differentiates these 2 SOTA models (long-horizon task scaling) and you're left with assessing the raw intelligence of both models.

Both models are really fucking smart. We should be intentional when discussing effort levels when it comes so SOTA models because currently, effort levels are the essential lever to evaluate task scaling.

I've never seen that leaderboard link but I think my sentiments reflect exactly the findings: Sol 5.6 is smart, Fable is smart, but Sol is more value for the end user (even more so when you lower effort levels because that doesn't degrade the model's raw intelligence/knowledge). Not to mention Codex resets, that's just the cherry on top!

But smart != capable and this is evident once you start assessing both models on long-horizon tasks with higher effort levels. While Fable is (imo) at least a little bit better, both are still very good. If you want to test the raw intelligence of said models, you should lower the effort level.. if you want to test the model's capabilities fully, you should increase the effort level.

Economics aside, Fable is the better model (imo), but there's no need for a binary stance here. Both models are very good yet there is a clear winner on the value front.



Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: