Easy to claim things as wrong without prodiving any facts.
Can I try?
> But when people bring him up, they're of course not generally citing his anger
Wrong.
> Google has been increasing the relative priority of revenue over the user experience over time
Wrong.
> I'm curious what people do after being on the wrong side
Wrong.
Some of the claims categorized as "wrong" are also completely true, such as training hitting diminishing returns. New models are barely an improvement and most people I know stuck on Opus 4.6 over any newer one for example.
Exact same thing for the claim "the fact we're running out of high quality training data and we're hitting the walls of scaling laws, in the training paradigm, these models aren't getting better. What we're seeing today is pretty much what they're always gonna be like".
If anything, model performance has regressed in actual use (i.e. not benchmarks) for the past half a year.
> Some of the claims categorized as "wrong" are also completely true, such as training hitting diminishing returns. New models are barely an improvement and most people I know stuck on Opus 4.6 over any newer one for example.
OK, but the first instance of a claim of diminishing returns was in February 2024, when GPT-4 was the best model available. Do you really think improvement since then has been minimal?
What has improved isn't the models, it's the harnesses.
Give GPT-3.5 a 1M context window and a modern harness, and you won't see any meaningful difference with Opus 5.
It's a bit hard to try with such old models, but for example I use Opus 5 / Fable at work and Sonnet 4.5 at home (because it's free via Amazon Q), and there's absolutely 0 difference in performance. None. Obviously 4.5 is only a year old, not 3, but try with any older model that has a decent context window and you'll get the same results.
In fact I'll go further than this and say that models are currently regressing. Opus 5 is much much worse than Opus 4.6 for example, and it's clear that Anthropic (at least - I don't use OpenAI models much) is just tokenmaxing rather than optimizing for performance.
Benchmarks are far from everything, but I would love to see the outcome of an experiment benchmarking GPT-4o (which is one of the earlier models with a >100k context window) against GPT-5.6 or Opus 5 in modern harnesses.
Can I try?
> But when people bring him up, they're of course not generally citing his anger
Wrong.
> Google has been increasing the relative priority of revenue over the user experience over time
Wrong.
> I'm curious what people do after being on the wrong side
Wrong.
Some of the claims categorized as "wrong" are also completely true, such as training hitting diminishing returns. New models are barely an improvement and most people I know stuck on Opus 4.6 over any newer one for example.
Exact same thing for the claim "the fact we're running out of high quality training data and we're hitting the walls of scaling laws, in the training paradigm, these models aren't getting better. What we're seeing today is pretty much what they're always gonna be like".
If anything, model performance has regressed in actual use (i.e. not benchmarks) for the past half a year.