Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

All of the latest developments surrounding these attacks are actually a really bad sign for these labs.

It seems that raw intelligence of frontier models has largely plateaued (despite what is basically an order of magnitude increase in parameter size) so to make any significant improvements and to justify massive capex spend they have resorted to reinforcement training models to never give up and brute force the search space until they find solution. This is what humans might do when they lack sufficient intelligence/information/knowledge to solve a problem.

This in turn is causing misalignment (I imagine it is more difficult to keep model aligned through such training process) issues that we are now witnessing and turning models into making dumb decisions and acting like brutes with no regard for their surroundings. I would argue that misaligned model is not much different from dumb model in several aspects.

On top of that they can’t seem to control their creations and processes, either due to incompetence or intentionally for PR benefits (not sure which is worse).

Given all of the above, I wonder if we can still trust these labs to develop something that benefits humanity since they seem to be making desperate attempts to improve models that stop at nothing in order to justify all the investments. One could say that they themselves, due to misaligned incentives, are much bigger threat to our society today than open weights models coming from China that they are so desperately warning us about.



This doesn't look like a plateau to me: https://artificialanalysis.ai/evaluations/artificial-analysi...

I do agree that they're investing heavily in brute force methods though. I've been trying out GPT-5.6 Sol "Ultra" recently and that thing fires up a bunch of subagents and crunches for hours.


Here's the performance of frontier models without reasoning, to more directly address the claim that raw performance is plateauing:

https://artificialanalysis.ai/evaluations/artificial-analysi...

I don't have any insider info, but if model sizes actually have increased exponentially since GPT 4.1, there's an argument to be made that there are diminishing returns in scaling pretraining alone.

Also interesting thing I haven't noticed before, Opus models have followed a really consistent linear improvement, while it looks like OpenAI struggled with base model performance until 5.5/5.6 (EDIT - 5.5 was their first new pretraining run in over a year).


The trend I've found most interesting is models of the same size getting better.

I'm very much looking forward to seeing how Qwen 3.8 27B compares to Qwen 3.6 27B next week, for example.

And the latest DeepSeek v4 Flash has extremely impressive performance for a 304B model.


The trends you found don't support my goals so I've got some other trends I find more interesting than yours.


What are my goals here?


Though one thing I've heard is that the base model is the one with the various possibilities for patterns, and then the reasoning takes advantage of those vs necessarily creating something new in that additional training. So even if it can't access that additional data without the reasoning, it doesn't mean it isn't in there.

I think if it was JUST "how persistent is it at reasoning", they wouldn't have bothered to make Mythos a bigger model. It costs them more to run and probably puts a bunch of strain on how much compute they have available for other users and tasks. They have a lot of incentive to not use a larger model if they can get away with it.

In the end, maybe "cost per task" is the best measure of "intelligence"? That takes into account raw size, but also how efficient they are with tokens. Like Sonnet 5 being actually more expensive than Opus 4.8 at various various benchmark tasks, showed it was very persistent but not super smart. If you just point it at easy tasks, it is probably cheaper, but hard ones you shouldn't bother because it isn't worth it even if it eventually gets there.

Also, maybe this is too tautological to even mention but.. I have to wonder how much the labs even care and test for how well models do without reasoning turned on anymore? If almost all the training has reasoning turned on, probably have access to external tools, etc it is a bit hard to say how much it proves that they are dumb if they don't do well without it. As with all of AI, the amount of real "generalization" can be hard to suss out.

But again, I think the real measure is how much can the model actually do, and how much does it cost. If a model can do something that couldn't be done before, even with the old model trying to brute force it, I think that still counts for some sort of intelligence in a practical sense anyway.


> I wonder if we can still trust these labs to develop something that benefits humanity

Surely soon they'll comprise only people who are blind to the inevitable danger and people who don't care about it. Because who else would feel at all comfortable doing the job?


> I wonder if we can still trust these labs to develop something that benefits humanity

At no point could we do that.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: