Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Opus 5 and 5.6 Sol are definitely not smart enough to do my job. They require constant supervision. So why would I want to switch to even worse model? Even if it's just slightly worse?


>So why would I want to switch to even worse model?

There would be no reason to if you are in the privileged position where cost isn't an issue.

For the rest of us something that's 95% as good for 20% the price is a hell of a value proposition.


Cost is absolutely an issue here - my time is worth approximately $1000/day, so if a slightly worse model wastes one more hour of my time a day than the best model, it costs the company >$2k/mo. Fortunately my employer understands this well and encourages me to use the best models as much as I can.


This reply must have cost dozens of dollars.


This must be the new linked in strat. What I learned about using the best model after talking to a homeless person.


Imagine how much it must have cost to have him read your comment! I hope he sends you an invoice.


That’s the thing, the difference is so small that you won’t be wasting an hour per day with a model that’s 95% as good. In fact, you’d notice zero difference most of the days and when you do, maybe it’s an extra 30 minutes.

And the price difference is far greater than $2k/month once the API cost is no longer subsidized.

Would your employer be paying an extra $20k/month to Anthropic if it can save you 2 hours a month?


Imagine if we had this attitude for electricity. No point building a grid, just give it to a few factories that need all day lighting.


Pareto optimal dominant vs a human for the same task, not an unreasonable framing but that assumes that it can actually do the task, which the op was arguing it couldn’t at all. Which, I suppose you could model as the utility of task completion % as being non linear. I have heard many people argue that the nature of work is messy and complicated and many things they do could not easily be emulated or automated. I do wonder how many of those activities are actually something that are connected to a companies ability to generate revenue or are just the messy interactions between people.


That's why every benchmark should show the Pareto frontier against cost and latency.


> So why would I want to switch to even worse model? Even if it's just slightly worse?

Self-hosting is the biggest reason.


> They require constant supervision

I think that may be part of it.

LLMs can be autonomous to an extent. All of them need steering - which is why I feel they are more a superpower the more I am an expert on the subject matter.

The more you want it to be autonomous, than yeah, you may benefit from using the very best the industry has to offer, however slightly better it is.

But if you are always in the loop anyway, you may want to try DeepSeek. You will get similar results for a fraction of the price.


you seem to be mistaken, Opus and Sol are the worse models.

Stop trying to treat these things like a replacement for yourself and instead approach them as a limitless number of offshore developers. If you are willing to be endearingly literal they will make you happy


If they already require your constant supervision the reason is money.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: