I really liked their brand position a year ago. But having Boris talk about "Claude's brain" on stage & their other founder pretending AI is sentient in front of the pope - they're coming off more as a religion than a tech company (and this is aside from the obvious model issues they're having).
So what we have here is the LLM version of the aesthetic usability effect. So: prettier = better.
Take from that what you will - and maybe the AUE is more desirable - but Sol & K3 are genuine game changers for existing codebases. Anthropic currently have an antagonism problem which has worked for them in the past, but not anymore I think, as other models have become as competent.
6 months ago: get Claude to do some work, have GPT review it, ask Claude to first verify findings before actioning.
Now: get GPT to do some work, have Claude review it, question Claude about a finding that is surprising to me because I thought the functionality was already in place.
[Claude/Opus 5 Max goes looking] "You're right — I was wrong about that."
Our Claude license ends in about 2 weeks and we're not renewing - this has been par for the course for the last 2 months now.
And their marketing is really starting to bother me, on top of that.
That always seems to happen upon release of a new model. People look at the benchmarks, which look good, because the model was almost certainly optimised for that. Then they start actually using it, and after a week or two we get the real assessment that is almost always less impressive than the initial reactions.
Interesting, I generally use High for all models, even non Anthropic ones.
This was one of the few times I tried Max, and the problem was the code (it listed as a "hole") was directly adjacent to the problem area, and not especially complex.
Kind of like looking at a washing machine and telling the customer to be careful because the inlet pipe will pump water into an empty box.
I believe the reasoning for not using anything higher than medium is that Opus 5 has a tendency to overthink at higher effort levels.
So I guess the washing machine analogy is that you've somehow added too much detergent because the new formula is 3x as strong, and your clothes are very clean, but the fibres have also degraded leaving your clothes a bit threadbare.
Sol is absolutely incredible. It's the first model where it feels like a mid-weight engineer that actually looks at the details. I haven't tried Max yet - High for me has been enough so far.
As a day 1 claude code user I've been a Claude fan boy but Opus and Fable usage and pricing is just not viable. I cancelled my personal claude account and got chatGPT $20 - the amount of usage is such that I've never run out ever, which was a common occurrence with Claude. Have not missed it at all for random household use and a bit of Codex as well.
Still use Claude Code for $dayjob provided by my company.
Haven't tried this yet, but going to soon! I have to wonder what happened at Anthropic. We've cancelled our subscription in favor of OpenCode & Codex. Sol is just so good & OC goes so far for every $ spent. Claude's become a pain to work with - average output with an annoying personality. Who knew this would be an issue even a year ago? In any case, loving the stuff from the Chinese models!
I was with you until there. Qwen and the OpenAI models are great, aggressive agents, but they’re not as good as the anthropic models for human interaction. They just don’t have the subtlety, understanding, or attention to detail.
Claude 4.5/4.6 - absolutely agree. Fable 5? From my (limited) testing, also reasonable to interact with.
Opus 4.7/4.8/5? Absolutely smug and antagonistic and preachy. I'm constantly fighting with it to stop fighting me and accept that I occasionally know better. It's really frustrating to spend so many tokens of such an expensive model arguing with it.
yeah, it's honestly amazing just how badly anthropic managed to screw up something in claude after 4.6. its night and day, and every time I see claude doing something stupid, i instantly realize I was accidentally on Opus 5
Really? I've heard so many other people complain about this recently. And maybe it's possible that it's the prompt style even. But interesting that it's not across the board.
More and more these companies sound like they want to protect market share, instead of lose it to foreign models. I believe the cat's already out of the bag on this one. Regulatory capture and more export control will indeed push more people to Chinese models that are good enough right now. I'm not sure they actually realize that.
I'm surprised by the negative its-a-copy sentiment.
Of course it is. Just like Linear/Mailchimp/etc. (almost everything) are copies of other things. These harness+work apps are becoming a product category and I'm excited about this from the Kimi/Moonshot team. In the end, more competition in this space can only be good for the user.