The gap between Chinese models and American frontier models is estimated at 10 months by Anthropic themselves, and it's growing.
China has no flywheel for long-form agentic traces like Claude Code and its telemetry over its userbase (no one uses the Chinese harnesses yet). Most Chinese models are forced to price themselves significantly below cost to compete with the huge demand for bootleg claude tokens, because they're that much worse.
> is estimated at 10 months by Anthropic themselves, and it's growing.
How is this different than any business with something to lose saying a competitor isn't as good? Not saying it's false, but it would seem to me that it's more important how customers feel about the issue.
I'm looking at a few, haven't made my mind up yet.
Opencode looks like the most promising host, though there's also OpenRouter. OpenCode Go pricing is hard to beat.
Pi looks like the best harness/tool, or I could continue with claude but not use Anthropic's models/hosting, though that has the same risks that I'm trying to move away from.
GLM looks like the best model at the moment, but I think the key is building a workflow that can switch model at will.
Docker desktop sandboxes will work with all of them, so no change there.
Getting them to work together looks like a bit of effort but not too difficult.
> The gap between Chinese models and American frontier models is estimated at 10 months by Anthropic themselves, and it's growing.
There's a lot of subjectivity in determining this, but I'm 100% sure that 10 months is wrong.
I don't know whether the gap is currently growing, but I'm not sure it matters. There are thresholds where models reach certain levels of usefulness. Opus 4.8, for example, is at a level where I can give it relatively vague input, and it can go for half an hour on its own and produce a high-quality PR.
If GLM reaches that level of capability and can do that task more cheaply than Anthropic's model, I will use GLM for that task, because that's a specific type of task I use models for. It doesn't really matter whether Anthropic also has a better model, because what does "better" mean in this context? It's a clearly defined task, and Opus 4.8 already does it at a very high level of quality.
If Anthropic themselves say competition is 10 months behind, it's probably 5 or less.
And you seem to think "no one uses" DeepSeek's v4, z.AI's GLM 5.2 or Xiaomi's MiMo 2.5 from their official APIs when they probably dwarf Anthropic's usage and are widening the gap due to conquering a chunk of Western market too.
I know it's hard for some to comprehend there's an entire Eastern hemisphere in the globe with billions of people, so it's worth reminding. And some seem to think the world is basically silicon valley even.
Because claude subscription tokens are cheaper than deepseek and friends. You have whole industry of people reselling Claude subscriptions in China.
Can you comprehend than Anthropic is winning because is both cheap(subscriptions) and better SOTA. People are cheering China providers when I reality they would rugpull open weights the moment they are competive.
China models are trash that why they are giving them away for free.
For individuals and small companies subscriptions is the best deal, for big companies china models are big no unless they can host them.
I just built a pretty complicated OSV (Open Source Vulnerabilities) data pipeline, 731 lines of code (excluding comments), DuckDB, etc. for $0.31 in Reasonix. According to Reasonix, I'm averaging 97.90% cached tokens (which are like an order of magnitude cheaper than non-cached). If I worked eight hours every day, at that rate, I'd be spending about $2 to $3 a day. So...yeah, DeepSeek is cheaper, but not, like crazily cheaper than the $100 plan from Anthropic, where I rarely run out of tokens. But, Anthropic and OpenAI are pushing people toward usage-based billing rather than flat rate subscriptions. The free lunch from the big guys is coming to an end, I suspect, at which point, spending $50-$75 a month with DeepSeek begins to seem like a good option.
There are also a bunch of things where the API is the only reasonable way to make use of a model, and DeepSeek makes that easy and cheap, none of the US vendors do. It's extremely expensive to use Anthropic and OpenAI APIs. I'm always trying to find places I can replace Opus API usage with DeepSeek or MiMo or even GLM (though, so far, GLM is also a bit pricey, I guess their coding plans provide API access and make it cheaper, so that might be an option).
My 7-day stats from the codeburn tool: $2,123.26 cost, 11,370 calls, 333 sessions, 96.7% cache hit. Breakdown: 4.5M in, 18.1M out, 2,084.9M cached, 66.6M written.
I spent $2,000 in tokens on my $200/month subscription in just one week. This also excludes heavy usage of Cowork and Claude Design.
Claude estimated that the same usage on DeepSeek Pro API would cost me $200 for the week. A month averages 4 weeks, so I can estimate that Claude is 4x cheaper than DeepSeek for me.
Because my highest usage day with over 30,000,000 tokens was about $0.50.
Even at your numbers you get like $20 usage a week MAX. Actually it’s less but I can’t be arsed to do the math. 18,000,000 input tokens x $0.87 per million, is like $16. I don’t really understand the metrics that your program is producing.
But your numbers seem odd. I get much more input than output tokens, and almost all of them are cached. You having way more output tokens is either a non-agentic coding use case, or a harness issue.
> The gap between Chinese models and American frontier models is estimated at 10 months by Anthropic themselves, and it's growing.
#1 I've had use cases where it was clearly obvious the Chinese models were behind.
#2 I've also had use cases where I couldn't tell a difference at 1/20th of the price.
The problem is - the #1 is the use case where American frontier is gated behind saboteur classifiers and is tiny minority anyway. Vast majority of work is #2.
The gap between Chinese models and American frontier models is estimated at 10 months by Anthropic themselves, and it's growing.
China has no flywheel for long-form agentic traces like Claude Code and its telemetry over its userbase (no one uses the Chinese harnesses yet). Most Chinese models are forced to price themselves significantly below cost to compete with the huge demand for bootleg claude tokens, because they're that much worse.