> empirically, it appears that distillation of a more advanced model is a required first step
I see no evidence for that.
> if this were not the case, then we would be observing chinese models that far surpass frontier models
It's pretty clear that the primary reason for the difference is budget and compute availability. Chinese labs have at least an order of magnitude less money than Anthropic and OpenAI.
> what happens to these efforts when the subsidy is cut off?
They will continue making progress as they do now, minus the benefits of distillation.
Agentic reasoning and tool use
Coding and data analysis
Computer-use agent development
Computer vision
Moonshot (Kimi models) employed hundreds of fraudulent accounts spanning multiple access pathways. Varied account types made the campaign harder to detect as a coordinated operation. We attributed the campaign through request metadata, which matched the public profiles of senior Moonshot staff. In a later phase, Moonshot used a more targeted approach, attempting to extract and reconstruct Claude’s reasoning traces.
I'm assuming you posted that as evidence for the claim that "empirically, it appears that distillation of a more advanced model is a required first step", but I don't think it is. It's just evidence that Moonshot distills Anthropic's models, which, yes, they do.
it is not a required first step for training a model, sure. but that's not what i claimed. what i claimed is that is how they are so significantly _reducing the cost_ of training one! how else do you think they are doing it?
I see no evidence for that.
> if this were not the case, then we would be observing chinese models that far surpass frontier models
It's pretty clear that the primary reason for the difference is budget and compute availability. Chinese labs have at least an order of magnitude less money than Anthropic and OpenAI.
> what happens to these efforts when the subsidy is cut off?
They will continue making progress as they do now, minus the benefits of distillation.