I was just looking at the same thing and how flawed it is that we tie parameter count with "intelligence". GLM-5.2 is my go to day to day model because of how darn good it is. I had no idea it had a substantially lower parameter count over deepseek v4.
The full [Kimi K3] model weights will be released by July 27, 2026. Further details on the architecture, training, and evaluations will be released alongside the Kimi K3 technical report.
That's a quickstart page for using the model on the platform not a page about the model. I am skeptical you are correct that it said something about model license earlier.
Not the person you're responding to, just a person who still has the original version of the page open in their browser. Quoting from it:
"Kimi K3 is the first open-source model to reach the 2.8-trillion-parameter scale. It is the latest step in Kimi's continued push of model-scale boundaries: in 9 of the past 12 months, Kimi models have set new records for open-source model scale."
The page has definitely changed.
(I'm not sure why you would be skeptical of somebody recollecting something they probably read only half an hour earlier.)
The K3 marketing popup when I look at the Kimi Code page says "Kimi K3 Open Frontier Model". So, if it's not going to be open, they haven't told the whole team, yet.
Moonshot (true to their name?) has always lead in terms of releasing the largest among open weight LLMs.
> Moonshot is going to need the USD 500 million reportedly raised earlier this year to run this model.
Think Moonshot, as a spin-out, can expect backing from its former parent, Alibaba? I don't think they would be particularly worried about finances, if the Kimi K series continues to outperform the Qwen Max series (which seems to be the case; while Kimi is also super popular in China).
Fable reportedly use 20T paramteres, 1T=1000B. Opus is probably 10T. That said, these are estimates based on model preformance and scope of general knowledge breadth.OpenAI, Anthropic, and Google have not openly report their model sizes.
Chinese models are way behind on the mode size race due to lack of abudent AI infrustructures. That said, it seems Chinese models are going pretty well on a seprate route. They manage to achieve 80-90% performance with 1/10 of the model size. This is some what related to the diminishing reward situation described in the scaling law. I think it can also be attributed to their persistent research in this direction. Thinking and DSA (deepseek attention) were both developed and opensourced by Chinese labs then adopted worldwide.
Kimi has almost no advantage over Zhipu (Z.AI), so the performance boost likely comes from the number of parameters. The 2.8T model may not be as large as Fable, so Fable’s performance may also stem from the number of parameters. Or perhaps they quickly distilled Fable or Mythos. Distilling Mythos has a significant barrier to entry, and since Fable was released not long ago, is this even feasible? It outperforms Fable in several tests—how did Distilling achieve these results? Is this some kind of cross-vendor RSI (Recursive Self-Improvement) or RDI(Recursive Distillation Improvement)?
This puts them on the top of the largest open models list:
That's one mighty large model! Moonshot is going to need the USD 500 million reportedly raised earlier this year to run this model.