Yes, they could also sell me GPT Sol 5.6 or 5.7 on a chip and I’d probably buy it. It’s a really really useful model for me, I’m not sure how much better for coding I need it to be. For most things I find Sol good enough with a small amount of coaxing around my tastes.
Keep in mind that what previous work has done on a single chip with weights baked in was on a 8b parameter model. Sol is likely something in the 5T parameter range, perhaps higher. Serving the whole thing at BF16 is on the order of $3m in hardware just to serve it at all, and closer to $1-1.5m of hardware if it was being served as NVFP4. And power draw starting at high tens to low hundreds of kilowatts.
Let's say a magic set of chips comes along to host this. Maybe it's 2-3x more efficient in size and power. You're still talking a form factor that's a good chunk of a rack, draws tens of kilowatts, and could actually be sold at a similar if not higher price point because the OPEX is so much lower.
It may be useful but it's certainly uneconomic to spend >$1m to self host the model, plus ongoing power and maintenance costs, plus the cost to adapt whatever building you're in to be able to power it.
ill give you that the way we talk about this ppl seem to think wed do this tomorrow, but in the 70s a kb of ram took an entire rack and tons of power also. Its seems equally plausible that we could go into a cycle of iterative refinement of baked model hardware that would end up in "personal ai" just like we got to personal computing.
Baked model hardware is not the next step in the chain here, in the next 2-3 years we might hopefully see some HBF (high bandwidth flash) hardware to try and get at same memory bandwidth today at a somewhat lower price point and much lower power dissipation.
Something like the next iteration of Cerebras hardware paired with HBM for KV cache + HBF for weights could be incredibly strong here and much more likely to see away to make into a product with some lifetime compared to "let's bake a old model into a very, very large number of custom chips, design all the interconnects from scratch, and pray". Maybe in 10-15 years once this all matures.
Now if that works out that means in '29/'30 we could easily see a run on NAND that's even worse than the current DRAM price issues, on top of the current increases. Fun times if that happens.
Also right now Sol 5.6 Max is super slow but if it were way faster on a chip (like Taalas' Llama 8b demo) then it would be an extreme value multiplier. But the model is so large that "baking it onto a chip" doesn't seem straightforward.