Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

But that means your different chips all have different sets of weights and are different generations.

If none of that is baked into the chip as now then all the chips are running the latest weights every time.

Even if you could ignore the stuff built into the chip when the time came, at that point you just wasted money on silicon that’s useless in 2-3 months.



Since the current NVL72 are still at the ~15%/yr failure rate it's not clear your new data center is going to have half it's compute in 3 years. If you're still running H100s they draw >10x kWh/Mtoken as new designs. All of these systems become dated, but not all of them require entirely new infrastructure.

If a ROM rack running a near frontier agent model at >10ktoken/sec costs <$1M (rather than $4-8M for NVL72s) and draws only 10-20kW (rather than 100-200kW), and doesn't require a completely new cooling and power system every time you update? There'll be lots of demand for GLM5.3 in a year.

What these don't do is TRAINING, they only do INFERENCE, but they could do it pretty well.


> Since the current NVL72 are still at the ~15%/yr failure rate it's not clear your new data center is going to have half it's compute in 3 years

This would be a stupidly bad failure rate, basically the worst business decision you could make, especially if you're somehow on the hook for eating those losses (which seems to be the implication?). Is there a linkable source on this?

The only thing I could find is SemiAnalysis claims that 15% of Blackwells end up RMAed[1]. That appears to be a total failure rate, though, and if you're RMAing them, you're getting replacements. So that appears to be a pretty different state of things.

[1] https://www.dwarkesh.com/p/dylan-patel#:~:text=GPUs%20are%20...


I'll note CoreWeave didn't commission the very first Blackwells until ~Jan 2025 so it hasn't been very long for RMAs. Furthermore, for training even when the NVL backplane can use the remaining GPUs a ~20% drop in performance means it isn't in the training cluster.

I've personally heard this from several sources in the data centers (installers, training, network). It's not uncommon for 10% of racks to fail on delivery. I hear that's improved somewhat from GB200 to GB300, but the number of FW updates from the time they ship, until they're commissioned is >>10. If an HBM or GPU or backplane supply/cooling fails, it is basically not swappable or repairable. You have a "dead" rack, and deliveries are on allocation so you don't get a replacement for months (eg some "RMAs" for early delivered parts in late 2025 are still dead racks 9 months later). "Tray" swaps are technically possible, but still quite rare, perhaps because debugging takes as much time as commissioning a new rack.

I don't want to out anyone, but these are similar comments:

https://www.linkedin.com/posts/neelmaster1_aiinfrastructure-...

https://www.hostzealot.com/blog/news/nvidia-gb200-nvl72-is-n...

https://introl.com/blog/gb200-nvl72-deployment-72-gpu-liquid...


They fail that often? Damn.


Doesn't matter if the chip is 100x-1000x more efficient and faster, and you can just make a new one for new weights. Imagine being able to run GPT Sol at 1k tokens/second a year from now, at a 100x lower cost per token than now. Would that be useful? Or Qwen 3.8 27B at 10k tokens/second. The super long thinking that makes qwen so effective would take a couple of seconds.


> Imagine being able to run GPT Sol at 1k tokens/second a year from now, at a 100x lower cost per token than now. Would that be useful?

Given the current rate of change, it would be hard to guess either way. By some measures the cost at fixed quality score goes down vastly faster than that:

  A similar trend is evident in the cost of models scoring above 50% on GPQA, a substantially more challenging benchmark than MMLU. There, inference costs declined from $15 per million tokens in May 2024 to $0.12 per million tokens by December 2024 (Phi 4).
- https://hai.stanford.edu/assets/files/hai_ai-index-report-20...

15/0.12 -> factor of 125 cost reduction in 7 months.

But that may well be an extreme case. To show how broad the range is, another quote from the same publication:

  Depending on the task, LLM inference prices have fallen anywhere from 9 to 900 times per year.


This is from 2025, how about the last 6 months?


You tell me. Most of these reports take that long to get published, or even longer. Sometimes I even see new-ish reports talking about 4o.


I’ll admit that it would be an interesting product, assuming you could buy the chips / cards and slot it in commodity hardware.


A model is not useless if it is not sota. Price, and speed are also important.

A 6 months old model that can run at 1/10th hardware and much faster too, can be much more capable than a sota model when you don't have unlimited budget.


Why would it be useless in 3 months?


Because on hacker news the only thing that matters is being in the current news cycle and not whether your business is profitable.


Imagine Anthropic gives you Opus of 6 months ago but at much higher speeds and much lower cost (that they might or might not pass on).

Would you use it?


Yes, this would basically obviate Sonnet and Haiku. If you consider them 1 and 2 generations behind, respectively (that's not really what they are), you can still get a ton out of those older chips. Not to mention people still use older Opus versions happily. (In part because they don't like the new Opus but still, the cost effectiveness is a huge boon.)


GPT 5.6 Luna is already rather fast at 300 tokens per second, performs better than Opus 4.6 from what I know, and is very cheap.

I don't know why anyone would use Haiku.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: