Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

The 20W number includes EVERYTHING else the brain does. The chips/models are literally only producing tokens. Let's see an LLM drive a robot harness and have the robot produce speech, as well as move through 3D space, keep track of metabolic needs, etc. etc. etc. before we compare efficiencies. That is even assuming the tokens are of equal quality. This comparison is currently Apples and Oranges.


Right, I can do the talked about ~3 tok/sec output and drive a car, hold my bladder, and eat chips at the same time.

Take that, Jalapeno!


Also, I can talk at a lot more than 3 tok/sec, its just going to be more gibberish (see thinking traces)


For what it's worth, LLMs don't really suffer from incontinence, so at least that part is pretty much a solved problem.


They sometimes leak their system prompt


So why are their water cooling systems filled with leak detectors? /s


kind of a moot point if you can't get your brain to not do everything else. I think it's a fun comparison, even if it's not a 100% equivalence.


> At low concurrency scenarios, Jalapeño demonstrates remarkable interactivity, hitting over 700 tokens per sec per user at concurrency 1 on the DeepSeek R1 model.

This is about the same rate you get out of Sol Ultraspeed.

Why do you think extra tool calls like that would be so unthinkable? It'd run circles around this, especially if the problem can be split up among a live-collaborating agent swarm, so that it's not a single user thing anymore, which is exactly what they have in the cooker with Astra.


They are not. If the robot speech is a tool call, then for a fair comparison we need to take the tool call scaffolding (and probably the reasoning too) into account. So rather than a sentence of 10 tokens worth of speech being the output, the raw token output would be maybe 10x or 100x that. Even more if we consider the management of other aspects of the robot embodiment (or we reduce the brain's 20W number to whatever is actually required to produce coherent speech, sadly it is all rather entangled so this is not so easy).


But there are already voice models that do a reasonable job at a fraction of the throughput available?

The real question is how expensive it is to coordinate between these different modalities, and I really don't see why it'd be all that much.

I half expect Boston Dynamics to show something like this off in Q4 or whatever.


I am not arguing that there are perhaps other models that can run at the same quality, can coordinate between the different modalities, but are way less power hungry. My point is exactly about the comparison between the token output of the LLM running on the jalapeno chip, and sneaking in the power "usage" of the brain in the "token output" of human speech.


And the brain is literally only producing electrochemical signals.

I don’t see how tokens can’t produce speech or track metabolic needs. You can talk to chatgpt can’t you? Or do you mean literally talking? Because that’s not a brain function, that’s the mouth, vocal chords, and lungs.


> I don’t see how tokens can’t produce speech or track metabolic needs.

It probably could, but the point is this would require additional tokens, blowing up the comparison. The token output of LLMs and "token output" of speech are simply at different abstraction levels. Hence my comparison to the LLM brain driving the robot harness to produce speech etc. This would be more comparable, and also look significantly worse than "only" the 22x less efficient number.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: