Our brain is more like a MoE model activating only a few neurons for specific activities making it more efficient unlike a dense model activating all the params.
Also brain produces quality tokens @ 3.3 tps instead of fast generating hallucinated tokens by certain models. Thus MTP can produce low quality tokens at 2x speed.
Also brain produces quality tokens @ 3.3 tps instead of fast generating hallucinated tokens by certain models. Thus MTP can produce low quality tokens at 2x speed.
Patience pays.