Not so in the domains of computer operation, programming, and mathematical reasoning where RL and self-distillation can be used. Presumably beyond that but those at least can be done at a large scale.
No doubt they'll be going to war too soon,-- they already are if you count propaganda!
> Any information an AI has access to has been extracted from text,
Modern LLMs are trained using reinforcement learning with the LLM operating a computer.
So the training process has a captive virtual machine and the LLM grinds away at it, running commands, changing configurations, invoking the complier, writing and running tests.
I read "extracted from text" as the original training model of downloading/torrenting huge amounts of data and training the model to reproduce it. That it only sees the "shadow" of the world reflected in text rather than interacting with the world.
On reread, maybe you meant that they only experience the world through the limited channel of text?
If so I fully agree*, although there is a significant portion of modern life which is pretty comprehensively captured in text... including programming!
I always understood platos cave as being more about observing vs interacting than it being about the observations being limited in a particular way. Maybe I always had it wrong. I think the interactive part is way more important than the sensory part. Blind people are still complete human beings, after all and they can't even see shadows.
[*] with perhaps a nit that the highest performing models are visual-llms that also can take images as input. But text and snapshots is still a pretty narrow gateway to the world. And AFAIK the visual part doesn't get the same degree of closed loop training that is now used for computer use.
No doubt they'll be going to war too soon,-- they already are if you count propaganda!