Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Is it because they will have to communicate back errors during training? I forgot that training these models is more of a global task than proteins folding. In that sense this is less parallelizable over the internet.


Yes, and also activations if your GPU is too small to fit the whole model. The minimum useful bandwidth for that stuff is a few gigabits...




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: