Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

It's not a compute problem. It's a knowledge problem. Even if you can process each parameter individually and re-link the model layers, you need enough information to know what each parameter is for which is necessarily more memory than the weights themselves. You can use the weights to know whether each parameter is useful for a given prompt, but that operation is a strict superset of just generating the answer. By the time you know which parameters are useful, you've already done all the work of generating your output tokens and the effort is pointless.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: