Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

"No evidence" -> Large lab saying that they use distillation when training their models


the initial post infers that distillation is all you need. it is not. in 2026 you need large scale distributed inference of gigantic models, rl envs and millions of dollars runs to get to something decent. if you think glm just has to sft on traces of claude to edge Mythos on some cyber benchmark you are fooled.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: