Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

My understanding is that tokenization is largely serial, so for a large initial prompt it can make up a large chunk of input processing time since after handing it off to the model inference it's (able to be) fully parallel across all tokens.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: