Hacker News
new
|
past
|
comments
|
ask
|
show
|
jobs
|
submit
login
fastball
12 days ago
|
parent
|
context
|
favorite
| on:
GigaToken: ~1000x faster Language model tokenizati...
Tokenization is <0.1% of the inference time for the first token in the same way it is <0.1% for the last.
help
marcelroed
12 days ago
[–]
Time to first token refers to the time until the model outputs one token, which includes the time to process the entire prompt (doing prefill). The GPU time per token is much lower when doing prefill, so the significance of tokenization is higher.
reply
dingdingdang
12 days ago
|
parent
[–]
Have you done preliminary numbers on replacing tokenizer on, say, llama-server?
reply
marcelroed
12 days ago
|
root
|
parent
|
next
[–]
Added numbers here:
https://news.ycombinator.com/item?id=49015014
reply
marcelroed
12 days ago
|
root
|
parent
|
prev
[–]
Running the numbers now
reply
Consider applying for YC's Fall 2026 batch!
Applications
are open till July 27.
Guidelines
|
FAQ
|
Lists
|
API
|
Security
|
Legal
|
Apply to YC
|
Contact
Search: