Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Much of this is why I stick to the rule of:

a) Don't quantize your KV cache

b) Don't run quantizations of the LLM that are worse than the best available Q8 (the largest possible file size unsloth GGUF for a given model like qwen 3.8 27B as an example). I would rather things go slowly but I have confidence that it's doing things more accurately.



Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: