Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

As for KV cache quantization, Q8_0 from llama.cpp / ik_llama.cpp should also work better than FP8 from vllm (see https://github.com/vllm-project/vllm/issues/33480#issuecomme...).


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: