Post
Daniël de Kok
danieldk.eu
did:plc:xbq7pvqlraojw25b27syamzp
So cool! transformers can now run GGUF models efficiently.
To bring the performance close to llama.cpp, it uses the underlying ggml kernels through our kernels library.
Awesome work by @marcsun.bsky.social et al.!
https://huggingface.co/blog/transformers-llama-cpp-quants
2026-09-22T12:51:48.462Z