Hugging FaceTransformers
Hugging Face Adds GGUF to Transformers — Research, Debug, Deploy on One Stack
Hugging Face adds native GGUF support to Transformers, letting Qwen3.5-4B quantized hit 98% of llama.cpp speed on M2 Max (70.4 vs 71.8 tok/s).
Sep 23·2 min read