Transformers can now run GGUF llama.cpp checkpoints directly using ggml Metal kernels on Mac
AIHugging Face's transformers library can now load the same GGUF checkpoints used by llama.cpp, with fast local inference on Mac powered by ggml's Metal kernels. The author, Georgi Gerganov, says the work brings those kernels into the transformers ecosystem to increase compatibility and performance.




