llama.cpp recently added DFlash support to its speculative decoding arsenal.
Original titlellama.cpp recently added DFlash support to its speculative decoding arsenal. Along with MTP, Eagle3 and various ngram-based techniques, t...
AISummary
Along with MTP, Eagle3 and various ngram-based techniques, the local model performance takes another step up. Special thanks to NVIDIA team and Ruixiang Wang specifically for leading this effort!
Source: Georgi Gerganov · x.comPublished · added here