llama.cpp adds ggml RPC for distributing inference across heterogeneous devices
Original titlellama.cpp can distribute inference on heterogeneous devices through the ggml RPC backend
AISummary
llama.cpp can distribute inference across heterogeneous devices through the ggml RPC backend, according to Georgi Gerganov. He says it is currently an advanced setting, but he expects it to become more accessible to regular users over time. A related post reports MiMo 2.6 Flash running across an RTX 6000 GPU and an M5 laptop over 10 GbE at about 40 tokens/sec.
Source: Georgi Gerganov · x.comPublished · added here