Decision models in llama.cpp are now available
The /v1/systemone endpoint is available in the latest llama builds. Use it to do Jev-style inference locally, efficiently and privately. Multiple open models are supported with more to come.
https://huggingface.co/blog/ggml-org/decision-models-in-llamacpp
