Skip to content
View original post on X: OpenBMB· 59/100AI score59/100

VoxWeft runs real-time interpretation locally on Apple Silicon using VoxCPM2

AISummary

OpenBMB highlights VoxWeft, an open-source simultaneous interpretation system for Apple Silicon built by developer @HenryZ30734018 on an MLX implementation of VoxCPM2.

The system turns live speech into translated speech on-device, with first audio streaming in about 170 ms on an M5 MacBook.

VoxCPM2 generates speech in 30 languages, supports direct language-pair interpretation, and clones a target voice from about 5 seconds of reference audio.

Post on XView on X
@OpenBMB

🎙️ Real-time interpretation, powered locally by VoxCPM2.

Developer @HenryZ30734018 built VoxWeft, an open-source simultaneous interpretation system for Apple Silicon. It uses an MLX implementation of VoxCPM2 to turn live speech into translated speech on-device, keeping audio private and responsive.

✨ Highlights:
⚡ VoxCPM2 streams first audio in ~170 ms on an M5 MacBook
🌍 Generates speech across 30 languages, supporting direct language-pair interpretation without a pivot
🗣️ Clones a target voice from ~5 seconds of reference audio for a consistent interpreted voice
💻 Runs in 4-bit quantization on MLX, making low-latency local speech generation practical on Apple Silicon

VoxWeft shows how VoxCPM2 can become the speech layer of a full real-time application: not just producing audio, but enabling private, multilingual interaction that stays on-device.

Try VoxCPM2 and see what you can build with it!

🔗GitHub:http://github.com/HenryZ838978/VoxWeft 🔗Devlog:http://github.com/HenryZ838978/VoxWeft/blob/main/docs/DEVLOG.md
🤗 VoxCPM2:http://huggingface.co/openbmb/VoxCPM2

Source: OpenBMB · x.comPublished · added here