Skip to content
Read the original: IndexTeam (Bilibili) · new models on Hugging Face·Published · 5d agoAI score20/100

IndexTeam releases NVFP4 quantized Index-Echo-S2TT-9B speech translation model

Original titleIndexTeam/Index-Echo-S2TT-9B-FP4

AISummary

IndexTeam published an NVFP4 (W4A4) quantized version of its Index-Echo-S2TT-9B speech-to-text translation model, quantizing only the text LLM backbone while keeping the audio tower and other components in BF16.

On an NVIDIA A100, perplexity rose from 3.4155 to 3.5113 (+2.81%), with zh->en and en->zh outputs semantically equivalent under greedy decoding. Full FP4 speedup requires an NVIDIA Blackwell GPU, while older GPUs get only memory reduction.

Read the original huggingface.co

Source: IndexTeam (Bilibili) · new models on Hugging Face · huggingface.co