Read the original: IndexTeam (Bilibili) · new models on Hugging Face· Published · added · 6d ago20/100AI score20/100
IndexTeam releases NVFP4 quantized Index-Echo-S2TT-2B speech translation model
Original titleIndexTeam/Index-Echo-S2TT-2B-FP4
AISummary
IndexTeam has published an official NVFP4 (W4A4) quantized version of its Index-Echo-S2TT-2B speech-to-text translation model on Hugging Face. Only the text LLM backbone is quantized, while the audio tower, connector, and speech-synthesis components remain in BF16. Perplexity rises 5.80%, from 4.8772 to 5.1599, on a fixed corpus, and full FP4 speedup requires an NVIDIA Blackwell GPU.
Source: IndexTeam (Bilibili) · new models on Hugging Face · huggingface.coPublished · added here