Index-Homura-9B-FP4 released with NVFP4 quantization for translation model
IndexTeam/Index-Homura-9B-FP4
AISummary
IndexTeam released Index-Homura-9B-FP4, an official NVFP4 (W4A4) quantization of the Index-Homura-9B translation model from the Index-Translate family. On a fixed corpus, perplexity rose from 2.5386 in BF16 to 2.6245, a 3.38% increase, and zh->en generations matched the original. Full FP4 compute acceleration requires an NVIDIA Blackwell GPU, while older GPUs get only weight-only memory savings and the FP8 build is recommended for them.
Source: IndexTeam (Bilibili) · new models on Hugging Face · huggingface.co