Skip to content
Read the original: FunAudioLLM (Alibaba Tongyi) · new models on Hugging Face· Published 32/100AI score32/100

PrismAudio Adds Reinforcement Learning to Video-to-Audio Generation with Chain-of-Thought Planning

Original titleFunAudioLLM/PrismAudio

AISummary

PrismAudio is a framework that integrates reinforcement learning into video-to-audio generation, using a Chain-of-Thought planning mechanism.

It builds on ThinkSound by splitting single-step reasoning into four CoT modules for semantic, temporal, aesthetic, and spatial dimensions, each with targeted reward functions.

Code, model weights, and datasets are released for research and educational use under the MIT License, and commercial use requires explicit author authorization.

Read the original huggingface.co

Source: FunAudioLLM (Alibaba Tongyi) · new models on Hugging Face · huggingface.coPublished · added here