Tri Dao Says Nonlinear RNNs Differ From Attention and Linear SSMs
Original titleNonlinear RNNs seem to do sth genuinely different from attn and linear RNNs/SSMs. By themselves they already do quite well w the right pa...
AISummary
Tri Dao says nonlinear RNNs seem to do something genuinely different from attention and linear RNNs or SSMs.
He reports they already perform well with the right parametrization, and adding just one nonlinear RNN layer substantially improves a transformer-Mamba/DeltaNet hybrid.
The post quotes the M²RNN paper, which introduces non-linear RNNs with matrix-valued states for language modeling, with links to the paper, code, and models.
Source: Tri Dao · x.comPublished · added here