flâneur — a map of the web's best reading

Hydra Part II - The Model | Goomba Lab

goombalab.github.io · 1,518 words · saved by 1 readers

In our previous post, we systematically compared various sequence models with different mixer matrices, and the quasiseparable SAM mixer emerged as the top performer. So, what exactly is it? Before diving into the details of quasiseparable SAM mixers, let’s briefly revisit some key findings from Mamba-2. Recently, Mamba-2 has shown that the mixer matrices of SSMs are inherently parametrized to one of the fundamental structured matrix classes – semiseparable matrices. Defintion of Semiseparable Matrices A lower triangular matrix is -semiseparable iff any submatrix from the lower triangle (on or below the diagonal) has a rank of at most . See (a) in the figure below. So why are SSMs semiseparable matrix mixers? Using our previously defined matrix mixer framework, we can represent SSMs as follows: where each matrix and vector . This decomposition shows that SSMs are indeed semiseparable mixers. [If you are not familiar with this concept, we recommend checking out this blog post for a gr

Hydra Part II - The Model | Goomba Lab Hydra Part II - The Model [ Paper ] [ Code ] Part I - Matrix Mixer Framework Part II - Hydra: The Model In our previous post, we systematically compared various sequence models with different mixer matrices, and the quasiseparable SAM mixer emerged as the top performer. So, what exactly is it? Recap: SSMs Are Semiseparable Matrix Mixers Before diving into the details of quasiseparable SAM mixers, let’s briefly revisit some key findings from Mamba-2 . Recently, Mamba-2 has shown that the mixer matrices of SSMs are inherently parametrized to one of the fund

Explore this link on the map →

related reading