Hybrid NoPE models pair local attention with global NoPE layers
Original titleHybrid NoPE models pair sliding window attention (SWA) or recurrent layers, which focus on nearby words, with global attention that uses ...
AISummary
Hybrid NoPE models combine sliding window attention or recurrent layers, which focus on nearby words, with global attention layers that use no positional encoding (NoPE). The post notes that NoPE layers receive no positional information yet can still learn long-range dependencies, and raises the question of how this works.
Source: Zyphra · x.comPublished · added here