MiTA Attention: Efficient Fast-Weight Scaling via Mixture of Top-ݑ� Activations
Roughly speaking, the central challenge in scaling the fast weights via MoE is to construct fast-weight experts from unstructured key-value pairs. For example, given the spatial and temporal locality priors in many modalities, a straightforward, hardware-friendly approach is to partition the sequence into contiguous, non-overlapping, fixed-size blocks, ...