Appen
SubQ-1.1-Small Technical Report on Sparse Attention
Pages
22
Time to read
50 mins
Publication
Language
English
Pages
22
Time to read
50 mins
Publication
Language
English
This technical report introduces SubQ-1.1-Small, a long-context language model that utilizes Subquadratic Sparse Attention (SSA), a mechanism designed to reduce the computational and memory costs associated with dense attention. The report outlines the challenges faced by existing AI systems in reasoning over complete artifacts, as they typically rely on retrieval and chunking strategies due to the quadratic scaling of dense attention. The authors present experimental results demonstrating that SubQ-1.1-Small achieves significant improvements in retrieval performance and context length generalization, showing strong results on the RULER benchmark and AutomationBench Finance. The findings suggest that efficient attention mechanisms like SSA not only lower inference costs but also facilitate broader experimentation in long-context training, making it feasible to conduct extensive evaluations on large-scale datasets. The report emphasizes the importance of whole-artifact reasoning in AI applications and the potential of SSA to transform long-context modeling.