We are thrilled to announce that our latest research, "Efficient Context Window Expansion via Dynamic Sparse Attention," has been accepted for a spotlight presentation at ICLR.
Our team developed a novel method to allow Large Language Models (LLMs) to process massive datasets—up to 1 million tokens—without the exponential increase in computational cost. By dynamically identifying "importance clusters" within the data, our model maintains 98% accuracy while reducing GPU memory usage by half.
Read the full pre-print: [Link to ArXiv]
Access the Weights: [Link to Hugging Face]
