HiLS-Attention
model Your tags
Your notes
"Hierarchical Sparse Attention Done Right: Toward Infinite Context Modeling" — Hunyuan attention research with an open 7B demonstrator. HiLS learns chunk selection end-to-end under the language modeling loss (rather than heuristic block scoring), claiming over 4x train-length extrapolation (RULER held at 128K from short-context training). Notably built by continued-pretraining 50B tokens on Ai2's OLMo-3-7B — a Chinese frontier lab publishing attention research on an American open base. Apache-2.0.
Model Details
License Apache 2.0
Base model olmo-3