Can we benefit from using streaming attention computation during pretraining?

Author: youyu-2024Created Dec 5, 2025Updated Dec 5, 2025

Hi, in your paper you discuss adding a learnable placeholder token (sink token) and adopting full attention computation. I would like to ask whether it is possible—and potentially beneficial—to use streaming attention computation during pretraining.

Source: mit-han-lab/streaming-llm