KDA · 54 LAYERS
A memory that doesn't grow
Kimi Delta Attention is linear attention: instead of keeping every past token, each layer folds the conversation into a recurrent state of fixed size.
Why it matters: Three quarters of the model carry the whole session forward in the same amount of memory on turn 5 as on turn 500. That's what makes long agent sessions practical on hardware you can buy.
- fixed-size state
- no KV cache
TOKENS IN