Kimi Delta Attention is the key architecture thread
The Kimi Linear paper introduces KDA and reports up to 75% KV-cache reduction and up to 6x decoding throughput at 1M context in the paper setting. The K3 page points to KDA as part of the model architecture.