llm
KV cache: how inference keeps attention fast
Why autoregressive transformers cache keys and values, the memory cost it implies, and the trade-offs in attention computation.
Why autoregressive transformers cache keys and values, the memory cost it implies, and the trade-offs in attention computation.