Aug 14, 2026 shared ai llm linear attention gdn Linear Attention Algorithms: LA, RetNet, Mamba2, and Gated DeltaNet A review of linear attention methods, covering LA, RetNet, Mamba-2, and Gated DeltaNet.
Aug 14, 2026 shared ai llm prefill decode cache Understanding Prefill, Decode, and the KV Cache in LLMs Hands-on example to understand the generation pipeline of an LLM.