FlashAttention pseudocode for computing tile statistics and updating the accumulated attention output

FlashAttention

Background Frequent HBM Access Creates an I/O Bottleneck in Standard Attention The authors focus on the I/O bottleneck in attention, pointing to the frequent HBM accesses in standard implementations as a major cause. To understand what is going on, we first need to look at the memory hierarchy found in GPUs and most other accelerators. Improving Actual Wall-Clock Time, Not Just Reducing FLOPs The authors note that previous work has tried to reduce the compute and memory complexity of attention, but many of those studies do not report improvements in actual wall-clock time. Their explanation is that these approaches tend to focus on reducing theoretical operation counts while overlooking the overhead of memory access. ...

September 18, 2026 · 13 min · Donghyung Ko

LightMem: Lightweight and Efficient Memory-Augmented Generation

This post summarizes LightMem: Lightweight and Efficient Memory-Augmented Generation. The easiest way to give an LLM agent memory of past conversation is to include the whole history in the prompt every time. That becomes harder to sustain as conversations grow. A long context triggers the “Lost in the Middle” problem, where the model ignores information buried in the middle, and memory systems that re-read the accumulated history on every turn pay for it with higher compute and slower responses. LightMem targets both problems at once. It’s a lightweight memory-generation system that cuts token usage to a fraction of what existing systems need, while outperforming them. ...

April 6, 2026 · 4 min · Donghyung Ko

GraphRAG

This post summarizes the Microsoft Research paper From Local to Global: A GraphRAG Approach to Query-Focused Summarization, also drawing on the video [Paper Review] GraphRAG by Seoul National University’s DSBA Lab. RAG works by building a trusted document collection ahead of time, then, when a question comes in, retrieving the relevant documents and handing them to an LLM as grounding for its answer. The basic pieces are indexing (chunking documents into a searchable form), retrieval (finding documents relevant to the question), and generation (producing an answer from the retrieved documents and the question). ...

March 13, 2026 · 7 min · Donghyung Ko