The notes get smaller
So, I have a habit of reading DeepSeek papers on the day they land, and I wrote about V4 in April when their trick was compressing attention so the model stopped comparing every token against every token that came before it, and about DSpark in July when their trick was turning the checking budget into a scheduling problem, and the house style I keep noticing is that they attack the same wall from a different side each time. The new V4.1 Flash paper is the most extreme version of that habit because this time they went after the shape of the model itself. I am no expert so please take this with a bag of salt, but I read this one twice.
The setup is the same as ever and everything downstream comes out of it, which is a lab with a fraction of the funding of the American frontier, a team many times smaller, no access to the best NVIDIA parts, and no giant data center to hide inside. And the thing they shipped anyway is a flash model that sits at the frontier on the benchmarks, generates more than two hundred tokens a second through their API, and arrives with a memory footprint four hundred and thirty seven times smaller than the first generation of the same family and about four times smaller than the version I wrote about in April. I keep having to say that number out loud because it does not survive being read silently. Its so cool!
To see why it matters you have to know what the model is doing while you talk to it, and the work splits into two phases which are reading and writing, where reading is called prefill and it pushes your prompt along with everything you attached to it through every layer of the network, and at each layer the model computes two sets of numbers for every token, called keys and values, and those get stored in the KV cache. In plain English; the model reads your prompt once and jots notes as it goes, the writing phase is it checking every new word against those notes, and that notebook is a load bearing wall of the whole generation which I had been treating as a side detail for years. Put that beside what we ask these models to do in 2026, which is to hold a million tokens of a codebase or a folder of documents or a week of an agent’s work in mind at once, and the notes stop fitting on the desk, because the fast memory inside a GPU is the high bandwidth memory sitting next to the processor and it is the most expensive real estate in computing. Once the cache outgrows it the notes get pushed down to solid state drives outside the chip, and every time the model wants to predict a word it has to drag a pile of notes back up through the motherboard while the processor waits, which is where the arithmetic that used to be the bottleneck stops being the bottleneck.
Which brings me to the first genuinely strange decision in the paper, which is that the layers are split into two halves with two different jobs, so the front half is an encoder that reads everything you gave it and builds what the paper calls a global KV cache, and the back half is a decoder that writes the answer, and while the reading is happening the back half is switched off and does nothing at all. The part that made me reread the section is how a decoder understands anything if it never read the input, and the answer is that it borrows the finished notes directly off the last layer of the encoder, which is a design that should visibly cost quality because half the model skipped the reading, and the engineering exists to keep that loss small enough that users stop noticing it (I am not sure if I would notice it either). What the decoder keeps for itself is the local picture, which is the immediate context of the sentence it is writing, and it gets that through sliding window attention, which stares very hard at the most recent tokens while never looking at the whole history. And the cleanest way I can hold both ideas at once is a firm analysing a thousand pages of accounts where the juniors read every page and hand up one dense summary, and the seniors work entirely from that summary except at the moment a line is about to be signed, when they put their glasses on and read the one page in front of them with great care. The summary is the global cache, the page is the local context, and the reading happens once and gets paid for once.
Halving the reading work is lovely and it leaves the memory problem exactly where it was, because a cache built once is still a cache, so the second move in the paper is about getting layers to share what they remember, and my April post described the first version of that idea where compressed sparse attention merged four tokens into one denser representation so the model stopped keeping four notes where one would do. CSA2 breaks the assumption that every layer needs its own copy, so a layer now runs in one of three modes. A layer in full mode does all the hard work, computing its notes from scratch and building an index across them, which is a table of contents the later layers use to find things without walking the whole pile. Reindex mode computes nothing new and reuses notes an earlier layer already made, while writing a fresh index over that same memory, which is how a pile of notes taken in chronological order gets reread thematically or by breakthrough, giving the model a second path into the same facts without paying for them again. And reuse mode creates nothing at all, living off notes and indexes that other layers already funded, so once enough layers sit in those last two modes the cache stops growing with depth, which is where a large share of that four hundred and thirty seven times comes from.
The rest of the compression is a gatekeeper, because the first layer of the decoder reads the entire global memory and builds a shortlist, and the number in the paper is that out of a million tokens it keeps roughly sixteen thousand, after which every later layer is forbidden from searching outside that pool, which is a decision that ought to be catastrophic and is defended on the grounds that the shortlist gets trained until it is good enough that the missing notes never come up. I have mixed feelings about that claim since every compression story makes it, and the honest version is that it is testable and for now the tests are holding, and a model that stops looking at everything is exactly what I would want an appendix to spend itself on.
The most counterintuitive page in the paper is the one about the sliding window, because having built all this machinery to avoid recomputing things, the team then deleted a whole cache on purpose. The local memory of a conversation was being written out to disk on every turn so the model would not lose the thread, and in a long session those writes fill the drives and turn into a queue the model ends up waiting in, so the fix was to stop saving any of it, which means that when a reply finishes its local memory evaporates and the next turn rebuilds the last hundred and twenty eight tokens from scratch. Every instinct says that has to be slower because you are doing work you already did, and the arithmetic says otherwise the moment you price both options honestly, since writing a note to a drive and fetching it back means carrying bytes across a motherboard while recalculating a hundred and twenty eight tokens is a rounding error on a modern GPU. The decision comes down to which resource is actually scarce, and the scarce one is always the wire. Why? Because everything you persist you will one day pay to fetch, and a cache is not free just because we call it a cache. That is the pattern I find most useful to steal.
Two smaller mechanisms hold up the rest of it, and the first is about traffic inside the chip, since every operation in a network produces intermediate numbers that the next operation needs, and the usual pattern writes those to memory and reads them back, which costs nothing once and becomes the entire budget when you multiply it by billions of tokens, so single pass MHC lines several of these operations up so they land together and the round trip disappears. The second is the one I have been telling people about because it is so easy to picture, which is a module the paper calls the engram holding a hundred and sixty eight billion parameters that live in ordinary server RAM down the hall from the GPU, where memory is cheap, and its whole job is holding static facts like dates and capitals and constants so the GPU is left free for the part where the model actually thinks. The image I have for it is a senior lawyer working a case with an assistant sitting beside him, doing the reasoning and the strategy while the assistant fetches the exact clause and the exact date the moment he asks, and if you tried to make that lawyer memorise the whole library you would be spending the most expensive mind in the building on filing.
Some of this release I have already written about, because DSpark is folded in here, and that is the speculative decoding work from July where a small fast drafter writes several words ahead while a memory trick keeps the draft coherent and a scheduler decides how many of them are worth checking, and it is the piece of the stack behind the output speed, which the paper measures above two hundred tokens a second. Seeing it turn up inside a release like this is its own satisfaction, because the thing I read about in isolation two months ago was never a side quest.
Then there are the numbers, and the one I keep coming back to is that the decode curve stays almost flat as the context window goes from four thousand tokens to a million, and every earlier generation of this family bends upward as you feed it more. In plain English; the model spends roughly the same energy per word whether you hand it a page or a codebase, so a curve that refuses to bend is the entire point of the exercise. The footprint numbers are the ones worth saying out loud, since the first generation needed about three hundred and ninety thousand bytes of cache per token, the version I wrote about in April wanted about three and a half thousand, and this one sits at eight hundred and ninety, which makes the story one about a compressed filing cabinet and a better filing clerk at the same time. On the benchmarks it scores seventy four point two where the frontier model everyone compares against scores seventy four, it tops the open leaderboards, and the price is where I stopped reading and started doing arithmetic, since the nearest alternatives are quoted at more than twenty times the cost per task or four times the latency, and the part of that comparison I care about is who can afford to run it. The previous pro model in this family is quoted at something like three hundred thousand dollars of hardware to run locally and this one lands near a quarter of that, which is still a house in most of my country, and the direction of travel is what makes it interesting, because at this slope the thing ends up on hardware nobody thinks of as a computer in a year or two and the argument about AI stops being an argument about datacenters.
The catch is in the paper and it is honest, which is that this model likes to think, and thinking costs tokens, so something cheap per token can still be expensive per answer if it wanders, and I have felt that in my own work where the bill rarely comes from the price on the label. The second thought I cannot shake has nothing to do with this lab, and it is about what papers like this do to the shape of the industry, because the story of the last three years was that capability lives in the datacenters and the datacenters live in America, and a curve flattened by a team without the best hardware is how a frontier turns into a commodity and a commodity turns into plumbing, and plumbing is the only kind of technology that ever reaches a school, a clinic and a shopkeeper’s monthly accounts. Nobody writes essays about the electricity supply and everything runs on it anyway, and each release like this makes the electricity a little more likely to arrive.
I closed the paper with the same vertigo I had in April, which is how much of the field’s progress keeps coming from the team with the least to spend, and a small gratitude that they publish the recipe along with the cake, because there is nothing in that for them except the standing that comes from being read. In a couple of years a version four hundred times smaller again will run on something we do not currently describe as a computer, and it will be close to free, and the argument we are having right now about who gets to use it will look the way arguments about long distance calls look today (or atleast it seems so to me). I put the paper down with a page of notes about things to try, which is the most I ever ask of anything I read, and I have not stopped thinking about a model that got faster by deciding to forget. So cool.
If I made mistakes in the mechanics help me correct them, and do tweet to me what you would run on your own machine if this fit in your pocket. I am @troysk704.