I spent last week watching my agent bill stay flat while my sessions got longer, and I read model benchmarks as a layman, not as a lab researcher, so if I mangle a cache number help me correct it.

I have been running DeepSeek 4.1 Flash heavily for about a month across a dozen projects, and it is super capable at orders of magnitude less money than the frontier models, and mid-session without looking at the model name I honestly cannot tell whether I am on DeepSeek or Opus, across conversations and work and speed, and I treat it like a frontier model because it behaves like one.

Its so sad to see how long we waited for someone to say the quiet part out loud.

I keep thinking about dates because we lived through this democratisation with hardware, which is why this feels familiar. The mainframe in the seventies belonged to the air-conditioned room, and then the Apple II in 1977 and the IBM PC in 1981 put a computer on a desk for something like fifteen hundred dollars, and then Jio in 2016 put data at something like three hundred and ninety nine rupees a month into phones costing less than seven thousand rupees in India today, and the machine stopped being a privilege and started being furniture. Models are walking the same road at ten times the speed, and the distilled Chinese models trail Anthropic and OpenAI by a month or two on paper and handle the same workloads in practice, which is exactly how Trevithick’s high-pressure engines embarrassed Watt’s monopoly once the patents expired, except the expiry now arrives every few weeks through weights and caches instead of courts. And yes, there is a theft debate underneath with Anthropic accusing Chinese labs of mining Claude back in February, and Anthropic itself standing accused of training on everyone else’s words, and I stopped following that trial because I am trying to get the most bang for the buck and the buck just collapsed.

The shift is not a benchmark chart and it is a change in manners, because once a model is good enough for unattended high-quality work, chasing the latest and greatest starts looking silly, like asking a math PhD to organise files on a desktop, and the fun demos keep coming while daily work quietly moves to the cheap seat. I run an OpenCode Go sub at $10 a month which makes DeepSeek basically unlimited, and I spin up mindless tasks without shame, with exploratory UI monkey testing and desktop reorganisations that cost $0.003 instead of $1, and I rarely cross $1 in expected cost even when a session runs most of the day, and I felt the same flip last month on Omarchy when a resolution switcher took about ten minutes for something like forty cents in tokens, which is less than the coke I was holding.

I lean on 4.1 Flash for complex planning and research too, and I call in Opus 5.5 for a final review which catches edge cases the main loop went blind to, and then hand execution of the fixes back to DeepSeek, with GLM sometimes playing the same fresh-eyes role, which tells me the expensive call is becoming a second opinion rather than the engine. The load-bearing number here is the KV cache shrinking roughly 437x since DeepSeek V1, and holding that cache in GPU memory is one of the biggest costs of running long coding sessions, which is how all-day sessions stay under a dollar. In plain English; the model remembers the whole day’s conversation in a notebook that suddenly fits in a pocket, so the GPU stops renting warehouses just to keep context alive, which burns less money alongside less water and electricity, and Opus 5.5 quietly picked up the same efficiency trick which means the magic is spreading through every lab.

On money, I am not nickel-and-diming here since work already pays for Claude and Cursor and the rest, and my argument runs on planning and sustainability alongside democratising access, because a dollar-a-day coding brain changes who gets to build, and a kid in Delhi or Bangalore priced out of stacked Claude Max subs can now run the same loop as a FAANG engineer with twelve load-balanced subscriptions. Most times the industry still behaves like price tags signal worth, with the biggest labs spending the most for the highest intelligence and shrugging when the bill lands on the climate or the economy, and that dog-eat-dog reflex is why people build wild rigs to juggle Max subs and then complain when the trough runs dry.

Wouldn’t it be nice if intelligence went the Jio way and stopped being a status purchase?

Honestly, the self-hosting math surprised me, since at Flash prices self-hosting never recoups costs when saving money is the goal, and the privacy case stays real, with these cache optimisations heading toward local runs that make waiting the rational move, since Flash is technically self-hostable even while practical local hosting lags by a season. Technology always works for good and the leader always claims to be right, so watch where the rent moves next, with tokens and agent seats and skill stores ready to become the new meter once weights go free.

My backlog is full of dusty tickets that never made a sprint, and the cheap brains finally make them worth doing. Start with one and keep tripping.

Send me the cheapest all-day session you have run this month and what it built. I am @troysk704.