Ai Engineering 4 min read

A Month With DeepSeek 4.1 Flash: 'I Couldn't Tell You If I'm Using DeepSeek or Opus'

A widely shared essay by the blogger Jono documents a month of heavy DeepSeek 4.1 Flash use across a dozen projects, finding it indistinguishable from Claude Opus 5.5 for most work at $10 per month, with the KV cache 437 times smaller than DeepSeek V1's.

A blog post titled “Why Isn’t The Industry Freaking Out About DeepSeek 4.1 Flash?” hit 648 points on Hacker News on October 7, and its core evidence is the mundane kind that lands harder than benchmarks: the author, writing as Jono, spent a month using DeepSeek 4.1 Flash across a dozen real projects and concluded that without looking at the model name, “I could not tell you if I’m using DeepSeek or Opus.” His OpenCode subscription costs $10 per month for effectively unlimited use; a trivial file-reorganization task cost $0.003 where it would have cost roughly $1 on a frontier API; and he rarely exceeds $1 in expected costs per session even when it runs most of a day.

The Claim, and Why It Is Credible

Anecdotes about cheap models being “good enough” are a genre; this one carries unusual weight for three reasons. First, the duration and breadth: a dozen projects over a month is a real workload, not a toy evaluation. Second, the author’s workflow uses the frontier model where it matters, keeping Opus 5.5 for occasional critical tasks like final code reviews because it catches edge cases, then letting DeepSeek execute the fixes, which is a division of labor rather than a blanket downgrade. Third, external grounding exists: the post cites Artificial Analysis comparisons of Claude Opus 5.5 against DeepSeek 4.1 Flash, and the specific technical claim, that DeepSeek shrank the KV cache roughly 437 times compared with its V1 model, is an architecture-level improvement consistent with the cache-efficiency leaps this season’s Chinese releases have been shipping.

The Economics Deserve the Freak-Out

The post’s economics section is where the industry implication lives. Chinese labs are, in the author’s assessment, “a month or two behind Anthropic/OpenAI” on capability but a factor of a hundred behind on cost for equivalent work. That combination breaks the assumptions underneath frontier-API pricing: the price war that Haiku 5.5’s $0.10 input tier escalated yesterday has been about small models, while DeepSeek 4.1 Flash makes the frontier-adjacent tier itself effectively free at the consumer level. The author also notes a wrinkle: self-hosting is no longer economical versus these subscriptions, because the cache optimizations that make hosted DeepSeek cheap are exactly what a self-hosted setup struggles to match, which inverts last month’s local-inference narrative for this workload class.

The Geopolitics Is the Subtext

The essay’s most pointed passage concerns the data-mining accusation, Anthropic’s claim that Chinese labs mined Claude’s training data, which the author notes while observing that Anthropic trained on others’ data too. That accusation, and the pricing pressure from labs accused of it, is the open wound in the American pricing strategy: matching DeepSeek’s price means matching the cost structure of labs operating under different constraints. The essay’s prediction, that cache optimizations will soon make capable models economical to self-host after all, is the reconciliation path, and this week’s local-inference releases are the early evidence.

What to Watch

Three things. First, whether the “frontier for final review, cheap model for execution” workflow becomes a documented pattern in agent frameworks, since it is already how experienced developers use these models and it changes what benchmarks should measure. Second, DeepSeek’s next release: if 4.1 Flash closes to within weeks of Opus-class quality at this price, the subscription pricing model itself comes under pressure. Third, the freak-out itself: the essay’s title is a question about everyone else’s composure, and the answer, judging by a market that repriced models weekly all month without flinching at Chinese releases, is that the industry stopped freaking out sometime around the third price cut. Composure and complacency are now indistinguishable, and only the next frontier-model comparison will tell them apart.

Get Insanely Good at AI

Get Insanely Good at AI

The book for developers who want to understand how AI actually works. LLMs, prompt engineering, RAG, AI agents, and production systems.

Keep Reading

Ai Engineering

How to Automate Google Pay Integrations With MCP

Connect your AI development environment to real-time merchant data and documentation using the new Google Pay and Wallet Developer MCP server.

Ai Coding

DeepSeek's 242k-Star Agent Harness Reaches the Desktop

DeepSeek Harness v0.2 preview shipped September 29 with official macOS and Windows desktop apps, in-app plugin management without Node or pnpm, and scheduled automations, maturing the MIT-licensed framework that has drawn 242k GitHub stars.

Ai Engineering

Claude Haiku 5.5 Collapses Small-Model Prices to $0.10 Per Million Input Tokens

Anthropic released Claude Haiku 5.5 on October 7 with GDPval 1620 against Haiku 4.5's 735, an adjustable effort setting, and pricing of $0.10/$0.50 per million tokens for prompts under 100k, roughly 90% cheaper than its predecessor.

Ai Engineering

The Creator of Redis Built an Engine to Run Frontier Models at Home

Salvatore Sanfilippo, the creator of Redis, released DwarfStar 4 (ds4), an MIT-licensed local inference engine purpose-built for huge routed MoE models like DeepSeek V4 on a 128GB Mac, using asymmetric 2-bit quantization and disk-based KV caching.

Ai Engineering

Simon Willison: Agents Need Default Hard Budget Caps on Everything

Simon Willison argued on October 3 that usage-based services should ship with default hard budget caps that cut off service when a monthly limit is hit, because AI agents make it trivially easy to deploy code that incurs real costs.