Why isn't the industry freaking out about DeepSeek 4.1 Flash?
- vishvananda - 4684 sekunder sedanThe reason people aren’t freaking out is because most people are using heavily subsidized subscriptions.
I tried the cheapest provider on openrouter and burned through $50 in a few days. Quality was ok, seems slightly above Luna quality perhaps? But that $50 is 1/4 of my codex subscription where I could have burned that many tokens or more using Astra within my weekly reset.
This won’t last forever but as long as the frontier labs are subsidizing this heavily the open models won’t matter.
- giancarlostoro - 10487 sekunder sedanCall me crazy but:
VRAM & Memory Requirements by Precision
• FP16 (Full Precision): Requires ~1,664 GB of VRAM (e.g., an 8x B300 288GB cluster).
• INT8 Quantization: Requires ~832 GB of VRAM (e.g., 8x H200 141GB).
• INT4 Quantization: Requires ~416 GB of VRAM (e.g., 8x A100 80GB)
VRAM aint cheap, Sam Altman ruined the cost of memory, Nvidia doesnt make enough consumer GPUs letting the market go insane over them, I still have friends on 1070s or 1070 TIs because GPUs have been severely overpriced for too long. I remember when a gaming PC was only $1000.
Even so why would anyone not sleep on a model they cannot run?
- gregwebs - 3322 sekunder sedanI have been using DeepSeek 4.1 flash intensively for over a month. If I run it all day long it costs $1-2. Its fast. Previously I was always quickly running up to my Claude/Codex 5 hour window (on the $20/month plan). The cost savings of DeepSeek is real as shown in this article and I am using subsidized plans.
DeepSeek is horrible at grilling sessions (the /grill* skills to make technical decisions). It doesn't know how to explain things. Maybe the skill could be adjusted. It also doesn't come up with as good solutions as Opus/Sol.
What I use it for is
Previously I planned with Opus/Sol/Astra and then I used DeepSeek for coding, and then reviewed with Opus/Sol/Astra. With the cost improvements to Opus/Sol I am trying to use them for coding instead now so there will be less back and forth review needed.* the orchestator of my coding workflows * the tester/verifier of code changes * the sub agent that explores code or does web searches * putting together code base research reportsThey are all working together in Pi using the extension @tintinweb/pi-subagents where my workflow skill is calling different subagents that use different models.
Luna is cost competitive, but doesn't score as well on intelligence. I do need the intelligence for most of what I use it for, so I am not motivated to use Luna. Haiku also doesn't seem like a competitive price/performance mix.
- p1necone - 7139 sekunder sedanI have a pretty large, complex project I've been building with heavy AI use (new language + compiler). I was following a 'strong model as orchestrator launching cheap models as implementers' pattern, but I recently trialled just using Deepseek-V4.1-Flash as the model for both layers because of the cost savings (with mimo v2.6 flash on code review agents for some decorrelation).
I was previously using GLM-5.3 as the orchestrator, after switching to DS anecdotally there was an unnacceptable quality loss, mostly around not taking all the relevant context into account when making decisions, pulling new design out of thin air without discussion too often, and being way too wordy and rambly in documentation despite prompting to avoid it. There's a lot of docs, rulings, core concepts, design philosophy to uphold and DS was just not cutting it.
However, it's perfectly capable of being the sole agent for all of my well specced implementation tasks. I've gone back to GLM as the orchestrator.
- mlinsey - 7937 sekunder sedanI'm paying for the heavily-discounted subscriptions, not the API rates. There isn't really a cost gap for me. DeepSeek doesn't have a subscription to compare to, but when I compared the GLM 5.3 usage I got from a $100/mo Z.ai subscription compared to Opus 5.5 on a $100/mo Claude subscription, there wasn't a big gap. And GLM 5.3 is very clearly not a frontier model (deepseek v4 seemed a lot closer, but I didn't use it enough to really say for my workloads).
I don't think those subscriptions nave negative contribution margins, either. I think we're seeing a lot of price discrimination by the big labs, and huge margins on their frontier models. The fact that they have been cutting prices to their second-biggest tier of models (Opus/Sol).
Open models catching up and collapsing these margins would worry me if I were a shareholder in the big labs, but as a user, I really doubt that the western labs have bigger environmental impact just because they have higher API costs, I think they have a ton of efficiencies they aren't sharing with customers yet because demand is so high.
- lmf4lol - 9130 sekunder sedanOh man. v4.1-flash has been an sbolute game changer for us. We run all our Personal Assistants now on flash (thinking high) by default and it works incredibly well. There is really no need for basic agentic tasks that might require Kimi K.3 or GLM-5.3 levels.
Once its gets juicier, we let flash launch specialized subagents with specific models. GLM-5.3 for coding or Kimi K.3 for research and critique.
But as a main driver. I love flash. And it brought our bill down by A LOT :D
- user43928 - 4512 sekunder sedanBecause DeepSeek is not "a month or two" behind as claimed in the article.
These open models still did not beat February's Mythos / Fable 5.
DeepSeek 4.1 Flash is behind GPT 5.6 Sol, and that one is left in the dust by the excellent Opus 5.5.
Rumors say Anthropic is holding in reserve the big improvement, Fable 5.5, for the IPO.
It's plausible that open models are 6 - 12 months behind, and there is no "good enough". As long as progress doesn't slow down, leading labs have nothing to fear.
- brunooliv - 199 sekunder sedanIt’s obvious: they train on prompts and store data when using through their official API. And for third party it’s just… not good. That’s it.
- ctolsen - 367 sekunder sedanNot sure "freaking out" is the word I would use, but it’s fairly obvious looking at OpenRouter usage that the price cuts on Luna a while back were in response to intense competition from dsv4.
So the industry is responding, where it matters. Which is on heavy API usage, not coding subs.
- arush15june - 5059 sekunder sedanI am 4.1 maxxing on commandcode GOAT Plan + api rates with oh my pi for the last 4 weeks, it's absolutely amazing and crazy fast, it's alright if it makes a mistake, I have enough time to iterate again, I have also added an advisor layer of mimo 2.6 pro which does make it a notch smarter. Getting haiku 5.5/sonnet5.5 to work on plans and letting 4.1 flash work through it is helping a ton too.
I am a big ChatGPT fan, all our team has ChatGPT Subs, but the TPS across all models including luna is just so damn slow.
Commandcode giving 60$ worth of Deepseek for 10$ is just genuinely goat.
And it never says no for cyber tasks so that's a big win
- hmontazeri - 9756 sekunder sedanI had the same experience using ds 4.1 last couple of weeks. It’s insanely good for the price. I’m doing mostly web dev it excels at everything I throw at it. The pricing is ridiculous. I canceled my gpt subscription and haven’t looked back hope the pricing stays like that. I almost never need a better model. I still keep my Claude 20$ sub for now but I feel like one more iteration and I won’t need even that anymore I hardly use it
- profsummergig - 484 sekunder sedanWhy isn't the author worried about sending her/his ideas to DeepSeek online (instead of hosting it and using it locally)?
- james2doyle - 3278 sekunder sedanBeen using Flash 4.1 via the ante harness to blast through a GBA recomp. The ante team has pushed hard to make Flash 4.1 perform well under it. So far, I've maybe spent $10 over the last 3 days. Its a real workhorse and works much better in this harness
- zug_zug - 6116 sekunder sedanI did a test on this a couple weeks ago. What I found was that the chinese models were far better than API rates, but about comparable on price vs the subsidized subcription model (chatgpt). Also it was my experience that codex completed tasks quicker.
That said, it's my best understanding that these american companies aren't profitable and will eventually raise rates (the old uber trick) so I'm keeping myself ready to switch when that day comes.
- simpaticoder - 9873 sekunder sedanThe question seems rhetorical but I think there are two reasons in some combination. First is it there is some awareness lag here. That lag can be on the producer and consumer side. Software enterprises are pretty slow to adopt new things and slow to try new things so they might only be aware of openai and Claude as options. Plus there are some scariness because deep seek is a Chinese model and therefore export restricted - never mind that there are American in European providers.
The other reason is more interesting. Maybe the frontier providers think that price performance is irrelevant in light of very powerful frontier models that can start the RSI loop and or a huge displacement of work and a winner take all economic situation. After all if frontier providers earn everyone's money then you won't have any money to spend on any model 100x cheaper or not.
- RGS1811 - 8008 sekunder sedanThis model finally got me off my Claude Max subscription. I’ve found it superior to Opus 5.5 in certain use cases, and certainly faster.
I’m convinced that I’ll have good enough inference on my laptop at reasonable speeds within the next year.
- nerdypepper - 1124 sekunder sedanhttps://tangled.org/astrra.space/ds4-recipe is an incredibly cool writeup on making deepseek v4.1 flash run really fast.
- wg0 - 9891 sekunder sedanWhile using DeepSeek v4.1 Flash I was architecting a system and I made a mistake of drawing the RPC boundaries at a wrong place that did cost me in so many ways.
I realized that mistake and guided DeepSeek where it should be.
Next I fired Fabble 5.5 set to high to check if the hype is real about Fabble. It exhausted 89% of quota and came up with NOTHING that DeepSeek hadn't flagged itself already in its notes.
- apitman - 5242 sekunder sedan> With my OpenCode Go sub of $10/month, DeepSeek is basically unlimited
My OpenCode Go monthly window was scheduled to reset this morning. It was sitting at 22% used despite me using DeepSeek V4.1 Flash heavily as my implementation agent the past couple weeks (I use gpt-6.1-sol high for planning/orchestration).
I had 1.5 hours left so I fired up first 10, then 20, and finally 50 concurrent subagents all working on reverse engineering C code from an old PC game. They found over 100 new functions.
This is the first workload I've found that could make a dent in my sub. It got my 5 hour window to 85% used, but sadly my monthly was still only at about 35% when it reset. So that cost maybe $2.
- swiftcoder - 9638 sekunder sedanI think the interesting provider to cross-check this assumption with here is Meta, who is clearly freaking out, and is currently providing Muse 1.3 even cheaper so long as you are willing to share data with them
- alex-moon - 7795 sekunder sedanI think because we're all just using it thinking we have found the "model for me" and never mentioning it to anyone because what would we say? It's good. It's a bit like the Logitech MX Master, as more and more people assumed they had found the ideal mouse for their purposes, it quietly became the professional standard through sheer adoption.
- aguilaair - 8724 sekunder sedanWhat about MiMo v2.6 Pro? It’s throughput is slower by default (UltraSpeed is faster than DS4.1F) but is above the pareto line, and cheaper.
see https://artificialanalysis.ai/models/releases/comparisons?co...
- - 4215 sekunder sedan
- _jayhack_ - 2655 sekunder sedanEnterprise is not freaking out because DeepSeek 4.1 Flash does not actually occupy a spot on the Pareto frontier for non-coding enterprise workflows. We see this at my employer, focused on non-technical knowledge work. Luna 6 and now Haiku 5.5 are both very competitive if not better on all axes that we care about
- wren6991 - 2859 sekunder sedanIt's a solid little model, and I appreciate DeepSeek's commitment to the bit in releasing a brand new pretrain, double the size, numerous architectural innovations as a ".1" release over the excellent DeepSeek V4 Flash.
- ne01 - 2440 sekunder sedanDeepseek V4.1 Flash is a hidden gem, really. Not to mention, you can easily get it through many providers that offer zero data retention and consistent speeds above 200 tokens per second!
- elmer2 - 8188 sekunder sedanDeepSeek isn't even on my mind. I use the frontier models and can get the best in the industry for a relatively cheap price.
- LeBit - 3100 sekunder sedanI have subscriptions to OpenAI and Claude but use DeepSeek 4.1 Flash for my coding agents.
It costs pennies and you got really great output.
The author is spot on.
- pants2 - 4031 sekunder sedanProbably because Luna is faster, cheaper, and approximately as smart
- potsandpans - 599 sekunder sedanI'm using it quite extensively in my PlayStation decompilation harness
- 0xbadcafebee - 605 sekunder sedan[delayed]
- jeffrallen - 781 sekunder sedanAlso, it is willing to do legitimate work I need done which other models flag as dangerous and refuse to do. (Software testing of a DHCP server to survive bad inputs.)
- browningstreet - 10219 sekunder sedanWhat would freaking out look like, or is this just a stupid bloggish title flourish?
Is OpenAI coming in $20B under a sign of "freaking out"?
- bitfilped - 878 sekunder sedanBecause in two weeks someone will be asking why I'm not freaking out about AlphaDolphins 0.3 Zip and then in a month FrozenMonkey 2.5 Artic.
- smallmancontrov - 10240 sekunder sedanThey might be. They would delay public admission as long as possible, because public admission would make stocks go down.
- f6v - 5561 sekunder sedanMy anecdotal experience is that I can’t even trust DS4Pro let alone Flash. I always have to have Sol reviewing the code.
- jbellis - 6424 sekunder sedanI built mjolnir in large part so I could have Opus manage DeepSeek Flash subagents. It's phenomenal and extremely light on the Claude tokens. https://github.com/BrokkAi/mjolnir/
And yes, Opus is enough smarter than DSF that it's worth the extra steps. This ranking is from live tickets, no contamination: https://slopcop.com/power-ranking
- booi - 10479 sekunder sedanBecause GLM 5.3 Flash is even cheaper?
- liuliu - 8442 sekunder sedanDeepSeek 4.1 Flash 0910 is perfect for M5 Ultra 256GiB. Running it fully resident in RAM, prefill at ~2500 tok/s and decode at ~40 tok/s. Probably tons of room to improve from there.
- wildster - 6614 sekunder sedanI like GLM 5.3 Flash, it seems good enough for coding features if you have a good structure and a good AGENTS.md
- try-working - 1278 sekunder sedanI have used over 40B tokens and spent over $800 on DeepSeek API over the past 30 days, mostly on V4.1 Flash.
It's good, and you can do most work with this. For complex software implementation you need to split your runs into various phases, build in verification, and use subagents so that work gets another audit and repair pass from the lead agent. You can do pretty much everything then. Frontier models can do without compelx workflows, that's the difference.
- xyzsparetimexyz - 7942 sekunder sedanThere was a moment 3 months back where the sentiment was that cheaper models were the way to go. Since then the pendulum has swung back.
- aszen - 6775 sekunder sedanBecause subscription plans are cheaper, only enterprises paying per tok pricing should be freaking out
- thefourthchime - 9819 sekunder sedanFor non-coding tasks it may be fine. But for coding, Opus 5.5 is just a completely another level than something like Deepseek 4.1 Flash.
Opus 5.5: TIME 9.3m COST / $1.99 / SCORE 99/100 https://jonclegg.github.io/pacman-bakeoff/#claude-opus-5-5
Deepseek 4.1 Flash: TIME 2.8m / COST $1.89 / SCORE 72/100 https://jonclegg.github.io/pacman-bakeoff/dev/#deepseek-v4.1...
- aussieguy1234 - 2876 sekunder sedanWhat blows me away about this model is it's speed.
It's way faster than Opus or any of the GPT models.
I have a coding harness which is opencode plus a few skills relevant to my workflow. Deepseek 4.1 Flash does very well in this environment. I haven't noticed much difference quality wise compared to Opus 5, which I use in my day job as my employer pays for it (although I'm considering using DeepSeek here too given how cheap it is).
- pianopatrick - 9244 sekunder sedanI was just using a bunch of models in Cursor to review a project. I went looking for DeepSeek and it was not one of the options.
Would be cool if they added it.
- tengbretson - 10469 sekunder sedanI don't know about "freaking out", but I'd say I'm having a good time here with DS 4.1 flash.
- gsky - 8843 sekunder sedanAmerica bans Chinese models sooner or later just the China banned American big tech
- hypfer - 9456 sekunder sedanIs it known why unsloth seems to not have touched DeepSeek 4.1 Flash?
- pizza234 - 7320 sekunder sedanPeople have been raving since forever about Deepseek, but if one looks at the CoT, it's evident that it's way way stupider than frontier models (there's a reason why it's cheap). It's laughable to compare Deepseek 4.1 with Opus 5.5.
I've benchmarked, rigorously, deepseek-v4-flash for programming and personal use, and it is definitely less smart than Qwen3.8-flash-next (which in turn, is not terribly smart).
Local models are also really slow, unless one spends insane amounts of money.
Having said that, Qwen3.8-flash-next is an impressive evolution; it reaches the small versions of the frontier models (like Sonnet) - but again, it's massively slower and not 100% reliable (including: stability).
- robertlane0 - 4230 sekunder sedanHonestly for me the intelligence gap between DS 4.1 Flash and Muse Spark 1.3 makes Muse more worth it for me, especially on a $10 OpenCode Go sub, with the caveat that everything I use it on is open source which makes the fact that I'm sharing it with Meta a little moot because it's already published permissively on GitHub anyways.
- - 8574 sekunder sedan
- kristianp - 8751 sekunder sedan> shrank the KV cache by roughly 437X
Can't you just say "shrank to 1/437th the size"? It's not that hard.
- MisterMunchkin - 8600 sekunder sedanI had it make 25 different things today and it cost $0.70
It’s disgustingly good value. I find it capable of doing anything I want.
Obviously can’t use it at work, but for home projects it’s awesome.
- anguralbanish2 - 9950 sekunder sedanI would love to get them more better, it's good not a bad thing.
- cactusplant7374 - 6251 sekunder sedanBecause engineers are lusting for 1000 tokens per second. You can only achieve something like that with OpenAI.
- pessimizer - 7381 sekunder sedanI'm no expert, but it think that it's the pricing on GPT-6 Luna. I'm also guessing that it's been underpriced just for this reason. I also don't think it's all that great, but it's definitely very cheap.
If it's underpriced, it's a loss leader to sell the other models, so it actually can't be too good.
I really put these things through their paces because I use them to review and work with new abstract game rules and models, so they're always flying blind. Luna misses the obvious (and more importantly, the clearly explained) consistently. My second prompt is listing all of the points in its first response, and saying "No, it doesn't work like that." The third prompt is picking out the two or three suggestions it made after correcting itself on all of the original points and saying "That's how it already works." The fourth prompt is "Now that we're done going over the rules, can we start?"
I actually feel like 5.6 Luna seemed better.
- sergiotapia - 7986 sekunder sedanIn my experience it just takes so much longer to arrive at "done" state for me. It thinks for soooooo long. I guess if you're running 12 sessions at once you don't really notice.
- AIblemblio - 9762 sekunder sedanNo they can't.
And as long as I pay as little for claude opus 5.5 i do right now, i'm using it.
But yes i'm glad that we have alternatives.
- m3kw9 - 5989 sekunder sedani thought 6.1sol copied the caching architecture so this isn't such a big deal no more
- doctorpangloss - 8481 sekunder sedanbecause it doesn't work very well?
if you have a legitimate coding application, it isn't very good. if you have some kind of inauthentic activity, which could be what it is trained for for all sorts of reasons...
- verdverm - 73824 sekunder sedanWhy would we freak out? The systems we use have always gotten better, faster, cheaper with time
- kydanet - 7368 sekunder sedan[flagged]
- oh_no - 2397 sekunder sedanAA shows Luna at 1/4 the price, 1 point behind on intelligence matrix with a 38.
Haiku 5.5 is 23% cheaper with a 4 point intelligence lead.
I'm on subscription usage so I can't compare Flash 4.1 to them directly but the OP has his head up his ass if he thinks Opus 5.5 is the best point of comparison. Why is anyone using Opus if the new Haiku is indistinguishable /s
Just absolutely terrible post, admits to using Opus for review but claims its intelligence isn't needed, why aren't you using Haiku or Sonnet then?
- CurbStomper4 - 5515 sekunder sedan[dead]
- distantsounds - 8317 sekunder sedanbecause we've all figured out that AI is just a huge grift?
- sroussey - 8994 sekunder sedanNot comparing to gpt-6-luna which seems comparable and priced well.
- wewewedxfgdf - 8206 sekunder sedanYou might also choose to pay money for a service that provides real value instead of actively choosing to support the Chinese deliberate effort to undermine this country.
Nördnytt! 🤓