DeepSeek V4 Flash 0731
- LaurensBER - 38516 sekunder sedanI've been using it extensively since the release and the best summary I can give is that it's good enough to use it for (almost) everything and cheap enough that the cost are irrelevant. I'm running it in Oh My Pi with a second instance running as "advisor" and even with 5-6 active sessions (effectively 12 streams) I'm struggling to spend more than 5 bucks per day.
OpenCode Go even has double limits temporarily so for 10 USD you effectively get 140 USD of tokens to spend. It would impress me if someone could burn that amount with "normal" usage. Even when running multiple sessions.
I have a Claude Max subscription but I've barely touched it, it just feels like a step back to have to think about limits and usage even though the models are stronger.
The beauty of intelligence at this cost (even if it's not SOTA) is that it opens a whole bunch of new use cases. Test failure in CI? Have the bot automatically propose a fix, its cheap enough that you can discard it w/h issues. Test coverage too low? Auto generate tests on CI for every pull-requests! Monitoring server logs, continuous security audits and investigating every received exception now becomes possible.
I'm thinking about having it automatically filter and re-rank my social media feeds so I can steer the algorithm instead of the other way around.
Perhaps other people (with enormous budgets) were already doing all of the above but for us this is a really exciting release!
- ak_t - 29308 sekunder sedanNote this is the 07/31 release of DSv4 flash and not the "preview" that they put out a couple months or so ago.
I've been running this model locally for a week, and the preview version before that. This updated one feels like a whole tier up. It's very capable for debugging and analyzing documents/data I upload.
The killer feature, IMO, is the speed. On 2x RTX Pro 6000 Blackwell, its ~8k tok/s prefill and ~250 tok/s on a single stream. I saw 1000 tok/s with ~64 concurrent streams on vLLM.
That's fast enough that you can interactively chat with it without switching tabs while you wait, and its a ~300B (13B active, hence the speed) model so the responses are also very good. It's actually more convenient now for me to direct 95%+ of my day to day usage to my local model, and only use Claude Fable for really big coding tasks.
Until this model was released, I was contemplating spending even more money on hardware to run GLM5.2 (~750B) at reasonable speeds, but I no longer feel that need. This is smart enough, and I think it only gets much better for local models from here.
- NoboruWataya - 30889 sekunder sedanMy Claude account was banned the other day. The only possible cause I can think of is that I tried to authenticate from the AI assistant in a JetBrains IDE and, not thinking, entered the details for my regular subscription rather than an API account. As soon as it became apparent that I needed an API account rather than a subscription, I just closed out of the tab. Nevertheless, about 20 minutes later I got an email saying my account was banned for a violation of the usage policy, and my appeal was rejected.
My initial thought was to sign up for ChatGPT, but I had $20 in OpenRouter so I've been trying out DeepSeek V4 Pro with Pi for the last few days and I gotta say, it's good enough for my use case. And even with paying for API usage rather than Claude's subsidised subscription, and with OpenRouter taking their cut, I will probably end up paying significantly less overall. And I really like the flexibility of being able to use whatever minimalist open source harness I want (and being able to switch providers easily, too).
(My demands probably aren't as high as many others' - I mostly use it for help with some hobbyist coding projects, and I tend to ask it questions about how to approach problems rather than just telling it to go off and code stuff for me.)
- nylonstrung - 28001 sekunder sedanCompared to the last Deepseek V4 Flash version I've had tons of issues with it getting in infinite loops and talking to itself without executing tool calls, wasting tons of tokens
This is on Pi agent, nothing fancy at all about my prompts or use case. Anyone else experiencing this?
I've also had it randomly go from talking about Rust to talking about the electric chair, controversies about D&D rules (both irrelevant and something I've never discussed) and it's completely blind to it in future prompts even when its pointed out and referenced directly
All this said its still worth it but the agentic performance has degraded in my experience at least
- modeless - 33954 sekunder sedanDeepSeek has announced an upcoming "significant increase" in price, so this line may have to move to the right soon. https://api-docs.deepseek.com/quick_start/pricing/
- 542458 - 38848 sekunder sedanKimi K3 was an interesting model only a month ago, and now we're looking at the same performance for 1/20th of the price. Wild how fast this is advancing.
- mosura - 36893 sekunder sedanI strongly recommend trying this for programming tasks.
It is strong (not Fable strong though) with a much better “persona” than Opus, and very different blindspots. If you flip between Claude and this you will find both catch the mistakes of the other before they get out of control.
On balance I actually prefer DeepSeek for programming now, because of the way it talks.
- andai - 30183 sekunder sedanThe recently announced they're raising their prices 10x right?
Which would put them... exactly where everyone else is on this graph.
Edit: I seem to have misunderstood the news. I thought the magical cache read pricing was going away (0.002) and they were going to be on par with everyone else (0.02). But I have no idea.
Edit 2: Apparently, neither do they!
>We plan to raise the overall pricing for DeepSeek API services in the near future, with a significant increase expected. Please plan your usage accordingly. The specific pricing plan will be subject to official notice.
- stuaxo - 11362 sekunder sedanSeeing everyone spend like 200USD a month seems kind of mad.
I have £20/month Gemini and £20 a month claude for a bunch of personal projects.
Yes I have to wait sometimes, it's probably a good thing.
- CharlesW - 38461 sekunder sedanLast weeks's discussion (591 points): https://news.ycombinator.com/item?id=49120299
- arjie - 34649 sekunder sedanIt's not frontier, but it's far past what we had at the beginning of the year. It's very usable. I get great instruction compliance, tool calling, and with a trivial workflows flow it has very good long-running performance as well.
- walrus01 - 33024 sekunder sedanOke of the great advantages of v4 flash 0731 is that even in the largest size unsloth quantized gguf, Q8 K XL, it will fit well within the resources of a 256GB DRAM server. If you have no gpu at all and are okay with setting up a workflow that handles slow token per second rate, give it a task and check back in 4-6 hours, it works great. And remember to give it more lengthy tasks to run overnight. Whatever workflow you set up, the idea is to keep it busy 24x7 doing different things in parallel.
- LeBit - 11474 sekunder sedanI always find it confusing that a meaningful volume of the comments are saying "this reached parity with SOTA models. Best $/task."
And a meaningful chunk of the comments are saying "this piece of garbage isn’t even at the level of gpt-oss 20B".
- jacquesm - 8481 sekunder sedanThis is the best model to come out since the beginning of open weights models for those working with classified data that you can not use hosted services for. I've been using it pretty much day and night since it landed and I'm nothing short of amazed. You'll need some pretty good hardware to run it though.
- kromem - 27250 sekunder sedanFlash is a delightful model and the start of intelligence at effectively insignificant cost.
From here on, it's going to become all about harnesses that best situate and organize swarm intelligence at scale.
- zacksiri - 6797 sekunder sedanI'm not sure about all these benchmarks, I did some very simple tests (I have my own benchmarks https://upmaru.com/llm-tests) and these models fail, not sure if it's the inference provider or the model. They seem to be optimized for benchmarks more than real use cases. Do anything outside their distribution (even if it's not complex) they fail.
I Compared Deepseek V4 Flash 0731 (low) to Gemini 3.5 Flash Lite (minimal) and GPT 5.6 Luna (no reasoning) and Deepseek V4 Flash 0731 gets it wrong alot, where as Gemini and 5.6 Luna just gets it done.
- SwellJoe - 34932 sekunder sedanDeepSeek is my cheap and cheerful Chinese model of choice for API use. Has been for a while, but now it's Flash instead of Pro. Even cheaper, and now better then Pro. I feel like most of the major Chinese models are benchmaxxed, they have weird quirks every time I use them (Qwen 3.8 Max doesn't check its work and leaves stuff broken, doesn't write tests unless prompted, etc., Kimi ends up being quite expensive and rarely better than GPT Sol or Opus 5), while DeepSeek models seem to be generally as good as the benchmarks indicate: Not the best, but stronger across the board than any model within an order of magnitude of its price.
- apitman - 24838 sekunder sedanThese are very interesting results, and honestly hard to believe, even as a big 0731 fan.
If I'm reading the chart correctly, a couple observations:
* deepseek-v4-flash-0731 max is better than kimi-k3 max
* glm-5.2 is dumber than a box of rocks (this must be on low reasoning or something, right?)
This is way more extreme than other results I'm seeing, like those from Artificial Analysis.
- g023 - 10159 sekunder sedanIt says 'yes' where the others say 'no'. Good enough for me.
- dools - 22167 sekunder sedanI have been using deepseek v4 pro almost exclusively. I was using Kimi a lot but it just nose dived. The decline started with the release of 2.7 and accelerated with the release of 3.
When I need vision capabilities I use GPT 5.3 codex and if deepseek can’t figure something out after a few goes I switch to GTP 5.5 or 5.6 (I’ve been giving Terra first bite recently and it does pretty well, and have used Sol a couple of times).
Using this regimen means I spend under $100 per month on inference and I work all day everyday with multiple agents running simultaneously all on API token spend not subscriptions.
- zmmmmm - 23509 sekunder sedanit's great but we need a multi-modal model of this quality and price to truly declare victory.
But it makes me quite curious, how a text-only model can do so well on ARC-AGI-2 being a set of visual puzzles? It would have to solve it entirely using text-only spatial reasoning about the grid (or maybe writing code?). I am curious if this is normal or do other models use their vision capabilities to solve the puzzles?
- surprisetalk - 38354 sekunder sedanThis reminds me of those pareto-style speedrun record charts when a new glitch is discovered.
[0] https://taylor.town/silver-landmines
When I see dramatic leaps like this, it tells me that the important hacks haven't yet been discovered.
- xyzsparetimexyz - 35926 sekunder sedanThat page needs a Pareto frontier display. But wow, it absolutely demolishes.
- Almondsetat - 18634 sekunder sedanIt's crazy to think V4 Pro still hasn't finished the post processing.
- minimaxir - 39234 sekunder sedanIt's always fun when Max reasoning is cheaper than High reasoning.
- nmitchko - 27357 sekunder sedanPerhaps it might be interesting: a latent thinking version is here https://huggingface.co/nmitchko/DeepSeek-V4-Flash-0731-Laten...
Does no thinking emissions for context saving.
- sourcecodeplz - 35870 sekunder sedanwow. i remember when GPT-5.2 (medium) was everyone's favorite.
ARC-AGI II:
- GPT-5.2 (medium) %26.7 ($0.759)
- DSV4-Flash (max) %61.4 ($0.04)
- gentlewater - 37744 sekunder sedanI’ve been refreshing hacker news constantly for a week now waiting for v4 pro, after they stated it would follow «soon». I have learnt «soon» is a matter of definition.
- seanmcdirmid - 23474 sekunder sedanNote they double the price if you use during peak time. However, they define peak time with respect to China, not Europe or the USA...so if you are out of Asia, I guess Australians might be impacted, and its still cheap anyways.
- KolmogorovComp - 30666 sekunder sedanLooking at the caching price of deepseek compared to its competitors, does it have a secret sauce or is it just subsidizing?
- h4fizwasabie - 4056 sekunder sedanive been enjoying this deepseek v4 flash 0731. i think its a great model, it helped me with finishing all of my abandoned projects.
- tarruda - 20842 sekunder sedanOne of the best things about this version is that it is trained in the codex harness. It feels just as good as OpenAI models in using codex tools, but extremely cheap and with 1M context
- - 33802 sekunder sedan
- w2seraph - 22924 sekunder sedanThere's no reason that LLMs should cost beyond grave digging sums when this one topples the charts it'll be over.
- tosh - 39422 sekunder sedanresults comparable to gpt 5.6 luna but cheaper
promising!
- harisamin - 33163 sekunder sedanI'm curious... is anyone using DeepSeek V4 Flash from HugginFace? Is the cost around the same as directly form DeepSeek or from Openrouter?
- evanjrowley - 30362 sekunder sedanThe benchmark performance tells me DeepSeek v4 Flash could be very cost-effective at playing SNES/Gameboy games.
- luyu_wu - 38526 sekunder sedanIt is wild that this a log scale of cost to me!
- - 38429 sekunder sedan
- simonw - 33300 sekunder sedanThat's a pretty great score for a model you can run on as (expensive) laptop.
- CrosswordPuzzle - 33317 sekunder sedanI'm really excited for where the open weight models go from here. I've had fun with just CPU inference on old servers that only have AVX1; here's hoping for commoditized TPU-like hardware!
- nikp123 - 33174 sekunder sedanI just used it for some Kubernetes + FluxCD tasks and oh my is it good.
- mycall - 32070 sekunder sedanI'm curious how much worse the 0731 quantizations do.
- johnmlussier - 32294 sekunder sedanBeen running it using Prime Agent and absolutely love it.
- clayhacks - 38333 sekunder sedanWhy wasn’t this run against ARC-AGI-3? Or did it fail to solve anything?
- Havoc - 38050 sekunder sedanThey did recently announce they're increasing prices though (got a mail yesterday I think), so not sure this analysis showing it as price outlier will last
- - 14167 sekunder sedan
- steadyw0 - 25495 sekunder sedanShould try it sometime
- m3kw9 - 16091 sekunder sedanThe token price seem to be jigged, how do you know if it's subsidized or temporary. Anyone can just lower the token price to get to the left.
- croes - 34812 sekunder sedanPrice raise incoming
- dcchambers - 37916 sekunder sedanThis latest DeepSeek is almost at the "too cheap to meter" level. That's going to be a larger unlock than models like Fable/Mythos that are way too expensive to justify, IMO.
What secret sauce do they have?
- system2 - 19182 sekunder sedanHow can OpenAI or Anthropic fight against these prices?! $0.14 input, $0.28 output. For 1M tokens...
- gxs - 20224 sekunder sedanOne thing that popped into my head is that this shows how committed they are to building something that scales across the world
China has zero energy concerns in terms of energy production - not literally zero, but they’d be able to prioritize other dimensions and not necessarily worry about efficiency
Here they are though releasing models that sip resources
- iagooar - 34657 sekunder sedanI love DeepSeek V4 Flash since the pre-0731, now even more. It is the first model that is truly too cheap to meter.
But I find it having a pretty significant problem with tool calling - no idea why, but tool calling with it is SLOW. As long as the model is reasoning, all good. But give it a bunch of tools and it becomes extremely slow.
Am I the only one experiencing this?
- casey2 - 26539 sekunder sedanFinally something that is breaking away from the pack. Interesting that max costs less than high. I still think, currently, TPS is more important than near frontier intelligence. Likely for reasons that LeCun outlined, maybe out of a billion prompts you will get value from that intelligence. When we have very fast models abstraction will work as that filter.
- esafak - 38312 sekunder sedanIt's serviceable but, like many Chinese models, it uses a lot of tokens to get work done.
- leizhou - 33812 sekunder sedanso cool. does it mean it can understand the verificated code
- WhitneyLand - 35452 sekunder sedanThe DeepSeek team is so strong, very impressive.
Imagine if they had GPU resources of western labs.
- artursapek - 12957 sekunder sedanThese prices are not real. They already said so.
- ClipBGNET - 10663 sekunder sedan[flagged]
- - 27903 sekunder sedan
- hnc3yfnu6f - 31308 sekunder sedan[flagged]
- antirez - 38461 sekunder sedanPrice is not a good meter. Active parameters per token are. Joule would be even better.
- muricula - 38567 sekunder sedanPrice is confounded by VC subsidies, economies of scale, and inference optimizations. I think a more interesting chart would be ARC AGI vs forwards pass flops or ARC AGI vs training tokens. Of course we don't have those numbers for the closed source models or even some of the open weight ones.
Nördnytt! 🤓