Gemini 3.8 Flash and 3.8 Flash Cyber
- simonw - 46712 sekunder sedanThe speed combined with the fact that this thing is really good at HTML JavaScript is pretty exciting.
Here's what I got for 1.8 cents and 13 seconds from the prompt "make me a cool thing in html":
https://gisthost.github.io/?6a77bc41a81718c6aaa10d4ab243c59f
Transcript here (it was part of a chat): https://gist.github.com/simonw/b6149a49d327164d67d62c3d12992...
- jampa - 48456 sekunder sedanI've been using Gemini 3.7 for my personal trip planning app. Across multiple benchmarks, it ranks higher on everything I tried:
- Real world knowledge (when a thing opens and closes, the geographic region, historical facts). It's also the best at taking a cluster of places and working out a visiting order.
- Photo ranking (which photo should be the hero). Gemini can tell whether a photo is of the thing or of the view from it.
- Document parsing (extracting the relevant trip info from PDFs).
If you use LLMs for anything other than coding, I definitely recommend not discounting Gemini like I did just because other models are more popular.
- mattlondon - 50484 sekunder sedanCurrently top at https://deepswe.datacurve.ai - beating Opus 5!
https://artificialanalysis.ai/models/gemini-3-8-flash shows an intelligence score of 59, the same as Opus 5 medium!
Wow - for a flash model this seems to benchmark powerfully. Remains to be seen what it is like to use.
- simonw - 49653 sekunder sedanPelicans (thinking effort high, medium, low): https://tools.simonwillison.net/markdown-svg-renderer?url=ht... - high cost 8.9742 cents
Here are the 3.7 pelicans for comparison: https://tools.simonwillison.net/markdown-svg-renderer.html?u... - high cost 8.4387 cents
(I think thinking level low is a regression on 3.8 compared to 3.7.)
- simonw - 48828 sekunder sedanThe most interesting thing about the Gemini models is still their multi-modal support: they accept audio and video input, OpenAI and Anthropic's flagships are still image-only.
Gemini Flash is also pretty cheap, so it's a great family for performing media analysis, like extracting structured data from images and video.
- brap - 41477 sekunder sedanPeople have been sleeping on Gemini lately but these last few Flash releases (which were very rapid) are damn good.
These sort of fast and cheap models are great for tasks that are verifiable and can be retried infinitely (like coding), you can basically get frontier results with a good harness (at a fraction of the time and money).
- EFLKumo - 44006 sekunder sedanSomething maybe unfamiliar with you: not about coding but writing. I've asked it to write an argumentative essay, which is a part of "gaokao" (China's university entrance exam), and its work is *extremely* impressive. speaks and writes like a real senior high school student, and the opinions unfold progressively with deep hierarchy. I don't know how the Gemini team reaches this because this kind of Chinese capability literally outperforms at least 2/3 Chinese students, no to mention those who speak Chinese. After all, the model speaks like a real humankind if you prompt it well. That's AGI guys
- mattlondon - 51007 sekunder sedanWow this comes after what - 3 or 4 weeks since 3.7 Flash, which was also 3 or 4 weeks after 3.6 Flash IIRC?
I eagerly wait more info but sounds like Deepmind without Demis calling the shots has been unleashed and are operating at full speed? Shocker!
At this point it is a meme of course, but where is 3.5 Pro :)
- a11r - 48748 sekunder sedanLooks like the strategy of regular updates with incremental improvements is working out well. Interestingly, the biggest jump in Artificial Analysis Intelligence Index score is for reasoning level Medium ( 3.7 was 51, 53, 57 for Low, Medium and High, 3.8 is 52,57, 59 respectively). I think scores at lower reasoning levels are more indicative of model capability since higher reasoning levels are focussed on benchmaxxing. We use the lowest reasoning level in production with good results.
- abixb - 40038 sekunder sedanI like Google's strategy here. These new Flash models of late (Flash 3.6, 3.7 and now 3.8) have obviously been distilled from a much larger unreleased model (Gemini 3.5 Pro, iirc from the rumors).
One aspect of model releases that don't get discussed as much are the cache invalidation (changes in underlying architecture, weights, or tokenizers); I assess Google seems to be squeezing the maximum out of the last 'Pro' version they released with 3.1 back in February.
Small models cataching up with their bigger siblings are fantastic news.
- kamranjon - 47665 sekunder sedanThey've interestingly left out any mention of speed.
I have been testing 3.7 flash against 3.5 flash and it seems to lose every time in overall latency. Every benchmark I've seen seems to suggest the opposite[1] - that 3.7 flash is significantly (at times 2x) faster than 3.5 flash - but I have never been able to prove this out in real world use cases.
Has anyone found their latency numbers to actually be accurate? Is this why they've toned it down in this release? For context, I'm testing larger generation payloads that take 8-10 seconds in 3.5 flash and 15-25 seconds in 3.7 flash. Lowest reasoning settings in both cases.
1: https://artificialanalysis.ai/?speed=intelligence-vs-speed&m...
- andai - 50705 sekunder sedanWait, I didn't realize 3.7 Flash was already beating Sol on a bunch of the benchmarks. Isn't it a way smaller models?
- asdaqopqkq - 2876 sekunder sedanGemini models look so good on paper by IRL dev and daily life usage totally make it seem like it's way behind Codex and Claude.
- j-bu - 47467 sekunder sedan"The knowledge cutoff date for Gemini 3.8 Flash is March 2026 – users can expect updated information for some domains while in others they may experience the model’s knowledge is limited to January 2025 (in line with the Gemini 3 Model Family)."
Kind of wild that they haven't (successfully) pretrained a base model since Jan-25.
- raincole - 46376 sekunder sedanI don't know if Google is having the worst marketing fumble or the most genius marketing one. Their "flash" models are very comparable to other companies' "pro" or "flagship" models. It seems to be a quite counterintuitive naming convention as it undersells the models.
Unless they have an even more powerful Gemini Pro in the oven...?
- xnx - 50641 sekunder sedanSeem like a great, no-compromise, upgrade over 3.7 which is already a bargain, fast, and doesn't have the brain-damaged writing style of Claude.
- lysecret - 38964 sekunder sedanAlso just want to let my appreciation here for 3.7 it’s cheap super fast super reliable incredible at information parsing eu host able (important for us) and perfectly integrated into gcp. Great job google!
- tagalog - 7721 sekunder sedanGemini flash seems to have been a bit of a sleeper. Somehow it's ended up as the most used LLM for my client document extraction work these past few months.
I have an eval harness that runs every Thursday to determine which models are the current best for a few different client workflows. And since May(?) flash has slowly been taking over more and more stuff to the point it is now 100% on 8 out of 11 document extraction flows with the other 3 being a Flash / Opus 4.8 mix for high value stuff where cost is less of a factor.
- meh2frdf - 49984 sekunder sedanThe flash models, for coding are reckless in my experience. I have a Ultimate subscription, get good quota, but still use Opus 4.6 as it's much more reliable if you manage the context window carefully.
- jerkstate - 47353 sekunder sedan3.7 flash was by far the best model for image recognition tasks according to my benchmarks. 3.8 flash didn't regress any candidates and improved some specificity (positive ID of common name vs species name of exotic fruit, correct identification of cast/replica of artifact and statue) but is still relatively weaker (26/30) on esoteric public figures (Korean beatboxers). I'm going to have to make my benchmark harder.
- throw10920 - 10896 sekunder sedanWe've gotten an unusually fast speed of Gemini Flash releases over the past few months. Is this Recursive Self Improvement, or Google just trying to distract from the fact that it's been a while since the last Gemini Pro release?
- throwa356262 - 45373 sekunder sedan
Then why even bother announcing this? Ordinary people can use K3 and GLM 5.3 or whatever drops next and avoid all this hassle."available to trusted defenders through our new Fairwind Program" - pampas - 34040 sekunder sedanGemini 3.8 Flash is top of the Redactle LLM benchmark but so was Gemini 3.7 Flash. Both one shot all puzzles in the evals though 3.8 is just a bit faster. It also does the evals cheaper and faster than almost all the other models I've tried.
- leopoldj - 48375 sekunder sedan
- adbachman - 47358 sekunder sedanStill zero on the felony bench.
Is this weakness in their training regimen the impact of operating under regulatory frameworks for too long?
- hmate9 - 48518 sekunder sedanIt is more expensive per task than 5.6-sol high: https://artificialanalysis.ai/models/gemini-3-8-flash#price-...
- gere - 44020 sekunder sedanI have mixed feelings about Gemini 3.7 Flash. I used it for a personal project in Java and it was ok: it was crazy fast and it reached the correct result, but the code quality was barely passable.
I also used it for a an app for my Garmin watch, and it wasn't good. The code was compiling, but functionality was totally broken and even with a lot of steering it wasn't able to make it work. GLM 5.3-flash instead was up for it and the code wasn't bad at all. I am curious to see if 3.8 is an improvement in this use case.
- AM1010101 - 47780 sekunder sedanSeems to do reasonably well in opencode according to artificial analysis. https://artificialanalysis.ai/agents/coding-agents
If I had to pay per token I would probably consider using this (they seem to be on the pareto of performance) but not being able to use opencode with a subscription is not really something I'm realistically going to do when claude and codex are around. Also never gotten along well with gemini-cli / antigravity-cli.
- buntp - 48391 sekunder sedanIt seems like this is one of the most powerful models for the price, really didn't see that coming from Google
- npn - 38744 sekunder sedanStill refuse to search internet for stuff it thinks does not exist lol.
And even when searching for internet, it still cannot suggest a up-to-date approach to the problem.
For example I'm using crystal, it recently revamped the concurrency/parallel model. Even using web search, gemini still does not aware of the new feature and still give the outdated code.
I'm sure my crystal usage is not the unique case here.
- f311a - 50069 sekunder sedanIs the google infra stable enough right now? At the start of the year, the flash model was unusable for a whole month via gemini CLI. They could not fix it for a whole month and I was a paid customer.
- maxnevermind - 5522 sekunder sedanJust tried Gemini 3.8 Flash on these 2 consecutive prompts at gemini.google.com:
1 what is tesla cybercab plan to address legal implications of accident that will happen? who is going to be responsible for them when they happen? are they covered by tesla insurance or some other insurance? are there any official plan/statements around that?
2 what was the name of the experiment they started in san antonio tx when some cars didn't have a driver? what was the results of it? did they expand the operations? it was much smaller than waymo, is it growing? how it is related to robotaxi?
It is not able to connect the dots that I keep asking about Tesla in 2nd prompt and spit out some unrelated stuff. Really? How it can be that bad? Gemini 3.1 Pro model works fine in this case btw. I thought maybe it is about knowledge cut over date and it doesn't know about those events from 2025 but it seems it has the knowledge up to March 2025. Top 10 in Intelligence on artificialanalysis ladies and gentlemen.
- pimeys - 45464 sekunder sedanIt's interesting that Deepseek models were missing in the comparison. I see Deepseek v4 Flash a direct competitor to Gemini Flash for text-based agentic work.
- akurilin - 13325 sekunder sedanCurious which model this can supplant as a clear winner on almost every metric. Sol? Looks like it's not quite there on a couple of benches, but I'm not clear how much they matter in practice.
- sfink - 46227 sekunder sedanFor my application, I'm still happily using gemini-2.5-flash and the only problem is when it reports being overloaded. It's for interpreting a downscaled phone camera photo of a hand-written shopping list on a whiteboard, and it works stunningly well. My handwriting sucks, too.
(I guess the only relevance here is that if your problem matches a model's strengths, then you can do fine with a model that is several generations out of date.)
- andreygrehov - 47267 sekunder sedanI don't use Gemini, but I thought `cool, let's give this new model a try`. Opened gemini.google.com, and I'm not even surprised. The drop down gives me the following options:
- Flash-Lite
- 3.6 Flash [new]
- 3.1 Pro
The above is why i don't use LLM products from Google. If the model is not available right this minute (heck, hours before the release!), then I'm not gonna bother getting back to it tomorrow, because tomorrow I'll be playing with the new model from OAI/Anthropic.
- aff-vasileva - 44064 sekunder sedanThe model seems fast enough to solve your problem before Google finishes explaining which of its three products you need to open to access it.
- 2001zhaozhao - 39745 sekunder sedanHow generous is the Google subscription quotas compared to Anthropic and OpenAI? This sounds like a really good potential model for high volume due to its speed and cost effectiveness.
(By high volume I mean things like "main app just updated with XYZ commits, please scan XYZ plugins and surface any compatibility issues")
- yipinwong - 39022 sekunder sedanAs a big proponent of GPT-5.6-Luna for the combination of speed/perf/(especially)price,
Flash 3.8 seems like where I can specify Flash3.8 as the coding model as part of agent workflow.
The video recognition is especially impressive as they got all of Youtube to train from.
- Def people who has to queue video recognition jobs to use the model.
- mowmiatlas - 48787 sekunder sedanWow fable5.1 was the first model to do what I actually told it and I couldn’t find any problems with it, excited to try this just a day later lol
- jetter - 36720 sekunder sedanCAD for 3D printing is finally becoming feasible with Flash 3.7 and 3.8. Exciting times. https://github.com/ModelRift/openscad-skill/
- henry-xli - 38021 sekunder sedanI can’t wait until waiting hours and spending a big chunk of your usage per task seems antiquated, and real-time iteration on massive code changes is the norm. This might just be the year of efficiency, that truly allows AI to be used to the heart’s content.
- robertwt7 - 14343 sekunder sedanthis is cool for all other non coding task. however I am still stuck on 3.6 flash on my gemini web as a plus user, can anyone else even access 3.7 flash in AU?
- thrdbndndn - 6840 sekunder sedanWhen can we use it in Gemini (web)?
It still uses 3.6 Flash for example.
- lpolovets - 45040 sekunder sedanI'm surprised the introductory 50% discount is good for 4 months. It seems like frontier models release new versions every 2-3 months, so raising prices in 4 months seems like a bad plan: you're effectively planning to charge users twice as much for a model that is no longer frontier.
- drivebyhooting - 16475 sekunder sedanI’ve used the Gemini flash, but then when I have soul ultra check its work, it found a bunch of cut corners and improper design.
As much as I like the speed and interactivity, I really don’t trust it
- wjellyz - 48048 sekunder sedanbeen absolutely loving 3.7 flash for coding. it feels very fast and quality is decent for implementing product features. usually use opus or sol for hardcore debugging.
- kelvinjps10 - 41658 sekunder sedanI think about Google is the value you get of their plans, for 5$ a month you get their ai plus model combined with 400gb you can share this with your family. The other ai companies don't provide family plans
- speak_plainly - 48139 sekunder sedanAfter struggling with Gemini for months, I think the trick to getting the most out of the model is writing a really solid personal intelligence/instructions prompt. The results are night and day in terms of performance.
- nharada - 41224 sekunder sedanMeanwhile I pay for Pro and still don't have access to 3.7?
- leumon - 50099 sekunder sedanSo 89.4% on Terminal Bench 2 but only 19.1% on Tbench 4. Opus 5 is 89.1%/51.8%.
- Galorious - 30734 sekunder sedanIs anyone here using using these models via google subscription (not api). I tried to in the past using gemini cli and then agy - headless invoked by codex and claude code, but they were so incredibly buggy that it stalled 1/2 times and I cancelled. Interested to know if that has changed!
- ddp26 - 37852 sekunder sedanThere must be a deeper read on why Google can rapidly ship better small models while being delayed months on the bigger model.
What's the simplest explanation?
- _aavaa_ - 45111 sekunder sedanDo they officially support you use their AI Pro subscription (or whatever the heck it's called this month, the one that gives you models in antigravity) in a 3rd party harness?
- pwython - 50323 sekunder sedanIs there any reason to even use 3.1 Pro now?
- jpau - 32410 sekunder sedanThe iteration cycle is becoming very quick. Gemini 3.8 Flash arrived just 20 days after 3.7 Flash.
Similarly Qwen3.8-Max was updated in just 30 days (to the 0902 release) and Muse Spark in just 28 days (to the 1.3 release).
A year ago iterative releases were every 3-6 months. At what point will they reach nightly candidates?
- kelvinjps10 - 50175 sekunder sedanI see benchmarks beating sol terra and sonnet. But is actually better? Has someone used it? I don't see actually much people that use Gemini for coding.
- satvikpendem - 50091 sekunder sedanIs the Gemini CLI still terrible compared to Claude Code and Codex? The harness the main thing holding back Google models as they could've been the best given all the advantages in compute capacity and training data they initially had, where now even the Google CEO said they're falling behind in agentic tasks, which is sort of a vicious cycle because RLHF relies on human usage.
- therealmarv - 43156 sekunder sedanOn my short tests: This model is amazing and the speed makes it feel like another sort of AI.
But it's bad at code reviews (maybe it's the harness agy cli?). Could not get it to same quality level on reviews like Opus, GPT 5.6, Grok. Even tried special code review skills but no luck.
- luciana1u - 38502 sekunder sedanFlash Cyber sounds like a villain from a 90s hacker movie and I'm here for it.
- sreekanth850 - 42166 sekunder sedanDear Google, Kindly make you chat window on the right side of vscode in antigravity extension, There is a reason others kept it like that. I can see the code and inspect the files changed while Agents keep working. its critical for me personally.
- newppc - 20983 sekunder sedanIf Google has the juice and wants to win, they need to start releasing world models.
- hmokiguess - 49971 sekunder sedan
- mrbonner - 40264 sekunder sedanI’m interested in a general knowledge model (closed or open weight) and not coding specific. I want to plan for travel and trip. Do you have one of your favorite HN crowd?
- TechRemarker - 47007 sekunder sedanHopefully before they release 4.0 Flash we will finally get Gemini 3.5 Pro.
- centaurz - 32998 sekunder sedanA company with 400+B revenue from software cannot build a usable command line cli for its vital AI model?
- algoth1 - 19714 sekunder sedanI asked gemini 3.8 high to review the site I'm working on for points of high cpu/ram consumption - it failed spectacularly and also halucinated the server i/o limits
- ldm0 - 35637 sekunder sedanIt’s strange that its score on Terminal‑Bench 4.0 is so low. They aren’t fast enough to benchmaxx that section.
- ASinclair - 50143 sekunder sedanFrom personal experience it feels much more capable than 3.7 Flash.
- sva_ - 50964 sekunder sedanmodel card https://news.ycombinator.com/item?id=49537354 (doesn't 404)
- - 48660 sekunder sedan
- johnnyApplePRNG - 14517 sekunder sedanWhy is it still such a bad coding agent? Does anybody have any insight?
I am continually impressed with Gemini's chat responses, which encourages me to test their agentic capabilities and... no... no... and no... every single time.
It's terrifying watching it, really.
- prometheus1992 - 48683 sekunder sedanGoogle keeps flashing everyone where everyone is expecting to get PRO'bed.
- im_soul - 40573 sekunder sedandisclaimer : Introductory price expires on December 31, 2026. Starting January 1, 2027, $1.50/1M input tokens and $7.50/1M output tokens will apply.
- atemerev - 47462 sekunder sedanEveryone is censoring models now with anything remotely resembling cyber or bio. I already have problems with my research in mathematical epidemiology because of that - both Sol and Fable simply refuse. They keep pushing people towards Chinese models that can be decensored.
- koalaman - 43832 sekunder sedanI use Gemini to make sense of things Claude says to me.
- lgl - 46610 sekunder sedanAm I the only only one thinking that Google might still "win" the AI race, despite the apparent gap?
They're apparently evolving slower than most SOTA models but "slow and steady wins the race" is probably still a thing.
And since Google doesn't depend exclusively on AI models, they can probably afford to "wait and see" where all this craze is heading.
- schmorptron - 25144 sekunder sedanA reminder that google is the only major lab without a meaningful opt-out of training on your data. The only way to opt out is to disable message history entirely, which seems like a darkest of dark patterns to get users to leave "opt in" to training on, because next to nobody wants to use it without message history.
- fitsumbelay - 50545 sekunder sedanshows up in /models though and encourages you to use it over 3.7 Flash I prefer this over reading specs: the "just show me" way
- realist_not - 51210 sekunder sedanAnyone has a cached page / mirror ? 404
- firemelt - 37424 sekunder sedanI wish google to thrive
- levelZero - 39499 sekunder sedanGemini 3.8 flash thinks Entoloma sinuatum is good to eat... Otherwise feels great
- dismalaf - 47090 sekunder sedanNice surprise. In a few of my own tests it seems maybe a tad slower than 3.7 (but still way faster than any other LLM I've used) and even smarter. With 3.7 I felt I could just not use 3.1 Pro at all and 3.8 seems even better.
- amazingamazing - 47802 sekunder sedanCould someone explain to me why it matters if google has the best model? Isnt the real metric cost per task?
- advenn - 50441 sekunder sedanBut where is Gemini 3.5 pro?
- - 51601 sekunder sedan
- casey2 - 4381 sekunder sedanMeh, not any noticeable improvement and unlike 3.7 high it eats all your credits, perhaps medium would be better
- mohamedkoubaa - 31906 sekunder sedanThe race to the bottom continues
- OG_BME - 51227 sekunder sedanWhat did it say?
- HardCodedBias - 38239 sekunder sedanI have to say:
The Google brand remains powerful on HN!
I’m shocked.
- alex1138 - 44173 sekunder sedanIt's a shame Google crams it ham-fistedly into search results and that Google has some of the reputation it has because I actually really enjoy Gemini and I don't even use it for the reason people often list which is that you can cross-reference it to stuff in your Google account
- dcchambers - 45191 sekunder sedanI would really love to be able to use these Gemini models in Opencode or Pi with my existing Google AI Pro subscription.
- eis - 45388 sekunder sedan3.8 uses nearly twice as many tokens as 3.7. One might be inclined to think that they just increased the thinking budgets...
3.7 used 64M on high: https://artificialanalysis.ai/models/gemini-3-7-flash 3.8 used 120M on high: https://artificialanalysis.ai/models/gemini-3-8-flash
Even their own chart showed more than 2x higher cost compared to 3.7: https://storage.googleapis.com/gweb-uniblog-publish-prod/ima...
- eis - 45435 sekunder sedan3.8 uses nearly twice as many tokens as 3.7. One might be inclined to think that they just uppsed the thinking budgets...
3.7 used 64M on high: https://artificialanalysis.ai/models/gemini-3-7-flash 3.8 used 120M on high: https://artificialanalysis.ai/models/gemini-3-8-flash
Even their own chart showed more than 2x higher cost compared to 3.7: https://storage.googleapis.com/gweb-uniblog-publish-prod/ima...
- dyauspitr - 47227 sekunder sedanWhatever they’re using within the Maps app is not good at all. I cannot just ask it for things conversationally like I do with ChatGPT. They really need to put a better model in there. I don’t even think it maintains context across two different queries within the same session. It’s not seamless and doesn’t just “get it” like ChatGPT does.
Yesterday I asked for food stop on my road trip 45 minutes from the current time and it gave me some options, but then I changed my mind and specifically asked for Asian restaurants and it completely forgot about the 45 minutes and gave me the closest Asian restaurant to me.
- FpUser - 47627 sekunder sedan>"safety performance" - this starting to get long in the tooth. Gemini cut programming session 3 times for "safety reasons" yesterday for mentioning image generation (I need to generate bunch of those for infinite zoom virtual training app experience). After I got creative and managed to trick it to answer t was of course because "think of a children"
And in my other app I was debugging and using OpenAI to optimize some path it cut me off numerous times because it did not like JIT functionality (this is my commercial business rule evaluation engine that compiles rules to executable code inside the app to increase performance using asmjit library)
I am basically paying for them to waste my tokens and time on these 2 tasks
- sergiotapia - 48320 sekunder sedanWho coined the phrase "cyber" for security related things lol. It's so 1999.
- barapa - 48419 sekunder sedanlove these flash models
- jdw64 - 48217 sekunder sedanThe biggest problem with Gemini is that its performance degrades the longer you use it for coding. Is it just me?
- yipinwong - 51045 sekunder sedan"Page not found"...
- Razengan - 28497 sekunder sedanWhat is with Google's dumb ass STILL refusing to respect the OS dark mode setting in fucking 2027??
- Mashimo - 51454 sekunder sedanIt's 404 now.
- zuzululu - 30288 sekunder sedannot really getting the excitement over this, its at opus 5 medium level, and opus 5 is not really the go to model , claude purists hate it
so its fast sure and decent at non coding usage but for developers nothing can really top sol or fable.
even grok 4.6 is so so and i would not choose 3.8 flash over it.
- yoga666 - 839 sekunder sedan[flagged]
- DrewKeller - 39041 sekunder sedan[dead]
- Helldez - 36732 sekunder sedan[dead]
- deanc - 49949 sekunder sedanAnd yet again another failed launch from Google. I pay for their AI plus Google one package to get more cloud storage (have no interest in their AI bundle but you have to pay). and all I see in the Gemini app is 3.6-flash
- mythz - 50414 sekunder sedanI'm trying it now for token heavy coding tasks, it's capable for many tasks but in noway compares to Claude/Sol - requires more prompts and the output isn't as good.
So just another mid-tier flash model, nothing exciting, but Antigravity has very generous quotas so it's a good workhorse model when your Claude/OpenAI subs run out.
And whilst it's a fast model, having to baby sit through and approve prompts every few seconds ends up making it slower than the Auto approve modes of Claude/ChatGPT - they definitely need an auto approve mode.
- - 46979 sekunder sedan
- tacomonstrous - 51010 sekunder sedanLooks like Google's given up on frontier models for external consumption?
- - 51153 sekunder sedan
- shuvrojit - 50382 sekunder sedanGemini is getting less useful with each update. I could edit a pdf with the 3-pro model before but 3.1-pro couldn't edit the given pdf nor it could generate one for me.
- coffeecoders - 49456 sekunder sedanOne place where I find the Flash models surprisingly bad is Google Search's "AI Mode".
A recent example - I searched for how to unsubscribe from Pearson emails. Google Search "AI Mode" confidently gave me a sequence of steps along the lines of Settings > Profile > Email preferences > Unsubscribe.
Of course, I looked for an unsubscribe link before asking Google. None of those options existed. The correct answer was there is no way to unsubscribe through the account, so I just blockthe emails instead.
I've run into this pattern quite a few times. AI Mode seems to make up things all the time.
- greenowl - 45992 sekunder sedanNot to rain on anyone's parade but I find it strange how excited and giddy people on HN get for any new X.X model releases. Pumping it straight to the top, clamoring to use it, check and compare benchmarks, bragging about it being your "daily driver"?
Are you people truly this excited about this crap? I mean I guess if you work for Google or Anthropic or whatever I could see it??? Otherwise, are these just bot comments?
Nördnytt! 🤓