GLM-5.3: Frontier coding with emergent cyber capabilities
- leobuskin - 22119 sekunder sedanI bought $18 GLM official subscription yesterday (5.2, but new model version was already leaking on some docs), set it up with Claude Code harness... and I’ve bumped to $80 plan almost immediately. It’s the first model that agreed on a proper security research (red team scenario), executed it seamlessly, including 0-days in WP plugins, RCE, 6.8 kernel exploit adaptation, etc - while playing against another GLM agent as a defender (following HF story)!
I understand that such models can be used by malicious actors, but it’s fair to have it publicly available (and play on your side in case of emergency). This is what changes the world in a better way, I think, not the guardrails.
- z4y5f3 - 35102 sekunder sedanApparently they are scanning OSS and popular software at scale and disclosing the vulnerabilities they found: https://cvd.z.ai/
Most of these are under embargo, but it seems there are a lot of CVE here from a wide range of popular software, many considered critical or high.
I understand the argument of "people are not actively looking", but isn't the cost for such a scan getting lower by the week, and Anthropic's Project Glasswing is supposed to find them quite a while ago?
- aliljet - 44767 sekunder sedanThis is absolutely still shy of Sol and Fable, but only just by a hair. Ridiculous results. There's still not a compelling economic reason to drop OpenAI courtesy of the ludicrous reset addiction that's taken place, but it feels like we're on the precipice.
How are you all toying with running this kind of thing in a mega quantized way locally? Two weeks out from released weights, but this is still just GLM 5.2 with post-training magic.
- hypfer - 43682 sekunder sedanI might be just reading my positive bias into that text, but is it possible that it is written less like SV marketing hype trash and more like researchers wrote it?
It does feel like it respects both me and my time.
Thank you, Z.AI. Amazing what difference it makes when the top of your org are actual university professors.
- aand16 - 42884 sekunder sedan> Mythos 5 remains well ahead at 181 and 247 tasks. The pattern across the three is consistent: the further up the exploitation chain a benchmark sits, the wider the remaining gap to the closed frontier. Capability is growing fastest exactly where we are furthest behind.
I appreciate they don't just take the opportunity to self-glaze.
- jjcm - 40117 sekunder sedanSame image->html test as I showed in the Gemini 3.7 flash thread. Note that GLM isn't multimodal, but it still was able to generate something similar-ish by writing a python script to inspect the image and extract elements from it.
Original images: https://image.non.io/neonRamenDesigns.webp
GLM 5.3 build: https://html.non.io/neonRamenGLM5.3
Opus 5 build for comparison: https://html.non.io/neonRamen
For having no vision, it did a tremendous job. I'm pretty impressed it was able to extract so much detail.
The Opus one is still significantly better, but that's to be expected since it's multimodal. Curious to see where a future version from Z.ai lands on this.
- - 1188 sekunder sedan
- bertili - 25354 sekunder sedanThis will be roughly on pair with Kimi K3, but using a third of its parameters.
Just 4 weeks ago the "Kimi K3 moment" was seen as a threat to Closed AI and in less than a month Z.ai have cut the parameter/RAM barrier to a third.
Congratulation to Z.ai and all the hard working Chinese researchers who are quitely boiling the frog.
- wxw - 44621 sekunder sedan> Scaling post-training is all we did for GLM-5.3.
Love this opening line. And wow, great results.
> As agent capability improves, much of the difficulty in scaling post-training moves from the model to the environment.
- jjice - 11796 sekunder sedanAm I correct in understanding that this is just 730B-ish parameters as an MOE? That sounds like incredible performance per parameter. The new Deepseek was also very impressive with its 280B or so. Plus the most recent 30B-ish Qwen and Muse.
I find the performance to size ratio of these models to be way more interesting, selfishly because it makes me bullish on what I'll be able to run on a machine I own over the next few years. The progress is just incredible.
- virgildotcodes - 44928 sekunder sedanOpenAI and Anthropic need to just go ahead and give people access to the cyber models.
Otherwise we have a world of attackers using open and closed source models against a much smaller group of maintainers that are likely heavily dependent on Anthropic and OpenAI and for whom it may not be a simple matter to just get approval to start using the open model flavor of the month.
- fcanesin - 11142 sekunder sedanGLM-5.3 is further proof that all >1T models are currently undertrained. I was looking at inteligence density ( https://www.pasteboard.co/6q2-5f92mtj9.png ) from recent open models (where parameters sizes are known) and taking DS-v4-flash as upper limit GLM-5.x can 3x its performance.
- vmware508 - 36713 sekunder sedanApple will release M7 MacBook Pros / Mac Minis next year, and they will be able to run free LLMs locally at native speed. All software developer notebooks will be replaced to run local models, saving a lot by cancelling Claude Code subscriptions. Developers win. Apple stocks will be rocketing. Everything else will go down. You're welcome.
- zmmmmm - 39989 sekunder sedanMissing multimodal again?
It is so valuable in practise to be able to have the models see screenshots - I guess if they aren't in the benchmarks then nobody will focus on it. But it completely nixes these for some of my main use cases.
- lazarus01 - 7893 sekunder sedanI’m using deepseek v4 flash to build a complex full stack production ai app and it’s a total beast.
I break out Claude when I hit some serious roadblocks, but that doesn’t seem to be happening much after the last deepseek flash release.
Deepseek prices just went up, but are still low.
I will def try GLM on my next project
- KronisLV - 38585 sekunder sedanTheir coding plan switched to credits, didn’t it? What are the rate limits like, compared to Anthropic or Kimi K3?
I remember trying their Coding Plan out before the change and the 5 hour limits felt too restrictive then even for light/medium work, especially cause of the whole peak and off-peak thing: https://blog.kronis.dev/blog/z-ai-s-glm-5-2-is-a-great-model...
Nowadays, I’d probably go with their Max plan if the rate limits are okay? Anyone using them now?
Oh also unrelated but ZCode was surprisingly good, which is surprising for a tool that came out of nowhere - even some of the critiques in my blog post have been patched out. Sadly they don’t support using Claude Code as an agent so can’t use it like Paseo or Kepler or Agent Orchestrator.
- Gecko4072 - 43938 sekunder sedanPeople familiar with the topic, how will models continue to get better? Post training it seems? Labs have already used up internet-scale data, so are there any limits to architecture improvements and post training or can we expect this trend to continue? ByteDance is training a 10T-parameter model. Here, GLM 5.3 outperforms models 3-4x its size of roughly 700B, so parameter count doesn’t seem to be a direct correlation anymore.
- anana_ - 44291 sekunder sedanWhat a week for AI model releases
- mraza007 - 43090 sekunder sedanSuch an interesting times we are in,
We just had amazing releases this past two months
kimi k3, glm5.3 qwen3.8 and now glm5.3
These open models are getting really good
- moinism - 31658 sekunder sedanGoogle: Here is the next iteration of our flash model series, with a discount. please use. thx.
Z.ai: Here is our next iteration, neck and neck with Fable/Sol. weights releasing in two weeks.
- jamesponddotco - 10172 sekunder sedanIs there a plan somewhere that gives access to Kimi K3 and GLM-5.3? I was thinking of testing both to run security reviews of my code.
I know OpenCode Go has both, but their limits seem kinda low, so I'm not sure how feasible it is to run such a task with them.
- CuriouslyC - 13178 sekunder sedanThese results look pretty good, given the smaller model size and the GLM family's historic robustness. Cheaper than Kimi and more robust than DeepSeek. The question in my mind is if you're going cheap, are you going to stop here or go all the way down to DeepSeek Flash?
- - 10399 sekunder sedan
- alienbaby - 24048 sekunder sedanOne htought I had; if The chinese allow unfettered access to cyber capabilties while th US does it's best to neuter it's model releases, from China's point of view they have the US all tied up in knots dealing with problems they don't give people the tools to solve. China giggles as it watches the US under threat from people using it's models. The US is restricting citizens from owning this particular kind of weapon, while China is handing it out to the wrolds citizens freely. It feels like the US would only come out worse overall?
- jameshart - 10444 sekunder sedanSo is ‘cyber’ just short for ‘cybersecurity’/‘cyberwarfare’ now? That is not what cyber used to mean…
This is like when ‘crypto’ started meaning cryptocurrency.
- ikari_pl - 9469 sekunder sedanSuch a smart model and didn't warn them how confusing the headline is to anyone who understands what "cyber" means?
- andai - 13387 sekunder sedanWe got nukes capable of having existential crises, before GTA 6...
- maxdo - 18058 sekunder sedanThey just ignore in their benchmarks opus 5 for some reason :) also grok 4.6 . I wonder why
- scottfits - 9474 sekunder sedanwhat i appreciate most about this post is the level of transparency in how they built and scaled an RL pipeline. my friends at the big labs are so cagey about everything, and Zai is just putting out a great crash course for free.
- newyankee - 45044 sekunder sedanA flood of releases today, really difficult to make out for someone who does not use or test all these models on complex real world use cases as to how people decide which ones to use (besides price)
- joshk401 - 43710 sekunder sedanLove these open source models keeping close source models honest.
- bertili - 43474 sekunder sedanMusk: Open Chinese models will rival Fable 5 in Q1 2027
JieTang (Founder of Z.ai): It won't take that long
- dimgl - 43100 sekunder sedanI was extremely impressed by GLM 5.2, although you could definitely _feel_ it was a bit behind Opus 4.8 at the time. Eager to see where GLM 5.3 is at.
- himata4113 - 7771 sekunder sedanThere goes the last argument that anthropic had. I think beyond this point we're entering the 'dark scary world' that dario predicted which in fact result in things going on as usual. Really, the amount of fear mongering is astonishing.
Hopefully they will drop it all together and focus on making models that are useful for everyone like their original mission was instead of playing games with politics.
- Havoc - 34711 sekunder sedanWohoo. Congrats to team. Been using 5.2 for a while for hobby use and it's been solid - smart enough for my needs & I'm on a grandfathered plan.
Nice to see a commit to open weights straight off the bat
- matheusmoreira - 24298 sekunder sedanMeanwhile, my OpenAI TAC application lingers in a total limbo. I suppose I'll switch to this at some point.
- Ruca_AI - 13124 sekunder sedanSame base model, this much improvement just from post-training is kind of insane.
Really curious to see how GLM-5.3 performs on messy, real-world repositories once the weights are released
- rob74 - 36272 sekunder sedanI'm not that up to date with the latest AI developments, but I noticed that this article seems to use "Cyber Capabilities" as a shorthand for the model's ability at cybersecurity tasks? Is that now an established expression, same as "crypto" now refers to cryptocurrencies rather that cryptography? Because "cybernetics" actually means something different (yeah, old man yelling at clouds, I know)...
- jadbox - 14053 sekunder sedanNo API yet? I don't see it on OpenRouter yet.
- maxloh - 44760 sekunder sedanNo Hugging Face link yet. I wish they would release it under a true FOSS license.
Kimi and QWEN are now moving on to a restricted-usage license, which, although is still better than the proprietary American models, is a step back from the open source Chinese LLM culture.
- quantumwoke - 44146 sekunder sedanFeels like Fable's edge ended up just being long horizon task scaling, which post-training seems to achieve as seen here. Wonder what the next frontier is? Improvement in specialised tasks or computer use?
- tw1984 - 43025 sekunder sedandario must be writing another angry essay arguing why his closed model AI is too dangerous to be used by others.
- kashif - 32369 sekunder sedanUnless its multi-modal and can deal with screenshots - its not really usable for a lot of coding use-cases.
- tmsh - 41815 sekunder sedanIs post-training magic just overfitting to benchmarks?
- postatic - 30516 sekunder sedanLook, GLM, Kimi, Deepseek and Qwen should just join forces and come up with THE model that will beat the frontier lab models even just for the benchmaxxing perspective - all just to create hype and chaos to derail the trillion IPO conversations surrounding OpenAI and Anthropic.
- peiyan_wang - 39081 sekunder sedanCan't wait to see it in practice.
- mostlyk - 44964 sekunder sedanIncredible numbers, will have to wait and see how it actually performs. The timing of GLM updates are always suprising
- adrian_b - 36787 sekunder sedan> The model weights of GLM-5.3 will be publicly available soon in two weeks.
- SwellJoe - 42921 sekunder sedanThey're taking security seriously with this one, with their own disclosure page, like Anthropic did for Mythos. https://cvd.z.ai/
- - 15815 sekunder sedan
- Jacopos311 - 27504 sekunder sedanThis looks very interesting indeed!
- yogthos - 15588 sekunder sedanI'm so glad I managed to get their subscription when it was on sale for 250 bucks a year back when it was 5.1. Back then it was just ok, but after 5.2, it's become my main workhorse. And 5.3 is looking fantastic.
- - 16120 sekunder sedan
- aizk - 41771 sekunder sedanThe model releases just don't stop!
- cubefox - 38788 sekunder sedan> Open Source: We will release the weights in two weeks after launch, once safety evaluation and hardening are complete.
What safety evaluation? What safety hardening? They already evaluated it and found it to be highly capable at exploiting security vulnerabilities. So we know it is not "safe", and they don't seem to plan to do anything against it. What could be more dangerous than hacking? Biological weapons research? I don't think Chinese labs are doing anything against this either.
- tw1984 - 44280 sekunder sedanjust imagine the world without these open weight models - we'd probably have to reverse mortgage our homes to pay for tokens to those trillion $ companies to have access to their models.
- ofjcihen - 25227 sekunder sedanThe capabilities of open models approaching or meeting that of SOTAs is good in every way except for our short-sighted economic reliance on their success (in the US at least).
- smurf9852 - 23010 sekunder sedan" a judge agent then attempts each task to verify that it is actually solvable "
I understand you need to verify the goal is achievable. But if the judge agent has the same goal as the training agent (solve), and both are of the same model, then aren't the judge and the training agent doing the exact same thing? What is the point then? Can someone explain this to me.
- petesergeant - 32821 sekunder sedanTheir own hardness (ZCode) seems to be a GUI, which doesn't work for me. They say they support other harnesses. However, it seems like I can inject the plan into other harnesses, like Claude Code[0]. Does anyone who's been using GLM models for a while have a strong feeling for if it does better in some harnesses than others, or should I just use my favourite harness?
- bsenftner - 24727 sekunder sedanSo, "cyber capabilities", whoa there horsey, what the fuck is that? Are we making up words or are you trying to court the black hat crowd?
- petesergeant - 32292 sekunder sedan[total rewrite: their subscription code is buggy. It takes a while for paid subscriptions to show up, and for upgrades to take effect. Original comment was whining about this]
- MrBuddyCasino - 42884 sekunder sedanAn I the only one who was disappointed with GLM 5.2 after all the hype? It was thinking forever and sometime just stopped mid task.
- steffi_oliver - 3881 sekunder sedan[dead]
- Satoshi_Bro - 15334 sekunder sedan[flagged]
- zaj00l - 11319 sekunder sedan[dead]
- jocelyner - 37824 sekunder sedan[dead]
- libertas_quae_s - 29962 sekunder sedan[dead]
- bigwheel - 23322 sekunder sedan[dead]
- felixlu2026 - 26064 sekunder sedan[dead]
- maxkim869 - 22542 sekunder sedan[dead]
- Culonavirus - 36486 sekunder sedan[flagged]
Nördnytt! 🤓