Claude Opus 5
- postalcoder - 10719 sekunder sedanI think the most important thing here is not absolute performance. It's that organizations now have access to a Fable-ish model without Fable's 30-day data retention requirement[0].
> "Consistent with prior Opus models, Opus 5 does not have data retention requirements for general access."[1]
On the Opus model release page, the reason why Fable doesn't have an ARC-AGI score is because of that retention policy[2].
0: https://support.claude.com/en/articles/15425996-data-retenti...
- jjcm - 6307 sekunder sedanDoing testing with it now, specifically for image->html conversion.
Previously Fable was the best at this, followed by Gemini 3.1 pro (a surprising #2, but Google has great vision models).
Opus' results seem to be more accurate than Fable, following the design source of truth better.
Example results:
Design source of truth: https://image.non.io/73e239a3-880f-4793-b65f-4810be2d9378.we...
Opus 5 build: https://html.non.io/solaraOpus/
Fable 5 build: https://html.non.io/solara/
Note the buttons - for fable they're pill buttons, opus got the rounded rectangle nature of them. Opus' images are closer to the source of truth as well (both LLMs were provided with image gen capabilities for the assets).
Running more tests now, but preliminary results are saying this is indeed better than Fable in some areas. Crazy.
- rb2e - 12050 sekunder sedanhttps://www.anthropic.com/news/claude-opus-5 - A blog post for those not wanting to go through a 190ish page pdf
- paxys - 11705 sekunder sedanLooking at all these releases it’s not a surprise that model routing is the fastest growing segment in AI right now.
There are 10+ LLM companies, each with dozens of models of different modalities, each model with multiple size variants, then different “thinking” levels, then agentic modes, “pro” modes, a “fast” option, standard vs flex vs batch execution. And of course each end combination has a different input/output/cache token price.
Companies that say “give me a prompt and I’ll route it to the most ideal and cost effective model and setting for you” are capturing a ton of value from a gap that model developers don’t seem to understand exists.
- deet - 1916 sekunder sedanI compared the writing style of Opus 5 vs Fable 5, and Opus 5 continues many of the "Claude-isms" of its 4.8 predecessor in a way that Fable broke away from.
Opus 5 still uses "carry the argument", "worth stating plainly", ", and the trap", "The X matters more", the use of "move"
We need an "annoying English" benchmark.
- Fable 5 Max: https://gist.github.com/deet/3d97f854b48eac6658d642fa18bb24d...
- Opus 5 Max: https://gist.github.com/deet/1a43693a732dfccb4d0d914bfc42692...
- nerdsniper - 11492 sekunder sedanEdit: It was pointed out to me that Opus 4.8 got "21%" for successfully fully completing ~1-in-5 tasks, but also got "55.7%" for obtaining significant partial credit on some of the ~4-in-5 tasks it could not fully complete.
---------------
Why does Anthropic say here that Opus 4.8 scored 55.7% on OSWorld 2.0 benchmark, but the paper published by the authors of OSWorld 2.0 say they achieved a benchmark of ~21% with Opus 4.8? [0]
That's a huge gap, considering that the paper was published just 2-4 weeks ago.
I understand that the benchmark authors have an incentive to publish lower numbers (to show that the benchmark has potential longevity) and that Anthropic has incentive to publish higher numbers, but the other models seem pretty inflated as well. The benchmark authors shows GPT-5.5 at 14%, and Anthropic shows GPT-5.6 Sol at 62.6%.
Is there any reasonable explanation for this? Do all the other benchmark numbers need to be sanity-checked as well? Are SOTA benchmarks really this difficult to get consistent, replicable results within a reasonable range of tolerance/variability? Can these benchmarks be compared from one paper to another, or are they only valid to compare intra-paper results?
- HyperL0gi - 11628 sekunder sedanIsn’t it just hilarious that a model that seemed so superior to Fable but didn't get doomsay marketing from Anthropic got released without any issues? In theory, this was supposed to be AGI level according to Anthropic, yet here we are, just a normal Friday.
- 6thbit - 10420 sekunder sedanTheir communication is confusing. They say "Opus 5 is not more capable overall than Fable 5", but their blog post proceeds to list how much better Opus 5 is than Fable 5 on __most__ benchmarks listed.
Then system card goes on to "Its AI R&D capabilities are comparable to those of Claude Mythos 5", which is supposed to be fable minus restrictions.
- Dibes - 11139 sekunder sedanI'm not sure what to make of this graph[0]. It shows medium as the most effective thinking mode by far for frontier code.
It's the only case that I saw going through the system card where more reasoning effort meaningfully negatively impacted the resulting eval. I know sometimes max efforts show a small dip, but this is substantial. I wonder why in the world that is?
- theHocineSaad - 6597 sekunder sedanOpus 5 is considered the most intelligent model[0], while it's half the price of Fable 5[1], and Anthropic is still positioning Fable 5 as the most capable model[2].
Is it because maybe Anthropic engineered Opus 5 to work well on benchmarks and didn't do the same thing to Fable 5, or is there another reason?
[0]: https://artificialanalysis.ai/#intelligence
[1]: https://platform.claude.com/docs/en/about-claude/pricing
[2]: https://platform.claude.com/docs/en/about-claude/models/over...
- acmnrs - 11674 sekunder sedanFrom the prompting guide<https://platform.claude.com/docs/en/build-with-claude/prompt...>:
> Claude Opus 5's default user-facing responses run longer than prior Opus models'.
The benchmarks do show Opus 5 as slightly more expensive than 4.8, although the scores are much higher.
This still feels like a step in the wrong direction, though, especially with OpenAI making so much progress with the efficiency of their models. Fable's token efficiency made it seem like Anthropic would start following OpenAI's approach but that doesn't seem to have carried over to their other models.
- not_a9 - 11736 sekunder sedan> Opus 5’s safeguards match those of Claude Fable 5’s, with one change: it now permits source-code vulnerability discovery at all access levels. This means that the model can support defensive cybersecurity work while still blocking vulnerability discovery in compiled binaries, which is more commonly used offensively.
Okay so it’s worse than Opus 4.8 for my purposes I guess?
- atraac - 11997 sekunder sedanGreat that there's a new model but they could fix their existing infra. We're considering dropping our Claude Team sub cause it's unusable recently. Constant bugs, dropped sessions, issues switching models, http errors. It's becoming ridiculous
- wuhhh - 1205 sekunder sedanIt really feels as though my 20 year career as a front end developer is coming to a very abrupt end; at least as I have know it these past two decades.
- abroszka33 - 11413 sekunder sedanWhat's the point of 150 pages description of a model that's going to be replaced in a couple months? Who even reads this? I know it's cheap to generate text with LLMs, but this is just noise at this point.
- ianberdin - 4775 sekunder sedanPelican svg: https://playcode.io/blog/macbook-svg-benchmark#model-claude-...
It creates the MacBook svg way better than 4.8, yet only fable can make it perfect without visual defects. Results similar to Kimi K3.
- ealready_value - 11311 sekunder sedanI've yet to understand why they call a 190 page PDF a "card". Calling something a card invokes a small, quick rundown of pertinent details, not every single possible detail.
- Sol- - 11705 sekunder sedanHow does it perform on HuggingFaceExploit bench? Suspiciously absent, so not sure if I can take the model seriously.
On a serious note, I hope they improved their extremely sabotaging and unspecific bio safeguards, which prevented Fable from being used in any codebase that ever so slightly grazed medical terminology or data and made me switch to 5.6 Sol.
- adamhowell - 533 sekunder sedanOpus 5 Pelican SVG: https://pelocan.ai/drawings/ese0s599
- thewebguyd - 11831 sekunder sedan> Opus 5’s safeguards match those of Claude Fable 5’s, with one change: it now permits source-code vulnerability discovery at all access levels. This means that the model can support defensive cybersecurity work while still blocking vulnerability discovery in compiled binaries, which is more commonly used offensively
Why can't they also allow Fable to do so also? Why is source-code vulnerability discovery limited to a lower capability model? If Fable and Opus have the same safeguards, except for this one change, I see no reason they can't also allow this for Fable.
- yewenjie - 10277 sekunder sedanWait, 30% on ARC-AGI-3! I definitely didn't expect that jump so soon. Are there any rumors of what they are changing in architecture that is leading to this?
- pyridines - 11351 sekunder sedanThe wording in this post seems much more... restrained? than usual. Maybe Anthropic is afraid of exaggerating the capabilities and consequences of their new models to avoid government scrutiny and sanctions.
> we’ve intentionally avoided training Opus 5 on cyber tasks [...] it remains substantially behind Mythos 5 on the exploitation of those vulnerabilities
I wonder if Anthropic would still intentionally nerf their models without the threat of government intervention.
- cebert - 11719 sekunder sedanI am very confused about what the difference between Opus 5 and Fable 5 is now. What is the purpose of having two models that are so similar? The main differences I see are cost and marginal capability, according to the Anthropic-provided benchmarks.
- hrpnk - 3783 sekunder sedanThe breaking changes vs. Opus 4.8 are interesting [1]
1. Thinking on by default: On Claude Opus 4.8, requests without a thinking field run without thinking; on Claude Opus 5, the same requests run with adaptive thinking.
2. Disabling thinking is capped at high effort: You can still turn thinking off with thinking: {type: "disabled"}, but only at an effort level of high or below.
[1] https://platform.claude.com/docs/en/about-claude/models/migr...
- artninja1988 - 11546 sekunder sedanThat's a crazy arc 3 score. What do people think of this? Are models actually developing fluid intelligence like what the creators claim to be measuring? Is it jus do to training for it? Is the benchmark flawed?
- guybedo - 7776 sekunder sedanLooking at intelligence vs cost:
- Opus 5 is 10% smarter than Grok 4.5 for 10x the cost. - Opus 5 is a bit smarter than Gpt 5.6 Sol for 2.75x the cost
ref: https://artificialanalysis.ai/?cost=intelligence-vs-cost-per...
- petilon - 6096 sekunder sedanThe naming system is so confusing. Is Opus better than Sonnet? Where does Haiku fit in? How can you tell from the name? I can't keep track of all these names or make guesses from the names. Suggestion for a better naming system: use the words "Pro", "Plus", etc.: Claude 5 Pro, Claude 5 Standard, Claude 5 Fast, Claude 5 Mini.
- trunnell - 3073 sekunder sedanThe chaos appears to be tamed for now.
From the system card [1]:
[1] https://www-cdn.anthropic.com/c5fbac3f0b1280a933ebd26d3cb8bb...The Fable cyber classifier we have previously discussed also applies to Claude Opus 5 , with one notable exception: for Claude Opus 5 , we’ve unblocked vulnerability finding in source code to help our coding customers develop more secure code. If you are a cyber defender and are experiencing blocks on Claude Opus 5 , we are also offering exemptions through our Cyber Verification Program, which will remove blocks to enable activities such as bug bounty hunting and vulnerability research and verification. Enterprise customers can also apply to join the Cyber Verification Program to have mitigations removed to enable penetration testing. - visiondude - 11437 sekunder sedanThe signal here is tokeneconomics are very real, price vs performance is starting to be a consideration even at the bleeding edge labs. maybe a subtle indication scaling is not all that is needed since if AGI was around the corner leading labs would still be incentivized to pour all resources into larger (smarter - or maybe not?) models
- vatsachak - 8927 sekunder sedanGPT 5.6 Sol is the first model I've used where I can trust it to add 100-500 lines of code maintainably.
It's great with Codex.
I still find that LLMs tend to not know how to compose larger ideas but on the scale of small ideas or short form well defined tasks like small scale debugging/performance engineering it's safe to say that they are now superhuman.
- ddxv - 11690 sekunder sedan"Cybersecurity. Opus 5’s cyber classifiers are proportionally less restrictive than those on Fable 5. They allow Opus 5 to find vulnerabilities in source code, but block “binary-based” vulnerability scanning (a method more likely to be associated with malicious actors), penetration testing, and exploit generation."
Nice of them to be more explicit for what is blocked. Will be interesting to see if this is true or not.
Also, a notable lack of mention of open source models. They only compare themselves to ChatGPT.
- modeless - 11103 sekunder sedanWow, 30% on ARC-AGI-3 for $20k total. Huge jump from GPT-5.6's 7.8% at $20k per task. I continue to believe ARC-AGI measures something different and important compared to other benchmarks.
- tekacs - 9175 sekunder sedanSomething fun: on our AWS Bedrock console right now, there's a 'NEW' model called 'anthropic.honey'. Wonder if that's the codename just for this one or in general?
- itissid - 2704 sekunder sedanI found opus 4.8 too agreeable and too wordy(as opposed to codex) and too agreeable. If you are reading documents generating by it was too much. TBH. Fable did a bit better on this. Anyone seen a marked difference with opus 5 on this?
- alasano - 11016 sekunder sedanHalf the price of Fable 5 and useable with 100% of your subscription means roughly 4x the usage using Opus 5, presuming similar token use for solving problems.
Not that they should get credit for giving you only 50% of your plan worth of Fable usage but still.
- luciana1u - 1355 sekunder sedan473 comments in 3 hours. people are speedrunning having opinions about it
- irthomasthomas - 8115 sekunder sedanChangelog - fixed issue where model acts like qwen when prompted in chinese
- albert_e - 11152 sekunder sedanJudging by the pace at which new models are released these days -- it feels like a Windows KB or VS Code patch release now.
Older models must be getting deprecated at the same (or faster) pace. So anything you built 3 months ago is probably going to break soon.
AI solutions need better insurance around model deprecation. Commercial API-only models that complete the full cycle from SOTA / gated-preview to unsupported and deprectated in a matter of months -- is no way to build serious software!
- lucamark - 8375 sekunder sedanBut why GPT 5.6 Sol is so behind on the benchmarks? In real-world projects, it is the best frontier model to me in terms of accuracy, speed and consistency. It can just be compared to Fable 5, but I prefer GPT 5.6 Sol because of inference speed.
I've never trusted on model cards though. I'm sorry.
- consumer451 - 4259 sekunder sedanI have a side project that I always run a simple security analysis prompt on in CC, at each model release. Obviously, Fable 5 would downgrade to Opus 4.8 on any such request.
Nothing since Opus 4.6 has found anything interesting. Just ran it using Opus 5, and it found a genuine issue that I verified. Neato!
- the_lucifer - 11807 sekunder sedanNoticed none of the comparisons mention Kimi K3. Is there a comparison chart?
- williamstein - 9789 sekunder sedan> This means that the model can support defensive cybersecurity work while still blocking vulnerability discovery in compiled binaries, which is more commonly used offensively.
Annoyingly, this is a concrete argument that open source software may be easier to attack.
- vinhnx - 7876 sekunder sedanFor anyone wanting a faster overview, I used NotebookLM to create a brief video summary after going through the system card and announcement blog using a cinematic video overview. Link: https://www.youtube.com/watch?v=SUFBhvQ2tY4. And a podcast companion: https://www.youtube.com/watch?v=nYZTW2snXow
- born-jre - 4762 sekunder sedanIs it me or these have gotten very boring. We have 5 more points on xyzbench or whatever .
- paxys - 5025 sekunder sedanIt’s funny to share benchmarks showing Opus 5 scoring better than Fable 5 across the board and then saying “but it isn’t actually better than Fable 5”. So then what’s the real definition of better? And why post all these numbers if even you don’t trust them?
- bottlepalm - 8539 sekunder sedanPage 151 of the linked system card - did Opus 5 get nerfed to prevent it being better than Fable? The graph makes no sense. Huge decline in coding performance at effort levels higher than medium.
- mulhoon - 5635 sekunder sedanAs a coder, I’ve had no desire to use Fable. In fact I switched from Opus models to sonnet 5 and haven’t noticed any drop in quality on large repos. It seems the gap at the top is very small and not hugely noticeable for backed/frontend. Has anyone else had this experience?
- skybrian - 9331 sekunder sedanLooks like the API price in tokens is same as previous Opus or Sol, double the price of Terra.
Maybe there’s a better comparison than cost per token, but it will be application-specific.
- - 7685 sekunder sedan
- skerit - 10540 sekunder sedanInteresting, they finally support `system` messages anywhere in a chat conversation:
> Mid-conversation system messages are available on the Claude API, Claude in Amazon Bedrock, and Google Cloud. > > This feature is available on Claude Fable 5, Claude Mythos 5, Claude Opus 4.8, and Claude Opus 5. No beta header is required. This feature is not available on Claude Sonnet 5; use the top-level system field instead.
For nearly all models EXCEPT Sonnet 5? That is weird. How old is Sonnet 5 really?
- beydogan - 4065 sekunder sedanmy early and non scientific feeling:
- it has this annoying Opus response style(since Opus 4.7) with bunch of very hard to interpret word salad
- on >xhigh it eats tokens like there is no tomorrow
I don't like it. Since Fable is unaffordable for anything meaningful, I'll stick with Sol for now. I was on Max 5x, saying hi to Fable costs %5 weekly.
- jatins - 11646 sekunder sedanBetter than Fable 5 on all but 3 evals.
Has Anthropic ever mentioned how do Opus and Fable differ? It used to be Haiku < Sonnet < Opus in terms of params. Where does Fable fit in this?
- 6thbit - 10291 sekunder sedan"although Opus 5 shows improvements in its ability to identify software vulnerabilities, it is substantially behind Mythos 5 in its ability to exploit them."
"Opus 5’s safeguards match those of Claude Fable 5’s, with one change: it now permits source-code vulnerability discovery at all access levels".
This is probably great news, but then again, where does this leave Fable as a choice?
- dehugger - 11425 sekunder sedanIs Fable 5 just Opus 5 with some additional long-context management modifications for extended self-directed work? Or are they actually truly different models?
- rad_val - 6560 sekunder sedanAfter Opus 4.8 intelligence really started to matter less and less for the programming tasks I have. If I have to handheld anyway, why would I wait more or pay more?
- vinhnx - 7346 sekunder sedanThe benchmark appears to have a mistake, as Opus 5 and Fable 5 score 53.4% and 53.5%, respectively, for the Agentic Coding row (FrontierCode v1.1). But Opus 5 is the highlight.
- markasoftware - 11128 sekunder sedanSoo most of the benchmarks are better than fable... Is this naming scheme just to avoid getting banned again?
- arjie - 5916 sekunder sedanI wonder when a model will be released that can work in a loop and port Qwen-3.6 27B to run on Tenstorrent P150.
- geooff_ - 11301 sekunder sedanFYI: `/model claude-opus-5` works to use it even through `/model` still tries to serve 4.8
- theplumber - 5716 sekunder sedanThe most important thing is it has the same drama queen mode on safety “guards” like Fable.
- destring - 11341 sekunder sedanGoogle is having their Meta moment where they failed to stay at the frontier
- bovermyer - 9467 sekunder sedanThis stood out to me as a little concerning:
> The model hallucinates factual claims slightly more than Opus 4.8, despite being more accurate overall.
- boc - 9814 sekunder sedanSeems really good so far using it in Claude Code CLI - it gave me a new flag when I asked a question:
"I don't have a reliable way to read that number, so I'd be guessing if I gave you one — and this is exactly the kind of question where a confident guess is worse than none.
What I can tell you is what I actually observe:"
I really like this update - gave me a clear sense of the facts but didn't give me a guess just for the sake of guessing.
One oddity is that it appears to only have a 200K context window right now via CC. Hopefully the 1M version will appear soon!
- tyre - 10441 sekunder sedanI'm interested in benchmarks for Claude Design. There is so much opportunity there and I hope they continue investing in it. It EATS tokens though.
- internet2000 - 6983 sekunder sedanKimi K3 already left behind in the dust. They can't keep getting away with it!!!
- shockembopper - 9038 sekunder sedanI wish these releases came out earlier in the day so I could try them during my work day instead of waiting until the next.
- alvis - 12253 sekunder sedanWhat really impress me is opus 5 is better in alignment than fable 5!
- briandoll - 12383 sekunder sedanVery interesting to see such a focus on cost for performance here
- m_w_ - 12184 sekunder sedanVery impressive headline benchmark numbers. I expected a step change, but not past Fable. That said - it all depends on whether the classifiers make the model unusable...
- korabs - 7592 sekunder sedanSo in benchmarks it's better than Fable?
But they say it's "almost as good as fable"
- pmg1991 - 11747 sekunder sedanSame cost as 4.8 but better that 4.8. Happy to get more efficient model. But is there any reason all companies are releasing models back to back after GLM 5.2.
- arrowleaf - 10738 sekunder sedanI can't find anything about whether this is zero data retention, or falls under their required 30 day retention like Fable and Mythos?
- twothreeone - 11988 sekunder sedanIt starts at page 148.
- skinfaxi - 11244 sekunder sedan
- - 11170 sekunder sedan
- 8note - 9856 sekunder sedanim excited that cad and object=>cad is getting into the test tasks
i guess the next stuff will be tool use for the rest of what cad does in assemblies and simulation?
itd be fun to try to set up a 3d printer as part of a feedback loop, and see what a model can build.
the automated test harness for physical stuff seems a bit beyond reach still
- jakeogh - 1951 sekunder sedanAnyone else not getting chain of thought? Opus 4.8 would show it to me, until around the time Fable came back. Now I dont see it with 4.8/5.0 or Fable. Not having it makes catching mistakes harder.
- arseniitrut - 2292 sekunder sedanatp, is it the end of fable 5 era?
- himata4113 - 12001 sekunder sedanRather interesting that this makes sonnet 5 look even worse! There is no reason to use sonnet over opus with low or no reasoning at all.
- whatever1 - 11033 sekunder sedanWhere does this leave Fable? I am confused.
- yusufozkan - 12257 sekunder sedan> arc-agi-3 30.2%
wow
- urams - 10855 sekunder sedanSo Opus 5 is basically "distilled" Fable? The benchmarks look often better than Fable.
- doctoboggan - 9718 sekunder sedanAccording to these charts I should switch from Fable to Opus in Claude Code now?
- stevefan1999 - 9311 sekunder sedanWhere's the reset...
- mcast - 12121 sekunder sedanInteresting timing to release this on the same day Jensen makes a statement on open source AI.
- arj - 7301 sekunder sedanOn a Friday, I'm out of tokens ;-)
- inshard - 11430 sekunder sedanArc AGI score is astounding
- 6thbit - 10012 sekunder sedanAnyone has an insight into how much money labs are putting into benchmarks?
Just Arg-AGI-3 is quoted above 20K USD and footnote says average of 5 runs (!!). Likely just a drop in the bucket to the training budget but still..
- spstoyanov - 11556 sekunder sedanSo same as Sol? I guess I’ll see which one is more token efficient.
- mkurz - 11500 sekunder sedanWhere is the pelican?
- - 8622 sekunder sedan
- taf2 - 11742 sekunder sedaneager to see how it benchmarks on https://deepswe.datacurve.ai/
- toephu2 - 8420 sekunder sedanHow does it score on DeepSWE?
- Eldodi - 11863 sekunder sedanModels benchmarks start to get saturated again!
- _pdp_ - 5476 sekunder sedanWake me when they deliver Opus 4.8 level performance for $5 per million tokens.
- throwaw12 - 11661 sekunder sedanis coding and engineering solved yet?
- abc42 - 7669 sekunder sedanAre we getting to singularity or something? This seems a bit crazy.
- zuzululu - 12008 sekunder sedanso almost fable 5 with 50% cheaper cost? sign me up
- - 11649 sekunder sedan
- hmontazeri - 10569 sekunder sedanHonestly if reached a level of coding that sonnet 5 is more than enough for my needs as assistant/agent I don’t need long Horizon stuff…
- mihau - 10620 sekunder sedan30% on ARC-AGI-3
- LoganDark - 7889 sekunder sedanThese cybersecurity safeguards are really annoying. There are ethical reasons to reverse-engineer and binary-patch software; for example Rewind got acquired by facebook and, as a gift to all their customers, implemented a killswitch in their software to ensure it will eventually stop functioning. I kept using a version without the killswitch, but the macOS 27 update killed it, and I needed binary patching to fix it. I should be allowed to repair software I purchased (I did purchase it like a month before they sold out), but unfortunately this overlaps significantly with cybersecurity.
- throwaway23597 - 9054 sekunder sedanThe truth for me at least is that these models became "good enough" around Opus 4.6. I feel like further capability improvements, "step changes" like we saw with agentic coding, aren't necessarily going to come from the model. I think the next crown goes to whoever can figure out the right scaffolding so that these models can be inserted into your organization.
Maybe I'm wrong and Opus 5 is a real unlock?
- simianwords - 9849 sekunder sedanMy thoughts: fable is the bigger model. Opus is distilled from it but since it is smaller it doesn’t need the online classifiers. Though benchmarks show Opus to be near Fable level, I think it’s nowhere near Mythos (fable without safeguards).
- sbochins - 10878 sekunder sedanQuick read is that this is more capable and cheaper than 5.6sol. Same price for input tokens and $5 cheaper per mil output tokens.
- alex1138 - 1867 sekunder sedanAm I misreading anything or are comparisons to Fable (and/or Mythos although AFAICT it was only a crackdown on Fable) always going to be a bit missing the mark now due to what the Trump admin did?
- ismailmaj - 6426 sekunder sedanI'd pay good money to see OpenAI "oh fuck" war rooms.
- alvis - 12621 sekunder sedanhere we go
- sudohalt - 8389 sekunder sedanAnthropic is no longer a good model company in my mind, they are optimizing for an IPO and padding themselves on the back for being the next Aristotle. They're so far up their behind they don't realize how s**y their products are, and their research team hasn't done anything ground breaking in probably over a year other than release "scary" reports.
- mrcwinn - 9009 sekunder sedanCan someone help me understand something? I thought Fable was such a miraculous leap forward in capability. But now it seems Opus is basically on par with it, and in some cases (computer use) far exceeds it.
- StrauXX - 11349 sekunder sedanThe benchmark table is manipulative, borderline lying through statistics. In every line the top performing cell is marked red. Except the line where Sol leads, there it is marked in gray.
- wyre - 11521 sekunder sedanIn the wake of OpenAI’s model hacking Huggingface it’s interesting how the first quarter is entirely about how good Opus 5 is at hacking and finding vulnerabilities in software.
- mnky9800n - 9466 sekunder sedanYay just in time for neurips lol
- justindotdev - 11705 sekunder sedan> . Opus 4.8 served as fallback on safety-classifier refusals for Opus 5 and Fable 5.
ffs just keep it man.
- - 425 sekunder sedan
- gorkemyildirim - 2311 sekunder sedan[flagged]
- vilmire - 8475 sekunder sedan[flagged]
- nee_oo_ru - 9770 sekunder sedan[dead]
- emunova - 6765 sekunder sedan[dead]
- Nevin1901 - 11874 sekunder sedanExcited to use it? Will we be seeing Haiku 5 next? /s
- datakan - 12223 sekunder sedan> Claude Opus 5 is not more capable overall than our most capable general-access model, Claude Fable 5
Ok then so what's the point?
- aleenz1102 - 11491 sekunder sedanthis claude fable & opus 5 should be cheaper and can compete in pricing with chatgpt latest models
- midnightbobarun - 11632 sekunder sedanIt looks great, and those coding benchmarks are impressive... now if only it didn't come out just days after I let my Claude subscription expire :')
- TheJCDenton - 8270 sekunder sedanI think it's the first time Anthropic release a model without any meaningful disruptions while doing it
Nördnytt! 🤓