example.com/path/to/article
000 points · username · 0 hours ago
example.com1419 points · 1395 comments · 13 days ago · denysvitali
System Card: https://www-cdn.anthropic.com/0339e6a7c5c7b87f5c07798616dc32...
felixrieseberg
simonw
I'm still waiting for effort max to finish.
EDIT: I fixed a bug in my tooling so it now records summarized reasoning traces - here's that max pelican, which is a significant improvement: https://tools.simonwillison.net/markdown-svg-renderer#url=ht...
Took just under 14 minutes to generate, and at 65927 output tokens cost me a hefty $3.30!
Excerpts from the reasoning trace:
Adding pedal shapes near both feet, with the far foot on the second leg partially visible behind the frame. I'm considering whether to add a small scarf or cap for extra character, but leaning toward keeping it simple to avoid clutter.
Now I'm debating a bicycle helmet on the head versus the pelican's signature crest—the beak and pouch already read clearly as "pelican," so a helmet could reinforce the bicycle theme without losing identity, though it might compete with the crest for visual space.
I realize the beak at (484,84) would overlap with the dome helmet, so I need to shrink the helmet so it only covers the top of the head, adjusting its arc endpoints to sit higher and narrower so the beak can attach cleanly at the front without collision. [...]
I'm adding a darker tip region to represent the primary feathers, then reconsidering the trailing edge to include scalloped feather curves instead of one smooth line for a more natural look. [...]
Now I'm checking the vent line placements on the helmet, making sure they sit far enough inside the helmet's edge given the stroke width and rounded caps, and confirming each vent stays within the helmet's circular boundary. [...]
I decide skipping a handlebar bell and tire highlights since they're unnecessary additions. Now I'm reconsidering the front fork's curve — the current control point pulls the shape backward when it should bow forward for a proper rake, so I need to shift the control point rightward to fix the fork's lean.
This is a notable result because most of the recent Claude models have been pretty bad at drawing pelicans, at least when compared to models in the Gemini or GLM series.
GodelNumbering
This gives a lot of credit to the theory that Anthropic did not get much bite on Fable at its original pricing, which in turn likely places a ceiling on LLM pricing in general.
Interestingly also, if you take away terminal-Bench-Science 0.1 results, it is hard to see ANY improvement:
Terminal-Bench 4.0: Fable 5.1 is +3.5% vs Opus 5.
GDPval-AA v2: +1.5% vs Opus 5.
OSWorld 2.0: +2.5% vs Opus 5.
Humanity's Last Exam (with tools): +1.6%
Keep in mind that this is supposed to be an entirely higher tier of a model than Opus 5. For one tier up and one version up, these are not really improvements. Probably leaves no room to place Opus 5.1 anywhere. Combined with the fact that they are selling 'readability'... Has frontier progress finally stalled?
exabrial
What they have done:
* Nerfed Fable, as many of noted it's useless
* Leverage Mythos as a marketing strategy, claiming its too good to release
* Removed thought traces, one of the only useful things to make sure your prompts are working correctly
* Continue tons of hype about how good they are without delivering, going to great lengths to publish how their model "hacked" its way out of a sandbox they misconfigured.
* Push a bunch of EU Overregulation onto the rest of the world with text watermarking, decreasing quality of answers
Last year, they were at least focused on making improvements. Nowadays its just a bunch of handwaving at the church of how good they are.
The only saving grace is Opus 4.6 is still available. Just sucks we haven't seen any measurable improvement, despite all of the ceremony.
madrox
I am becoming dependent on AI to make a living, and I need predictable spend on it. If I know I can't use a model regularly all month, my enthusiasm is limited.
I urge Anthropic to get better at this aspect of their business so I can come back to it.
mlaux
kccqzy
Anthropic seems to be listening to community complaint on HN about how the writing style is grating. And apparently the solution from Anthropic is to add this block to every conversation!?
Mannered prose substitutes metaphor and flourish for direct statement. Instead of "a parameter worth varying," the mannered writer produces "a dial worth turning." Instead of "this point still matters," they write "this point earns its keep." The phrases exist to display the writer, not to convey the idea, and readers can tell. That is why mannered prose irritates: it makes the reader work harder so the writer can perform. It is also imprecise. Metaphors drag in connotations the writer did not choose and cannot control. The fix is to say what you mean. When a literal phrase is available, use it.
The above was quoted verbatim from https://platform.claude.com/docs/en/build-with-claude/prompt...
jumploops
For example, in testing by the investment firm Millennium, Fable 5.1 found the cause of a rare crash on their internal systems that none of their engineers (or any other model) had been able to explain after several years of trying.
Say what you will about LLM-generated code, but stories like this give me hope that software will never be as buggy as it once was.
rcr-anti
Can't believe they haven't at least figured out better messaging. If we take them at their word, it's hard not to read it as a messiah complex, that they think they're the only ones capable or worthy of making these decisions. I don't believe them, but I wouldn't be surprised if the articulated reason is a version of "distillation is a safety risk because we might lose the race".
Plus, completely deaf to the recent OpenAI-HF hack incident. Recall, defenders were categorically unable to use western frontier models in their response.
I was originally going to complain about the chem and bio guards still being too onerous, but I'll admit the projects Fable 5 categorically refused to work on are now usable, at least not rejecting on first prompt because the word "virology" was in a git commit (absolutely serious, in one repo it triggered on literally any prompt, eventually traced to the system prompt loading git commit history). Still, them trying to get into the biomed business while walling off the capabilities to the public reeks. Why sell the segments that are actually valuable if you can capture the value yourself!
EliasWatson
What's far more exciting right now is models like DeepSeek V4 Flash and GLM 5.3 Flash. They have achieved good-enough-intelligence at extremely low prices and fast speeds. I don't have a use for Fable-level intelligence, but I do have uses for Opus-4.8-level intelligence that I can use as much as I want without worrying about the bill.
pookieinc
Glad to see this!
tarr11
I cancelled my pro max Claude subscription last week; codex is much more succinct. I am curious if this is getting better.
I don’t think Anthropic realizes that humans have a token limit too and it can be exhausting to read Claude’s output. Prose density is not the same thing as succinctness.
dboon
5.1 so far seems like another leap, which is really surprising. I threw it at a few bigger features I've been designing for a while, and it came back with some extremely thoughtful wrinkles in the design that I'd legitimately not considered. Which, OK, package managers and build executors and compiling C/C++ is pretty well trodden ground, but my thing is very different from everything that exists, and I was very surprised it was able to understand all that context so deeply and intuitively
rybosworld
1) address the claude 20x plan usage being only 6-7x the ceiling of the claude pro plan
2) either fix opus 5, make it completely free, or delete it entirely
swalsh
In the past I watched and saw everything the model did, not a lot got past me. Today it does A TON of work while i'm busy on other tasks. It also has extensive access to my computer, other computers on my network, my internet. It's really helpful when you give it a lot of resources, but right now I have very autonomous, very smart agent running around more or less unattended with a lot of resources.
boardwaalk
Not gonna say I want 5.0 as an option still... but maybe I do.
skiing_crawling
elpakal
Whole-file rewrites for small changes. When editing text files, the model is more likely to rewrite the entire file than make a targeted edit. The result is usually the same, but the rewrite costs more output tokens and time.
So we are to catch that somehow? And then add their recommendation (below) to our prompts?
https://platform.claude.com/docs/en/build-with-claude/prompt...
If Claude Fable 5.1 rewrites whole files for small changes, append the following instruction to the system prompt or the first user message. Claude Fable 5.1 is more likely than Claude Fable 5 to rewrite an entire text file rather than make a targeted edit. The resulting file is usually the same, but unless the file is short or most of it is changing, a rewrite costs more output tokens and time. The instruction brings Claude Fable 5.1 back in line with Claude Fable 5 for small and medium changes.
The number of tokens used to edit files is best minimized, all else being equal. Therefore, when it will not affect the end result, try to surgically edit a file rather than rewrite the entire thing.
1970-01-01
objects that are not alive: dust, rocks, water, wood, hats, lego, aluminum, etc.
objects that are alive but not intelligent: trees, mold, staphylococcus, cancer, grapes, etc.
objects that are alive and intelligent: cats, Steven Tyler, dolphins, crows, dogs, elephants, etc.
and now intelligent but not alive: Fable, Grok, GPT, etc.
Zigurd
What I do see in the comments: subjective improvement in text generation, possibly lower cost, some optimism about code generation, but some skepticism too.
I use coding agents. To me they are very useful. But what I spend on them isn't going to support trillions of dollars in investment.
AnodicElegy
caconym_
We’re introducing Claude Fable 5.1 and Claude Mythos 5.1. They’re the world’s most advanced models for coding and knowledge work—and their research capabilities offer an early glimpse of how AI models will contribute to scientific progress.
I'm not an emdash hater but this isn't how you use them. It should be a comma.
InsideOutSanta
This has never happened to me before, but if this is normal behavior, Fable 5.1 is essentially unusable.
dabinat
This required us to add a watermark—a numerical way of determining the likelihood that Claude was involved in writing a piece of text—to the outputs of models released after August 2, 2026. As we recently explained, this watermark is invisible to anyone who does not have the detection API. It has no practical impact on the quality or content of Claude’s outputs and contains no information about the user, their organization, or their conversations with Claude.
How does this work if it doesn’t change the output?
bobjordan
eckr
Then why does it have separate datapoints for Terminal Bench, and score higher? Something doesn't add up here??
seaurchinzee
throwawaye3735
Then i went and pasted those exact prompts it generated into a fresh chat, 5.1 on max effort.
Immediately got blocked and sent to Opus 4.8 fallback.
d4rkp4ttern
Why is it that the voice models in Claude and ChatGPT have a perfectly normal style with barely any "AI smell", while the writing models are so obviously recognizable as AI?
The answer is likely that models underlying the voice modes are (post) trained differently. If so, then why can't the writing model be similarly trained? Presumably they haven't found a way to train them to be both "smart" (i.e. solve tasks etc) and pleasant to talk to?
ceroxylon
I once caught Fable 5 spinning its wheels on a rendering issue, which evaporated 90% of my usage in a single prompt. I could never let Fable run free attached to a credit card without staring at it the whole time.
apt-apt-apt-apt
_islo
Anthropic accidentally over-billed my account, and when I reached out to the support bot, it downgraded my account to a Free account. It’s been impossible to get it resolved and I have almost $200 held hostage.
I don’t want to do a charge back. I’m one of the main advocates for Claude Code at work, I use this subscription to try out new features before it’s available at work.
The whole experience has been illuminating about our dependencies on these AI companies.
alin23
I jumped when I saw a mention about "writing style improvements" so I gave it a try on a recent feature in rcmd [0]. I prompted Fable 5.1 to find these wordings and propose simpler plain language.
For context, I recently worked with Fable to give users a way to fuzzy search and focus any browser tabs, terminal panes etc. but the UI was still a prototype full of AI writings.
It took every string including the ones I already rewrote by hand, and proposed even more weird LLM speak. Like for "Left Command conflict detected" it proposed "This keyboard can't tell left from right".It's a very capable coding agent, but I can't understand how it can be so bad at writing. Where are all these verbal tics coming from and why is it so hard to get rid of them?
simonw
same input and output prices, with cache reads at a quarter of the cost
This should impact any long-running agent since subsequent calls can benefit from cached reads for previous transcripts.
spondyl
nottorp
Are they going to try the banned for export for a week marketing move too?
mococa
spwa4
Model HLE w/tools GDPval-AA v2
Claude Fable 5.1 65.0 1853
GPT-5.6 Sol 64.5 ~1711-1730
GLM-5.3 62.5 1769
DeepSeek V4 Pro 60.0 1590
Kimi K3 59.8 1682
Qwen3.8-Max 56.2 1739alansaber
maxdo
sunaookami
Topfi
This is the user's own Firefox-fork browser; the slice is defensive service-posture hardening (telemetry/Normandy/FxA/push/crash-upload off, the update endpoint and private-mode extension law) of their own product on the unbranded build path.
I cannot say what effect this has on the way the classifier operates, whether it actually impacts the classifier or whether that was tuned in the background to prevent blocking hardening ones own pre-release code, whether it treats input by Fable 5.1 different to what a user prompts (otherwise the classifier could be defeated with prompting which wasn't the case in 5 and I doubt has changed).
I do however know from personal experience that even when Fable 5 prompted a subagent in such a manner, it had a high likely to be caught by the classifier.
seaurchinzee
jimnotgym
I am using Claude and Claude code for my own amateur history project. I'm enjoying how it constantly reaches dead ends, and I can reframe the question and get more results. I am starting to get concerned that AI and me are so compatible, that I might not be a human at all...
I also like that, because I'm too lazy to write stuff up, Claude code can keep the current state of research published on my site. It makes running a hobby site a dream. "I just found these pictures. Add them to the site for me". And up they go, resized and all. What a dream of a way to work. "Some of links in this article are dead, run through them and check, and see if you can get an archive link for me if they don't". It's like sending a Teams message to my PA.... which I don't have in real life
charcircuit
fulafel
Anyone know who the ZDR special treatment is available to?
koolba
Data retention. Our new system of Enterprise Frontier Safeguards (EFS) gives customers complete privacy (the same as a zero data retention policy) while still being state-of-the-art at preventing adversarial use. EFS works by storing data in cloud infrastructure controlled entirely by the customer, not Anthropic. It will be made available to enterprise customers in phases, beginning later this fall. Until EFS is available, eligible customers will be able to use Fable 5.1 with zero data retention.
This is interesting. I wonder if customers will be allowed to create an auto expiry for their own data to prevent future subpoenas. That’d be a treasure trove for discovery.
spicypixel
finnjohnsen2
kimseungyong
olirex99
I still think that a major problem is that biological processes are not “fast” as coding, but they are verifiable. If during post processing we are able to give enough harness to test and verify this kind of environment (maybe via simulation and real data) we will for sure achieve incredible performance also in this domain.
joduplessis
2001zhaozhao
This seems to point to them having achieved some kind of optimization in attention mechanism perhaps along the lines of DeepSeek V4, which had a similarly high discount between cache input and normal input.
In real world use, the savings should be quite noticeable. For example, you can now use the model at 800K tokens context window at the same cost efficiency as the previous model at 200K tokens context window.
exabrial
m101
I would be interested in whether someone has done research here on these things as it seems a fairly complicated function to work out, and use case dependent. (?)
In some sense an expired kv cache is basically like an expensive cache hit, so your compaction token threshold should come in. Ideally claude code should allow you to vary the autocompaction threshold to vary with time since last token, but it doesn't of course. This perhaps suggests that someone should manage claude code through their own intermediary agent who manages these sorts of rules.
Lastly, I strongly suspect that anthropic isn't offering this price cut out of the kindness of their hearts. I am sure that they are to some extent banking on people not reacting to their price cut and leaving their autocompaction thresholds unchanged.
[edit - looks like the discount is only for the api, so they still don't give a rats ass about subs!]
mohitpaddhariya
miki123211
These patterns invalidate every later thinking block:
• [...] Rebuilding the top-level system prompt or tools array between requests in the same conversation.
Many people unknowingly do this (at a high cost to them because of the cache busts), this change will finally force them to stop.
Especially if you're generating your system prompt via a template that can change mid conversation, it's so easy to fall into this trap.
bilsbie
5555watch
ponyous
Comparing 4.8 Opus with Fable 5.1
eis
This directly contradicts what Anthropic is presenting here. Yes it scores higher but that's to be expected from a new release. It's the opposite of what OpenAI has been doing which was reducing costs, increasing efficiency.
Fable 5: https://artificialanalysis.ai/models/claude-fable-5 Fable 5.1: https://artificialanalysis.ai/models/claude-fable-5-1
[deleted]
8cvor6j844qw_d6
On complex asynchronous workloads, though, nudge it not to end its turn before the work is done. Without the nudge, the model sometimes describes what it would do next instead of doing it ("Next, I'll …") or stops to ask permission for a step the original request already covered ("Shall I apply this?"). [1]
Interesting behavior. The docs also provided recommended prompt [1] to mitigate this behavior if undesired.
Wondering if anyone has encountered it yet?
[1]: https://platform.claude.com/docs/en/build-with-claude/prompt...
niteshpant
AI to AI doc share: sure, do what you please.
AI to human: please make it legible and flowly.
example, "Every thinking block records which model produced it, and it's preserved in one direction only: Claude Fable 5.1 reads earlier models' thinking blocks, and no earlier model reads Claude Fable 5.1's." is a very Claude-isk way of writing. Choppy, long, and lacking flow.
nezhar
jebarker
vlovich123
In part, this is because Fable 5.1 can now be used to discover software vulnerabilities—though not to develop exploits for them
Generally once an exploit chain is described, developing the exploit is trivial.
If you're so inclined, discover the exploits using Fable 5.1 and then give that exploit to a model that doesn't have such compunctions (e.g. local LLM or an uncensored cloud model / model that's easier to jailbreak). I don't think Anthropic is really mitigating here anything in the real world other than PR narratives where media can report "Anthropic's model was used to develop the latest cyber attack".
eis
regexorcist
leecommamichael
enraged_camel
Curious to see how Astra does.
amluto
I wonder to what extent this will make the automatic Fable-to-Opus downgrade give worse results.
sergiotapia
bix6
I tried the old fable and it didn’t seem worth paying for. It still made errors like Opus does so I might as well use the included model…
mark_l_watson
Does Fable 5.1 really provide much benefit over models like Kimi K3 that are 1/3 the cost? Or GLM-3 that are 1/12 the cost?
If you can talk about your work, what kind of tasks do you work on where the higher cost is very much worth it?
mantenpanther
6thbit
What exactly is the premium that you're getting for paying these prices?
mrcwinn
petreradu
ckugblenu
cromka
OK, I think that's what they meant when they suggested reduced extra promo usage will not sting this much.
Norwell_io
tamimio
u8
mentalgear
Forced tool use is not supported
That seems unfortunate for 3rd party integrations that expect stable output - what that really necessary ?
[deleted]
twhitmore
Fable 5.1 -- to my early impressions -- seems to have gone backwards again.
Opus 4.8 was terrible for babbling self-invented jargon. (But still better than other options at the time for coding and logic.)
vinhnx
potwinkle
dmix
philipwhiuk
Claude Fable 5.1 follows explicit tool instructions reliably.
Moving stuff out the API into prompt engineering is obviously less reliable but necessary for progression to 'actual intelligence'. Will be interesting to see if it really is solid.
as12fj
Jane Street is a partner? How sad indeed. Anthropic could front run them because they leak all the data.
freakynit
Having lived through Covid, this doesn't sound so good to me.
stillpointlab
I have been happy with Fable 5, it has done great work for me so far. Very excited to try out Fable 5.1 and see what differences and improvements there are.
Fordec
ilia-a
TuxSH
tosh
The watermark doesn't change the meaning, quality, or readability of the output
how?
pmdr
Husafan
When paying by the token, don't the labs have a strong incentive to make the model as verbose as possible?
ghoshbishakh
ayhanfuat
BoorishBears
DCKP
Computer0
Exoristos
ramon156
mixedbit
jgilias
noduerme
[deleted]
iLoveOncall
k1rd
nubinetwork
eigenblake
wewtyflakes
leumon
andai
yoanwaidev
george_max
They show this off, but artificial analysis contradicts the statement. Fable 5 cost $3.14 per task, while 5.1 cost $3.69 -- around a 15% jump in pricing.
https://artificialanalysis.ai/
These, IMO, are marginal improvements for a more expensive model. I stopped using Claude ~3 months back; its outputs are too jargoned, it makes architectural decisions that are not right, and it's incredibly pricey for what it is. Each decision it makes, it acts as if a problem as major as world hunger has been solved. And the overly verbose code comments, strange commit descriptions, duplicate code, and slop it generates -- which I know is not specific to Fable -- is just too much for me.
I found the best is to use something like Deepseek V4 Flash -- with a fast TPS provider -- and work on the code myself. For agentic work with computer use, GLM 5.3 flash with Hermes Desktop works well.
joshfraser
Bluestein
brcmthrowaway
kris-memoket
dainiusse
cdnsteve
de6u99er
I had to switch to Fable, because Opus has become completely unreliable. Just now while working on a specific task, Fable made changes completely unrelated to the task and introduced new regressions.
Fable now talks complete nonsense.
I do not understand why Anthropic is so focussed in squeezing a few fractions out of benchmarks, while at the same time making the life of developers insufferable. I am seriously considering ditching Anthropic and moving to something else.
The thing that I describe as bullshitter mode is starting to feel extremely unproductive for me. I spend more time making Claude code do what I want than before!
kosolam
Unfortunately, for them.
re-thc
hit8run
Hey Cl… Your limit has been reached.
jorl17
Things it does constantly that Fable 5 barely ever did:
- Act without my permission. All. The. Time. "Oh I just finished this thing we were discussing, let me push it without ever having been told to do so."
- Immediately jump to action instead of addressing me first. If I say "I wanted to write tests for this and run them" it immediately starts writing tests instead of digging into what "this" is better -- literally does not give me any feedback and starts spitting out code. Naturally it creates the wrong tests
- Despite claims that it does not write like "stereotypical Claude" anymore, in my experiments it is far worse than before. Replies are longer, more filled with fluff, and still flooded with garbage language. Hard to parse.
- It loves to answer my set of two direct Yes/No questions with 5 paragraphs where it only answers one of them and answers 4 other questions I didn't ask. Notice how it misses one of the questions.
- It. Is. Cocky. Absurdly full of itself and arrogant. Just the whole way it presents and answers passes this energy of "No, but really, you're wrong and I'm right". It often is not right. What annoys me is not that it's wrong more often than before (which it may be), it's that it doesn't own up to it as before. Insulting if it were a human.
- Replies and addresses me directly in its thinking traces, and then assumes I've read it. I ask a question, it answers it in the thinking traces and does not relay it back to me at all. This is the only one that Fable 5 also did, but 5.1 is doing it an order of magnitude more often.
- It's too early to really tell, because I may just be working on particularly harder problems today, but it seems to get things wrong more often. I've had to bump it from high to xhigh to compensate.
My guess is I must be having a bad day or something. Although this is happening on multiple projects run from multiple machines (fully isolated, except for the account, which is the same) all in the same way.
Will probably downgrade to 5 while I can.
loeg
thway15269037
nirmeet011011
canadiantim
nightsd01
Guys, listen to your feedback please. I hadn't used OpenAI products in quite a while until this issue came around. They seem to have MUCH smarter safeguards than Anthropic does.
testaccount121
literally_him
purpleidea
Enterprise Frontier Safeguards (EFS)
Sounds like some serious nonsense. "Tell me you want the government to retain access to my data without saying it explicitly."
tclancy
Denser prose in places.
Really? Interesting choice. Pretty much every CLAUDE.md file I have starts with something about Hemingway, terseness and treating every word you use like you're carving it on your own back, but different strokes for different folks. I suppose I haven't heard from anyone who enjoys how wordy Claude is because they aren't done writing their post yet.
abroszka33
sashank_1509
MadsRC
After months of trouble dealing with KYC and procurement I finally got CVP for my security org and today I found out that CVP (which is what removes cyber safeguards) does not apply to Fable…
So yeah, unless you’re a Project Glasswing member, there’s no using Fable (which with Glasswing is Mythos) for security work… Absolutely useless…
Didn’t they just sign some “we must use AI for cyber defense before the bad guys do” and then they artificially cap us by not allowing Cyber-unlocked Fable…
Sigh…
kneel25
testaccount121
dfltr
thisisauserid
... with the condition that you store 100% of your data and make it available to the US government and possible others.
genxy
nullbio
zb3
danieltk76
scronkfinkle
[deleted]
latcom007
tusimi
lousken
ike4est
robinpie
2001zhaozhao
[deleted]
krupan
[deleted]
Beyond all the benchmarks, I think Fable 5.1 is a big improvement in writing style. It sounds a lot less stereotypically like other Claude models, has (imho) a much more natural style, and responds to my style instructions more reliably. More work to be done (and we will!) but reading better prose makes me so much happier.
Another point I expect not to get much attention until it all happens at once is science. People have been correctly excited about the many "sudden" breakthroughs LLMs are making in Maths, but some of the science benchmarks make me believe we'll soon see similar developments in other scientific domains. Fable 5.1 more than doubled Fable 5's Terminal-Bench-Science [1] score, which I think is meaningful.
[1] https://github.com/harbor-framework/terminal-bench-science