
Anthropic accidentally exposed internal assets through a CMS misconfiguration, revealing that it was developing two unreleased models: Claude Mythos and Capybara. Cybersecurity stocks fell shortly afterwards.

That last detail is the tell. A leaked model name does not usually move a sector; a leaked capability does. What surfaced with the names was enough to suggest Mythos was not an incremental release, and the market reaction arrived before most of the details did.
More information has since emerged about Mythos, via Project Glasswing and researchers working through the exposed material:

The coding jump is the kind of progress the field has learned to read: 80.8% to 93.9% on SWE-bench Verified is a large step, but it is a step along a curve everyone has been watching, and the remaining headroom on that benchmark is small enough that it will stop being informative soon.
Autonomous zero-day discovery is a different category of claim. Finding and exploiting an unknown vulnerability is not a benchmark task with a graded answer key; it is open-ended search against real software, and it is the part of offensive security that has stayed stubbornly human. A model that does it at machine scale changes the economics on both sides at once. Defenders get a tool that can audit a codebase continuously rather than in quarterly bursts. Attackers get the same tool. Which side gains more depends almost entirely on who has access first and under what controls — which is exactly the question a leak, rather than a release, leaves unanswered.
The stock move suggests the market read it the same way: not as a better coding assistant, but as a capability that touches the value of incumbent security products.
Everything here comes from exposed material and third-party analysis, not from a launch. Benchmark figures attached to unreleased models are typically internal numbers, run under conditions nobody outside the lab can inspect, and they have a habit of moving between a leak and a launch. Capybara remains close to a name and nothing more — no capability claims of substance have surfaced alongside Mythos's.
There is also no released safety documentation to weigh against the capability claims. For a model whose standout skill is autonomous exploitation, the deployment terms — who gets access, with what monitoring, under which usage policy — matter at least as much as the SWE-bench score, and those are precisely the details a CMS misconfiguration does not leak.