a lot landed this week, but most of it pointed the same way. the price of frontier capability fell, pushed down by a chinese open model large enough to run on your own hardware and matched within days by the labs that had been holding the line. then, in a quieter corner, a model found something real: a weakness in a cryptography candidate that two years of human review had missed.


1. anthropic ships opus 5 with a dial for cost against capability

on july 24 anthropic launched claude opus 5, the newest version of its heavyweight model — smaller than fable 5 but cheaper and less restrictive, and beating it on several of the benchmarks in the announcement. anthropic stressed the model is "much stronger at verifying its work and iterating carefully until it succeeds," and shipped a beta "automatic fallbacks" feature that reroutes safety-blocked api requests to a weaker model instead of returning an error; its safety classifiers are expected to trigger 85% less often than fable 5's. the following week, on july 30, openai cut gpt-5.6 luna's price by 80% and terra's by 20% (cnbc) as buyers grew cost-sensitive — a blunter version of the same competitive pressure.

read the source at TechCrunch →

2. moonshot open-sources kimi k3, the largest open-weight model yet

on july 27 beijing's moonshot ai released the full open weights — roughly 1.5 tb — for kimi k3, a 2.8-trillion-parameter mixture-of-experts model (104b active params, 896 experts) with a 1m-token context, plus a 47-page technical report and self-host tooling (vllm, sglang). it ships under a custom "open weight" license, not true open source: firms over $20m/yr running it as a "model as a service" must cut a separate deal with moonshot, and very large deployments must display "kimi k3." it can run on clusters of consumer rtx 5090s but is really aimed at well-resourced orgs. separately — and not in this venturebeat piece — a white house official (ostp's michael kratsios) alleged moonshot distilled anthropic's claude to build k3 and accessed restricted nvidia gb300 chips in thailand, threatening investigations and sanctions; beijing's commerce ministry shot back that washington is practicing "ai hegemonism" and threatened countermeasures (reuters).

read the source at VentureBeat →

3. a claude model found a real cryptographic weakness, and a candidate standard was pulled

anthropic reported that a gated model, claude mythos, working semi-autonomously, found a new key-recovery attack on hawk, a signature scheme in the third round of nist's post-quantum standardization. it halved the scheme's security, dropping one variant's attack cost from about 2^64 to 2^38, in roughly 60 hours and $100k of compute, after the design had survived two years of human review. hawk's author announced he was withdrawing it. no deployed system was touched, and anthropic disclosed to the authors and nist first, but the shape is new: a general model producing a genuinely novel result in a hard, adversarial field rather than reciting a known one.

read the source at Anthropic →


✶ one good thing

the harness, not the model. openai showed that gpt-5.6's mediocre scores on the arc-agi-3 agent benchmark were largely an artifact of the test's generic harness, which discarded the model's reasoning after every step. rebuilt to let the model keep its reasoning across turns, the same model rose from 13.3% to 38.3% while using roughly six times fewer tokens. a clean idea to sit with this weekend: much of what we call a model's ability is really a property of the loop we put it in, and memory of its own thinking is often the piece that was missing.

read it →


pick one of these to read slowly this weekend. back next friday. — jordi