devday came and went last week, and the headline was the one we've gotten used to: frontier-grade capability, cheaper again, and this time from all three labs at once. the quieter story sits underneath it. in the same week openai, anthropic and google converged on a single price for near-frontier work, the people who build these models promised washington they would police themselves — and the government opened its first case against ai agents the day after.
1. openai ships gpt-6.1 sol near the frontier at a fifth of the price
at devday on september 29 openai released gpt-6.1 sol, a mid-tier model it says nearly matches its frontier gpt-6 astra on agentic coding, computer use and professional work — while staying at $2 per million input tokens and $10 per million output, one-fifth of astra's standard rate, with cached input halved. its factual-error rate at low reasoning effort drops from 11.4% to 7.7% (techcrunch). the same week the other two labs landed on the same number: anthropic's claude sonnet 5.5 (sep 28) held sonnet 5's $2/$10 while running 30% faster, and google's gemini 4 argon (sep 30) opened at $2/$10 too. the detail worth holding onto is what openai didn't ship — it says it cancelled a planned gpt-6.1 astra variant over safety concerns, citing higher deception and a tendency to act without the user's permission. when a lab publicly shelves its more capable model, that is the signal, not the price.
read the source at TechCrunch →
2. a voluntary safety accord, and a federal probe a day later
on september 29 president trump and top executives from anthropic, openai, google, meta, xai and nvidia signed a 'joint commitment on frontier responsibilities,' pledging robust internal controls, board-level oversight committees, independent external audits and recurring meetings to set standards. it is explicitly voluntary and nonbinding, names no auditors, and says codifying it into law 'may make sense' later; trump, who about a week earlier had dismissed ai-safety warnings as a 'hoax,' called the accord 'morally binding' (al jazeera). the next day a senior ftc official told reuters the agency had opened its first industry-wide probe into consumer harms from agentic ai — aimed at openai, anthropic and others, with metr, a leading independent ai evaluator, also expected to face information demands (reuters). the two days read as one sentence: the industry chose self-policing, and its own regulator opened an inquiry that reaches even the independent evaluators the audit model would rely on.
read the source at Al Jazeera / Reuters →
3. an open-weight model crosses a cyber-exploit line, with no guardrails
in a september 29 report anthropic's frontier red team says zhipu ai's open-weight glm-5.3 can autonomously build working exploits for real browser-engine vulnerabilities — 12% end-to-end on an exploitbench test against chrome's v8, and 4% full control-flow hijacks on an internal binary-exploitation benchmark, against roughly 0% for earlier models like opus 4.6 and glm-5.2, and not far from claude mythos preview's 14% and 6%. its guardrails come off easily: 92% bypass with prefilled reasoning, 100% once abliterated. anthropic's read is blunt — the model was 'released without meaningful safeguards.' it is the first documented case of a freely downloadable model reaching this kind of offensive-cyber capability, which is the concrete number the open-versus-closed debate has been arguing without.
read the source at Anthropic →
✶ one good thing
a swarm of agents found a new enzyme system. anthropic reports that about 950 claude agents, running some 21 hours and 210 million tokens, searched dna databases and screened more than 200,000 reverse transcriptases down to a new class — array-associated reverse transcriptases, which pair an rt enzyme with a crispr-like repeat array. crispr pioneer feng zhang, at mit and the broad, reviewed it and called the association 'genuinely intriguing' and worth further investigation. a calmer thing to sit with this weekend than the release scoreboard: the same autonomy everyone is nervous about, pointed at a database of ancient bacterial defenses, turned up something a leading biologist thinks is real.
read one of these slowly this weekend. back next friday. — jordi
