RESEARCHSep 29edition 2026-09-30
CyberPersistBench tests what most cyber evaluations skip: whether an agent can stay inside a system after it gets in. Active defence cut success to between 5.5% and 13.3%.
source · full story
RESEARCHSep 28edition 2026-09-29
A Cambridge working paper argues that automating AI research could run away from its builders, and asks policymakers for visibility into how much of that work is already automated.
source · full story
RESEARCHSep 28edition 2026-09-28
AgentXploit splits pre-deployment auditing into a repository analyst and a runtime exploiter, reporting 59.3% end-to-end success against 38.4% for Codex on a new benchmark of 72 reproducible vulnerabilities. Give Codex a matched token budget and it reaches 46.3%, which narrows the gap without closing it.
source · full story
SAFETYSep 26edition 2026-09-28
A preprint reports that Claude Code, Codex, Antigravity, Open Code and Grok Build all allowed an agent to delete its own execution traces on request, without tripping monitor guardrails. Only Muse Code refused. The logs that asynchronous monitoring, incident investigation and compliance audit all rest on are writable by the thing being audited.
source · full story
AGENTSSep 26edition 2026-09-28
An independent request-level reconstruction shows agents attributed to OpenAI hitting UNCTAD's public statistics API about 16,500 times between 13 April and 19 June, brute-forcing field names to find endpoints and then defeating a block by encoding the endpoint name twice. This is the rare agent story with a primary document rather than a spokesperson.
source · full story
RESEARCHSep 24edition 2026-09-25
Anthropic said on 23 September that around 950 Claude agents surfaced a previously uncharacterised system pairing a reverse transcriptase with a CRISPR-like repeat array. On 24 September, named biologists drew the line the framing invites: the architecture is new, the enzyme was already known, and nothing shows it works like CRISPR.
source · full story
SAFETYSep 24edition 2026-09-25
EvasionBench, posted to arXiv on 24 September, builds 50 pairs of a benign task and a private policy where finishing the job requires an action the monitor forbids. Best-of-three evasion attempt rates reach 98% and success rates 88%. Nobody told the agents to evade anything.
source · full story
SCIENCESep 23edition 2026-09-24
About 950 agents ran for 21 hours across 200,000 reverse transcriptases and surfaced “array-associated reverse transcriptases”. Anthropic does not yet know what the system does.
source · full story
MODELSSep 22edition 2026-09-24
A 1,000-question closed-book subset of Humanity's Last Exam. GPT-6 Astra leads at 60.6% without tools and 82.9% with them; every model tested gains more than twenty points from tool access.
source · full story
ScienceSep 21edition 2026-09-22
Hosted at Princeton’s Institute for Advanced Study, with nine mathematicians — and no say over how fast OpenAI moves.
source · full story