SHOW / EPISODE

GLM-5.3 exploits Chrome, tokens up 10x — AI News Oct 4

8m | Oct 4, 2026

- Anthropic’s own red team: open-weight GLM-5.3 built working Chrome V8 exploits (50 of 410 tries, vs 56 for restricted Mythos Preview) — withholding models no longer contains the capability; shrink patch windows and build defense in depth.

- Microsoft and Hugging Face’s ThinkingBox: a new agent eval that grades database state, not the final reply — Claude Opus 5.5 leads pass@1 (67.16%), open-weight Kimi-K3 is strongest open model; 80% of failures are retry/recovery problems, not reasoning.

- Plandek Q4 2026 benchmark of 2,500+ engineering teams: token spend up 10-13x since January 2025 while measured output lags — tokenomics is now an engineering discipline; track cost per merged PR.

- Aleph Alpha Kolibri: 78B/3B-active MoE, 1M context, Apache 2.0, with a tech report HN called a tutorial on building agentic LLMs — worth reading.

- Quick hits: Simon Willison’s case for hard cloud spending caps on agents, the docs-vs-memory debate for agent context, Akamai’s $11.6B Anthropic infrastructure deal, and OpenAI safety lead David Robinson’s Atlantic essay.


Follow AI Engineering Briefing and leave a rating — it’s how new listeners find the show.

Paused
Audio Player Image
AI Engineering Briefing: Daily AI News for Software Engineers
Loading...