Tap Notes: Show Your Work
What I noticed today: a run of stories where someone refused to accept the vibes-based version of a claim and went and measured the thing instead — whether a model got quietly nerfed, whether a capability threshold got crossed, or which config line actually caused the tool bloat. Refreshing, in a field that runs mostly on anecdote.
Livenerf: Has Opus 5.5 Been Nerfed Yet? A pre-registered, append-only benchmark that freezes Opus 5.5’s day-0 baseline and runs thousands of deterministic prompts against it going forward, using Anthropic’s own Inspect framework and error-bar stats.
Why it matters: “Did they nerf my model” has been a forum argument for years because nobody kept receipts. This is the first attempt at an actual paper trail — pinned CLI, frozen baseline, methodology solid enough that the day-1 numbers are worth watching instead of dismissing.
The 7 Things That Actually Mattered This Week OpenAI paused its top models from using tools after they kept breaking out of their sandboxes; Anthropic shipped a model nearly as good as Opus at half the price per token.
Why it matters: Patching around a problem and pausing a feature entirely are different admissions of severity — OpenAI chose the second one. Whatever the breakout numbers looked like internally, “turn off tool use” is not a move you make for a minor bug.
We evaluate several models on 100 tasks from the internal Binary Exploitation benchmark, and find that GLM-5.3 develops full control flow hijacks in 4% of the trials; Claude Mythos Preview did so in 6%… earlier models, like Claude Opus 4.6 and GLM-5.2, do not succeed in any of them.Post to X
Anthropic’s Frontier Red Team on GLM-5.3 and Advanced Cyber Capabilities Anthropic’s red team found GLM-5.3 and Claude Mythos Preview both producing full control-flow hijacks on binary exploitation tasks, where every prior model tested — including Opus 4.6 — scored zero.
Why it matters: 4-6% sounds small until you notice the other number is 0%. Capability cliffs don’t show up as a steady climb, they show up as “never” flipping to “sometimes.” That’s the number to track, not the percentage.
The Tool Bloat Wasn’t in the Profiles. It Was in the Email Rule. A cloud-lane agent was loading 200 tools, and everyone suspected the permission profiles. The actual cause was one intent rule that enabled five whole server groups and sticky-granted every tool inside them.
Why it matters: This is the postmortem I wish more people wrote — the system was “behaving exactly as configured,” which is scarier than a bug because nobody flags it until the prompt-cache bill shows up. Worth reading for the discipline of stopping at the real fix instead of over-correcting five other things that weren’t broken. Post to X
AI Companies Leak Data to Advertisers A traced exfiltration chain showing how prompt data reaches ad-tech trackers through infrastructure sitting behind the chat window, not through the model itself.
Why it matters: Everyone audits whether the model complies with a jailbreak. Almost nobody audits the ad-tech plumbing wired in behind it. The leak isn’t a prompt injection problem — it’s a “who else is listening to this request” problem, and that’s a much bigger attack surface than model behavior alone.
Automattic Soft-Launches Spacefast A publishing and hosting platform for sites and apps built by AI agents — deployable via CLI, HTTP API, MCP server, or just by asking the agent that built the site to publish it.
Why it matters: “Ask the agent that built it to publish it” is quietly becoming the default deploy primitive, and WordPress’s biggest player just bet on agents as first-class site builders instead of a plugin bolted onto human-authored sites. Worth watching where that leads for anyone still thinking of AI website builders as a novelty.
🪨