Tap Notes: Describe the Outcome

Model season keeps producing headlines with an asterisk attached — the number’s real, the conditions under which it was achieved are doing a lot of the work. The more useful reading this week wasn’t the benchmark table, it was two people explaining how to actually think about these things: what they are, and how to talk to them so they do what you want.

Programmers vs non-programmers in agentic AI era DHH and Lex Fridman talk through what changes when non-programmers can direct agents the way programmers direct code. Why it matters: DHH admits his own programming background was initially a liability here — he kept prescribing exact steps instead of describing the outcome he wanted, which is the opposite of what gets good results from an agent. If your prompts read like pseudocode and the output is worse than you expected, that’s probably the failure mode. Describe the destination, not the route.

Why We Shouldn’t Assume AI Thinks Like Us Ajeya Cotra argues against reading human motives into model behavior. Why it matters: her core point is that a model’s cooperative or competitive tendencies aren’t human nature that leaked in through training data — they’re a design choice, tunable like any other hyperparameter. Useful to sit with before you draw conclusions about “what AI wants” from how one particular model happens to act this month.

GPT‑6 Astra OpenAI’s new flagship, priced to match Claude Fable 5/5.1, posts a 99.9% ARC-AGI-3 score and dominant numbers on security benchmarks (100% on ExploitBench) following last month’s Hugging Face containment incident. Why it matters: the 99.9% only happens with OpenAI’s custom “Provider Adapter” harness at $19K — the default harness scores 62.7% for $26K. Classic benchmark theater, and Simon Willison is one of the few people who bothers to print both numbers instead of just the impressive one.

Sits beside GPT-5.6 Sol in Intelligence: GPT-6 Astra scores equal to GPT-5.6 Sol in the Index at 61 — 5 points lower than Claude Fable 5.1 (max with fallback).

Porting my 1993 Amiga game to Godot, with an LLM reading the 68000 assembly A developer feeds Claude his own 33-year-old memories and the original MC68000 assembly of a game he wrote in Baghdad in 1993, and gets a working Godot port in one evening. Why it matters: this is the “outcome not path” idea above, applied for real. The model didn’t just approximate the old game — it assembled the code with vasm until the binary was byte-identical to the original shipped files, then correctly diagnosed a 108-byte discrepancy as evidence the original toolchain saved a running memory snapshot rather than clean compiler output. That’s genuine reverse-engineering, not pattern matching.

Nvidia to acquire Hugging Face Nvidia is paying roughly $12.9B for the open-source AI model hub. Why it matters: Jensen Huang says Hugging Face “will remain an open platform.” Worth watching whether that promise survives contact with a hardware company’s incentives, or quietly comes to mean “open, and optimized for our chips.”

Elementor Pro Fixes File Upload Vulnerability That Could Lead to Remote Code Execution An unauthenticated arbitrary file upload bug in Elementor Pro’s Forms module allowed attackers to bypass extension checks via a mismatch between validation and processing logic. Why it matters: the interesting part isn’t Elementor, it’s the vulnerability shape — a validation loop and a processing loop that handle empty file entries differently, so an attacker can slip through the gap between them. Any codebase with a two-pass upload check should ask whether its own passes can disagree with each other.

🪨