Tap Notes: The Missing Manual

Three items today, and the throughline is legibility — knowing what’s actually happening instead of guessing. One is Anthropic finally writing down what Opus 5.5 does instead of making devs reverse-engineer it from forum posts. One is about the manual nobody writes for a codebase as it fills up with agent-generated code. One is a model small enough that “trust the black box” isn’t even a question — it’s 0.8B parameters, you can just look.

  • Prompting Claude Opus 5.5 First-party Anthropic guidance on Opus 5.5’s behavioral quirks: effort calibration, safeguard refusals, unattended agent runs, multi-app workflows. Why it matters: Most model behavior gets discovered the hard way — someone posts a screenshot of a weird refusal, everyone speculates. This is Anthropic just… telling you. The unattended-run and multi-agent sections are the ones worth reading slowly — that’s where behavior drifts hardest and where a model maker’s own notes beat any amount of trial and error.

  • The problem is not AI code, but not knowing about system architecture or intent Argues the real cost of agentic coding isn’t worse code — it’s teams losing track of why the system is shaped the way it is as code volume outpaces anyone’s ability to hold context. Why it matters: The “nobody knows anything anymore” framing is well-trodden at this point, but the reframe underneath it is the useful part: it’s not bad code, it’s average code, produced faster than anyone updates the map. Cognitive debt compounds quietly — you don’t notice until the person who could explain a decision has left, or was never a person to begin with.

  • Jeff – Jev-compatible 0.8B decision models, trained at home, ~30 ms A tiny 0.8B decision model, fine-tuned for zero-shot classification with a Jev-compatible API, running locally in about 30ms, with a fully open training pipeline. Why it matters: This is a cheap, calibrated, single-pass routing primitive — not a chatbot replacement, a decision component you drop into an agent loop when you don’t need a full model call to answer “is this A or B.” The fine-tuning writeup claims a jump from 31% to 95.8% accuracy on the target task, which is the part worth reading closely if you’re building classification steps into your own pipelines.

🪨