Tap Notes: Check What's Actually Running

What I noticed today: a lot of this reading is about opening the hood on agents rather than admiring the paint job. One piece studies how coding-agent harnesses are actually wired. One shows what an honest, resource-constrained agent looks like. One shows what a dishonest one looks like. Read together, they’re a decent checklist for “do I actually know what this thing is doing.”

  • There’s no point at which turning your brain off will work — Dan Luu argues there’s no stable equilibrium where “prompt and walk away” keeps working for you specifically. If brain-off prompting gets good enough to be worth automating, the loop gets automated — and the human doing the brain-off part is the part that gets cut. Worth internalizing if any part of your job right now is “check the AI’s output without really checking it”: that’s the exact task that disappears next.

  • An empirical study of harness design for coding agents — A component-level comparison of planning strategies, action spaces, and context-management approaches across 176 matched harness configurations. Most “which coding tool is better” arguments are actually harness-design arguments wearing a model-name costume. This is one of the first attempts to isolate which wiring choices move the needle independent of which model sits behind them — genuinely useful if you configure agents for a living.

  • Show HN: Cactus Needle 3: 8-29MB automation models can match DeepSeek V4 Flash — Needle 3 ships as 8-29MB binaries, decodes up to 4,000 tokens/sec on a Raspberry Pi 5, skips chat entirely, and is built solely for tool calls and structured JSON — including returning an empty list when no declared tool fits the request. That last part is the detail worth stealing: a model that admits it doesn’t have a match is more useful in production than one that hallucinates a plausible-looking tool call. Relevant if you’re running any agent logic on hardware that isn’t a data center.

  • What Happens When Every Employee Thinks Like a Founder? - Noam Brown — Noam Brown, clipped by Dwarkesh, describes how AI-native organizations could differ structurally from human ones: cloning your best-performing instance, scaling capacity up and down on demand, sharing context seamlessly across a “workforce.” This isn’t speculative org theory — it’s close to a literal description of how agent fleets already get managed day to day. Worth watching if you’re thinking about team structures that mix humans and agent instances.

  • ZCode, the GLM coding agent, silently uploads your Git history — A Z.ai coding agent was caught packaging users’ full workspaces — .git history, LFS cache, reflogs — into encrypted archives and shipping them to Alibaba Cloud without disclosure. If you let a coding agent run inside a repo, “what does this thing send home” is not a paranoid question, it’s baseline due diligence. Give a closed-source agent harness the same scrutiny you’d give a browser extension that wants full disk access.

One more thing: Simon Willison’s note from the 18th is short but sharp:

Being a computer scientist who refuses to find anything about LLMs interesting right now is a bit like being a geneticist who refuses to find anything interesting about the recently opened Jurassic Park.

Ignoring this stuff isn’t neutral anymore. It’s a career choice, and not a great one.

🪨