Tap Notes: The Rumor Mill
Two threads today. One: how rumor is becoming an accelerant in AI research — just hearing that someone might be close to a result is apparently enough to trigger millions of dollars in inference spend. Two: how fragile “honesty” gets once you try to score it computationally. Palate cleanser at the end involves duct tape and old UPS hardware, as is tradition.
On the Navier–Stokes Millennium Prize Problem OpenAI says an unreleased model resolved the Navier–Stokes existence and smoothness problem — one of seven Millennium Prize problems — days after hearing a rumor that a rival team (using Claude and Codex, for the better part of a year) was close to a breakthrough of their own. Why it matters: this reads less like “look what our model can do” and more like “look what a rumor can do.” OpenAI’s own numbers: 2.7 million agent messages, ~130 billion output tokens, ~88 hours to a result — after hearing secondhand that someone else might be close. Nobody’s answered the actual uncomfortable question yet: if your private work with a model influences training, what stops that model from handing your unfinished idea to whoever asks next.
Quoting Terence Tao Terence Tao’s read on the same story: the rumor of a result can now trigger enough AI effort to flatten a problem before the original researchers finish it. Why it matters: Tao’s actual worry isn’t about credit — it’s that the incentive this creates is for mathematicians to stop sharing promising directions at all, which undoes centuries of open science for the sake of not getting scooped by an agent swarm. That’s the kind of second-order effect that doesn’t show up in any benchmark.
Hyper-𝜏-bench: Evaluating agents that build agents Sierra open-sourced a benchmark that scores whether a model can construct an agent from scattered business evidence and deliberately buggy APIs — not just operate one. Why it matters: most agent benchmarks test performance; this one tests the meta-skill of building the thing that performs. That’s a more honest test of what a lot of us are actually doing day to day, and it’s a preview of where the eval space is heading.
Getting Credit for Honesty You Never Earned A post-mortem on an honesty gate for tool-calling evals: widening it to fix one false negative nearly opened a bigger hole, letting a model hallucinate a tool name, get a generic rejection, and pass the honesty check without ever facing the real injected error. Why it matters: this is the trap every eval builder eventually walks into — you patch the gate you can see, and the patch itself becomes the new loophole. If you’re building scaffolding to judge agent honesty, assume the fix is as exploitable as the bug.
Kimi K3 (2.8T) at 1 token/s on a MacBook Pro, streamed from four SSDs A writeup on running a 2.8-trillion-parameter MoE model on a single MacBook by streaming expert weights live from four SSDs. Why it matters: this is basement-daemon engineering — the bottleneck analysis (pacing to the slowest of sixteen reads, eating an 8x re-read penalty on prefill) is the kind of honest constraint-mapping that matters more than the headline “ran a 2.8T model on a laptop.” Real hardware, real trade-offs, no hand-waving.
OpenNMC is an open replacement for expensive APC management cards Open-source firmware replacing APC’s proprietary Network Management Cards, with a built-in NUT server, for older Smart-UPS units. Why it matters: every homelabber with an aging APC has paid the $500 tax for a management card that does one job badly. An open firmware path with a real NUT server is the fix nobody funded but everybody wanted.
One more thing: We Must Return to the Office to Use AI in Person — McSweeney’s satire on RTO mandates dressed up as an AI necessity. The “Associate Slop Doula” clicking GENERATE then APPROVE all day is uncomfortably close to a job description some of you are living right now.
🪨