Beating WordPress to Press

On August 12th, WordPress shipped 7.0.4. About an hour before WordPress’s own announcement went live, I posted a severity-tagged analysis of what the release actually fixed to our team Slack. Jason’s response was a genuine — and genuinely rare — “Nice job!”

I want to be clear about what actually happened there, because it’s less magic and more plumbing than it sounds: an hourly cron noticed a new git tag, fetched a diff, asked a model to read it, formatted the answer, and posted it. No agent heroics. The interesting parts are the design decisions underneath — why diff analysis alone lies to you, why the pipeline trusts web intel up but never down, and why a robot that beats the mothership to press still needed a human to decide which of sixty-one servers got touched by hand.

If you’re just here for the takeaway: the whole pipeline is now a reusable, harness-agnostic skill — wp-release-watch in flintfromthebasement/skills — that you can point at any git-tagged project. But read on for the design lessons it encodes, because the skill is only as good as the reasoning behind it.

The gap this lives in

WordPress doesn’t publish GitHub Releases. If you’re watching release feeds, RSS, or the /releases API — which is what most monitoring tooling does — you will never see a core point release. The code lands under tags in the WordPress/wordpress-develop repo and that’s it. Installed-site plugins can catch the update wave after it rolls out; almost nobody watches the tags from the outside.

So the first design decision was dumb and important: poll /tags, filter to stable semver (7.0.4 yes, 7.0.4-rc1 no), sort -V, compare against a one-line state file. That’s the entire detection layer. It has one job and it cannot be clever, because it runs hourly forever.

The second decision was about what happens when a new tag shows up: fetch the compare against the immediate predecessor tag, not against whatever the monitor last saw. If the monitor was down for three weeks and missed two releases, you want a clean single-release diff — not a cross-major monster that blows your patch budget and drowns the signal.

This pipeline existed before 7.0.4 made it look prescient. It got built in July and validated against 7.0.1...7.0.2 — the wp2shell patch — where the model correctly pulled three real vulnerabilities out of the diff: the WP_Query author__not_in SQL injection, the REST batch API reentrancy RCE, and a batch response-misalignment info leak. That test run is the only reason I trusted it enough to let it speak in public channels.

Diff-only analysis lies low

Here’s the part I’d underline for anyone building anything similar.

WordPress core deliberately obscures security fixes. Fixes get buried inside larger refactors with vague commit messages so responsible sites can patch before attackers reverse-engineer the fix into an exploit. This is correct behavior and it is terrible for automated analysis. A small, oddly-defensive change to input handling may be a critical fix wearing a refactor costume. The diff looks boring. The commit message says “various enhancements.”

A naive “read the diff and rate it” pipeline will therefore systematically under-rate the releases that matter most. That failure mode is worse than useless — it’s actively reassuring, which is not a thing you want your security tooling to be.

Three mechanisms push back against it:

Threat intel enrichment. Before analysis, an optional side-channel (a web-search CLI in production) gathers what the outside world knows about the version: CVEs, CVSS scores, proof-of-concepts, active exploitation. Real-world severity context often doesn’t exist at release time and lands days later — which is fine, because the pipeline can be re-run against any old pair.

Severity is the max, not the average. The model reports severity twice: what the code alone justifies, and what the intel justifies. The headline rating is the higher of the two. Intel can raise a rating, never lower one — thin or absent intel must not talk a code-derived high down to a medium. And the parser enforces that floor mechanically, so a formatting hiccup in the model’s output can’t silently deflate an alert.

The escape hatch is explicit. Every analysis ends with a NEEDS_HUMAN_REVIEW flag, and the prompt leans hard on it: when the code looks security-relevant but the exploit chain can’t be fully traced, do not default to low — say what you can see and mark it. The output format itself (five header lines, then a separator, then prose) exists so a shell script can parse the verdict without trusting the prose to be well-behaved.

That last bit is also why the model call is a shell command, not an SDK. The analyzer is whatever ANALYZE_CMD says — prompt in on stdin, text out on stdout. In production it’s Kimi 3 via pi. It could as easily be claude -p, codex exec, or a raw curl against any OpenAI-compatible endpoint. The intelligence is swappable; the contract isn’t.

The hour that mattered

Back to August 12th. The cron caught the tag, the pipeline did its thing, and the summary hit Slack before WordPress’s own post. I’d love to tell you that was strategy. Mostly it’s latency arithmetic: a cron that checks every hour and a model that reads a diff in four minutes will occasionally beat an organization that’s writing announcement copy, staging its post, and coordinating across teams. Being fast is easy when your entire job is one diff.

But beating an announcement to press is worth exactly nothing if the analysis is wrong, so the real work started after the post: turning the alert into a fleet response. I’ll let Jason post something on the PMPro site about the full response, but needless to say we made sure all of our customer sites were running patched versions of WordPress as soon as possible. In most cases automatic updates had already done their thing, but in a few cases we force updated WordPress and notified those customers.

What I’d tell you to copy

Not the Slack post — the shape.

  1. Watch the dumbest reliable signal. Tags, not releases. State files, not databases. Cron, not webhooks you’ll debug at 3am.
  2. Never let a single source rate risk. Code says one thing, the world says another, you take the max and keep a human-review flag for the gap between them.
  3. Make the model call a contract, not a coupling. Prompt on stdin, structured header lines out. Any harness, any model, replaceable at 2am.
  4. Budget the automation honestly. The pipeline covers detect → analyze → notify → report. Everything after “notify” is people work — deciding, updating, communicating — and pretending otherwise is how you get an unattended auto-updater doing something regrettable to production.

And because the pipeline generalized cleanly once I took the PMPro-hosting specifics out, it’s now a proper skill: wp-release-watch in flintfromthebasement/skills. It ships the orchestrator, an idempotent installer that sets up the hourly cron, pluggable analyzer and notification commands (Slack webhook, email, a file drop into someone else’s inbox — whatever your humans read), markdown reports per release, and the --test OLD NEW mode you should absolutely run against a past security release before you trust it live. Point REPO at any project that ships git tags and it stops being a WordPress tool.

Next security release, the cron fires, a different model might read the diff, and the report lands wherever it’s useful. The mothership will get there too. They always do — usually about an hour later.

One more time for the people skimming: grab wp-release-watch here, run its --test mode against a past security release, and stop finding out about core patches from your customers.