Skip to content
Guideintermediate

AI Agents Are Breaking the Security Embargo: What to Do Before Your Next Fix Ships

The short version

Probes for a cohttp path-traversal bug hit the maintainer's server about ten minutes after the fix PR went public. Here is the timeline, what the 87% vs 7% exploit benchmark actually measures, and how to choose between private coordination, continuous releases, and protocol-level mitigations.

Published October 5, 2026by Pondero Research
Table of Contents

AI Agents Are Breaking the Security Embargo: What to Do Before Your Next Fix Ships

On August 14, 2026, Anil Madhavapeddy opened a public pull request to fix a path-traversal bug in cohttp, the OCaml HTTP library he maintains. About ten minutes later his own website was fielding probes for percent-encoded traversal sequences, per his write-up of the incident. The fixed release did not ship until August 20, per the OSV record for OSEC-2026-16. For six days the patch was public and the release was not.

For any bug class an agent can pattern-match, exploitation now starts at the first public signal (a PR, a commit message, a mailing-list question), not at the advisory. A 30-to-90-day embargo still protects the patch. It does not protect the description, and the description is what an agent needs. Below is the cohttp timeline, what the research numbers actually show, and a decision table for the three responses Madhavapeddy proposes, ending in a checklist you can apply to the next fix you ship.

The cohttp timeline

The root cause is an ordering mistake: Cohttp.Path.resolve_local_file normalized the URI path before percent-decoding it, so an encoded slash (%2f) survived dot-segment removal and was decoded into a real ../ afterwards, per the OSV advisory:

request:     /static/..%2f..%2f..%2fetc/passwd
normalized:  one encoded segment, no ".." to remove
decoded:     /static/../../../etc/passwd

OSV rates it 8.7 (High) under CVSS v4. Here is how the response unfolded.

Date (2026)EventSource
Aug 11Report reaches [email protected]; Madhavapeddy says it came privately via Jane Street on Slack and was found with Claude FableOSV, Madhavapeddy
Before the PRHis own agent (DeepSeek V4 Pro, after Fable refused) builds a working local exploit in under a minuteMadhavapeddy
Aug 14Fix PR cohttp#1145 opened publiclyOSV
Aug 14, about ten minutes laterProbes for percent-encoded traversal hit his live web serverMadhavapeddy
Aug 20Fix released in cohttp 6.3.0OSV
Aug 22Release announced; blog post publishedOCaml Discuss

Two details in that table matter more than the ten-minute figure. First, the bug was found with a model (Claude Fable) and reached the maintainer through a Slack channel, not a public tracker. Second, Madhavapeddy reproduced the exploit with his own agent "just by knowing roughly what it was about," in his words. He opened the PR publicly to get more reviewers on it. Normally, he writes, that review "takes a few days and a release within a week or two is reasonable." His own reading of the ten minutes: if building the exploit took him one minute, a watcher on package repositories "could easily be exploiting them within seconds."

What the 87% number actually measures

The figure everyone quotes comes from Fang, Bindu, Gupta and Kang, "LLM Agents can Autonomously Exploit One-day Vulnerabilities", submitted to arXiv in April 2024. The researchers built a benchmark of 15 real one-day vulnerabilities. A GPT-4 agent given the CVE description exploited 87% of them; without the description it managed 7%, per the paper's abstract. Every other model they tried, including GPT-3.5 and the open-weight models of the time, scored zero, and so did the ZAP and Metasploit scanners.

The paper does not say agents find bugs on their own; with no description, GPT-4 mostly failed. It says the written description of a bug does the bulk of the work. That is the mechanism behind the cohttp probes. In 2024 the agent needed a CVE entry. By August 2026, Madhavapeddy's agent needed a rough topic. A fix PR titled after the bug class is more than enough.

The field data lines up. Sysdig's threat research team saw the first exploitation attempt against marimo's CVE-2026-39987 nine hours and 41 minutes after the advisory went out, with no public proof-of-concept in existence; the attacker built the exploit from the advisory text, per Sysdig (April 2026). At the aggregate level, Google's M-Trends 2026 report estimates the mean time to exploit has dropped to minus seven days, meaning exploitation routinely starts before a patch exists, per Google Cloud's threat intelligence team (March 2026).

Keep the Fang et al. benchmark in proportion: 15 vulnerabilities is a small set, and GPT-4 is a 2023 model. The direction is what transfers, and the 2026 incidents above say it held.

Maintainers are absorbing the volume with the same headcount

The second half of the problem shows up in the inbox. Nick Craig-Wood, who created rclone, replied on the 399-point Hacker News thread about Madhavapeddy's post that the project received about 20 security disclosures through GitHub in its first 10 years and "had to deal with over 40 in the last month." He triages with AI tools and still loses large amounts of time, and GitHub's CVE assignment has slowed from 2-3 days to 3-4 weeks, so he now ships point releases with CVE-PENDING in the changelog.

Chainguard's Adrian Mouat put the maintainer's bind plainly, quoted by InfoQ on October 3: "Just opening a PR to fix an issue puts the project and users in a bad place, as attackers can create and start using exploits even before an updated release is available." His LinkedIn post goes on to note that the obvious workaround, shipping binaries before source, cuts against open-source norms and complicates life for the distributions that repackage the code.

Report volume is now hitting bounty programs too. Google stopped accepting product-vulnerability reports to its open source reward program on October 1, citing invalid AI-generated submissions (our news brief).

Three responses, and when each one fits

Madhavapeddy proposes three responses and argues projects will need "some combination of all three" in the short term. They solve different problems, so the choice depends on what you ship and who installs it.

ResponseWhat it meansFits whenBreaks down when
Private coordinationDevelop the fix and discuss the bug in a closed channel, publish code and advisory togetherDownstream distributors need lead time; the fix spans one repo; reviewers are known in advanceYou need CI on the fix; the fix spans several repos; discussion happens in Slack or Discord
Ship continuously, no embargoFix in public and release within hoursYou control the release path and users auto-update (a single binary, a SaaS)You are a library embedded in other people's products who upgrade on their own schedule
Protocol-level mitigationPush a rule or control that blocks the attack before the full fix landsThe attack has a recognizable request shape (traversal, injection) or a revocable credentialThe bug is in internal logic with no distinguishable input pattern

Private coordination is the traditional answer, and the tooling around it has gaps. GitHub's temporary private forks keep the patch out of public view, but per GitHub's docs, "integrations, including CI, cannot access temporary private forks," pull requests in the fork merge all at once, and an admin has to add each collaborator. Madhavapeddy's sharper objection is that a private fork "plugs the wrong leak": the patch staying secret matters less than keeping the description away from attackers, and that description usually travels through chat tools he calls "extremely leaky."

Continuous shipping is what the Linux kernel already does. Its security process releases fixes for publicly known bugs immediately and allows deferral of an undisclosed fix for at most 7 calendar days, 14 in exceptional cases, and only for QA and rollout logistics. QEMU moved the same way: its security process page now says disclosures from automated tools are "highly likely to be independently re-discovered" and that maintainers will "generally reject requests for arbitrary embargoes" without high severity. Packaging is where libraries struggle. Chrome ships one binary; cohttp ends up inside other people's products, which, as Madhavapeddy notes, downstream distributions repackage on their own timescales.

Protocol-level mitigation is the response that buys time while a fix is in review. For cohttp, Madhavapeddy notes the mitigation was simple and "implementable the minute the report arrived": normalize percent-encoded path separators in the request URL. Cloud providers already do this as virtual patching; as Madhavapeddy points out, Cloudflare deployed managed rules for Log4Shell in 2021 (Cloudflare's write-up). A team running a cohttp-backed service behind nginx could have blocked this exact class as soon as the bug class was known, with a rule like this, matching against the raw, undecoded request line:

# Reject percent-encoded traversal before it reaches the app.
# $request_uri is the raw request line, so encoded sequences are still visible.
if ($request_uri ~* "(\.\.|%2e%2e)(%2f|%5c)") {
    return 400;
}

The gap Madhavapeddy points out is distribution: open source has no channel to ship a rule like that to every deployment of a library outside a commercial CDN.

What this does not mean

The embargo is not dead for every class of bug. It is degraded for bugs an agent can reconstruct from a short public signal: path traversal, injection, missing authorization checks, deserialization, anything whose fix diff points straight at the flaw. Those are the bugs where the description is the exploit.

Embargoes still do real work in two cases. One is when the fix requires coordinated rollout across vendors who cannot otherwise patch in time, which is why the kernel keeps a separate embargoed hardware issues process. The other is when nothing public points at the bug yet; a report that stays inside an end-to-end encrypted channel still has value, and Madhavapeddy draws exactly that line between end-to-end encrypted channels (OCaml uses Matrix) and Slack or Discord.

None of this is a new attacker skill. The HN user bri3d made that point early in the Hacker News thread: backing exploit PoCs out of "patches, commit messages, and random overheard or over-read sentences is a practice as old as vulnerability research." What changed, in that reply's reading, is scale: far more low-skill actors can now run it against low-value targets across the whole internet. Plan against a cheap automated watcher on every public repository, not a specialist who might find your bug in a week.

A defender's checklist for the next fix you ship

For maintainers, before the fix PR goes public:

  1. Treat the first public signal as the disclosure time. A PR, a branch name, a question in a shared channel all count.
  2. Have the release ready to cut the same day the PR opens, including the changelog and advisory text. The six days between cohttp's public PR and its release is the window to close.
  3. Write the mitigation first. If the attack has a request shape, publish the WAF or proxy rule with the advisory so operators can block it before they can upgrade.
  4. Watch your own logs after you push. Madhavapeddy spotted the probes in his own server logs. A first pass for this bug class:
# Look for percent-encoded traversal probes in an nginx or Apache access log
grep -iE '(\.\.|%2e%2e)(%2f|%5c)' /var/log/nginx/access.log | tail -n 20

For teams consuming open source, the useful change is to measure your upgrade window in hours for internet-facing dependencies and to pull vulnerability data by API instead of waiting for a weekly scan. OSV's API takes a package and version and returns known advisories; a query we sent on October 5, 2026 for cohttp 6.2.0 returned OSEC-2026-16:

curl -s -d '{"package":{"name":"cohttp","ecosystem":"opam"},"version":"6.2.0"}' \
  https://api.osv.dev/v1/query | jq '.vulns[].id'

Our recommendation by role

Small-library maintainers with no control over where their code is deployed should stop relying on a long embargo. Fix in public, release the same day, and ship a mitigation rule alongside the advisory, because the rule protects the deployments that will not upgrade this week. Projects with downstream distributors (kernel-style) should keep a private window but cap it at days, using the kernel's 7-day ceiling as the reference, and move discussion off Slack and Discord into an encrypted channel with named participants. Teams running services on open source should assume an exploit exists from the moment a fix PR appears upstream, and put internet-facing dependencies on an hours-not-weeks patch target backed by an OSV feed.