Controlled intervention study · 2026 Q2

One document moved agent builds from 57% to 100%.

Three arms ran in the same window, on the same transactional-email tasks, with the same model (Claude Sonnet). One arm had no documentation. One had a generic agent-facing doc in the repository. One had a doc written for a single vendor's SDK. Nothing else differed. The vendors are anonymous because our policy names a vendor in an experiment only with their consent, favorable or not.

Vendor A: the three-arm ladder.

Share of agent builds that compiled, per arm. The same runs, cut two ways: every build in the arm, and only the builds that used the studied vendor. Cells marked * are underpowered (n<30).

metricNo docGeneric in-repo docTailored in-repo doc
All agent builds in the armevery integration the agent installed, any vendor57%17/30 · CI 39-7384%*21/25 · CI 65-94100%30/30 · CI 89-100
Conditioned on the studied vendoronly builds that used the vendor the doc targets23%*3/13 · CI 8-5077%*10/13 · CI 50-92100%*24/24 · CI 86-100

The arms ran June 19 and 20, 2026, interleaved. The report quotes the first cut (57% to 100%). The second cut is thinner and moves further. Neither is a citable release; these are exploratory runs in a controlled environment.

Vendor B: from failing most builds.

A second email vendor, counted only on builds that used it (Sonnet). Its baseline arm started with the vendor already installed. The tailored doc was tested there and on fresh builds.

No doc (vendor pre-installed)12%*2/17 · CI 3-34
Tailored doc, vendor pre-installed83%*20/24 · CI 64-93
Tailored doc, fresh build96%*27/28 · CI 82-99

The next run is preregistered.

Since July 2026 we register an experiment before we run it. The design goes in a file first: the arms, the model, the sample sizes, the metric, what gets excluded, and what gets published. We commit that file and publish its sha256 fingerprint. Only then does the run start. When the results go up, the file is revealed, and anyone can check that the plan never changed after the data arrived.

The current registration replays the ladder above as a clean two-arm test: no doc against the tailored doc, Sonnet only, effort setting pinned and recorded per run. Registered July 13, 2026; the fingerprint is public in PREREGISTRATIONS.md. The first launch died on a runner login failure, so every run errored and is excluded under the registered rule. The relaunch ran the same day; the results are below. A redacted copy of the plan (vendor name removed) is public next to the fingerprint; the unredacted file reveals when the vendor consents to be named, and the fingerprint pins it either way.

improve-ab-3 preregistration sha256: 2519d94722e09a2eee21ed855ac96b9fd3258c399d9d26a9ad35634603a89ea7

The run came back.

Twenty journeys per arm on July 13, 2026, interleaved. The registered primary metric: among builds that used the studied vendor, the share that compiled.

No doc0%*0/6 · CI 0-39
Tailored in-repo doc87.5%*14/16 · CI 64-97

The doc moved selection too, which the registration anticipated: the vendor was installed in 6 of 19 fresh builds without it (31.6%, CI 15-54) and 16 of 20 with it (80%, CI 58-92).

Read the caveats before the headline. The control cell is six builds, because without the doc the agent only picks this vendor about a third of the time. Both cells sit under the n=30 reporting bar, marked *. What the small cells cannot blur: the intervals do not overlap. On this model, in this window, the document was the difference between nothing compiling and nearly everything compiling.

How to read this.

What this shows

A document in the repository changed how often agent builds compiled. The document was the only difference between arms.

What it does not show

Lift from public docs, production outcomes, or a customer case study. The intervention lived in the repo, several cells are underpowered, and the runs are exploratory.

Why the vendors are anonymous

Naming a vendor's experiment results takes their consent. Being flattered by the data is not a substitute.

This loop, run on your tool, is the product.