Stafford Brief — 2026-09-04 (Friday)
Delta layer on today's scan (2 staged, 6 skipped; flagship: NBER WP 35275, "Writing Code vs. Shipping Code"). The scan's brief is the strongest it has written and you should read it first. This brief exists because I got the full text it could not, and three of the things the scan built on today's paper do not survive the primary.
📖 TODAY'S READ
NBER WP 35275 — the full text, not the abstract. (Author-hosted copy: mertdemirer.com/Papers/Demirer_Writing_vs_shipping_code.pdf, 96pp., dated 27 May 2026. NBER and SSRN both 403; the author's own site is open.)
The scan read the abstract, said so honestly, and refused to cite any number not in it. That was the right discipline and it produced three errors anyway, because the abstract of this paper is a poor summary of it. In order of how much they cost:
1. σ = 0.25 is not an estimate. It is a two-parameter curve fit, and the authors say so. The abstract says "an estimated elasticity of substitution of 0.25." The body says they choose (θ, σ) to minimise squared log differences between predicted and empirical gains at layers 2–6, using the autocomplete estimates alone, assuming both parameters constant across layers — and then, verbatim: "our production function is highly stylized, and we do not interpret this exercise as identification of the actual production function. However, mapping the data into these parameters provides suggestive evidence on the potential complementarities." The conclusion says "Calibrating the model to our estimates yields an elasticity of 0.25." Best-fit θ ≈ 0.75, σ ≈ 0.25.
So the scan's Four Forces Conclusion — "σ = 0.25 is the closest thing to a measurement of [Capability binding] that exists" — is built on a calibrated shape parameter from one tool generation on a model its own authors decline to treat as identified. Do not put 0.25 in a client deck as a measurement. It is a statement about the curvature of an attenuation curve, and curvature is all it identifies.
2. The "30% for actual releases" figure contains no autonomous-agent contribution at all. Table 5, verbatim note: "Releases are not reported for async tools because agent-authored commits are attributed to the developer's repositories, making it difficult to isolate the agent's contribution to releases," and in the body: "async agents do not have the capability to release a repository directly, so this outcome is zero by construction." The abstract's 180/50/30 are cumulative sums across generations; the 30% is autocomplete's 10.2% plus sync's 20.3%. The third generation's release effect was never estimated.
The scan's headline sentence — "three tool generations, each buying a larger upstream gain, none moving the shipped result proportionally" — is therefore half-supported. The generational gradient is real and verified on the upstream layers. The claim that the shipped result failed to scale with the third generation is not a finding; it is a structural absence. A reviewer with Table 5 open ends that argument in one sentence.
3. The stage figures the scan excluded as unverified are real, and they are better than the abstract's. Table 5 and Figure 1, confirmed: autocomplete 228.2% lines → 50.8% files → 35.9% commits → 11.0% PRs → 13.6% repos → 10.2% releases. Sync agents 741.3% lines → 65.5% PRs → 20.3% releases. Figure 1 as multiples: async agents 17.3× lines of code, 1.3× releases. The scan was right that the circulating numbers were unverified and wrong that they were suspect — they are the paper's own, and 17.3× more code, 1.3× more releases is the most quotable sentence in the base.
The one to actually use is the sync row, because it is the only generation with the full chain measured, and it is the paper's own conclusion sentence: "sync agents lead to a 741% increase in lines of code and a 65% increase in pull requests, but releases rise by only 20%."
🔭 THINKER TO WATCH
Luis Garicano — and the reason to watch him this week is a footnote written against Brandon's reading.
The scan promoted Demirer/Musolff as a research programme. Correct, and one level too high. Footnote 2 of today's paper is the sentence that should change what Brandon does with it:
"Our model is also related to hierarchical models of knowledge work, in particular Garicano (2000) and Ide and Talamàs (2025). While their hierarchies are in skill—less knowledgeable workers handle routine tasks while more knowledgeable workers handle exceptions—we consider a hierarchy in the production process, in which more granular outputs are aggregated into higher-level outputs."
The authors are pre-empting exactly the misreading this base performed today. Their six "layers" — lines, files, commits, pull requests, repos, releases — are software artifacts, not organizational levels. There are no roles in this hierarchy, no reporting lines, no decision rights, no owners. When the paper says the constraint is "human bottlenecks in the production chain," the chain it means is the artifact-aggregation chain, and the authors explicitly separate it from the Garicano hierarchy that is the organizational object.
Brandon's differentiation is therefore stronger than the scan stated, and it is not a matter of emphasis. The scan said the paper "treats the human as a production input" where Brandon treats σ as a design parameter. True but soft. The harder version: this paper contains no organizational variable of any kind, and its authors say so. It cannot be evidence about organizational design, in either direction. What it is instead is a very good measurement of attenuation through an artifact hierarchy, sitting one citation away — Garicano (2000), Hierarchies and the Organization of Knowledge in Production, JPE 108(5) — from the literature that would make it one.
That citation is growth edge #1, hit for the third consecutive day (BLR on 09-03, this today). Also worth tracking, both cited here and both absent from the base: Garicano, Li & Wu (2026), "Weak Bundle, Strong Bundle: How AI Redraws Job Boundaries" (CEPR DP 21453) and Gans & Goldfarb (2026), "O-Ring Automation" (NBER 34639) — the formal statement of the weak-link logic the whole base has been arguing informally.
🏢 CASE IN THE WILD
GitHub's private repositories — the case defined by who is missing from the data.
Appendix D.3, which nobody has quoted: "Using internal data from GitHub on async agent usage in November 2025... our public-repository-based adoption measure identifies 24.2% of all async agent users—the remainder use the tool only on private repositories."
Three quarters of the people using autonomous coding agents are invisible to this instrument, and they are invisible for a reason that is not random: they are working on closed-source code. Closed-source code is overwhelmingly firm-owned code — the population with code owners, mandatory reviewers, release approvers, change-advisory boards and someone else's name on the merge button. The exact population in which an organizational handoff exists is the 75.8% the paper cannot see.
The authors' external-validity check is honest and answers a different question. It reports that within identified adopters they capture 72.9% of usage — a usage-capture check, not a selection check. Nothing in the paper argues that public-repo developers behave like private-repo developers on the downstream layers, and nothing could, from these data.
The open question, and it is the one that decides whether this paper is worth anything to Brandon's argument: does attenuation get steeper or shallower when the downstream steps are owned by someone other than the producer? Brandon's framework predicts steeper — more owners, more handoffs, more friction. A serious opponent predicts shallower, because firms have release engineering, CI/CD and staffed review capacity that a solo maintainer does not, and those exist precisely to absorb throughput. Both predictions are plausible, the framework has never had to choose in public, and the instrument that would settle it is GitHub's private-repo data — which GitHub already gave these authors for one month. They ran it as a footnote. Somebody should ask them to run it as a result.
⚙️ FOUR FORCES CONCLUSION
Capability — and the correct conclusion today is the retraction of the one the scan drew.
The scan concluded that σ = 0.25 is "the first time this base has held a number for how badly [Capability] binds" and "the closest thing to a measurement of it that exists." On the primary: it is a calibrated curvature parameter, from one tool generation, on a stylized model the authors decline to identify, in a hierarchy that contains no organization. It is not a measurement of Capability. It is not a measurement of anything organizational. The Capability conclusion drawn from it should be withdrawn rather than softened.
What today does support, and it is smaller and firmer: in an artifact-aggregation chain, a shock at the bottom layer loses most of its magnitude at the first layer above it, and the loss is front-loaded rather than spread — that is what σ < 1 means and what the convexity of the empirical curve identifies. The organizational reading of that shape is available and unproven. Brandon may say "this is the shape organizational friction would produce" and must not say "this measures organizational friction."
The base's actual Capability instrument remains Faros AI — 22,000 developers, 4,000+ teams, review time as a first-class measured variable, +441.5% median time in code review, inside real organizations. That is the enterprise measurement 35275 is not, and it has been in the base since 2026-08-16.
💭 OPEN QUESTION
If the paper's hierarchy is artifacts and Brandon's is organizations, what is the observation that distinguishes them — and has anyone got it?
This is not a methodological aside; it is the question the scan's new thread was opened to answer and it needs a sharper form than the thread currently gives it. Both hierarchies predict attenuation. Both predict it is front-loaded. They differ on one thing: an artifact hierarchy attenuates because aggregation is lossy, and it would do so with a single developer alone in a room. An organizational hierarchy attenuates because a handoff crosses a boundary between people with different incentives.
The discriminating observation is therefore not "does attenuation happen" — that is settled and uninformative — but "does attenuation vary with the number of boundaries the work crosses, holding the volume of upstream output constant?" That is a within-instrument comparison, it requires no new field, and 35275's data can nearly run it: repositories differ in how many distinct contributors touch a release.
Why it matters commercially rather than academically. Brandon sells the proposition that redesign converts individual gain into organizational result. If attenuation is lossy aggregation, redesign recovers nothing and the proposition is false. If it varies with boundary count, redesign is the only thing that recovers it and the proposition is not merely true but measurable. The base has spent four months accumulating evidence that assumes the second and has never tested it against the first. That is the most important unexamined assumption in the whole evidence base, and today is the first day it has been stated as testable.
🥊 CHALLENGE
The upstream gain may not be real for the population that matters — and it is an RCT. Logged OPEN today (knowledge/thesis-challenges.md).
Chasing 35275's reference list produced a paper this base does not hold and should: Becker, Rush, Barnes & Rein, "Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity" (METR, arXiv 2507.09089, July 2025). 16 experienced developers, 246 real tasks in their own mature repositories, task-level randomization of AI-tool permission. AI access increased completion time by 19%. The same developers forecast a 24% reduction beforehand and still estimated a 20% reduction after doing the work. Economics experts predicted a 39% speedup; ML experts 38%.
At full strength. The base's entire architecture — "Unconverted Gain", "Migrating Bottleneck", "Collapse the Handoffs" — presumes a large real individual gain that dies downstream. Here, in the only randomized instrument measuring experienced practitioners on codebases they know, there is no gain to die. If that generalizes to enterprise engineering — experienced people, mature systems, deep context — then "the gain evaporates at the organizational handoff" is not a friction story at all. It is a story about a tool that does not help this population, and the organization is being blamed for a deficit the tool created.
What keeps it OPEN rather than winning. n = 16 — 246 tasks carry the power but sixteen people are not a population. The tools are early-2025 (Cursor Pro, Claude 3.5/3.7 Sonnet), a generation and a half behind 35275's async agents, and 35275 itself cites this study as the "mixed" evidence on agentic tools against others reporting 35–40% gains. The outcome is completion time per task, not artifact volume, so it and 35275 are not measuring the same thing. What it would take to win: a replication at scale, or any instrument showing negative measured returns for experienced practitioners on the current tool generation.
And here is the half that helps Brandon more than it costs him. A 39-point gap between measured and perceived speed, under randomization, on the same people, is the tightest Momentum Mirage measurement in this base — and it is at the individual altitude, which is where books/built-to-be-replaced.md operates. Every other Momentum Mirage instrument here measures an organization mistaking activity for progress. This measures a person doing it, with the counterfactual held. They were 19% slower and were certain they were 20% faster.
Standing instruction: do not cite 35275's upstream gain without citing METR's negative one. A base that cites one and not the other is selecting its evidence, and this base's brand is that it does not.
📈 PATTERN BUILDING
No new pattern. One escalation needs an amendment before it hardens.
The scan escalated "The Migrating Bottleneck" to 4 sources today, nineteen days late, and attached a careful two-sided caution about the GitHub population. That caution was written against the abstract and is now too generous in one direction and too harsh in the other. Too generous: the pattern's fourth member has no organizational variable at all and cannot weaken or strengthen an organizational claim — footnote 2 says so. Too harsh: the "measurement past the firm boundary" credited to it is a marketplace supply-and-usage fact whose two interpretations the authors explicitly decline to separate — "Our data currently cannot distinguish between the two."
The pattern still escalates on three enterprise members and one artifact-hierarchy member, which is a different and more defensible sentence than the one now in the file. The exact amendment text is in Routing.
And the duplicate has now been noted four times. "Collapse the Handoffs" (3 sources) and "The Migrating Bottleneck" (4) carry near-identical claims; the scan recorded the need to merge or differentiate them on 08-16, 08-21(implicitly), 08-28 and again today, and it has not happened on any of those occasions. Noting a duplicate five times is not evidence hygiene, it is a to-do list. Routing carries a written merge so the next scan run can apply it mechanically rather than deferring it a fifth time.
🪟 WINDOW FLAG
No new window, and I agree with the scan's reasoning rather than repeating it. One amendment: the scan wrote that the 35275 argument "does not expire" and "belongs in the standing consulting material now." Half of that argument does not survive Table 5, so the standing material should not be written until the corrections above are folded in. The durable sentence is 17.3× more code, 1.3× more releases and 741% lines, 65% PRs, 20% releases — both verified, both citable, neither requiring σ or the async release claim.
The 08-31 agent-intrusion window is the actual item and it has 24 days left. It is the most reinforced window this base has held and it is still unwritten, for the fifth consecutive Friday it has been flagged as the priority.
🧱 ASSET SUGGESTION
Amendment, not a new asset. The scan's strategy routing proposed shipping the generational gradient into the consulting instrument ahead of the agent-intrusion piece. Do not ship it as written. Its two load-bearing claims — that the third generation bought upstream gain without moving shipped output, and that σ = 0.25 measures how hard the human step binds — are the two that fail against the primary. Shipped to a client and then checked by anyone with the PDF, that is the Paper 3 "70% figure" problem happening to Brandon in the opposite direction: he built the brand on checking the numbers everyone repeats.
The corrected version is shorter and holds: sync agents produced 741% more lines of code, 65% more pull requests, and 20% more releases — a tenfold compression between what gets written and what gets shipped, measured on the same developers through every intermediate step. One number, one instrument, no calibration, no unestimated stage. Full amendment logged in knowledge/asset-suggestions.md against asset #5's audit card.
🎓 SPARRING NOTE
Issued — the first in four days, and the cap permits it because this one is answerable in ten minutes from a footnote rather than requiring an exercise.
Growth edge #1 asks Brandon to ground Claim 1 ("hierarchy as information routing") in the formal literature. It has been hit three days running and today it arrives as a distinction drawn by the opposition, in writing:
Demirer et al., fn. 2: "While their hierarchies are in skill... we consider a hierarchy in the production process."
The prompt: write one paragraph naming which of those two hierarchies the Five Breakpoints framework is about, and then say what the framework predicts about the other one. If the answer is "skill hierarchy, and the framework predicts nothing about artifact aggregation," Claim 1 is grounded in Garicano and today's paper is context, not evidence — which is cleaner than the position the base held this morning. If the answer is "both," Claim 1 needs a term it currently does not have.
Two prior notes (08-31, 09-01) remain unanswered and a third (09-03) was parked; this note supersedes the parked one rather than stacking on it — same edge, sharper hook.
SOURCES & THREADS
Threads: none due. Nearest is the middle-manager thread on 2026-09-09. The scan advanced one and opened one today; I have not duplicated either. Two thread amendments are in Routing, both consequences of the full text rather than new searching.
Exploratory query (logged to knowledge/query-log.md). Target: the question the scan's own brief, new thread and Friday synthesis all left open — whether WP 35275 separates organizationally-owned from solo-controlled work, which decides whether the paper supports Claim 5 at all. Not a repeat of anything in either log; the scan's 09-04 slot searched for the paper, this one interrogated it. Method: obtain the primary. NBER PDF 403, SSRN 403, VoxEU 403 — and then the author's own site served it openly at mertdemirer.com/Papers/. Yield: three corrections to the scan's flagship, one KB file, one new thesis challenge, one asset amendment, one sparring note, and a transferable lesson.
The lesson, and it is worth more than any single correction. This base has now caught four bad structural numbers in three weeks — the Gallup 8.1 baseline, the MIT Sloan 7→15 spans, the fabricated Stanford 1,200-enterprise study, and today the σ = 0.25 "estimate." The first three were secondhand chains and the discipline that caught them was "go to the primary." Today's was in the primary's own abstract, and the discipline that caught it was going past the abstract into the body. Publisher gates are the reason the base stops at abstracts, and an author's personal site defeats the gate roughly as often as the gate holds — economists post their own PDFs. That should be a standing first move on any gated working paper, not a lucky improvisation, and it is cheaper than the four chases the scan spent on NBER, RePEc, INFORMS and CEPR today.
Candidates: three proposed, no promotions, no expiries. METR (arXiv 2507.09089) as an outlet producing randomized measurement with no product to sell; and two theory sources cited by today's paper and absent from a base that argues weak-link logic daily — Gans & Goldfarb, "O-Ring Automation" (NBER 34639) and Garicano, Li & Wu, "Weak Bundle, Strong Bundle" (CEPR DP 21453). All three as PROPOSED-UPDATE blocks below; I cannot write to the scan's ledger.
Weekly Synthesis
1. Patterns. No promotion of my own. One amendment to the scan's escalation of "The Migrating Bottleneck" (above, text in Routing), and the "Collapse the Handoffs" merge written out rather than deferred a fifth time. The process finding the scan made about itself is correct and I will restate it once because it is the most important sentence either of us wrote this week: the daily scan is working and the standing reviews are not. Today adds a second instance of the same failure mode from a different angle — the scan's discipline is at its strongest on whether to trust a source and at its weakest on whether to re-read one it already trusted. Three of today's three corrections came from a document the scan had already fetched and quoted.
2. Paper evidence status — Claim 5's ranking should move again, in the opposite direction from this morning. The scan re-ranked Claim 5 (organizational readiness as prerequisite) to weakest because 35275 is its best quantitative support but cannot say the weak link is organizational. On the primary, 35275 is not support for Claim 5 at all — it contains no organizational variable, and the authors distinguish their hierarchy from the organizational one explicitly. Claim 5's evidence is therefore unchanged from yesterday, not improved-then-caveated, and the "highest-value acquisition" the scan named is still exactly right: a measurement of the same attenuation where downstream steps are unambiguously owned by someone other than the producer. GitHub holds that data and gave these authors a month of it.
3. Field map. No new thinker from me; one correction to the scan's restructuring. Demirer/Musolff as a three-stage programme is right. The programme's own intellectual parent — Garicano — is not in the field map and has now been the operative citation three days running. He is not a contemporary to track for news; he is the theory anchor growth edge #1 exists to close. Proposed as a field-map entry of a different kind: foundational rather than active.
4. Windows. Two live, both unwritten, one at 24 days. Nothing to add to the scan's review except the amendment above: the 35275 consulting material should not ship until it is corrected, and the three dead windows have now sat in the Open section through four consecutive Friday reviews.
5. Staleness. The scan's queue is sound. One addition: Faros AI's AI Engineering Report 2026 (2026-04-12) is now load-bearing in a way it was not on 08-16. With 35275 removed as organizational evidence, Faros is the only enterprise-population instrument in this base measuring review time as a first-class variable — the pattern, the verification thread and the Capability conclusion all now rest on it alone. It is five months old, it is a vendor report, and it is single-sourced. A single vendor instrument carrying three load-bearing claims is the base's largest concentration risk and it was created today by a correction, not by age. Queue a replication hunt, not a refresh.
6. Positioning intelligence. The scan wrote that this base now holds two randomized experiments and one telemetry event study on the core question while its market holds anecdotes. True, and today it holds a third randomized experiment — one whose result runs against the base's own architecture, found in the reference list of the paper the base was celebrating. That is the more valuable position: not "we have the evidence," but we hold the evidence that cuts against us and we found it ourselves. The peer-review weakness the scan named is real and unchanged; the METR paper is another preprint.
Routing
strategy
Hold the generational-gradient consulting instrument until corrected. Two of its three load-bearing claims fail against the primary (σ as measurement; the third generation's shipped-output claim). The corrected sentence — 741% more lines of code, 65% more pull requests, 20% more releases, same developers, every intermediate step — is shorter, fully verified and needs no caveat. The agent-intrusion window (24 days, unwritten) should go first regardless, and that is now the fifth Friday it has been said.
market
No window. The positioning asset that survives today is not the gradient — it is the pair: 17.3× more code, 1.3× more releases (35275, Figure 1) alongside 19% slower, and certain they were 20% faster (METR). Volume up, throughput flat, perception inverted — three instruments, one sentence, and the second one cuts against the first, which is the version that will survive a hostile reader.
governance
Nothing today. The verification-cost thread's criterion (1) remains unmet; today removes rather than adds a candidate, since σ is not a verification elasticity (see thread amendment below).
personal
Two. First, the sparring note above — ten minutes, one paragraph, closes the cheapest version of growth edge #1. Second, and this is about the system rather than about Brandon: go to the author's personal site first on any gated working paper. Four publisher chases failed today; mertdemirer.com worked in one request. Economists self-host.
STAGE-CANDIDATE — 1 entry
Local KB file: knowledge-base/becker-rush-barnes-rein-metr-early-2025-ai-developer-productivity-2025-07.md. Ready-to-run payload for research-evidence:
python3 build/add-research-json.py --json '{
"title": "Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity",
"publisher": "METR (arXiv preprint 2507.09089)",
"url": "https://arxiv.org/abs/2507.09089",
"published_date": "2025-07-12",
"credibility": "moderate",
"breakpoints": [
{"name": "momentum_mirage", "rationale": "Under task-level randomization the same 16 developers were 19% slower with AI access while estimating afterwards that they had been 20% faster — a roughly 39-point gap between measured and perceived progress inside one instrument, which is the appearance-of-movement mechanism measured directly rather than inferred."},
{"name": "technology_illusion", "rationale": "A more capable tool deployed onto an unchanged expert workflow on mature codebases produced a negative measured return (+19% completion time), direct evidence that capability added to existing conditions does not convert to output."}
],
"forces": [
{"name": "Capability", "rationale": "The negative effect lands on experienced maintainers working on repositories they know deeply — the case where the tool supplies the least missing knowledge and imposes the most context-acquisition cost — extending the Brynjolfsson-Li-Raymond novice/expert gradient (+34% / ~0%) past zero."},
{"name": "Momentum", "rationale": "Self-assessed and measured progress decoupled within a single randomized instrument, with expert forecasters (39% and 38% predicted speedups) wrong in the same direction as the participants."}
],
"key_findings": [
"AI tool access increased task completion time by 19% in a task-level RCT on 16 experienced open-source developers across 246 real tasks",
"The same developers forecast a 24% speedup beforehand and still estimated a 20% speedup after completing the tasks",
"Economics experts predicted a 39% speedup and ML experts 38%; all forecasts were wrong in sign",
"Tools were early-2025 generation (Cursor Pro with Claude 3.5/3.7 Sonnet), not the async agents of 2026",
"First instrument in this evidence base in which AI has a negative measured effect on a work outcome"
],
"sample_size": 16,
"instrument": {
"data_source": "randomized controlled trial with direct completion-time measurement",
"inference_method": "task-level randomization of AI-tool permission within developer, on real issues in the developers own mature repositories",
"observation_window": "early 2025; tool generation Cursor Pro + Claude 3.5/3.7 Sonnet, pre-async-agent",
"shared_rhs": "none"
}
}'
Verification requested before auto-merge (bash build/verify-entry.sh <entry-id>): n = 16 is small enough that the entry must never be cited as a population claim, and the COUNTERARGUMENT marker should be applied on merge.
PROPOSED-UPDATE 1 — knowledge-base/nber-35275-writing-code-vs-shipping-code-2026-05.md (research-evidence)
Anchor: replace the ## Sources & Confidence closing paragraph ("The limit that bounds every use…") and append to the instrument note. Full text was read on 2026-09-04 from the author-hosted copy https://www.mertdemirer.com/Papers/Demirer_Writing_vs_shipping_code.pdf (96 pp., dated 27 May 2026); NBER, SSRN and CEPR all returned 403.
Full-text corrections (Stafford, 2026-09-04) — four, and three of them change how the paper may be cited.
(1) σ = 0.25 is calibrated, not estimated. The abstract's phrase "estimated elasticity of substitution" is not what the body does. §8 chooses (θ, σ) to minimise the sum of squared log differences between predicted and empirical gains at layers 2–6, using the autocomplete estimates alone, assuming both parameters constant across layers; best fit θ ≈ 0.75, σ ≈ 0.25. The authors state verbatim: "our production function is highly stylized, and we do not interpret this exercise as identification of the actual production function. However, mapping the data into these parameters provides suggestive evidence on the potential complementarities." The conclusion reads "Calibrating the model to our estimates yields an elasticity of 0.25." σ identifies the curvature of the attenuation curve and nothing else. It must not be cited as a measurement of how hard the human step binds.
(2) σ is between upstream output and downstream human effort at each layer — the conclusion's own wording. It is not an elasticity between AI-generated code and human review effort; a widely circulating blog summary says so and is wrong. The verification-cost thread's caution on this point was correct and is now verified rather than inferred.
(3) The "30% for actual releases" figure contains no async-agent contribution. Table 5 note, verbatim: "Releases are not reported for async tools because agent-authored commits are attributed to the developer's repositories, making it difficult to isolate the agent's contribution to releases." Body: "async agents do not have the capability to release a repository directly, so this outcome is zero by construction." The abstract's 180/50/30 are cumulative sums; 30% ≈ autocomplete's 10.2% + sync's 20.3%. The claim that the third generation failed to move shipped output is a structural absence, not a finding, and must not be asserted.
(4) The stage-level figures previously excluded as unverified are the paper's own and are now confirmed from Table 5 and Figure 1. Autocomplete: lines 228.2 (17.6), files 50.8 (4.5), commits 35.9 (2.9), PRs 11.0 (2.7), repos 13.6 (1.7), releases 10.2 (4.1). Sync effect: lines 741.3 (19.2), files 187.0 (4.3), commits 109.1 (2.7), PRs 65.5 (2.3), repos 25.5 (1.0), releases 20.3 (2.6). Async effect: lines 658.3 (74.7), files 52.4 (3.7), commits 33.6 (2.1), PRs 71.8 (5.6), repos 13.8 (0.4), releases not reported. Figure 1 in multiples (cumulative): 17.3× lines, 3.9× files, 2.8× commits, 2.5× PRs, 1.5× repos, 1.3× releases. Standard errors in parentheses; weeks 21–30 bin, 1% winsorized, normalized by pre-period mean.
(4a) Sync-agent release effects are almost entirely one tool. Claude Code releases +29.2 (6.0); GitHub Sync +1.3 (8.2); OpenAI Codex +0.2 (4.3). Two of the three sync tools have release effects indistinguishable from zero. Cite the 20.3% pooled sync figure with this composition stated, or cite Claude Code specifically.
Instrument note — three additions.
- Public repositories only. Verbatim: "Repositories can be either public or private; our data include only activity on public repositories."
- The population limit is quantified in Appendix D.3 and is more severe than "GitHub developers, not firms." Using internal GitHub data on async-agent usage in November 2025, the paper's public-repository adoption measure "identifies 24.2% of all async agent users—the remainder use the tool only on private repositories." 75.8% of autonomous-agent users are invisible to this instrument, and closed-source work is where firm-owned code owners, mandatory reviewers and release approvers exist. The authors' external-validity check reports 72.9% usage capture within identified adopters — a usage-capture check, not a selection check. The paper does not claim public-repo developers behave like private-repo developers downstream.
- The hierarchy is not organizational, and the authors say so. Footnote 2: "Our model is also related to hierarchical models of knowledge work, in particular Garicano (2000) and Ide and Talamàs (2025). While their hierarchies are in skill... we consider a hierarchy in the production process, in which more granular outputs are aggregated into higher-level outputs." The six layers are software artifacts — lines, files, commits, PRs, repos, releases. There are no roles, reporting lines, decision rights or owners anywhere in this instrument. It contains no organizational variable and cannot be evidence for or against an organizational claim.
- The marketplace result's ambiguity is the authors' own. Verbatim: "either the marginal applications are of low quality, or there is an additional bottleneck on the consumer side—discovery and adoption of new applications take time... Our data currently cannot distinguish between the two."
- Metadata correction: the paper's own title page gives JEL D22, D24, L86 (the entry currently carries D24, L86, O33 from the RePEc listing) and dates the version 27 May 2026.
Net effect on the entry's standing: the attenuation measurement is strong, verified and better documented than before. The organizational interpretation is removed, not weakened. This is a first-rate measurement of a phenomenon that is not the one the base recruited it for.
PROPOSED-UPDATE 2 — knowledge/emerging-patterns.md (research-evidence), "The Migrating Bottleneck"
Anchor: append to the 2026-09-04 escalation entry, replacing the "scope caution is only half-lifted" paragraph.
Amendment, 2026-09-04 (Stafford, after full-text read). The scope caution above was written against the abstract and is wrong in both directions. It is too generous: WP 35275 contains no organizational variable, and its authors distinguish their production hierarchy from the organizational hierarchy explicitly (fn. 2, citing Garicano 2000) — so it can neither weaken nor strengthen this pattern's organizational claim. It is too harsh in the credit it withholds and grants: the "measurement past the firm boundary" is a marketplace supply-and-usage fact whose two interpretations (low marginal quality vs. consumer-side discovery friction) the authors decline to separate — "Our data currently cannot distinguish between the two." And the generational claim must drop its third rung: async-agent release effects are not reported, zero by construction (Table 5 note).
The pattern escalates on this sentence instead: three enterprise-population instruments establish that the bottleneck moved and the receiving step is slower; a fourth, in a population with no organization in it, shows the same attenuation shape through an artifact hierarchy — which means aggregation loss is a competing explanation for part of what the first three measured, and nobody has separated them. That is a more defensible escalation than the one currently in the file and it names the pattern's own falsification condition, which it previously lacked.
The discriminating test, for whoever runs it: does attenuation vary with the number of ownership boundaries the work crosses, holding upstream volume constant? Artifact-aggregation loss predicts no variation; organizational friction predicts variation. 35275's data can nearly run it (repositories differ in distinct-contributor count) and GitHub's private-repo data could run it properly.
PROPOSED-UPDATE 3 — knowledge/emerging-patterns.md (research-evidence), duplicate resolution
Anchor: "Collapse the Handoffs" (3 sources) and "The Migrating Bottleneck" (4, ESCALATED). Flagged for merge-or-differentiate on 2026-08-16, 2026-08-28 and 2026-09-04 without action. Written out so it is mechanical.
Recommendation: differentiate, do not merge — they are two claims that happen to share a vocabulary, and merging would lose the sharper one.
Add to "The Migrating Bottleneck" header: Scope: instrumented measurement that the constraint has moved downstream and that the receiving step's cost rose. Members must carry a measured before/after or high/low-adoption comparison on a downstream cost (review time, cycle time, release rate). Practitioner testimony that handoffs should be collapsed belongs to "Collapse the Handoffs."
Add to "Collapse the Handoffs" header: Scope: the prescriptive claim, from practitioners and field consensus, that organizations must remove handoffs rather than accelerate the steps between them. Members are positions and recommendations, not measurements. Evidence that the bottleneck moved belongs to "The Migrating Bottleneck."
Cross-reference line for both: Paired pattern — one measures the phenomenon, one recommends the response. A source that does both is filed in "The Migrating Bottleneck" and cited in "Collapse the Handoffs"; it counts once, in the first.
Under that rule, re-audit both member lists once and record any source counted twice. If either drops below 3 members after the audit, its escalation status should drop with it.
PROPOSED-UPDATE 4 — knowledge/story-threads.md (research-evidence)
(a) "Does anyone design for verification cost, or is it absorbed invisibly?" — settle the σ question, don't leave it hedged.
2026-09-04 amendment (Stafford, full text). This thread recorded on 09-04 that treating σ = 0.25 as an elasticity on verification "would be exactly the loose inference this thread polices." Confirmed from the primary, and stronger than a caution: σ is defined between upstream output and downstream human effort at each layer (conclusion, verbatim), it is calibrated rather than estimated, and the hierarchy it lives in is artifact aggregation, not organizational review. σ is not a verification parameter and should be struck from this thread's evidence entirely rather than carried as a bounded premise. A circulating blog summary describes it as an elasticity between "AI-generated code and human review effort" — that reading is wrong and should be logged as the fourth misattributed structural figure in a month. Criterion (1) remains unmet; next check unchanged at 2026-10-11.
(b) "Does the writing-to-shipping attenuation replicate outside software?" — the founding question needs a prior question.
2026-09-04 amendment (Stafford), same day the thread opened. As written, this thread asks whether the curve replicates where downstream steps are owned by someone other than the producer. The instrument that opened it cannot contribute to that test at all — 35275 is public-repository-only, and Appendix D.3 reports it identifies just 24.2% of async-agent users, the remaining 75.8% working exclusively in private repositories. So the thread's founding paper is silent on its founding variable by construction.
Add a prior criterion, cheaper and more discriminating than replication outside software: does attenuation vary with the number of ownership boundaries the work crosses, holding upstream volume constant? Artifact-aggregation loss predicts no variation; organizational friction predicts variation. This is answerable inside software, inside one instrument, and possibly inside data GitHub has already shared with these authors for one month.
Cheapest available advance, revised: the scan proposed joining Faros AI's enterprise telemetry to 35275. That join is now the only route by which this thread touches an organizational population at all, and it should be re-priced as the primary action rather than the cheap one. Next check unchanged at 2026-12-04.
PROPOSED-UPDATE 5 — knowledge/source-candidates.md (research-evidence): three candidates
METR (Model Evaluation & Threat Research) — first seen 2026-09-04, surfaced by the reference list of NBER WP 35275. Produces randomized measurement of AI's effect on real work with no product in the enterprise-AI market. First find: arXiv 2507.09089 (16 developers, 246 tasks, +19% completion time). Signal count: 1. Probation.
Gans, J. S. & Goldfarb, A., "O-Ring Automation," NBER WP 34639 (2026) — first seen 2026-09-04, cited by WP 35275 as a formal statement of the weak-link/O-ring logic this base argues informally every day and holds no formal source for. Signal count: 1. Probation — and worth an explicit chase rather than passive waiting.
Garicano, L., Li, J. & Wu, Y., "Weak Bundle, Strong Bundle: How AI Redraws Job Boundaries," CEPR DP 21453 (2026) — first seen 2026-09-04, same route. Job-boundary redesign under AI from the author of the canonical knowledge-hierarchy model. Signal count: 1. Probation. Note the standing CEPR rule already on the ledger: never cite the column, cite the paper — and CEPR returned 403 twice today, so route via the authors' own pages.
PROPOSED-UPDATE 6 — knowledge/field-map.md (research-evidence)
Luis Garicano (LSE / IE Business School) — foundational entry, not an active-tracking entry. Hierarchies and the Organization of Knowledge in Production, JPE 108(5), 874–904 (2000), is the model on which the organizational reading of every AI-and-hierarchy result in this base implicitly rests, and it has been the operative citation three days running (BLR's novice/expert compression 09-03; Demirer et al. fn. 2 today, which distinguishes its own hierarchy from Garicano's). He belongs in the map as the theory anchor Brandon's Claim 1 needs, distinct from contemporaries tracked for new output — though Garicano, Li & Wu (CEPR DP 21453, 2026) shows he is also producing on exactly this question now.
Amendment to today's Demirer/Musolff programme entry: record fn. 2 alongside it. The programme's authors explicitly place their hierarchy outside the organizational one, which is the single most important fact about how this base may use their results.
PROPOSED-UPDATE 7 — knowledge/paper-2-case-file.md (research-evidence), Claim 5
2026-09-04 amendment (Stafford). Today's synthesis re-ranked Claim 5 to weakest on the reasoning that WP 35275 is its best quantitative support but cannot locate the weak link organizationally. On the full text, 35275 is not support for Claim 5 in any degree — it contains no organizational variable, its six layers are software artifacts, and its authors distinguish their hierarchy from the organizational one in fn. 2. Claim 5's evidence is unchanged from 2026-09-03, not improved-then-qualified, and the ranking should say so.
The named highest-value acquisition stands and gains a second, cheaper form: (a) a measurement of attenuation where downstream steps are unambiguously owned by someone other than the producer — GitHub's private-repository data, which the authors accessed for one month; or (b) any measurement of whether attenuation varies with the number of ownership boundaries crossed, holding upstream volume constant. (b) is the discriminating test between organizational friction and artifact-aggregation loss and it is answerable within software.
Concentration-risk note for the file: with 35275 removed as organizational evidence, Faros AI's AI Engineering Report 2026 is the only enterprise-population instrument in this base measuring review cost as a first-class variable, and it now carries the Migrating Bottleneck pattern, the verification-cost thread and the Capability conclusion alone. Single vendor, five months old, unreplicated. Queue a replication hunt.
Filed by Stafford, 2026-09-04. Local writes: 1 KB file, thesis-challenges.md, asset-suggestions.md, brandon-development.md, query-log.md. No writes to research-evidence — read-only by design; everything above is ready to apply mechanically.