Stafford

Asset ledger

Posts, essays, and instruments Stafford has proposed — with status and full reasoning.

1. The Evidence Audit — "Four Papers, One Method" DECLINED

The argument, exactly: Everyone repeating "AI flattens organizations" is citing what looks like four independent confirmations — Babina et al. (NBER 31325), Ewens & Giroud (NBER 34162), Wang/Feng/Sun (arXiv 2511.02099), Alekseeva et al. (SMJ 2026). They are one data genre (organizational structure inferred from résumé job titles or job postings — no org chart, reporting line, or payroll record anywhere), mostly pre-generative observation windows (three of four end 2018–2022), and in two cases literally one shared variable (Ewens & Giroud's AI-adoption measures are Babina et al.'s). The one administrative-data instrument covering the generative period (Visier, 7.7M employee records) finds consolidation-flattening in only ~10% of companies and measures no AI variable at all. The claim that survives: firms measured as AI-adopting flattened, by résumé-inferred hierarchy, in the pre-generative era — and nobody has yet measured whether flattening works.

Asset type: Substack long-form + a living instrument-audit table on research.workthatholds.com (data source · inference method · window · shared right-hand side · what it can/cannot support), updated as new instruments arrive. The dashboard is the moat: nobody else writing commentary has a maintained public evidence base to anchor a standing audit page. The essay is the announcement; the page is the durable asset that earns return visits and citations.

Why this form: it's measurement criticism, not a theory fight — verifiable from public papers, hard to rebut, and it extends the brand Paper 3's August update established (the person who checks the numbers everyone repeats: first the 70%, now the flattening consensus). It also pre-softens the empirical ground for asset #2.

Feeds: the Résumé-Panel Consensus pattern entry (emerging-patterns.md) is ~80% of the draft already. Shan & Zhu (SSRN 6456498, quarantined) becomes content: "a fifth instrument arrived; here's the first question the audit asks of it."

Effort/risk: low/low. Venue: workthatholds.com; tighter version pitchable to MIT SMR as "what the AI-flattening evidence actually shows."

2. The Position Essay — "The Flattening Bet" (formerly The Coordination Mirage) DRAFTED

The argument, exactly (the case file's unused academic-strength core): A serious literature now argues AI's organizational meaning is coordination compression — Sequoia/Block's "From Hierarchy to Intelligence" as the practitioner statement, Farach's coordination-compressing-capital model and Klein & Wieczorek's "Headless Firm" as the formal versions. Grant it its strongest form: hierarchy really did carry routing, routing really is collapsing, structures really are thinning. The mirage is the inference that cheap coordination solves the organizational problem. Three evidence classes say the binding constraint migrates rather than disappears: (1) the deployment–impact wall (88% deploying / 81% no bottom-line gains; Accenture's realized-value line falling while investment rises); (2) organizational context as a measured moderator — in the instruments where AI productivity is estimated in a production function, the organizational variable does the work while the raw AI term goes weak (Wang/Feng/Sun; BIS training complementarity; BCG's talent-pillar tripling while tech scores barely move); (3) the compression models' own fine print — Klein & Wieczorek relocate the constraint to verification, which scales with throughput and is an organizational property; McElheran et al. (Census) find a third of the adoption J-curve's dip is firms dismantling their own management practices. Coordination got cheap; interpretation, incentive alignment, verification, and momentum did not — and machine speed makes their absence bind faster and less visibly. Close on the falsifiable fork (the Regime Fork thread criteria): what evidence would prove coordination-primacy right — spans of control, manager wage premium, coordination-tooling quality predicting outcomes better than any readiness measure. Invite it.

Asset type: the seventh-paper-grade Substack long-form — practitioner-legible, citations load-bearing, opponents named and steel-manned. Not an SSRN working paper yet: the Claim 1 theory lineage (Galbraith's information-processing view, Garicano knowledge hierarchies) is an open growth edge (brandon-development.md #1), and an academic posting invites exactly that referee ambush; the essay format takes the position publicly while the case file keeps accumulating armor.

Why this form: it's the one argument in Brandon's corpus that engages his strongest opposition on their terms, it's genuinely absent from the published series (The AI Mirror asserts; this defends), and stating the win-conditions publicly converts the thesis from consulting frame to testable claim — the best available answer to the unfalsifiability challenge short of a readiness instrument.

Sequencing: after asset #1 (the audit becomes this essay's footnote armor). Both are argumentative value, not time-windows — the near-term window item remains the Accenture piece (publication-windows.md, ~Oct 5), which shares the J-curve/McElheran material.

Feeds: paper-2-case-file.md (the machine-speed evidence engine), Regime Fork thread, Conditional-Productivity and Unconverted-Gain patterns, thesis-challenges.md (coordination-primacy, CONTESTED).

3. The Missing Instrument — measuring "AI multiplies broken processes" PROPOSED

The argument, exactly: the most-repeated claim in Brandon's market is that AI deployed on broken processes makes things worse. The scan's 2026-08-30 exploratory query hunted for any study measuring a pre-deployment organizational condition against a post-deployment outcome and returned a clean null: the claim is asserted everywhere and measured nowhere. That absence is simultaneously (a) the open lane the scan's brief flagged, (b) the answer to the standing unfalsifiability challenge (thesis-challenges.md), and (c) growth edge #2 in brandon-development.md (the readiness instrument). The asset is not an essay first: it is a small prospective instrument — a breakpoints-based pre-deployment scoring of organizations about to deploy AI (the will-it-hold rubric already does the scoring half for transformations), tracked against 6-12 month outcomes. Even a dozen scored cases would be the only dataset of its kind.

Asset type: instrument + tracked cohort first (extend the will-it-hold forecast log with an AI-deployment variant); the essay writes itself after the first resolutions.

Feeds: exploratory-null of 2026-08-30 (scan query-log row 1), unfalsifiability challenge, will-it-hold rubric and forecast log, The Flattening Bet's "outcome column is empty" section.

Status note: requires Brandon's call on scope — this is a practice-building asset, not a post.

Update 2026-08-30 (scheduled run) — the method exists; the ambition needs pricing. Stafford's exploratory query looked outside AI, at implementation science, where readiness measurement is a mature discipline. Two findings change this entry.

The design is off the shelf. Helfrich et al. 2011 (knowledge-base/implementation-science-helfrich-orca-predictive-validity-2011-08.md) is a usable template: baseline instrument administered before a known, named upcoming change (respondents must know what is coming or the scores measure nothing), outcome measured 6–9 months post-baseline as an effect size, hierarchical linear model with the site as the unit, controlling for programme and for whether the intervention was actually received. Port it directly: breakpoint scores at baseline before a named AI deployment, outcomes at 6–12 months, site-level model. The will-it-hold rubric supplies the scoring instrument; this supplies the measurement design. That removes the "how would we even do this" objection entirely.

The sample arithmetic is the real constraint. Helfrich wanted ~30 sites for 90% power to detect R² ≥ 0.21. Twelve scored cases will not predict; they will illustrate. Both are worth doing and they are not the same asset, so split the entry: 3a — the scored cohort (start now, describe it honestly as a forecast log, value is the practice and the case material, no predictive claim), and 3b — the predictive study (needs ~30 sites, a real outcome definition, and probably a partner with access to deployments; the thing that would actually answer the unfalsifiability challenge). Announcing 3b's claim on 3a's sample is the failure mode; see the 2026-08-30 update in thesis-challenges.md.

And the framing got better. Miake-Lye et al. 2020 found no gold standard readiness instrument after 20 years of the field trying. That reframes the pitch: this is not "Brandon builds a measure," it is "the measure nobody has built, in the one setting where its absence is currently expensive." It also makes a defensible short essay on its own — the readiness measurement gap is 15 years old and AI just made it urgent — which is publishable before any cohort resolves, and which stakes the claim while 3a runs. Update 2026-08-30 (22:20 UTC) — the standalone essay in the last update is now a much stronger piece, and it has a title.

The evening update proposed "the readiness measurement gap is 15 years old and AI just made it expensive" as publishable before any cohort resolves. Tonight's thread chase gives that essay its spine, and it is better than the version pitched three hours ago:

  1. The definitive test was never reported. Helfrich's 2011 prospective ORCA study — 53 VA facilities, effect-size outcome, 90% power — has no results paper in PubMed, Europe PMC or Semantic Scholar, and no Helfrich-authored paper cites the protocol. The 2009 development paper had promised it: "Criterion validation using implementation and quality-of-care outcomes is the next phase of our work."
  2. The construct's architect concedes it. Weiner (ORIC's developer), systematic review 2020: "the question of whether readiness for implementation matters remains open" — with TCU-ORC rated "minimal" on predictive validity after 55 tests.
  3. It is still true in 2025. Caci et al., May 2025: 40 of 46 studies measure readiness once.
  4. And the sharpest paragraph is Noe 2014: ORCA predicted whether staff believed their facility was meeting the need, and predicted nothing about whether services were implemented.

Why this is a better essay than the 21:55 version. It is no longer an observation about a gap. It is a narrative with a missing document at the centre of it — a funded, designed, protocolled study that never reported — which is the same investigative register as Paper 3's retirement of the 70% figure, and the second instance is what makes that a practice rather than an incident. It also arrives at a conclusion Brandon can own: the reason "get ready before you deploy" is unfalsifiable advice is that the discipline built to test it never finished the test.

Hard constraints, and the first would sink the piece if broken. Do not claim the Helfrich study was run and found a null — nothing shows the analysis was completed, and file-drawer inference is precisely the reasoning this base refuses elsewhere. Do not tag Noe as Momentum Mirage; the resemblance is an argument, not a mapping, and the design is cross-sectional. Do not let the piece conclude that readiness does not matter — the honest claim is that nobody ran the test properly.

Sequencing note. This now competes with the Accenture window piece for the next writing slot. The Accenture piece should still go first: it is time-priced against the next Pulse of Change wave (see tonight's reclassification note), and this one is not time-priced at all — a fifteen-year-old gap does not close in October.

Status: PROPOSED (essay); feeds and is fed by 3a/3b above.


4. The Proxy-Statement Instrument — the ex-ante measure is already filed PROPOSED

The argument, exactly: Asset 3b — the prospective study that would answer the unfalsifiability challenge — is currently priced at roughly thirty sites plus a partner with access to live AI deployments, which is why it has not started. That price assumes the ex-ante organizational variable has to be collected. For the incentive layer it does not: it is filed, dated, public and machine-readable. Equilar's 2026-02-06 roundup (scan ingest, 2026-08-31) shows named public companies disclosing AI metrics in executive incentive plans at stated weightings — Qorvo 20%, Recursion ~16.7%, Juniper 10%. Cohen, Kadach, Ormazabal & Reichelstein (Journal of Accounting Research 61:805–853, 2023) already built precisely this design at international scale for the previous non-financial metric to enter executive pay, and it cleared a top-tier accounting journal. The design ports directly: code every disclosed AI metric in a defined universe (S&P 500 or Russell 3000 proxies) as activity-defined — exploration, deployment, "win the AI opportunity" — versus result-defined, at time T; measure firm AI-related outcomes at T+2. Nobody has to grant access, and the universe is the index rather than a convenience sample.

Why this is worth a slot rather than a footnote in #3. It answers three standing items at once. It is the first version of the readiness-instrument idea whose sampling frame is defined rather than assembled, which is the objection that has sunk 3a's predictive claim. It supplies the prevalence measurement the "Paid for Deployment" pattern explicitly says it is watching for — "any instrument reporting the share of a defined universe that discloses an AI metric." And it is the only route the base has to resolving the soft-metric objection logged in thesis-challenges.md today on AI evidence rather than on ESG analogy.

What it is not, and this bounds the claim. A disclosed incentive metric is not a Five Breakpoints readiness score. It measures one breakpoint (Momentum Mirage, with an Incentive Fragmentation edge) at one altitude — the compensation committee — and says nothing about strategic clarity, process friction or capability further down. It is a proxy for the framework, not the framework, and describing it otherwise would repeat exactly the over-claim the 2026-08-30 update warned about. Positioned honestly it is still the strongest thing available: the first measurement of whether the incentive layer's stated AI intent predicts anything.

Sequencing and dependency. Blocked on one cheap step: read Cohen et al.'s outcome variable. If their ESG outcome is rating-based, the design transfers but the precedent's headline result does not, and the essay writes itself differently. Do that before scoping. UNBLOCKED 2026-09-01 — the outcome variable was read, and it is better than the good case.

Update 2026-09-01 — the blocker is cleared and the design gets sharper, larger and more defensible than it was when proposed. Cohen et al.'s primary outcome is not rating-based: it is ∆CO2, Scope 1 emissions in tons from Trucost — an externally measured physical outcome (they also run ratings and financial performance as secondary outcomes). Three consequences for this asset, all favourable.

  1. The precedent's headline result transfers, and it is the result this asset most needed. Their Table 8 does exactly the activity-versus-result decomposition this asset proposes, on 21,715 firm-years: the generic "any ESG metric" indicator is a null on emissions (−0.07, t = −0.85), while the result-naming carbon metric is −0.77 at the 1% level. The coding scheme proposed here — activity-defined versus result-defined — is not a novel construct that a reviewer can call arbitrary. It is the published decomposition of a top-three accounting journal, and it produced the paper's only significant outcome coefficient. Cite Table 8 Column (2) as the template.
  2. The design should now be three-way, not two-way. Cohen et al. run three dependent variables — physical outcome, commercial rating, financial performance — and the interesting result is that they disagree with each other. Port that: code the AI metric at T, then measure at T+2 against (a) a substantive AI outcome, (b) whatever "AI maturity" scoring exists commercially, and (c) financial performance. If the activity metrics move the scores and not the substance, that is the Momentum Mirage result, measured, on AI, in public documents. A two-way design would have missed the finding that makes the ESG paper worth citing.
  3. Build in the cost measurement, because it is the non-obvious half. Their Table 10 shows the carbon metric — the one that worked — carrying the sample's largest negative stock return (−0.079, t = −2.66) and a negative ROA coefficient. Result-defined metrics are expensive to accept, which is a candidate explanation for why boards write activity metrics instead, and it is a far more interesting finding than "boards draft badly." The AI version of this question is directly answerable from the same public data.

The one genuinely hard problem this update surfaces, and it should be solved at scoping rather than discovered at analysis. Carbon had Trucost — an independent, standardized, cross-firm measurement that predated the metric. AI has no Trucost, and the missing outcome measure is this asset's central design risk, not a detail. Fix the T+2 outcome definition before any coding begins, exactly as the 2026-08-30 discipline required for asset 3b, and be willing to conclude that the honest version of this study is smaller than the ambitious one. This is also the substance of the new measurability-confound challenge in thesis-challenges.md, so the asset and the challenge now share a critical path: whoever defines the AI outcome variable resolves both.

Revised status of the dependency: no longer blocked; the remaining prerequisite is Brandon's call on the T+2 outcome definition and the universe (S&P 500 versus Russell 3000).

Feeds: Equilar KB entry and the "Paid for Deployment" pattern (research-evidence); knowledge-base/cohen-kadach-ormazabal-reichelstein-esg-pay-international-2023.md; the soft-metric objection and the unfalsifiability challenge in thesis-challenges.md; Paper 3 mechanism 3; growth edge #2 in brandon-development.md.

Effort/risk: medium/low — coding proxy disclosures is real work, but it is desk work on public documents with no access dependency and a published template.

Status: PROPOSED — unblocked 2026-09-01 (Cohen et al. read in full; see update above). Needs Brandon's call on two things only: the T+2 AI outcome definition, and the universe (S&P 500 versus Russell 3000). Of the four assets in this ledger, this is now the one with the fewest open dependencies and a published methodological template.

Update 2026-09-02 — three changes from the Wednesday falsification sweep, two of which make this asset better and one of which prices it honestly.

  1. The coding scheme needs restating, and the published-template argument gets stronger, not weaker. The 09-01 update claimed Cohen et al.'s Table 8 is "the published decomposition" of activity-defined versus result-defined metrics. Appendix B shows it is not: their taxonomy is by subject matter, and grammar varies inside it (carbon "kg CO2e/tonne", safety "DART incident rate per 100 full-time employees", diversity "percentage of women among the SMP" are result-worded; compliance "continue to assess human rights, bribery and corruption and other related risks" and governance "establish standalone corporate governance and risk procedures... that build trust" are activity-worded). What Table 8 licenses is a topic-specificity template — name the outcome you are measuring, or don't — and that is the comparison this asset should port, because it is the one with a significant coefficient behind it. The activity/result grammar coding remains worth collecting as a second, genuinely novel axis; it is simply not the one with a JAR precedent. Stating it correctly at scoping is cheaper than being corrected at referee.

  2. The three-way outcome design should become four-way, and the fourth arm is free. Table 9's carbon row lands at Refinitiv +0.001 (t = 0.07) — a precise zero on the disclosure-based score — Sustainalytics −0.583 (t = −2.12) and KLD +6.660 (t = 10.82), with Appendix A defining both of the latter so higher = better. The metric that verifiably cut emissions moved the three appearance measures in three different directions. The design lesson: do not pick one commercial "AI maturity" score, pick every one that exists and treat their disagreement as a finding rather than as noise. Vendor divergence on the appearance measure is itself the Momentum Mirage result, and it costs one extra join.

  3. The central design risk is now measured, and it is bigger than the 09-01 update said. That update named the missing AI outcome variable as the asset's key risk. Today sharpens it: carbon is the only domain in the entire paper with a physical-outcome test. Every other category — safety included, despite DART rates being as externally counted as Trucost tons — is tested solely against ratings and financial performance. So the precedent supporting "result-defined metrics move real outcomes" is a one-domain finding, and the one domain is the one that happened to have an outcome vendor. Do not describe the ESG literature as having established this generally; it has established it once. The honest pitch for this asset is stronger for saying so: the ESG precedent proved this in the single domain where somebody had already built the measurement — and the AI question is whether anyone will build it before the incentives are written.

One new and cheap prerequisite, added to Brandon's two open calls. Table 3 Panel A carries an eleventh and twelfth category the paper never estimates — Self evaluation (scores defined and measured by the firm), 884 firms, and External evaluation (scores defined and measured by external parties), 97 firms. That split is the measurability axis, already constructed, never regressed. Before scoping this asset, it is worth one email to the authors asking for that specification. If self- and externally-measured scores behave differently, this asset's outcome-definition problem is half solved by an existing dataset; if they behave alike, the asset's novelty claim gets sharper. Either answer is worth more than a week of coding.

Status: PROPOSED — unblocked, and now with a corrected template, a four-arm outcome design, and one cheap author query that should precede scoping.


5. "The Human-in-the-Loop Control That Isn't" — the autonomy contrast, with two experiments behind it PROPOSED

The argument, exactly (two sentences): Nearly every agent-governance framework in the base asserts "human in the loop" as a control, and two randomized field experiments in the same function now show the phrase means nothing without three further specifications — a supervisor was assigned to every AI-eligible chat in the Alibaba experiment and the customer rating still fell −0.412 on the chats the AI touched, because oversight quality is conditional (it held on technical escalations and failed on emotional ones) and no dashboard the organization owned would have shown it. The essay states the specification a real human-in-the-loop claim has to meet — which classes of failure the human is expected to recover, what evidence exists that they can, and what the review is measured on that is not throughput — and grounds it in the autonomy contrast the base can now draw for the first time.

Why now, and why it is stronger than a one-paper post. The value is not one experiment; it is that the base now holds both arms of the autonomy axis in one function. Alibaba (agentic, handoff mid-failure) degraded the touched-work quality; Brynjolfsson–Li–Raymond (assistive, human-in-the-loop throughout) improved customer NPS and productivity, with gains concentrated in novices. Same function, opposite sign, differing on autonomy. The consulting-usable object is the three-question audit; the intellectual spine is that "human in the loop" collapses two designs — assistive and agentic-with-handoff — that behave oppositely, and calls both a control. That is a distinction nobody in the governance-commentary market is drawing, and it is measurement-anchored on both sides.

Asset type: Substack long-form + a reusable three-question audit card for the consulting practice (deployable this week per the scan's 09-03 strategy routing, independent of the essay's timeline). Not time-priced — the scan's own market read is that the Alibaba find is fourteen weeks old and sitting unread, so this is argument value, not a window.

Feeds: knowledge-base/alibaba-agentic-ai-human-in-the-loop-field-experiment-2026-05.md (research-evidence); knowledge-base/brynjolfsson-li-raymond-generative-ai-at-work-2023-04.md; the "Does anyone design for verification cost" thread (nine cost instruments, none testing whether the checking works); the governance-survey prevalence (86% piloting / 46% in production / 20% mature governance) as the "asserting the control without the evidence" population.

Bounds on the claim, all binding. Both experiments are single-function (customer support) and single-setting; the three-question audit generalizes as a framework, not as measured fact outside customer service — say so. Alibaba is an unreviewed preprint with an Alibaba co-author (conflict runs against the finding) and a 2024 system; BLR is a 2020–2021 GPT-3-era assistive tool. Neither is a 2026 frontier-agent claim, and the essay's power is the contrast structure, not either coefficient standing alone.

Effort/risk: low/low — both sources are in the base, the argument is a contrast rather than new empirics, and the audit card is a page.

Status: PROPOSED — needs Brandon's call on whether the three-question audit ships to the practice now (per scan 09-03 strategy routing) ahead of the essay.

Update 2026-09-04 (Stafford) — a correction that must land before the audit card ships, plus a second exhibit that makes the card better.

The scan's 09-04 strategy routing proposed shipping the WP 35275 generational gradient into the consulting instrument ahead of the agent-intrusion piece, on the strength of "three tool generations, each buying a larger upstream gain, none moving the shipped result proportionally." Two of that argument's three load-bearing claims fail against the paper's full text and must not be shipped.

  1. σ = 0.25 is a calibration, not an estimate. Two free parameters (θ ≈ 0.75, σ ≈ 0.25) fit to the autocomplete attenuation curve alone, held constant across layers, on a model the authors describe verbatim as "highly stylized" and explicitly decline to treat as identification — "mapping the data into these parameters provides suggestive evidence." It identifies the curvature of an attenuation curve. It is not a measurement of how hard the human step binds, and in a client deck it is exactly the kind of number Brandon's own Paper 3 retired the 70% figure for.
  2. The third generation's release effect was never estimated. Table 5, verbatim: async release effects are "not reported" because "async agents do not have the capability to release a repository directly, so this outcome is zero by construction." The abstract's 30% is autocomplete's 10.2% plus sync's 20.3%. A structural absence is not a flat result, and one reader with the PDF open ends that argument.

The version that holds is shorter, needs no caveat, and is the paper's own conclusion sentence: sync agents produced 741% more lines of code, 65% more pull requests, and 20% more releases — a tenfold compression between what gets written and what gets shipped, measured on the same developers through every intermediate step. Add Figure 1's cumulative multiples if a second beat is wanted: 17.3× lines of code, 1.3× releases. Both verified from Table 5 / Figure 1. One composition note to carry if the 20% is used: it is almost entirely Claude Code (+29.2, SE 6.0), with GitHub Sync (+1.3, SE 8.2) and OpenAI Codex (+0.2, SE 4.3) indistinguishable from zero.

And the card gains a second exhibit that is better than the first. The three-question audit's premise is that oversight quality is unmeasured and assumed. METR (arXiv 2507.09089) supplies the measured version at the individual altitude: 16 experienced developers, 246 randomized tasks, 19% slower with AI — and they estimated afterwards that they had been 20% faster. Expert forecasters predicted 38–39% speedups in the wrong direction. That is a 39-point gap between measured and perceived progress under randomization, and it converts the card's third question — what is the review measured on that isn't throughput — from a governance prompt into a demonstrated failure: the people doing the work could not tell, and neither could the experts.

Bounds, binding. n = 16 and early-2025 tooling (Cursor Pro, Claude 3.5/3.7 Sonnet); cite it as what happened to sixteen experienced maintainers under randomization, never as developers are slower with AI. And it must travel with 35275 rather than instead of it — the standing instruction logged in thesis-challenges.md today is that the base does not cite one of a paired result without the other.

Sequencing unchanged and now firmer: the 08-31 agent-intrusion window (24 days, unwritten, fifth Friday as the flagged priority) goes first. This amendment removes the only argument for jumping the queue.

Update 2026-09-05 — the audit becomes four questions, and the fourth is the only one with an equation behind it.

The three questions as drafted — which failure classes must the human recover, what evidence exists that they can, what is the review measured on that isn't throughput — are diagnostic, defensible, and all three are Brandon's assertions. Bartolucci & Vivo (arXiv 2605.27202) supply the missing fourth and it is derived rather than asserted. Their review threshold is π⋆(θ) = θ/(κK): the congestion cost of reviewer time, over verification skill times K, the cost of an escaped error. So:

Q4. What does it cost this organization when a bad AI output gets through — and who specifically bears that cost?

Why this is the strongest question of the four. If the answer is "nothing," or "someone three steps downstream," then K ≈ 0 and π⋆ → ∞. The review is not degraded under load; it was never going to happen at any load, and the reviewer is behaving correctly. That converts "human in the loop" from a weak control into an inoperative one, and it is a finding deliverable in a first meeting with no measurement at all — every other question in the audit requires the client to know something about their own deployment.

And it is the cheap one. All three parameters sit in the same expression, but θ and κ are bought with headcount and training while K is set by a decision about accountability. Nothing in the eleven verification instruments across six functions, the Five Eyes prerequisites or the FRC guidance assigns a cost to an escaped output — they all assign a person. The essay's spine sharpens accordingly: the human-in-the-loop control has three parameters, your governance framework specifies one of them, and the cheapest of the other two is the one nobody has ever set on purpose.

Bounds, unchanged and one added. The paper contains no data — figure values are "a representative calibration" and "illustrative levels," and the "Empirical calibration" section is an agenda. Use the functional form and the definitions; never cite a number from it. The audit card is a framework claim, not a measured one, and Q4's force comes from the client's own answer, not from the model.

Feeds: as above, plus knowledge-base/bartolucci-vivo-queue-and-ai-variance-wedge-2026-05.md (research-evidence, staged 09-05).

Revised status: PROPOSED — the audit card is the only asset in this ledger shippable this week without a decision from Brandon, since it ships independent of the essay. Four questions, one page.