Stafford

Knowledge base

13 entries owned by this repo (the volume layer lives in research-evidence).

Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity (METR)
METR (Model Evaluation & Threat Research) — arXiv preprint 2507.09089 · published 2025-07-12 · flagged 2026-09-04

Tags: paper-2, practitioner, Momentum Mirage, Technology Illusion, Capability, Momentum, COUNTERARGUMENT Source: METR (Model Evaluation & Threat Research) — arXiv preprint 2507.09089 URL: https://arxiv.org/abs/2507.09089 Date published: 2025-07-12 Date flagged: 2026-09-04

Summary

Joel Becker, Nate Rush, Elizabeth Barnes and David Rein ran a randomized controlled trial on 16 experienced open-source developers across 246 real tasks drawn from mature repositories the developers already maintained (average ~5 years of prior experience on those projects). Randomization was at the task level: each issue was randomly assigned to allow or disallow AI tool use. The tools were the early-2025 generation — Cursor Pro with Claude 3.5/3.7 Sonnet — not the async agents of 2026.

The measured result runs opposite to every prediction attached to it. AI tool access increased task completion time by 19%. The developers themselves had forecast a 24% reduction before the study, and after completing the tasks still estimated a 20% reduction — they were slower and believed they had been faster. Outside forecasters were further off in the same direction: economics experts predicted a 39% speedup, machine-learning experts 38%.

The authors are careful about scope. They state that "the influence of experimental artifacts cannot be entirely ruled out," while arguing the effect's robustness across their analyses suggests it is not primarily an artifact of design choices. They do not generalize the finding to all developers, all tasks or later tool generations.

Five Breakpoints Intersection

  • Momentum Mirage: The gap between measured and perceived speed is roughly 39 percentage points inside a randomized design on the same people — developers who were 19% slower reported being 20% faster. This is the appearance of progress diverging from actual movement, measured at the individual altitude rather than inferred from an organizational dashboard, and it is the tightest instance of the mechanism anywhere in this base.
  • Technology Illusion: A more capable tool deployed onto an unchanged expert workflow on a mature codebase produced negative measured returns, which is direct evidence that capability added to existing conditions does not automatically convert to output — here it subtracted.

Four Forces Intersection

  • Capability: The effect is conditional on where the worker already sits — this population is experienced developers on codebases they know deeply, exactly the case where the tool's contribution to missing knowledge is smallest and its context-acquisition cost is largest. Read against Brynjolfsson, Li & Raymond (+34% for novices, ~0 for experts), the two instruments describe one gradient that continues past zero into negative territory for the most expert case.
  • Momentum: Self-assessed progress and measured progress were decoupled within a single instrument, which is the Momentum force failing at the point of measurement rather than at the point of effort.

Relevance to Brandon's Positioning

This is a counterargument to the base's own architecture, not to Brandon's thesis, and the distinction is the whole value of the entry. The "Unconverted Gain", "Migrating Bottleneck" and "Collapse the Handoffs" patterns all rest on one premise: the individual-level AI gain is real and large, and it dies somewhere between the person and the organization. This RCT reports that for experienced developers working on mature codebases — the population most enterprise software work actually consists of — the upstream gain is not merely unconverted, it is negative. If that holds, "the gain dies at the handoff" is the wrong story for a material share of the population, and the honest framework claim is narrower and better: AI's organizational return depends on where the worker sits, and there exists a regime where deployment costs output outright.

It also gives Brandon the cleanest Momentum Mirage citation he has ever had, and it is at the individual altitude where his book manuscript operates (books/built-to-be-replaced.md). Every existing Momentum Mirage instrument in this base measures an organization mistaking activity for progress. This measures a person doing it, under randomization, with the counterfactual held. The sentence — they were 19% slower and were certain they were 20% faster — does work in a paper, an essay and a client room, and it needs no organizational inference to land.

Weigh it against NBER WP 35275 with the tension stated rather than resolved: 35275 finds large positive task-level effects on commits and lines of code; METR finds negative effects on completion time. They are not directly contradictory — different outcomes (volume of artifacts versus time to close a task), different populations, different tool generations, and 35275 itself cites this paper as the "mixed" evidence on agentic tools. But a base that cites 35275's upstream gain without citing METR's negative one is selecting its evidence.

Sources & Confidence

  • Becker, Rush, Barnes & Rein, "Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity," arXiv 2507.09089, v1 2025-07-12, v2 2025-07-25 — https://arxiv.org/abs/2507.09089Moderate. Abstract and record read directly; figures (16 developers, 246 tasks, +19% completion time, −24% pre-forecast, −20% post-estimate, 39%/38% expert forecasts) taken from the paper's own abstract. Unreviewed preprint. METR is an AI-evaluation nonprofit, not a commercial vendor with a product in this market.
  • Corroborating citation: Demirer, Musolff & Yang, NBER WP 35275 (May 2026), which cites this study by name as finding "that agentic coding tools reduce experienced developers' productivity by 19% in an RCT" while other studies report gains of 35–40% — High (read from the primary PDF, author-hosted copy).

Instrument note (recorded before the finding, per CLAUDE.md): data-source genre randomized controlled trial with direct time measurement; inference method task-level randomization of AI-tool permission within developer, on real issues in the developers' own repositories, with screen-recorded/self-timed completion; observation window early 2025, tool generation Cursor Pro + Claude 3.5/3.7 Sonnet — pre-async-agent; no shared right-hand side with any other entry in this base. This is the base's first instrument in which AI's measured effect on a work outcome is negative.

Three limits that bound every use, and the first is severe. n = 16 developers. 246 tasks give the estimate its power, but sixteen people is not a population and the paper must never be cited as "developers are slower with AI" — the claim it supports is these sixteen experienced maintainers, on their own mature codebases, with early-2025 tools, were slower. Second, the tool generation is a year and a half behind the frontier the market is currently arguing about; 35275's async-agent results are a different technology. Third, open-source maintainers on repositories they personally own are not employees inside a firm — the same population limit that binds 35275 binds this, and for the same reason.

Filed by / Routing

Filed by: Stafford (daily brief, 2026-09-04) — surfaced through the reference list of NBER WP 35275 during full-text verification; not present in the 494-file base. Routing: strategy / personal

Generative AI at Work
NBER Working Paper 31161 (released April 2023); published *The Quarterly Journal of Economics* 140(2):889–942 (2025) · published 2023-04-23 · flagged 2026-09-03 · paper-2 · Technology Illusion · Momentum Mirage · Capability · COUNTERARGUMENT

Tags: paper-2, Technology Illusion, Momentum Mirage, Capability, COUNTERARGUMENT Source: NBER Working Paper 31161 (released April 2023); published The Quarterly Journal of Economics 140(2):889–942 (2025) URL: https://www.nber.org/papers/w31161 Date published: 2023-04-23 Date flagged: 2026-09-03

Summary

Brynjolfsson, Li & Raymond study the staggered rollout of a generative-AI conversational assistant to 5,172 customer-support agents at a Fortune 500 firm that sells business-process software; the majority of agents work from offices in the Philippines, with smaller groups in the US and elsewhere. The tool — built on a GPT-family large language model, fine-tuned on the firm's own resolved conversations — surfaces real-time response suggestions to the agent, who remains the decision-maker on every chat. Access was onboarded team-by-team over time (roughly 5% of workers had access in October 2020, growing to ~70% by January 2021), which the authors exploit as staggered treatment. Outcomes are read from platform logs and surveys: resolutions per hour (productivity), resolution rate, and net promoter score (NPS), a post-chat customer-satisfaction survey.

Access to the assistant raised productivity — issues resolved per hour — by 14% on average, with a 34% improvement for novice and low-skilled workers and minimal effect on experienced, highly skilled workers. The authors' stated mechanism is that the model disseminates the tacit best practices of the most able workers to everyone else, moving newer workers down the experience curve faster. AI assistance also improved customer sentiment (NPS), raised resolution rates, and reduced attrition — the retention gain driven specifically by newer workers. The assistant is squarely assistive, not agentic: it recommends, the human executes, and the human stays in the loop on every conversation.

The paper carries its own longer-run caution, which matters for the framework more than the headline does: top workers increasingly adhere to AI recommendations even when those recommendations are worse than what the worker would have done unaided, a deskilling / judgment-erosion signal the authors flag explicitly. The vintage is material — this is a 2020–2021 deployment of a GPT-3-era assistant, released as NBER WP 31161 in April 2023 and published in the QJE in 2025 — so it is argumentative and structural value, never a claim about 2026 frontier systems.

Five Breakpoints Intersection

  • Technology Illusion (counter-evidence, and this is why the entry is a COUNTERARGUMENT): the framework's Technology-Illusion reading holds that technology deployed onto an unchanged role produces the appearance of gain without the substance; here an assistive tool deployed onto the existing agent role, with no workflow redesign, raised resolutions/hour 14% and improved externally-measured customer NPS — a real, substantive, quality-positive result. The framework does not get to dismiss this; it has to absorb it (see Positioning), and the honest absorption is that the tool itself did the organizational work — it codified and disseminated top-agent tacit knowledge, i.e. it was a capability intervention, not "tech on broken conditions."
  • Momentum Mirage (a genuine edge inside a pro-tool result): the authors' finding that top workers increasingly follow AI suggestions even when those suggestions are worse is a measured instance of the appearance of a best-practice signal displacing the substance of expert judgment — the framework's own mechanism, arriving as the paper's cautionary tail rather than its headline.

Four Forces Intersection

  • Capability: the entire effect structure is a Capability story — the gain is 34% for novices and ~0 for experts because the tool substitutes for missing capability where capability is lowest and adds nothing where it is already high; the AI is functioning as a real-time knowledge-transfer/coaching layer, not as autonomous throughput.

Relevance to Brandon's Positioning

This is the single most consequential gap-fill available to the base right now, and it does three distinct jobs. (1) It is the base's missing anchor experiment. Across 492 KB files the entire canonical AI-field-experiment literature is absent — this paper, Dell'Acqua et al.'s "Navigating the Jagged Frontier," Peng et al.'s Copilot RCT, Noy & Zhang (Science) — and the 2026-09-03 paper-2 case file explicitly calls the 2026-05 Alibaba paper "the first entry in the case file where an AI deployment was randomly assigned." That framing was true only because BLR was never ingested; Brynjolfsson himself is already in the base as co-author of the Stanford Enterprise AI Playbook, which makes the omission clearly accidental. (2) It is the assistive comparator the 2026-09-03 scan brief wished for — same function as Alibaba (customer support), opposite sign on quality, differing on exactly one axis: BLR is assistive-in-the-loop and improved NPS; Alibaba is agentic-with-handoff and cut the customer rating −0.412 on the chats the AI touched. The pair isolates what autonomy costs with function, workforce and log-based measurement approximately held constant — the within-territory contrast the case file named as its highest-value acquisition. (3) It is a standing counterargument to the redesign-prerequisite claim (Claim 3 / Claim 5) that the 2026-09-02 falsification sweep looked for and missed — "measured AI returns without workflow or role change" — and it is the most-cited such result in existence. Logged in full in thesis-challenges.md; the crux is whether "the AI as knowledge-dissemination layer" counts as the redesign the framework requires, which is a real definitional question about the framework's scope rather than a point to wave away.

Sources & Confidence

  • Brynjolfsson, Li & Raymond, "Generative AI at Work," NBER WP 31161 (2023-04-23) / QJE 140(2):889–942 (2025). High — top-five economics journal, NBER peer-track, randomized/staggered design with administrative outcomes. URL: https://www.nber.org/papers/w31161
  • Sample and window read from the working-paper full text (CIRJE mirror of the 2024-11-18 draft): 5,172 agents; ~5% access Oct 2020 → ~70% Jan 2021; outcomes resolutions/hour, resolution rate, NPS. (An SSRN summary lists 5,179 agents; cite the paper's 5,172.)

Filed by / Routing

Filed by: Stafford (daily brief, exploratory slot, 2026-09-03) Routing: strategy (the assistive-vs-agentic autonomy contrast is the consulting-usable object); market (positioning — the base can now state the autonomy-cost comparison the market states with anecdote); personal (the base's missing-experiment gap is structural and should be closed deliberately)

Executive Compensation Tied to ESG Performance: International Evidence
*Journal of Accounting Research* (Cohen, Kadach, Ormazabal & Reichelstein) · published null · flagged 2026-08-31 · paper-2 · Momentum Mirage · Strategic Disconnection · Commitment · Purpose

Tags: paper-2, Momentum Mirage, Strategic Disconnection, Commitment, Purpose Source: Journal of Accounting Research (Cohen, Kadach, Ormazabal & Reichelstein) URL: https://onlinelibrary.wiley.com/doi/10.1111/1475-679X.12481 Date published: null Date flagged: 2026-08-31

Entry status — rewritten 2026-09-01 on a verified full-text read. This supersedes the 2026-08-31 abstract-only version, and it reverses that version's conclusion. The 08-31 entry was tagged COUNTERARGUMENT and stated that this paper cut against the framework. It does not. The full text was obtained from the University of Mannheim institutional repository (madoc.bib.uni-mannheim.de/63834/1/ESGPay.pdf, 60pp., text-extractable) after Wiley, SSRN, Taylor & Francis and the ECGI PDF all failed on 08-31. The abstract aggregates a result that the paper's own Table 8 decomposes, and the decomposition runs the other way.

Citation correction: the article is 61(3): 805–853, June 2023 — not 61(4), as the 08-31 entry recorded. Working paper: ECGI Finance WP N° 825/2022 / CEPR DP17267; SSRN 4097202; manuscript cover-dated December 2022. Date published remains null because only the issue month, not an exact publication day, could be verified from a primary.

Summary

Cohen, Kadach, Ormazabal and Reichelstein study the incorporation of ESG metrics into top-executive compensation contracts across international public firms. The instrument is ISS Executive Compensation Analytics (ECA) for contract terms, Datastream/WorldScope for accounting and market data, FactSet/LionShares for institutional ownership, Trucost for carbon emissions, and Refinitiv, Sustainalytics and MSCI (KLD) for commercial ESG ratings. Sample selection runs from 53,565 ECA firm-year observations to 35,076 (6,262 firms) with Datastream and FactSet coverage, and to 22,603 firm-year observations covering 4,395 firms in 21 countries where non-missing Trucost data is required. The observation window is 2011–2020. By 2020, 1,198 firms — 31% of sample firms that year — had adopted ESG Pay.

The headline sentence in the abstract is that "the adoption of ESG variables in managerial performance measures is accompanied by improvements in ESG performance and meaningful changes in the compensation of executives." Section 6 shows what that sentence is made of, and it is not one result but three, against three different dependent variables.

Emissions (Table 8, N = 21,715, firm and year fixed effects, standard errors clustered at firm, R² = 0.15). The dependent variable is ∆CO2, "the year-to-year change in the firms' direct GHG emissions (measured in tons of CO2 equivalent)" — Scope 1, chosen because "these are emitted by the firm itself rather than parties along the firm's supply chain." This is an externally measured physical outcome, not a rating. Column (1) regresses ∆CO2 on ESG Pay, an indicator for incorporating any ESG metric: the coefficient is −0.07 (t = −0.85), not significant. Column (2) replaces that single indicator with indicators for each of nine metric categories (corrected 2026-09-02 — the 09-01 entry said "ten"; the estimated categories are the nine specific indicators of Table 3 Panel A, and the two score categories defined in the same panel are never estimated. See the 2026-09-02 update below). Only one is significant: Carbon emissions, −0.77 (t = −2.88), significant at the 1% level. Other environmental variables −0.11 (t = −1.02); Safety and security −0.05 (t = −0.34); Diversity and inclusion −0.04 (t = −0.15); Employee satisfaction and development +0.14 (t = 0.91); Corporate culture +0.06 (t = 0.56); Compliance −0.11 (t = −0.78); Governance −0.06 (t = −0.46); Other +0.02 (t = 0.16). The authors state the contrast themselves: "while the coefficient on ESG Pay is not statistically significant, when we focus on emission-specific components of ESG Pay (Column (2)), the coefficient on Carbon emissions is negative and significant, which is consistent with the notion that introducing emission-specific metrics in top executive compensation contracts induces emissions reduction."

Ratings (Table 9) and financial performance (Table 10). Against ∆ESG Rating, ESG Pay is −0.001 (t = −0.17) on Refinitiv — a precise zero — +0.233 (t = 1.95, 10%) on Sustainalytics, and +1.004 (t = 2.42, 5%) on KLD. The authors note the divergence is unsurprising given documented disagreement across rating vendors (Berg et al. 2022). Against financial performance, ESG Pay is −0.003 (t = −0.94) on ∆ROA and −0.032 (t = −1.82, 10%) on annual stock return; the carbon-specific metric carries the cost — ∆ROA −0.015 (t = −1.89) and Return −0.079 (t = −2.66, 1%), the largest and most significant negative in the table. The authors' caveats are explicit: "the evidence presented here is mainly descriptive, which cautions against making strong causal claims"; "the results in Tables 8-10 are not statistically strong"; and "the interpretation of Table 9 depends on one's priors on the quality of ESG ratings as measures of ESG performance." The authors read Tables 8–9 together as supporting the view that "ESG Pay strengthens a firm's pledge to improve ESG performance" and as difficult to reconcile with a pure window-dressing account.

Five Breakpoints Intersection

Momentum Mirage — this entry is now evidence for the framework's inference, which is the reverse of how it was filed on 08-31. (Axis corrected 2026-09-02: what Table 8 isolates is topic specificity — whether the contract names the outcome being measured — not activity-versus-result grammar. Appendix B shows grammar varying inside the categories rather than between them. The support survives on the corrected axis; see the 2026-09-02 update.) The base reads an incentive metric that states no result condition as rewarding the appearance of progress rather than progress. Table 8 tests the adjacent proposition on the previous non-financial metric to enter executive pay, holding firm, year, specification and dependent variable constant, and varying only whether the incentive names the specific outcome: the generic "any ESG metric" indicator moves actual emissions by a statistically indistinguishable-from-zero −0.07, while the carbon-specific metric moves them by −0.77 at the 1% level. Paired with Table 9 — where the generic metric is associated with rising scores on two of three commercial ratings — the shape is the one Momentum Mirage names: the measure of appearance moves while the measure of substance does not.

Strategic Disconnection — second and weaker, tagged on the decomposition rather than on a single claim. (Counts corrected 2026-09-02.) Eight of the nine estimated categories name a subject rather than the outcome measured, and none of them is associated with movement in the one physical outcome the paper measures. That is the framework's claim that a vaguely stated intent produces the appearance of alignment without the substance of it, observed in the most binding statement of intent a public company files.

A discipline that must travel with both tags, because it is the honest limit of the result. Eight of the nine estimated categories are being tested against ∆CO2, an outcome most of them were never intended to move — a diversity metric failing to reduce emissions is a mismatched dependent variable, not a failed incentive. (Corrected 2026-09-02; this discipline now also voids the "safety behaved differently from carbon" argument, since the safety null is one of those eight — see the 2026-09-02 update.) The load-bearing comparison is the narrow one: ESG Pay (any metric) versus Carbon emissions, same outcome, same firms, same specification. Only that comparison licenses the Momentum Mirage reading, and it must be stated in those terms or a reviewer will correctly take the wider version apart.

Four Forces Intersection

Commitment. The framework defines Commitment as decision-makers acting consistently with stated direction under pressure, and this paper prices the pressure. The metric that moved emissions is the same metric associated with the largest negative stock return in the sample (−0.079, t = −2.66) and a negative ROA coefficient (−0.015, t = −1.89). A result-defined commitment cost the firms that made it; the generic commitment cost nothing and delivered nothing measurable.

Purpose. Clarity of intent is the only variable that differs between Table 8's Column (1) and the significant coefficient in Column (2) — the same firms, the same contracts vehicle, the same fixed effects, differing in whether the stated objective names a quantity a third party measures.

No Capability or Momentum tags — neither construct is measured anywhere in this paper.

Relevance to Brandon's Positioning

The soft-metric objection logged on 2026-08-31 is defeated by the paper that raised it, and Brandon should know that before he writes a word about compensation committees. Yesterday's position was that the base had tagged Momentum Mirage from metric grammar without ever testing whether metric grammar predicts anything, and that a 2023 JAR paper appeared to say it does not. The full text says the opposite: on the one comparison that isolates grammar, the metric that named a result moved the result and the metric that did not, did not. This converts the base's largest outstanding threat to the Momentum Mirage mapping into its first piece of empirical support — from a top-three accounting journal, 4,395 firms, 21 countries, on a physical outcome measured by a third party.

What it does for Paper 3, mechanism 3. Designed to Stall asks that "a meaningful share of executive compensation, not a token slice, is tied to specific multi-year transformation outcomes." Mechanism 3 has never had evidence attached. It now has a near-analogue with a published result: the word doing the work in Brandon's sentence is "specific," and Table 8 is the measurement that shows it is the operative word rather than a stylistic one. Equilar's AI disclosures show the market adopting the vehicle Paper 3 prescribed while inverting the content — and this paper supplies the reason that inversion matters rather than leaving it as an assertion.

The cost line is the part Brandon should not drop, because it is what makes the argument non-obvious. A boardroom writing "win the AI opportunity" instead of a result-defined metric is not necessarily incompetent. Table 10 says the result-defined metric was the expensive one. Activity metrics are popular because they are free — they impose no measurable cost on the executives who accept them and no measurable movement on the outcome. That reframes the diagnosis from a drafting failure to a revealed preference, which is a harder and more interesting claim, and it is the claim the framework should make.

The residual objection, which survives and is now the live one. Carbon worked because CO2 is countable in tons by Trucost. The result-defined/activity-defined distinction may be partly confounded with whether the domain admits third-party measurement at all — and in 2026 AI has no Trucost. If no firm can yet write a defensible AI outcome metric, activity metrics are evidence of an attribution problem rather than an organizational failure, and the framework would be diagnosing a measurement limit. That is logged as the successor challenge in knowledge/thesis-challenges.md.

Transfer discipline, stated so it does not get lost. The window is 2011–2020, entirely pre-generative. This is an analogue for AI metrics, not evidence about them. The authors call their evidence "mainly descriptive" and "not statistically strong," and both phrases must travel with any use.

Sources & Confidence

  • Cohen, S., Kadach, I., Ormazabal, G., & Reichelstein, S., "Executive Compensation Tied to ESG Performance: International Evidence," Journal of Accounting Research 61(3): 805–853, June 2023; DOI 10.1111/1475-679X.12481. Full text read from the Universität Mannheim repository copy (madoc.bib.uni-mannheim.de/63834/1/ESGPay.pdf, 60pp.). — High (peer-reviewed, top-tier accounting journal; authors at San Diego State, IESE, IESE/CEPR/ECGI, and Mannheim/Stanford). Coverage is now full-text: sample, specification, outcome variables and all coefficients cited here were read directly.
  • Dell'Erba, M., & Gomtsyan, S., "Regulatory and Investor Demands to Use ESG Performance Metrics in Executive Compensation: Right Instrument, Wrong Method," Journal of Corporate Law Studies (2024); authors' summary, Harvard Law School Forum on Corporate Governance, 2024-12-02. — Moderate (peer-reviewed law journal, normative argument carrying no data). The Harvard Law Forum is the venue of the summary, not the publisher of the research — do not upgrade credibility on the venue.
  • Berg, Koelbel & Rigobon on ESG rating divergence — cited by Cohen et al. as the reason their three rating results differ; not independently read here. — Unverified in this base.
  • Unverified and not to be cited: the "3% in 2010 to over 30% in 2021" ESG-KPI adoption series, which surfaced in a search summary on 08-31 with no locatable primary. Nothing in the full text supports it; the paper's own adoption figure is 31% of sample firms in 2020.

Filed by / Routing

Filed by: Stafford (scheduled daily brief, 2026-08-31; rewritten on verified full text 2026-09-01) — exploratory slot, not a scan find. Routing: strategy — resolves the soft-metric objection and attaches the first evidence to Paper 3 mechanism 3; governance — the proxy statement as an ex-ante, publicly filed organizational variable (asset #4).


Update — 2026-09-02 (Wednesday falsification sweep, third pass: Table 3, Appendix B, Appendix C, Appendix A)

What was checked and why. The 09-01 entry closed with a residual objection — the measurability confound — and named a cheap test for it: read Appendix B's worked examples of each metric category and Table 3's taxonomy, on the theory that "if the safety metrics are worded as programme activity rather than as injury-rate targets, this challenge weakens sharply on evidence already downloaded." That test was run today against the same Mannheim PDF. It came back the other way, and it produced three findings that change how this paper may be cited. None of the 09-01 numbers is wrong; what changes is what the decomposition is a decomposition of.

1. The safety category is result-defined and externally measured — so it cannot weaken the confound, and the null attributed to it never tested anything

Appendix B's worked example for Safety and security is "Days Away/Restricted or Transfer (DART) incident rate per 100 full-time employees" (New Jersey Resources Corporation, 2019). DART is an OSHA-defined, mandatorily reported, cross-firm-standardised rate — an injury-rate Trucost. So safety metrics in this sample are not activity-worded; they name a counted result in a domain with pre-existing third-party measurement infrastructure. The route by which the confound might have been weakened is closed.

And the safety coefficient adjudicates nothing at all, for a simpler reason. The −0.05 (t = −0.34) is measured against ∆CO2. Safety metrics have no reason to reduce carbon emissions. That cell is a mismatched-dependent-variable null exactly like the other seven, and any argument resting on "safety behaved differently from carbon" — including the one recorded as limiting evidence in knowledge/thesis-challenges.md on 09-01 — is invalid and has been struck. The paper never regresses safety metrics on injury rates, though DART data are as externally measured as Trucost tons.

2. The decomposition is a subject-matter taxonomy, not a grammar taxonomy — and this narrows the 09-01 reading

Appendix B shows metric grammar varying freely inside the categories rather than between them. Four categories are worded as counted results: carbon — "Greenhouse gas emissions intensity at gold producing operations measured in kg CO2e/tonne" (AngloGold Ashanti, 2020); safety — the DART rate above; diversity — "Percentage of women among the SMP (Senior Management Position)" (BNP Paribas, 2020); employee satisfaction — "Internal promotion rate in global leadership" (Adecco, 2020). Three are worded as activity or aspiration: compliance — "FY2021 actions and targets (continue to assess human rights, bribery and corruption and other related risks)" (Sandfire Resources, 2021); governance — "Establish standalone corporate governance and risk procedures at the company following internalization that build trust, create long-term securityholder value and align with company values" (Waypoint REIT, 2020); corporate culture — "Colleague Culture & Engagement survey" (Lloyds Banking Group, 2020), which names an instrument and no threshold.

Consequence for the load-bearing comparison. ESG Pay is an indicator for holding any ESG metric; Carbon emissions is an indicator for holding a carbon metric. The contrast between −0.07 and −0.77 therefore isolates topic specificity — whether the contract names the outcome you are measuring — not activity-versus-result grammar. The 09-01 entry was right that the narrow comparison is the only licensed one; it named the wrong axis. The framework's inference survives this and should be restated on the axis actually tested: an incentive that names the specific outcome moves that outcome; an incentive that gestures at the category does not. Qorvo's "exploration and deployment of AI tools" fails on both axes at once, which is why the reading still holds — but a reviewer who checks Appendix B will find the grammar claim unsupported by this table, and the sentence must be written so that it does not depend on it.

3. The variable that would settle the measurability confound exists in this paper, was constructed, and was never estimated

Table 3 Panel A defines eleven categories in two families, not nine. Alongside the nine specific indicators (# firms: carbon 172, other environmental 652, safety and security 744, diversity and inclusion 250, employee satisfaction and development 771, corporate culture 519, compliance 259, governance 397, other 161) sits family (b) Scores, split precisely on who does the measuring:

  • Self evaluation"scores defined and measured by the firm"884 firms. Worked example: "Combination of 3 criteria: (1) Diversity and equal opportunities; (2) Strengthen our People and the Digital Transformation of the Company; (3) Ethics and Good Governance" (Enagas SA, 2020).
  • External evaluation"scores defined and measured by external parties"97 firms. Worked examples: "Inclusion over the three-year period 2020-2022 in the DJSI, FTSE4GOOD, and CDP Climate Change" (Italgas SpA, 2020); "Bloomberg ESG disclosure score" (Newmont, 2020); "MSCI ESG rating" (Standard Bank, 2020); "Maintain citation in Bloomberg 'Gender-Equality Index'" (Scentre Group, 2021).

That split is a measurability axis with grammar approximately held fixed — both families say "achieve a score" — and it is exactly the variable the measurability confound requires. Neither category appears in Table 8 column (2), nor in Table 9 columns (2), (4) or (6), which contain only the nine specific indicators; and a full-text pass over pages 1–45 finds no discussion of either category anywhere in the body. They appear only in Table 3 and Appendix B. Table 3's own note (2) records a sample restriction on the score counts — "Restricted to the companies that use distinctive environmental metrics in the compensation contract" — which may be the reason, and no inference of suppression is available or intended: the fact is only that the estimate does not exist. The confound is therefore not resolvable from this paper's published tables, but it is resolvable from this paper's data — by a regression the authors already have the variables to run.

Appendix C sharpens the same point with a worked contract. Schneider Electric's 2020 LTI ties 6.25% to CDP Climate Change scored "0%: C score; 50%: B score (25% at B-); 100%: A score (75% at A-)" and 6.25% to DJSIW scored "0%: not in World; 50%: included in World; 100%: sector leader." These are threshold-bearing, result-defined metrics whose "result" is a third-party rating rather than a physical quantity — the cell that separates measurable by an outsider from physically counted. It is populated, and it is untested.

4. Table 9's carbon row — the substance and the appearance come apart, and the direction is not the flattering one

The carbon metric, the only contract term in this paper shown to move a physical outcome (−0.77 on ∆CO2, 1%), lands on the three rating vendors as: Refinitiv +0.001 (t = 0.07), a precise zero; Sustainalytics −0.583 (t = −2.12), significant at 5%; KLD +6.660 (t = 10.82), significant at 1%. Appendix A defines both Sustainalytics and KLD so that "a higher score indicates better ESG Performance," and Refinitiv's score as "based on the self-reported information." So the coefficients are directly comparable in sign, and they disagree: the metric that verifiably cut emissions is associated with no movement at all in the disclosure-based score and a decline in one of the two performance scores.

Two readings, and the entry commits to neither. Either the raters are poor measures of the substance they claim to score — the authors themselves defer to Berg, Koelbel & Rigobon on vendor divergence, and say plainly that "the interpretation of Table 9 depends on one's priors on the quality of ESG ratings" — or firms concentrating incentive weight on carbon lose ground on the other pillars a composite rating aggregates. What is not available is a story in which appearance and substance track each other. For the framework this is the more interesting half of the paper and the 09-01 entry did not have it: the appearance measure and the substance measure, on the same firms in the same years, moved independently and in one case oppositely.

A hard constraint on using it. Do not state this as "ESG ratings are wrong." The composite-shift reading is live, the KLD column rests on only 1,351 observations against 19,252 (Refinitiv) and 17,148 (Sustainalytics), and the authors' descriptive-evidence caveat governs here as everywhere else in Section 6.

What this does to the entry's conclusions

  • The 09-01 defeat of the soft-metric objection stands. Nothing found today disturbs Table 8's central comparison or its sign.
  • The axis of that defeat is corrected from activity-versus-result grammar to topic specificity. Any published use must be written on the corrected axis.
  • The measurability confound is not weakened; it is sharpened and stays OPEN, and its cheap resolution route is now known to be a regression the authors did not publish rather than a study nobody has run.
  • The transfer to AI is bounded more tightly than the 09-01 entry allowed. Carbon is the only domain in this paper with a physical-outcome test; every other category is tested solely against ratings and financial performance. "Result-defined metrics move real outcomes" is a one-domain finding — and the one domain is the one that happened to have Trucost.

Sources for this update — all primary, all from the same verified full text: Table 3 Panel A (p. 46), Table 8 (p. 53), Table 9 (pp. 54–55), Appendix A variable definitions (p. 37), Appendix B (p. 38), Appendix C (pp. 39–40) of madoc.bib.uni-mannheim.de/63834/1/ESGPay.pdf. Confidence: High.

Unpacking organizational readiness for change: an updated systematic review and content analysis of assessments
BMC Health Services Research (Miake-Lye, Delevan, Ganz, Mittman, Finley) · published 2020-02-11 · flagged 2026-08-30 · paper-2 · COUNTERARGUMENT

Tags: paper-2, COUNTERARGUMENT Source: BMC Health Services Research (Miake-Lye, Delevan, Ganz, Mittman, Finley) URL: https://pmc.ncbi.nlm.nih.gov/articles/PMC7014613/ Date published: 2020-02-11 Date flagged: 2026-08-30

Summary

Miake-Lye et al. systematically review and content-analyse organizational readiness-for-change (ORC) assessments in implementation science, covering 29 uses of readiness assessments drawn from 27 publications. They extract 1,370 individual survey items and map each to the Consolidated Framework for Implementation Research (CFIR). 897 items (68%) map to the CFIR inner setting domain, concentrated in readiness for implementation (n=220), networks and communication (n=207), implementation climate (n=204), structural characteristics (n=139), and culture (n=93).

Their conclusion is about operationalization, not validation: available instruments "predominantly focus on contextual factors within the organization and characteristics of individuals," and the item-level specificity they find suggests instruments must be tailored to the scenario in which they are fielded. On the state of the field they are explicit: "No gold standard exists within the realm of organizational readiness for change assessments," and "it remains unclear how best to operationalize readiness across varied projects or settings."

The review deliberately does not assess whether these instruments work. The authors state they "did not conduct a quality assessment of the included studies, since our analysis was not focused on the validity or robustness of study findings," and place the outcome link in future work: "Work testing the relationship between organizational readiness for change and implementation outcomes will help to better specify the underlying mechanisms of readiness."

Five Breakpoints Intersection

No direct Five Breakpoints mapping — this is a measurement source, not an organizational case. Logged because it establishes the field-level status of the construct on which Claim 5 (readiness as prerequisite) depends: after roughly two decades of instrument development, readiness measurement in its most mature discipline has no gold standard and no established link to implementation outcomes.

Four Forces Intersection

No direct Force implication — measurement-methodology entry. Note for later use: the two-factor structure that dominates this literature (commitment and efficacy, per Shea et al.'s ORIC) is a partial, health-sector-specific analogue of Commitment and Capability, with no counterpart to Purpose or Momentum. That non-correspondence is itself an argument, not a mapping.

Relevance to Brandon's Positioning

This is the source that turns the Unfalsifiability-risk challenge (knowledge/thesis-challenges.md, OPEN) from a structural critique with no evidence attached into a challenge with a citation behind it. Implementation science — a field with journals, funding, and a 20-year head start on readiness measurement — has produced dozens of instruments and, on its own account, no gold standard and no demonstrated link from readiness scores to implementation outcomes. That materially strengthens the objection: the difficulty is not that nobody has tried to make readiness predictive, it is that people have tried for two decades and have not closed it.

It cuts the other way too, and this is the more useful half. The absence is field-wide, not specific to the Five Breakpoints, so the critique "your framework is unfalsifiable" applies with equal force to every readiness construct in circulation, including the ones consultancies sell. And it prices asset #3 honestly: a prospective breakpoints instrument is not a weekend project that will settle the question — it is the thing an entire field has not managed, which is exactly why doing it first would be worth something. Stronger than anything in the base on this point; the base previously held no source at all on readiness-instrument validation.

Sources & Confidence

  • Miake-Lye IM, Delevan DM, Ganz DA, Mittman BS, Finley EP. "Unpacking organizational readiness for change: an updated systematic review and content analysis of assessments." BMC Health Services Research 2020;20:106. doi:10.1186/s12913-020-4926-z — High (peer-reviewed systematic review; quotes verified against the PMC full text).
  • Corroborating, cited within this entry: Shea CM, Jacobs SR, Esserman DA, Bruce K, Weiner BJ. "Organizational readiness for implementing change: a psychometric assessment of a new measure." Implementation Science 2014;9:7. doi:10.1186/1748-5908-9-7 — High. The ORIC authors themselves: "Although ORIC shows promise, further psychometric assessment is warranted. Specifically, the measure should be tested for convergent, discriminant, and predictive validity." (Verified against PMC3904699.)

Filed by / Routing

Stafford, 2026-08-30 (scheduled run, exploratory query slot). Audience: strategy. Routing: STAGE-CANDIDATE for research-evidence; evidence anchor for the Unfalsifiability-risk challenge and for asset #3 scoping.

Organizational Readiness for Change: A Systematic Review of the Healthcare Literature
Implementation Research and Practice (SAGE) · published 2025-05-15 · flagged 2026-08-30 · paper-2 · COUNTERARGUMENT

Tags: paper-2, COUNTERARGUMENT Source: Implementation Research and Practice (SAGE) URL: https://pmc.ncbi.nlm.nih.gov/articles/PMC12084713/ Date published: 2025-05-15 Date flagged: 2026-08-30

Summary

Caci and colleagues (University of Zurich and collaborators) systematically reviewed the healthcare literature on organizational readiness for change (ORC), covering 46 studies. The review is the most recent authoritative statement on whether the construct has been linked to what it is supposed to predict, and the answer is that it largely has not been tested.

The design breakdown is the finding. Of 46 studies, 40 measured ORC only once. Nineteen measured it before implementation only. Only three measured it at all three timepoints (before, during, after). The authors state: "scholars continue to measure ORC either retrospectively or at baseline only, without prospectively linking ORC to implementation outcomes," and "the limited number of studies linking ORC to implementation identified with this SLR confirms the previously critiqued shortage of studies examining ORC prospectively."

On the state of the claim itself: "claims about its importance often have not been based on the use of nuanced theories or empirical prospective designs clearly linking ORC to implementation outcomes."

Instrument note (required for structural/empirical instruments): this is a systematic review of studies, not a primary instrument. Its unit is 46 published studies of ORC in healthcare; its inference method is structured extraction and design classification; its observation window is the published literature to the 2024 search date. It measures the state of a literature, not the state of any organization.

Five Breakpoints Intersection

No direct Five Breakpoints mapping. Logged as measurement evidence bearing on Claim 5 and on the standing unfalsifiability challenge.

Four Forces Intersection

No direct Force implication — measurement-methodology entry.

Relevance to Brandon's Positioning

This is the source that makes the gap current rather than historical, and that distinction is load-bearing.

The readiness-measurement argument assembled on 2026-08-30 rests on sources from 2009, 2011, 2014 and 2020. The obvious rebuttal to a fifteen-year-old critique is that the field has moved on. Caci et al. closes that escape route with a review published fifteen months ago: as of 2025, 40 of 46 studies still measure readiness once, and the prospective link is still described as a critiqued shortage rather than a solved problem.

That converts the argument from "this was unresolved a decade ago" to "this is unresolved now," which is the version that can be published. It is also the version that prices asset 3b honestly: the design Brandon would be running is not a catch-up exercise, it is a design the specialist field has still not executed at scale, in a setting — AI deployment — where the outcome is commercially legible in a way that clinical implementation outcomes are not.

Caution against over-reading. A shortage of prospective studies is not evidence that readiness fails to predict. It is evidence that the question is open, which is exactly what Weiner et al. 2020 concluded and exactly what Brandon should say. The dishonest version of this material is "readiness measurement doesn't work"; the honest and more useful version is "nobody has run the test properly, and the tests that came closest split perception from implementation."

Sources & Confidence

  • Caci L, Nyantakyi E, Blum K, Sonpar A, Schultes M-T, Albers B, Clack L. "Organizational readiness for change: A systematic review of the healthcare literature." Implementation Research and Practice. 2025;6:26334895251334536. doi:10.1177/26334895251334536 — High (peer-reviewed; PMC full text PMC12084713; quoted passages verified).

Filed by / Routing

Filed by: Stafford (scheduled run, 2026-08-30, 22:20 UTC — story-thread chase, not the exploratory slot) Routing: strategy — establishes the gap as current (2025), which is what makes it publishable.

Predicting implementation from organizational readiness for change: a study protocol
Implementation Science (Helfrich, Blevins, Smith, Kelly, Hogan, Hagedorn, Dubbert, Sales) · published 2011-08-12 · flagged 2026-08-30 · paper-2

Tags: paper-2 Source: Implementation Science (Helfrich, Blevins, Smith, Kelly, Hogan, Hagedorn, Dubbert, Sales) URL: https://pmc.ncbi.nlm.nih.gov/articles/PMC3157428/ Date published: 2011-08-12 Date flagged: 2026-08-30

Summary

Helfrich et al. set out a protocol to test whether a readiness instrument can prospectively predict implementation success. The instrument is the Organizational Readiness to Change Assessment (ORCA), built on the PARIHS framework: three primary scales (evidence, context, facilitation) and 19 subscales.

The design is the part worth having. ORCA is administered at baseline to 208 respondents across 53 facilities drawn from four partner projects testing evidence-based clinical practice changes; each partner project sets its own baseline timing so respondents complete the survey knowing what change is planned. Follow-up covers 158 respondents across 33 facilities, with implementation outcomes measured 6–9 months after baseline. The planned analysis is hierarchical linear modelling with the implementation outcome expressed as an effect size (Cohen's h) as the dependent variable and ORCA scales as independent variables, controlling for partner project and intervention receipt; with 30 anticipated sites they compute 90% power to detect R² ≥ 0.21.

The authors state the field's position on prior instruments directly: "few have undergone rigorous validation, notably to demonstrate the ability to prospectively distinguish successful change efforts from those that will fail."

Five Breakpoints Intersection

No direct Five Breakpoints mapping — this is an instrument-design source. Logged because it is the closest existing methodological template for the prospective breakpoints instrument in asset #3: baseline scoring before a known change, a 6–9 month outcome window, effect-size outcomes, and multilevel modelling with the site as the unit.

Four Forces Intersection

No direct Force implication — methodology entry, no organizational outcome data reported in the protocol.

Relevance to Brandon's Positioning

Two uses, both concrete. First, as a design template: asset #3 (the missing instrument) does not need to invent its method. Score organizations on the breakpoints before a known AI deployment, fix an outcome window at 6–12 months, express the outcome as an effect size, model at the site level. The will-it-hold rubric already does the scoring half; this supplies the measurement half, and the sample arithmetic is sobering in a useful way — a properly powered version of this design wanted ~30 sites, which sets the honest floor for what a Brandon-run cohort would need before it claims prediction rather than illustration.

Second, as a date stamp on the gap. The sentence "few have undergone rigorous validation, notably to demonstrate the ability to prospectively distinguish successful change efforts from those that will fail" was published in 2011. Miake-Lye et al. (2020) find no gold standard nine years later, and the ORIC authors flag predictive validity as untested in 2014. Three points on a fifteen-year line, all saying the same thing. That is the spine of the argument that the readiness-measurement gap is structural rather than a lane nobody noticed.

Caveat that must travel with this entry: this is a protocol, not results. Do not cite it as evidence that readiness predicts (or fails to predict) outcomes — cite it as a design and as the 2011 statement of the gap.

Update 2026-08-30 (22:20 UTC) — the results paper was chased and does not exist

The open follow-up flagged earlier today has been run to ground. No results paper from this protocol has ever been published. Searches conducted: Europe PMC on the full author string AUTH:"Helfrich CD" (91 records, 2006–2026 — after 2011 his output moves to burnout, PACT/medical home, de-implementation, cardiology and EHR transition, with no ORCA criterion- or predictive-validity results paper anywhere in it); the Europe PMC citation list for PMID 21777479 (29 citing articles, none a results paper, none by Helfrich); the Semantic Scholar citation graph for doi:10.1186/1748-5908-6-76 (~68 citing papers, no Helfrich-authored citing paper at all — authors do not cite their own protocol when the results paper does not exist); and targeted searches on Cohen's h, 53 facilities, hierarchical linear modelling and ORCA. The only later ORCA output from this group is conceptual: Helfrich et al., "Mapping the organizational readiness to change assessment to CFIR," Implement Sci Commun 2021;2:14, doi:10.1186/s43058-021-00121-0 — a mapping exercise, not a validation.

What makes the null worse than an absence. The instrument-development paper had already promised this exact work two years earlier. Helfrich CD, Li YF, Sharp ND, Sales AE, Implement Sci 2009;4:38 (PMID 19594942): "This analysis does not address the validity of the instrument as a predictor of evidence-based clinical practice" and "Criterion validation using implementation and quality-of-care outcomes is the next phase of our work." Promised in 2009, protocolled in 2011, never delivered. Seventeen years.

Confidence: high that no results paper is indexed in PubMed, Europe PMC or Semantic Scholar. Residual risk, stated honestly: results could exist in a non-indexed VA internal report, a dissertation, or a conference abstract. No trace of one was found, and this entry should say "not published" rather than "does not exist."

How this may and may not be used. It is legitimate to say the field's most rigorously designed prospective test of readiness was funded, protocolled and never reported. It is not legitimate to infer the study was run and found a null — file-drawer inference is exactly the reasoning this base refuses elsewhere, and there is no evidence the analysis was ever completed.

Sources & Confidence

  • Helfrich CD, Blevins D, Smith JL, Kelly PA, Hogan TP, Hagedorn H, Dubbert PM, Sales AE. "Predicting implementation from organizational readiness for change: a study protocol." Implementation Science 2011;6:76. doi:10.1186/1748-5908-6-76 — High (peer-reviewed protocol; design details and quote verified against the PMC full text).
  • Related, not separately filed: Helfrich CD et al. "Organizational readiness to change assessment (ORCA): development of an instrument based on the PARIHS framework." Implementation Science 2009;4:38 — High (instrument-development paper; not read in full).

Filed by / Routing

Stafford, 2026-08-30 (scheduled run, exploratory query slot). Audience: strategy. Routing: STAGE-CANDIDATE for research-evidence; method template for asset #3.

Providing Culturally Competent Services for American Indian and Alaska Native Veterans to Reduce Health Care Disparities
American Journal of Public Health (APHA) · published 2014-09-01 · flagged 2026-08-30 · paper-2 · COUNTERARGUMENT

Tags: paper-2, COUNTERARGUMENT Source: American Journal of Public Health (APHA) URL: https://ajph.aphapublications.org/doi/10.2105/AJPH.2014.302140 Date published: 2014-09-01 Date flagged: 2026-08-30

Summary

Noe, Kaufman, Kaufmann, Brooks and Shore adapted the Organizational Readiness to Change Assessment (ORCA) — the PARIHS-based readiness instrument developed by Helfrich et al. — for a survey of 27 Department of Veterans Affairs facilities in the Western Region, fielded in 2011–2012. The study's purpose was programmatic rather than psychometric: the authors wanted to understand what shaped the delivery of culturally competent services to American Indian and Alaska Native veterans, a population with documented access disparities.

The result that matters for the readiness-measurement question is stated plainly in the abstract: "Several ORCA subscales (Program Needs, Leader's Practices, and Communication) statistically significantly predicted whether VA staff perceived that their facilities were meeting the needs of AI/AN veterans. However, none predicted greater implementation of native-specific services."

The 2025 systematic review by Caci et al. singles this study out in exactly those terms: "The authors found that no ORCA subscale predicted implementation of native-specific services."

Design limit, and it bounds the claim. This is a single cross-sectional survey administered in 2011–2012, not a prospective baseline-to-outcome design. The readiness measure and the outcome measure were collected in the same instrument at the same time. It is therefore a concurrent null, not a prospective one, and must never be described as a failed prediction over time.

Five Breakpoints Intersection

No direct Five Breakpoints mapping. Logged as measurement evidence bearing on the unfalsifiability challenge (stafford-research/knowledge/thesis-challenges.md) and on the Claim 5 evidence gap.

The temptation here is to tag Momentum Mirage on the grounds that the instrument predicted the appearance of performance while failing to predict the substance. That tag is deliberately withheld. The finding describes the behaviour of a survey instrument, not the behaviour of an organization undergoing transformation, and mapping it across that gap would be exactly the speculative tagging the operating rules forbid. The resemblance is real and is recorded under Positioning, where it belongs as an argument rather than as evidence.

Four Forces Intersection

No direct Force implication — measurement-methodology entry.

Relevance to Brandon's Positioning

This is the most damaging single data point in the readiness-measurement literature, and it is damaging in a way that is unexpectedly useful.

The standing unfalsifiability challenge says the Five Breakpoints framework risks being a taxonomy of autopsies unless readiness can be measured ex ante and shown to predict outcomes. The obvious defence has been that nobody else has built such an instrument either. Noe et al. sharpen that defence into something with more teeth and more danger at once: when a readiness instrument was pointed at both a perception outcome and an implementation outcome in the same organizations, it predicted the perception and not the implementation.

The resemblance to Momentum Mirage is worth stating in any piece Brandon writes on this, precisely as an observation and not as evidence: the measurement literature that would validate the framework's central prerequisite appears to reproduce, inside its own instruments, the failure the framework names — a measure that tracks how well people believe things are going and does not track whether anything happened. If that pattern held prospectively it would be the strongest available argument that readiness constructs measure organizational self-image. It has not been shown prospectively, and Brandon should not claim it has.

The correct use of this source is defensive and framing, not evidentiary. It says: the hole in readiness measurement is not that nobody tried, it is that the one time somebody checked both outcomes, the instrument split them. That is a reason to build the measure carefully with a hard outcome definition — asset 3b — not a reason to believe readiness does not matter.

Sources & Confidence

  • Noe TD, Kaufman CE, Kaufmann LJ, Brooks E, Shore JH. "Providing culturally competent services for American Indian and Alaska Native veterans to reduce health care disparities." American Journal of Public Health. 2014;104(Suppl 4):S548–S554. doi:10.2105/AJPH.2014.302140. PMID 25100420 — High (peer-reviewed, APHA). Abstract verified; quoted sentence taken verbatim from the published abstract.
  • Corroborating characterisation: Caci L, Nyantakyi E, Blum K, Sonpar A, Schultes M-T, Albers B, Clack L. "Organizational readiness for change: A systematic review of the healthcare literature." Implementation Research and Practice. 2025;6:26334895251334536. doi:10.1177/26334895251334536 — High.
  • Confidence in the finding: high. Confidence in its evidentiary weight against prospective prediction: low — the design is cross-sectional and cannot bear a prospective claim.

Filed by / Routing

Filed by: Stafford (scheduled run, 2026-08-30, 22:20 UTC — story-thread chase, not the exploratory slot) Routing: strategy — feeds the unfalsifiability challenge and asset 3b's outcome-definition requirement.

Measuring Readiness for Implementation: A Systematic Review of Measures' Psychometric and Pragmatic Properties
Implementation Research and Practice (SAGE) · published 2020-06-01 · flagged 2026-08-30 · paper-2 · COUNTERARGUMENT

Tags: paper-2, COUNTERARGUMENT Source: Implementation Research and Practice (SAGE) URL: https://journals.sagepub.com/doi/full/10.1177/2633489520933896 Date published: 2020-06-01 Date flagged: 2026-08-30

Summary

Weiner and colleagues systematically reviewed the measures used to assess organizational readiness for implementation in mental and behavioural health settings, rating each on psychometric and pragmatic properties. Nine readiness measures were identified. The review's central negative finding is stated directly: "Also striking is the lack of evidence for two psychometric properties of importance to implementation scientists: predictive validity and responsiveness." Fewer than half the measures had been tested for predictive validity at all, and "of these, few were tested for associations with implementation outcomes."

The verdict on the field's most-tested instrument is the sharpest passage. Of the TCU Organizational Readiness for Change measure: "Although the TCU-ORC has exhibited statistically significant associations with implementation outcomes in some studies, it has not in others and no discernible pattern emerges across multiple studies." And: "Despite 55 tests of association between individual TCU-ORC scales and various outcomes, the predictive validity rating of the measure remains 'minimal.'"

The review closes the question it opened: "Thus, the question of whether readiness for implementation matters remains open."

Who is writing this matters. Bryan Weiner is the developer of ORIC, the readiness instrument most widely used in implementation science, and the author of the field's most-cited theory of organizational readiness for change. This is the construct's principal architect reporting that the construct's practical claim is unestablished.

Scope limit: the review covers mental and behavioural health settings. ORCA is not within its scope.

Five Breakpoints Intersection

No direct Five Breakpoints mapping. Logged as measurement evidence against the load-bearing premise of Claim 5 — that organizational readiness is the prerequisite — and as the strongest available statement of the field-wide gap behind the standing unfalsifiability challenge.

Four Forces Intersection

No direct Force implication — measurement-methodology entry. Noted for later use: ORIC's two-factor structure (change commitment and change efficacy) remains a partial analogue of Commitment and Capability with no counterpart to Purpose or Momentum, and this review does not disturb that.

Relevance to Brandon's Positioning

This supersedes Miake-Lye et al. 2020 as the authoritative statement of the readiness-measurement gap, and it is a harder source to argue with. Miake-Lye said no gold standard exists. Weiner says something narrower and more damaging: measures exist, some have been tested extensively, and the tests do not produce a pattern. "No gold standard" invites the reply that the right instrument has not been built yet. "Fifty-five tests, minimal rating, no discernible pattern" does not.

For Brandon, the value is in the attribution. The single most effective sentence available on this material is that the developer of the field's most-used organizational readiness measure published a systematic review in 2020 concluding that whether readiness matters remains an open question. That sentence does three things at once: it establishes that the gap is field-wide rather than a weakness of the Five Breakpoints framework, it is unrebuttable by appeal to authority because the authority is the one saying it, and it sets the bar that asset 3b would have to clear to be worth doing.

It also raises the cost of over-claiming. If the framework's readiness rubric is ever pitched as predictive, this is the literature it will be measured against, and this literature has a documented record of instruments that did not predict.

Sources & Confidence

  • Weiner BJ, Mettert KD, Dorsey CN, Nolen EA, Stanick C, Powell BJ, Lewis CC. "Measuring readiness for implementation: A systematic review of measures' psychometric and pragmatic properties." Implementation Research and Practice. 2020;1:2633489520933896. doi:10.1177/2633489520933896. PMID 37089124 — High (peer-reviewed; SAGE full text; quoted passages verified against the full text).

Filed by / Routing

Filed by: Stafford (scheduled run, 2026-08-30, 22:20 UTC — story-thread chase, not the exploratory slot) Routing: strategy — the citation of record for the readiness-measurement gap; supersedes Miake-Lye as the lead source.

2026 IBM CEO Study: CEOs Are Reshaping C-suite Roles for the AI Era
IBM Institute for Business Value · published 2026-05-04 · flagged 2026-08-29 · paper-2 · Technology Illusion · Capability

Tags: paper-2, Technology Illusion, Capability Source: IBM Institute for Business Value URL: https://newsroom.ibm.com/2026-05-04-ibm-study-ceos-are-reshaping-c-suite-roles-for-the-ai-era Date published: 2026-05-04 Date flagged: 2026-08-29

Summary

IBM IBV's annual CEO study surveyed 2,000 CEOs and senior leaders across 33 geographies and 21 industries (fielded February–April 2026). Headline structural findings: 76% of organizations now have a Chief AI Officer (up from 26% in 2025), 79% of executives report decentralizing decision-making as AI takes a larger role, and CEOs expect 48% of operational decisions to be made by AI without human intervention by 2030 (vs. 25% today).

The study's core organizational claim: 83% of CEOs say AI success depends more on people's adoption than technology, yet only 25% of the workforce regularly uses AI — while 86% of leaders believe employees have the necessary skills. Organizations that redesigned five core business areas (technology, finance, HR, operations, cross-functional collaboration) are reported as four times more likely to have delivered on business objectives.

Qualifies under Category 7 (leader-defined instrument): it samples decision-makers about their own behavior and beliefs, which is exactly what makes the 86%-believe vs. 25%-use gap evidentiary — it measures the leaders' perception error directly, rather than relying on self-reported success.

Five Breakpoints Intersection

  • Technology Illusion: 83% of CEOs conceding that AI success depends more on adoption than technology — while pouring investment into CAIO roles and tooling — is the C-suite's own admission that technology deployed onto unready organizational conditions doesn't deliver.
  • Momentum Mirage (weak signal, untagged): the 76% CAIO adoption figure alongside only 25% workforce usage suggests structural announcements outrunning actual behavior change; not tagged because the study doesn't connect the two directly.

Four Forces Intersection

  • Capability: the 86% of leaders believing employees have necessary AI skills against 25% regular usage is a direct measurement of capability mis-assessment at the top — the force is implicated through leaders' inability to gauge it.

Relevance to Brandon's Positioning

Re-find note (2026-08-30): already in the v1 base since 2026-05-05 as the founding member of "The 86% Belief Gap" pattern and a member of "The Measurement Mismatch" (knowledge/emerging-patterns.md); this entry is the transferred repo's physical record of the source. The marginal contribution below stands.

Support for paper-2 Claim 5 (readiness as prerequisite): the 4x delivery multiplier attaches to organizational redesign, not to AI capability. Caveat for the case file: "delivered on business objectives" is self-reported by the same leaders — this is a leader-defined instrument, not an outcome-defined one, so it supports Claim 5 at moderate strength despite the High-credibility publisher. The 86-vs-25 gap belongs to "The 86% Belief Gap" pattern (see knowledge/emerging-patterns.md), relevant to Claim 4.

Sources & Confidence

  • IBM Institute for Business Value, 2026 CEO Study (n=2,000) — High (publisher), with the noted self-report limitation on the outcome multiplier.

Filed by / Routing

Stafford v2, 2026-08-29. Audience: strategy.

AI Agents at Work 2026: Securing the Agentic Enterprise (Okta)
Okta (commissioned survey), surfaced via Cloud Security Alliance blog · published 2026-05-30 · flagged 2026-08-29 · paper-2 · practitioner · Strategic Disconnection · Technology Illusion

Tags: paper-2, practitioner, Strategic Disconnection, Technology Illusion Source: Okta (commissioned survey), surfaced via Cloud Security Alliance blog URL: https://www.okta.com/newsroom/articles/ai-agents-at-work-2026-agentic-enterprise-security/ Date published: 2026-05-30 Date flagged: 2026-08-29

Summary

Okta-commissioned survey of 292 executives (CEOs, CIOs, CTOs, VPs) and 492 knowledge workers on AI and AI-agent usage. The design — asking executives and workers the same questions — produces paired perception measurements: 90% of executives express confidence in their organization's visibility into AI tools, while 52% of employees admit using unapproved AI tools; 95% of executives believe employees use AI responsibly within guidelines, while 54% of unapproved-tool users shared internal messages, 45% disclosed sensitive HR information, and 39% uploaded confidential documents; 65% of executives call their AI policies "very clear," while 57% of knowledge workers say policies are unclear, hard to find, or nonexistent.

On agent governance specifically: 88% of organizations report suspected or confirmed AI-agent security incidents, yet only 22% treat AI agents as independent identity-bearing entities and only 34% apply security controls to agents equivalent to those for human employees.

Five Breakpoints Intersection

  • Strategic Disconnection: 65% of executives rating AI policy "very clear" against 57% of workers finding it unclear or nonexistent is a paired measurement of the gap between stated direction and operational reality — the illusion of alignment, quantified within the same organizations.
  • Technology Illusion: 88% reporting agent incidents while only 34% apply human-equivalent controls shows agentic technology deployed on top of governance conditions everyone can see are broken.

Four Forces Intersection

No direct Force implication — breakpoints-only entry.

Relevance to Brandon's Positioning

Re-find note (2026-08-30): the v1 base ingested this study on 2026-05-30 (okta-ai-agents-governance-gap-may2026.md, not transferred) for the 92%/22% agent-identity-coverage gap, supporting Claims 3–4. This entry's marginal contribution is the paired exec/worker perception items (90/52 visibility, 65/57 policy clarity, 95% responsible-use belief vs. admitted leakage), logged as a candidate member of the 86% Belief Gap pattern in knowledge/emerging-patterns.md — flagged for Friday-synthesis review, not counted unilaterally. Weaknesses to respect: vendor-commissioned research with a product to sell (identity governance), modest samples, and self-report on the worker side too. Use it as convergent corroboration, never as the load-bearing citation.

Sources & Confidence

  • Okta, "AI Agents at Work 2026" (n=292 executives, 492 knowledge workers) — Moderate (vendor-commissioned; directional value from the paired-perception design).
  • Cloud Security Alliance blog (2026-07-17, Harish Peri, Okta SVP) — Moderate; surfacing route, author is the vendor's own executive.

Filed by / Routing

Stafford v2, 2026-08-29. Audience: governance / strategy.

AI Supercharges Scientific Output While Quality Slips (Kusumegi et al., Science)
Science (Vol. 390, Issue 6779), via ScienceDaily summary · published 2025-12-18 · flagged 2026-08-29 · paper-2 · Momentum Mirage

Tags: paper-2, Momentum Mirage Source: Science (Vol. 390, Issue 6779), via ScienceDaily summary URL: https://www.sciencedaily.com/releases/2025/12/251224032347.htm Date published: 2025-12-18 Date flagged: 2026-08-29

Summary

Kusumegi, Yang, Ginsparg, de Vaan, and Stuart (Cornell / UC Berkeley) analyzed more than 2 million papers posted January 2018–June 2024 across arXiv, bioRxiv, and SSRN, using a detection model trained on pre-2023 human-written text to identify likely LLM usage, then tracked publication rates and journal acceptance before and after adoption.

Output effects are large: LLM users posted roughly one-third more papers on arXiv, with increases exceeding 50% on bioRxiv and SSRN. The quality signal inverted: in human-written papers, high writing complexity correlated with journal acceptance; for AI-flagged papers, even those scoring high on writing complexity were less likely to be accepted. The surface features that historically signaled quality no longer discriminate once AI produces them cheaply.

Five Breakpoints Intersection

  • Momentum Mirage: the inversion of the complexity–acceptance relationship for AI-flagged papers is direct measurement of appearance decoupling from substance — output volume rose 33–50% while the traditional quality signal stopped predicting real acceptance.

Four Forces Intersection

No direct Force implication — breakpoints-only entry.

Relevance to Brandon's Positioning

Best evidence yet in the base for paper-2 Claim 4 (failure is harder to detect because AI produces the appearance of execution) — and it's outcome-defined in the Category 7 sense: the comparison group is defined by a measured outcome (journal acceptance), not self-report. The scientific-publishing setting is a limitation (knowledge work with an unusually clean external quality gate); the claim needs a corporate analogue where no such gate exists — which is precisely where Claim 4 predicts the danger is worst. Argue it that way: where science has journals to catch the mirage, firms have nothing.

Sources & Confidence

  • Kusumegi et al., Science 390(6779), 2025-12-18, NSF-funded — High.
  • ScienceDaily summary (read source) — Moderate as a rendering; verify exact figures against the journal article before quoting in a draft.

Filed by / Routing

Stafford v2, 2026-08-29. Audience: strategy.

AI Exposure and Organizational Structure
SSRN working paper (Guohou Shan, Feng Zhu) · published 2026-03-22 · flagged 2026-08-29 · paper-2

Tags: paper-2 Source: SSRN working paper (Guohou Shan, Feng Zhu) URL: https://ssrn.com/abstract=6456498 Date published: 2026-03-22 Date flagged: 2026-08-29

Summary

Shan and Zhu study whether AI exposure changes the two fundamental dimensions of organizational design — hierarchy depth and managerial span of control — exploiting staggered AI exposure across 7,681 U.S. firms between 1990 and 2020 in a difference-in-differences design. Finding: firms exposed to AI reduce the number of hierarchical ranks while simultaneously increasing managerial spans of control — systematic organizational flattening.

Read via abstract and secondary coverage only (SSRN full text returned 403 at ingest); methodology details beyond the design and sample are unverified. Flag for a full-text read before citing in a paper draft.

Five Breakpoints Intersection

No direct Five Breakpoints mapping. Logged as core paper-2 mechanism evidence: large-sample causal-design support for hierarchy responding to information-processing cost changes (Claims 1–2), which is the premise the breakpoints argument builds on.

Four Forces Intersection

No direct Force implication — mechanism-evidence entry. (Abstract-only read; no Forces tags per hard rule 2.)

Relevance to Brandon's Positioning

Corrected 2026-08-30 after the v1 import. The day-0 claim that this filled an empty Claim 1 slot is void: the v1 base already holds Babina et al. (NBER 31325), Ewens & Giroud (NBER 34162), Wang/Feng/Sun (arXiv 2511.02099) and Alekseeva et al. (SMJ 2026) on this ground — and its ESCALATED "Résumé-Panel Consensus" pattern requires recording any structural instrument's data source, inference method and window BEFORE its finding. Here only the window is known (1990–2020, entirely pre-generative); the data source is unestablished (403 at ingest). Consequences: (1) this paper is QUARANTINED as evidence until a full-text read establishes its data genre — if résumé/postings-based, it is a fifth member of one method-family, not independent confirmation; (2) it must never be cited alongside Babina et al. or Ewens & Giroud as if independent until that is settled; (3) if it turns out to use administrative or payroll data, it would be the pattern's second break condition met and worth far more than a confirmation. The Claims 1–2 framing (routing dissolves, alignment unmeasured) survives either way, but as an argument, not yet as this paper's evidence.

Sources & Confidence

  • Shan & Zhu, "AI Exposure and Organizational Structure," SSRN 6456498, posted 2026-03-22 — Moderate, quarantined (abstract-only; data genre unknown; see corrected Relevance section — confidence revisited after full-text read).

Filed by / Routing

Stafford v2, 2026-08-29. Audience: strategy.

Five Breakpoints — Source Article (Paper 1 of the Why Change Fails series)

Title: The Five Breakpoints Leaders Miss Status: Published 2026-04-28 — https://www.workthatholds.com/p/the-five-breakpoints-leaders-miss Tags: practitioner, foundation Text note (2026-08-30): the article text below is the pre-publication draft; the published version carries minor wording differences (e.g. "were not designed" for "were never designed") and the shorter title. For direct quotation in anything public, quote from the published URL; for framework definitions and tagging, this file remains the working reference.

This is Brandon's primary intellectual work. Everything Stafford does is grounded in this framework. Read this file before analyzing any content for Five Breakpoints intersections.


Full Article Text

The Five Breakpoints Leaders Miss in Transformation How leaders can design against predictable failure patterns and recognize them early

Organizations rarely fail to transform because they lack activity. More often, they fail because they were never designed to withstand the predictable pressures that large-scale change creates. The roadmap exists. Governance is in place. Funding is approved. Milestones are green. Yet leaders can still sense that the business is not moving the way it should. Teams are active without converging. Pilots are underway without scaling. Progress is being reported, but confidence is not increasing.

One executive described it to me this way: "It feels like youth soccer where everyone is kicking at the ball at the same time, but the ball still doesn't move."

That pattern shows up across cloud modernization, AI adoption, operating model redesign, customer transformation, and broader business change. The surface details vary. The underlying failure patterns do not.

Across large-scale change efforts, the same five breakpoints tend to appear again and again. They are not random setbacks. They are predictable failure patterns that organizations often fail to design against, then fail to recognize early enough once they begin to surface.

Leaders do not need another slogan about change. They need a way to design against these breakpoints from the outset and identify the warning signs before drift becomes expensive.


1. Strategic Disconnection

The first breakpoint appears when a transformation is launched with broad intent but without enough precision to keep the organization aligned under real operating pressure. Leaders believe the strategy is clear, but the organization is operating from multiple versions of the outcome.

A healthcare organization launched a cloud transformation with visible executive support and strong early agreement. But each major stakeholder quietly interpreted the effort through the lens of their own function. Security processes would remain intact. Change windows would remain intact. Review paths would remain intact. No one openly resisted. No one had actually committed to the same destination. A year later, the organization had new cloud platforms and an essentially unchanged operating model.

This is one of the most expensive failure patterns because it does not look like conflict at the start. It looks like consensus. Leaders hear the same language repeated back to them and assume alignment exists. In practice, teams fill in the blanks with their own definitions, priorities, and assumptions about what will or will not change.

The issue is not inspiration. It is precision. A goal like "become more digital" or "use AI to improve customer experience" sounds aligned in a kickoff meeting and fragments in execution. A goal like "reduce customer resolution time by 40 percent within two quarters" gives teams a shared outcome they can translate into decisions, tradeoffs, and accountability.

When strategic disconnection sets in, organizations do not usually stop moving. They simply begin moving in slightly different directions.

Self-check: If you asked ten leaders to describe the primary outcome of the transformation, would their answers actually match?


2. Incentive Fragmentation

The second breakpoint appears when a transformation depends on cross-functional cooperation but was never designed to align the incentives of the people whose decisions determine whether it moves. What looks like support early on often collapses once real tradeoffs appear.

A financial institution had a major migration effort underway with one deadline no one could ignore: datacenter leases were expiring. The VP of Data had aligned teams, built the business case, and set the timeline. The CISO attended the planning meetings and raised no visible objection. Months later, he revealed that he had engaged a separate consulting partner, defined a different set of security requirements, and that nothing would move until his scorecard was satisfied.

The cloud was not the problem. The program stalled because one executive with veto power had little reason to optimize for migration success. His job was to ensure security was never the cause of failure. Migration speed was not his metric.

This is where many leaders misread commitment. They interpret attendance as buy-in. It often means awareness. Sometimes it means observation. In this case, it was reconnaissance.

Incentive fragmentation is especially dangerous because teams often look productive while the system is quietly pulling them apart. One leader is measured on speed. Another on cost containment. Another on risk reduction. Another on quarterly output. Everyone works hard. The enterprise does not move coherently.

The issue is not whether leaders verbally support the change. It is whether the system makes it rational for them to prioritize it when tradeoffs appear. If a transformation succeeds while a powerful stakeholder's metrics stay flat or worsen, leaders should not be surprised when resistance shows up in delay, exceptions, side processes, or alternative standards. Those responses are not random. They are the natural consequence of a system that was never aligned to move together.

Self-check: If the transformation succeeds but a leader's metrics do not improve, will that leader still treat the work as a priority?


3. Process Friction

The third breakpoint appears when leaders set a faster ambition without redesigning the operating model required to support it. The strategy changes. The machinery does not.

A retail organization invested in cloud capabilities that could technically provision a working application environment in hours. But launching an application still required sequential handoffs across operating system, network, storage, identity, database, application, backup, monitoring, and security teams. Each group had its own queue. Each worked on its own timeline. No one owned the full journey from request to result. The cloud could move in hours. The organization still moved in weeks.

This is where many transformation efforts lose credibility with the people doing the work. Leaders announce a new ambition, but the underlying machinery remains unchanged. The handoffs are the same. The approvals are the same. The decision rights are the same. The dependencies are the same. Teams are told to move faster inside a system designed to prevent speed.

What makes this breakpoint so persistent is that organizations often overestimate capability by looking at talent rather than flow. They have skilled people. They have good technology. They may even have committed leaders. What they do not have is a delivery system that allows those ingredients to produce consistent business movement.

When process friction dominates, the strategy is not actually being executed. It is being negotiated one handoff at a time.

Self-check: If you mapped how work actually flows today, would it look like the formal process or the workaround your teams built to survive it?


4. The Technology Illusion

The fourth breakpoint appears when leaders invest in a new capability without designing the surrounding behaviors, workflows, and decision norms required to make it valuable.

An enterprise software company invested in a new sales analytics platform. The data was better. The dashboards were better. The engineering work was solid. Months later, teams were still relying on legacy reports and spreadsheets. Not because the new system was inaccurate or harder to use. Because the old process gave people more room to tune the story, soften the numbers, or avoid difficult conversations. The technology was ready. The organization was not.

This pattern is common in AI, analytics, workflow automation, and broader modernization programs. Leaders invest in the visible artifact and underestimate the operational and behavioral change required to make it valuable. They hope the tool will force the organization forward. More often, the organization absorbs the tool into its existing habits and continues operating much as before.

Technology can accelerate aligned execution. It cannot create alignment on its own. That is why so many well-funded initiatives look compelling in demonstrations and disappoint in practice. The technical capability becomes visible before the operating discipline required to use it.

This is not a technology failure. It is a leadership design failure. The organization was asked to use a new tool inside an old system.

Self-check: If you removed the technology tomorrow, would teams still agree on how the work should happen?


5. Momentum Mirage

The fifth breakpoint appears when a transformation is launched with early visibility but without enough built-in reinforcement to sustain progress once leadership attention shifts.

A semiconductor company launched a digital transformation with strong executive attention and promising early wins. The first quarterly review looked strong. Pilots were working. Teams were engaged. Progress was visible. Then margin pressure pulled the executive sponsor into other priorities. Governance meetings remained on the calendar. Status reports continued. But decision velocity slowed, obstacles stayed unresolved longer, and visible reinforcement faded. No one cancelled the initiative. No one needed to. The organization simply stopped feeding it. By the next planning cycle, the transformation was still alive in presentations and largely dead in practice.

This breakpoint is easy to miss because activity continues after momentum has weakened. Meetings happen. Slides update. The work still has a name. What disappears is the force that turns progress into confidence and confidence into adoption. People begin reading silence as a signal. They return to local priorities. The old system starts reclaiming ground.

Momentum is not enthusiasm. It is the organization's ability to convert intent into visible progress, then reinforce that progress fast enough to sustain belief.

Self-check: What would happen if you stopped actively managing this transformation for 90 days?


The Four Underlying Forces

The five breakpoints are visible signs that the transformation was not designed strongly enough in one or more of four underlying forces:

  • Purpose — Is the outcome clear enough that teams would define success the same way?
  • Commitment — Have priorities, incentives, ownership, and resources been aligned strongly enough to hold when tradeoffs appear?
  • Capability — Does the organization have the operating ability to execute and scale?
  • Momentum — Is there enough visible progress, reinforcement, and feedback to sustain movement after launch?

Breakpoint → Force mapping:

  • Strategic Disconnection → weak Purpose
  • Incentive Fragmentation → weak Commitment
  • Process Friction → weak Capability (flow)
  • Technology Illusion → weak Capability (readiness)
  • Momentum Mirage → weak Momentum

The Diagnostic Questions

Purpose: Are we aligned on the outcome, and would teams describe success the same way? Commitment: Have we aligned incentives, ownership, and priorities strongly enough to survive conflict? Capability: Does our operating model allow the work to move at the speed the strategy now requires? Momentum: Have we built enough visible progress and short enough feedback loops to sustain confidence and action after launch?