Stafford

← 2026-09-01 · all briefs · 2026-09-03 →

Stafford Brief — 2026-09-02

Delta layer. Today's scan ran at 12:19 UTC and its brief stands on its own — the Stanford Playbook, the sponsorship taxonomy, the Readiness Tax escalation, the EU AI Office RFIs, the three early thread advances. None of that is repeated here. This brief is what today's Wednesday falsification sweep did to the ledger, and it is mostly a report of my own errors.


📖 TODAY'S READ

Cohen, Kadach, Ormazabal & Reichelstein, Appendices A–C and Table 3 — not the tables everyone reads. (Journal of Accounting Research 61(3): 805–853, June 2023; full text at madoc.bib.uni-mannheim.de/63834/1/ESGPay.pdf.)

Yesterday this paper defeated the soft-metric objection and became the base's first empirical support for reading an incentive metric's wording as a breakpoint. The ledger closed with one cheap follow-up: read the appendix, check whether safety metrics are worded as programme activity, and the successor challenge weakens. I ran it. It went the other way on every count, and it caught two errors of mine.

The safety metric is "Days Away/Restricted or Transfer (DART) incident rate per 100 full-time employees" (New Jersey Resources Corporation, 2019) — OSHA-defined, mandatorily reported, cross-firm standardised. An injury-rate Trucost. The escape route is closed. Worse, the safety coefficient never tested anything: the −0.05 is measured against ∆CO2, and safety metrics have no reason to cut carbon. It is a mismatched-dependent-variable null exactly like the other seven. The argument I recorded on 09-01 — "if the confound were the whole story, safety should have behaved like carbon" — is invalid and has been struck from the ledger.

And the axis of yesterday's win was misnamed. Appendix B shows metric grammar varying inside Cohen et al.'s categories, not between them. Carbon ("kg CO2e/tonne"), safety (the DART rate), diversity ("percentage of women among the SMP") and employee satisfaction ("internal promotion rate in global leadership") are all crisply result-worded. Compliance is "FY2021 actions and targets (continue to assess human rights, bribery and corruption and other related risks)"; governance is "Establish standalone corporate governance and risk procedures... that build trust, create long-term securityholder value and align with company values." Their decomposition is by subject matter. The −0.07-versus-−0.77 contrast isolates topic specificity — whether the contract names the outcome being measured — not activity-versus-result grammar.

The defeat stands; the sentence Brandon may publish changes. Not "metrics that name a result move it." Rather: an incentive that names the specific outcome moves that outcome; an incentive that gestures at the category does not. Qorvo's "exploration and deployment of AI tools" fails on both axes at once, which is why the Momentum Mirage reading of it survives — but argue it from the grammar axis using this table and the first reviewer who opens Appendix B takes it apart. The take: yesterday's win was real and I filed it under the wrong heading, which is the more dangerous of the two failure modes because nothing about it looks wrong.

🔭 THINKER TO WATCH

Stefan Reichelstein — Universität Mannheim and Stanford GSB, co-author of today's paper. (Affiliations as recorded in the KB entry's verified author line; specific titles and centre roles not checked today and deliberately not asserted.)

Relevant now for a structural reason rather than a topical one. The decisive evidence for Brandon's incentive claim did not come from the organizational-design literature. It came from accounting. The field map holds economists, strategy scholars, consultancies and practitioners, and — as far as I can find — no performance-measurement theorist at all. That is a gap, not an oversight to shrug at: the question "does writing this metric into a contract change what the organization does" is accounting's home question, it has been asked there for fifty years, and the base has been answering it from first principles.

Brandon's differentiation is clean and worth stating precisely. Reichelstein's tradition asks whether a measure is informative — does it load on managerial action, at what cost, with what distortion. Brandon asks something the tradition mostly assumes away: what the organization does when no defensible measure exists yet. That is the whole of the AI case in 2026, and it is why the measurability confound is a genuine objection rather than a quibble. Proposed as a field-map candidate, not a promotion — one find.

🏢 CASE IN THE WILD

Schneider Electric, 2020 compensation contract (Cohen et al., Appendix C) — an organization that solved the threshold problem by outsourcing it, and the AI version of that choice is being made right now.

The annual incentive: 40% organic sales growth, 30% adjusted EBITA margin, 10% cash conversion, and 20% "Schneider Sustainability Impact." The LTI is the interesting half — 40% adjusted EPS, 35% relative TSR across two benchmarks, and 25% split evenly across four third-party ESG assessments, each with an explicit payout scale: DJSIW at "0%: not in World; 50%: included in World; 100%: sector leader"; CDP Climate Change at "0%: C score; 50%: B score (25% at B-); 100%: A score (75% at A-)"; Euronext Vigeo and FTSE4GOOD likewise.

The mechanism, and it is not the one it looks like. This is not a vague metric. It is a hard, threshold-bearing, externally adjudicated target — and the board got there by delegating the definition of the outcome to four rating vendors. Faced with a domain where it could not write its own defensible measure, the organization bought one. Whether that is Momentum Mirage or its cure depends entirely on whether the vendor's scale tracks the substance — and today's Table 9 says, on the one occasion the substance was independently counted, the vendors disagreed with each other about it.

The open question, and it is live for Brandon's clients this quarter. Commercial "AI maturity" scores already exist and boards are already being sold them. When an organization cannot define its own AI outcome and rents a vendor's index instead, has it built accountability architecture or bought Goodhart's law with a payout curve attached? The ESG case ran this experiment for a decade and the base now holds the result; nobody has pointed it at AI.

⚙️ FOUR FORCES CONCLUSION

Purpose, and the correction sharpens it rather than softening it.

The Force is defined as clarity of intent — the organization knowing what it is trying to achieve and why. Yesterday's reading located the evidence in the wording of the objective. Today's relocates it to something harder and more useful: the specific outcome named. Table 8 holds firms, contract vehicle, specification and dependent variable constant, and finds that the contract naming carbon moves carbon while the contract naming "ESG" moves nothing. Clarity of intent, measured, turns out not to mean stating the goal well — it means stating which quantity you will be held to.

That is a materially better claim for paper-2 than the one it replaces, because it is harder to satisfy rhetorically. A board can write a well-worded aspiration in an afternoon. Naming the quantity commits it to an argument about what it will count, who counts it, and — per Table 10 — what it is willing to lose to move it.

💭 OPEN QUESTION

When the base says "nobody has measured this," how often does it actually mean "nobody published the table"?

Today's finding was not that the measurability confound is unanswerable. It is that the variable answering it was built, counted, and never regressed. Cohen et al.'s Table 3 defines eleven metric categories, not nine: alongside the nine subject-matter indicators sit *Self evaluation — "scores defined and measured by the firm", 884 firms — and External evaluation — "scores defined and measured by external parties", 97 firms.* That split is the measurability axis with grammar approximately held fixed. It appears in Table 3 and Appendix B and nowhere else in the paper — not in Table 8, not in any of Table 9's six columns, not anywhere in the body across a full-text pass of pages 1–45. Table 3's note records a sample restriction that may be the whole explanation, and no inference of suppression is available or intended. The fact is only that the estimate does not exist, in a dataset where it could.

This is the second instance in four days. On 08-30 the base established that Helfrich's prospective ORCA study — the definitive test of whether readiness predicts anything — was designed, funded, protocolled and never reported. Two of the most load-bearing "nobody has measured this" claims in this evidence base turn out, on inspection, to be specifications that exist and were not published.

Why this matters more than it looks. The base's standing discipline is aimed at fabricated numbers — four broken chains in three weeks, a rule that a structural figure without a locatable primary should be assumed broken. That discipline finds numbers that were never real. It has nothing pointed at the opposite failure: results that are real, were computed, and were left in a file. Those are invisible to citation-chasing, because there is no citation. The question for Brandon is whether the honest description of several gaps in his case files is "unmeasured" or "unreported" — and they are not the same gap, because the second one can be closed with an email.

🥊 CHALLENGE

The measurability confound — OPEN, unweakened, and now cheaper to settle than a study.

Stated at full strength, unchanged: grant that Cohen et al. show contracts naming a specific outcome move it while contracts naming a category do not. That may not be about the contract at all. Carbon worked because CO2 is counted in tons by Trucost — an independent, standardised, cross-firm measurement that existed before the metric was written. In 2026 AI has no Trucost. If no company yet can write a defensible AI outcome metric, then Qorvo's "exploration and deployment of AI tools" is evidence of an attribution problem, not an organizational failure, and the framework is diagnosing a measurement limit as a breakpoint.

What today did to it. The one cheap route to weakening it is closed — safety metrics are result-defined and OSHA-measured, so measurability and result-definition co-occur there too, and the paper never regresses safety metrics on injury rates despite DART data being as externally counted as Trucost tons. And a bound emerged that the 09-01 entry did not have: carbon is the only domain in the entire paper with a physical-outcome test. Every other category is tested solely against ratings and financial performance. So "result-defined metrics move real outcomes" is a one-domain finding — and the one domain is the one that happened to have an outcome vendor. That is a real constraint on how far this paper can be carried toward AI, and it should be stated by Brandon before it is stated at him.

What it would take for the objection to win (unchanged): evidence that within externally measurable domains, result-defined and activity-defined metrics perform alike — grammar adding nothing once measurability is held fixed.

What changed is the next action. Not a new study. Ask the authors for the Table 8/9 specification with the two score dummies included. Self-measured versus externally-measured scores, 884 firms against 97, grammar approximately fixed, variables already constructed. One email either resolves the objection or sharpens it. Failing that, the free move: write the framework's metric claim on the topic-specificity axis only, which costs nothing and is true today.

Also today, and it cuts toward the framework: Table 9's carbon row lands at Refinitiv +0.001 (t = 0.07) — a precise zero on the disclosure-based score — Sustainalytics −0.583 (t = −2.12) and KLD +6.660 (t = 10.82), with Appendix A defining both of the latter so that higher means better performance. The one contract term shown to move a physical outcome moved the three appearance measures in three different directions, including downward on one. No reading is available in which appearance and substance track each other. Caveats that must travel: KLD rests on 1,351 observations against 17,148 and 19,252; the authors defer to Berg, Koelbel & Rigobon on vendor divergence and call all of Section 6 "mainly descriptive." Do not state this as "ESG ratings are wrong."

📈 PATTERN BUILDING

New pattern, logged at two sources and deliberately NOT escalated: "The Unrun Specification."

The shape: the study that would settle a load-bearing question in this base was designed or its variables constructed, and the analysis was never published — so the gap is invisible to citation-chasing, because there is no citation to chase.

Why this is not yet three sources, and the restraint matters. The base's four broken chains are a different pattern — fabricated or unsourceable figures, which this one is the mirror of. Folding them together would inflate both to five and would be exactly the loose aggregation the rules forbid. Promotion condition, written into the proposal: a third case where a specification that would answer a question in these case files is shown to have been constructed or protocolled and not reported. The two members also differ in one respect that must be recorded — Helfrich's was a whole study, Cohen's is one specification inside a published paper — and if the third member does not resolve that ambiguity the pattern should be split rather than escalated.

What it means for the thesis. Nothing directly. What it changes is method, and that is worth more: before writing "nobody has measured X" into a paper, check whether somebody measured X and did not report it. Two for two so far.

🧱 ASSET SUGGESTION

None new. Asset #4 (the Proxy-Statement Instrument) was materially revised today rather than added to — corrected template, four-arm outcome design, and one cheap author query that should now precede scoping. Full spec in knowledge/asset-suggestions.md; summary in Routing. Opening a fifth entry on a day that corrected the fourth would be ledger inflation.

🎓 SPARRING NOTE

None, deliberately. Two notes were issued this week already — 08-31 and 09-01, both on growth edge #3 — and neither has a recorded response. The cap is two a week and a third would be noise on top of an unanswered prompt. The standing item for the first /coach session is unchanged and now has a third question attached: today's correction is a worked example of the skill in edge #3's second half — before citing a decomposition as evidence for a claim, check that the decomposition's own axis is the axis of the claim. Brandon got the conclusion right on 09-01 and the reason wrong, from a table he had read. That is the exact failure a falsification habit exists to catch, and it caught it in one day.

SOURCES & THREADS

Threads: none due, none chased, and that is the honest report. Nearest next-check is the middle-manager thread on 2026-09-09; the scan advanced three ahead of schedule this morning (verification-cost, middle-manager, EU AI Act) and there is nothing left for this repo to add today without manufacturing it.

Exploratory query — Wednesday falsification sweep (logged to knowledge/query-log.md): are Cohen et al.'s safety-and-security metrics worded as programme activity or as injury-rate targets, and does the taxonomy contain any cell where third-party measurability and result-defined grammar come apart? Aimed at the framework's own metric inference at full strength, and at the single named cheapest open item in the challenge ledger rather than at a case-file gap. Not a repeat of the scan's same-day falsification sweep, which targeted the redesign-prerequisite claim and returned null. Yield: 1 KB file updated with three corrections, 1 challenge sharpened and kept OPEN, 1 DEFEATED challenge's axis corrected, 1 new pattern proposed at 2 sources, asset #4 revised, 1 field-map candidate. No new source was needed — the entire sweep ran against a PDF already in hand, which is the cheapest yield this repo has produced.

Verification chase (not the exploratory slot): the McKinsey operating-model article is not down. We are blocked, and the chase should stop. The scan recorded HTTP 503 on every route this morning and logged it as an active chase. Today mckinsey.com returns Akamai 403 Access Denied over HTTP/2, HTTP/1.1 with a browser UA, and the r.jina.ai text proxy — and a deliberately fabricated path under the same directory returns the identical 403. The block is indiscriminate: a real article and a nonexistent one are indistinguishable from here, so no agent-run retry can ever confirm this piece exists. web.archive.org is blocked by this environment's egress policy and archive.org's availability API returned 429. Disposition proposed: close the standing chase, reassign to Brandon's own browser, and hold the 63%/21% and 79%/51% figures as unverified — a potential fifth broken chain, not a confirmed one. Left as an open agent chase it will burn a slot every day and never resolve.

Candidates: one proposed — Stefan Reichelstein (field map, 1 signal). No promotions, no expiries from this repo.


Routing


PROPOSED-UPDATE — knowledge/emerging-patterns.md (new pattern, 2 sources, NOT escalated)

The Unrun Specificationrecurrence 2, logged 2026-09-02, deliberately not escalated.

The pattern: the analysis that would settle a load-bearing question in these case files was designed, or its variables constructed and counted, and then never published — so the gap is invisible to citation-chasing, because there is no citation to chase.

Members. (1) Helfrich et al. 2011 (knowledge-base/implementation-science-helfrich-orca-predictive-validity-2011-08.md, Stafford 2026-08-30): a prospective ORCA study across 53 VA facilities, outcome as Cohen's h, hierarchical linear model, 90% power at R² ≥ 0.21, promised in the 2009 development paper as "criterion validation using implementation and quality-of-care outcomes is the next phase of our work." No results paper in Europe PMC, in the protocol's 29 citing articles, or across ~68 Semantic Scholar citations, and no Helfrich-authored citing paper at all. No file-drawer inference is available — nothing shows the analysis was completed. (2) Cohen, Kadach, Ormazabal & Reichelstein, JAR 61(3) 2023 (knowledge-base/cohen-kadach-ormazabal-reichelstein-esg-pay-international-2023.md, Stafford 2026-09-02): Table 3 Panel A constructs and counts Self evaluation (884 firms) and External evaluation (97 firms) — the measurability contrast the measurability-confound challenge requires — and neither category appears in Table 8, in any of Table 9's six columns, or anywhere in the body (full-text pass, pp. 1–45). Table 3's note (2) records a sample restriction that may be the entire explanation; no inference of suppression is intended.

Explicitly distinct from the broken-chain rule. The base's four broken chains are fabricated or unsourceable figures — numbers that were never real. This is the mirror: results that are real, were computed or designed, and were not reported. Merging the two would produce a spurious five and would be exactly the loose aggregation the tagging bar forbids.

Promotion condition: a third case in which a specification answering a question in these case files is shown to have been constructed or protocolled and not reported. Unresolved ambiguity to settle at promotion: member (1) is a whole study, member (2) is one specification inside a published paper. If the third member does not resolve which construct this is, split the pattern rather than escalate it.

Method implication (the reason to log it at 2): before writing "nobody has measured X" into a paper or a case file, check whether somebody measured X and did not report it. Two for two so far, and the second one is answerable by email.

PROPOSED-UPDATE — knowledge/paper-2-case-file.md, Claim 5 / incentive evidence

Correction to the 2026-09-01 evidence note on Cohen et al. (axis, not sign). The entry records Table 8's −0.07-versus-−0.77 contrast as isolating activity-defined versus result-defined metric grammar. Appendix B (read 2026-09-02) shows grammar varying inside the categories rather than between them — carbon "kg CO2e/tonne", safety "DART incident rate per 100 full-time employees", diversity "percentage of women among the SMP" are result-worded; compliance "continue to assess human rights, bribery and corruption and other related risks" and governance "establish standalone corporate governance and risk procedures... that build trust" are activity-worded. The decomposition is by subject matter; the contrast isolates topic specificity — whether the contract names the outcome being measured. Result, sign and evidentiary weight unchanged; the claim's wording must change. Also correct the count: Table 8 column (2) estimates nine categories, not ten, and Table 3 defines eleven. And add the bound: carbon is the only domain in the paper with a physical-outcome test — every other category is tested against ratings and financial performance only, so this is a one-domain result.

PROPOSED-UPDATE — knowledge/field-map.md (candidate, 1 signal)

Stefan Reichelstein — Universität Mannheim and Stanford GSB (affiliations per the paper's author line; titles unverified and not asserted). First seen 2026-09-02 via Cohen, Kadach, Ormazabal & Reichelstein, JAR 61(3) 2023. Logged for the discipline, not the individual: the decisive evidence for the framework's incentive claim came from managerial accounting, and the field map currently holds no performance-measurement theorist. Differentiation to hold: that tradition asks whether a measure is informative about managerial action; Brandon asks what an organization does when no defensible measure exists yet — which is the whole of the AI case and the substance of the measurability-confound challenge. Promotion bar: a second on-thesis find from performance-measurement or managerial-accounting scholarship.

PROPOSED-UPDATE — knowledge/source-candidates.md, McKinsey operating-model chase

Close the chase opened 2026-09-02; reassign off-pipeline. Re-run 2026-09-02 from a different session: mckinsey.com now returns Akamai 403 Access Denied (not the morning's 503) over HTTP/2, HTTP/1.1 with a browser UA, and the r.jina.ai text proxy — and a deliberately fabricated path under the same directory returns the identical 403, so the edge block is indiscriminate and cannot distinguish a real article from a nonexistent one. web.archive.org is blocked by egress policy; archive.org's availability API returned 429. No agent-run retry can resolve this. Hold the reported ~2,000 executives / 63%-vs-21% / 79%-vs-51% figures as unverified — a potential fifth broken chain, not a confirmed find — and route the check to Brandon's own browser. Leaving it as a standing chase spends a slot per scan for a guaranteed null.

STAGE-CANDIDATE — supersedes the 2026-09-01 payload, which must NOT be run

⚠️ The 09-01 payload carries two defects fixed here: the strategic_disconnection rationale and the first key finding say "ten metric categories" (Table 8 column (2) estimates nine; Table 3 defines eleven), and the momentum_mirage rationale asserts the grammar axis that Appendix B shows is untested. If 09-01's has already been staged, re-run this one and verify the breakpoint rationales are replaced rather than appended — the duplicate path patches missing fields, and an appended rationale would leave the wrong axis in the index. published_date stays null: the issue is June 2023, but no exact publication day is verifiable from a primary.

python3 build/add-research-json.py --json '{
  "title": "Executive Compensation Tied to ESG Performance: International Evidence",
  "publisher": "Journal of Accounting Research",
  "url": "https://onlinelibrary.wiley.com/doi/10.1111/1475-679X.12481",
  "published_date": null,
  "credibility": "high",
  "breakpoints": [
    {"name": "momentum_mirage", "rationale": "Holding firms, contract vehicle, specification and dependent variable constant and varying only whether the contract names the specific outcome being measured, an indicator for incorporating any ESG metric is a null on Scope 1 emissions (-0.07, t=-0.85) while the carbon-specific metric cuts them (-0.77, t=-2.88, p<0.01); and in Table 9 that same carbon metric moves the three commercial rating vendors in three different directions (+0.001 Refinitiv, -0.583 Sustainalytics at 5%, +6.660 KLD at 1%, both latter scored higher-is-better), so the appearance measures and the one independently counted substance measure do not track each other."},
    {"name": "strategic_disconnection", "rationale": "Eight of the nine metric categories estimated in Table 8 name a subject rather than the outcome being measured, and none is associated with movement in the only physical outcome the paper measures — vague intent producing the appearance of alignment in the most binding statement of intent a public company files."}
  ],
  "forces": [
    {"name": "Purpose", "rationale": "Topic specificity is the only variable distinguishing Table 8 Column (1) from the single significant coefficient in Column (2) — same firms, same vehicle, same fixed effects, differing in whether the stated objective names the quantity the firm will be held to."},
    {"name": "Commitment", "rationale": "The carbon metric that moved emissions is also the metric carrying the sample largest negative stock return (-0.079, t=-2.66) and a negative ROA coefficient (-0.015, t=-1.89): a result-defined commitment cost the firms that made it, while the generic commitment cost nothing and delivered nothing measurable."}
  ],
  "key_findings": [
    "Table 8 (N=21,715 firm-years, firm and year FE, SEs clustered at firm): the ESG Pay indicator for any ESG metric on change in Scope 1 CO2 is -0.07 (t=-0.85), not significant; the carbon-specific metric is -0.77 (t=-2.88), significant at 1% and the only significant coefficient among the nine estimated metric categories",
    "Table 3 Panel A defines eleven metric categories, of which only nine are ever estimated: Self evaluation (scores defined and measured by the firm, 884 firms) and External evaluation (scores defined and measured by external parties, 97 firms) appear in the taxonomy and in Appendix B but in no regression in the paper",
    "Table 9: the carbon metric that cut physical emissions is +0.001 (t=0.07) on Refinitiv, -0.583 (t=-2.12, 5%) on Sustainalytics and +6.660 (t=10.82, 1%) on KLD, with Appendix A defining both Sustainalytics and KLD so that a higher score indicates better ESG performance — the three appearance measures disagree about the one substance measure",
    "Appendix B shows metric wording varying within categories rather than between them: carbon is kg CO2e/tonne, safety is the DART incident rate per 100 full-time employees, diversity is percentage of women among senior management, while compliance is continue to assess human rights, bribery and corruption and governance is establish standalone corporate governance and risk procedures that build trust",
    "Table 10: the carbon-specific metric carries the largest negative financial coefficients in the paper — annual stock return -0.079 (t=-2.66, 1%) and change in ROA -0.015 (t=-1.89) — so the metric that moved the outcome is the one that cost its executives",
    "Adoption reached 1,198 firms, 31% of sample firms, by 2020; authors caveat explicitly that the evidence is mainly descriptive and that the results in Tables 8-10 are not statistically strong"
  ],
  "sample_size": 22603,
  "instrument": {
    "data_source": "ISS Executive Compensation Analytics (contract terms, hand-coded metric taxonomy), Trucost (Scope 1 CO2 emissions), Refinitiv / Sustainalytics / MSCI-KLD (ESG ratings), Datastream-WorldScope (accounting and market), FactSet/LionShares (institutional ownership). Selection: 53,565 ECA firm-years to 35,076 (6,262 firms) with Datastream and FactSet coverage, to 22,603 firm-years and 4,395 firms in 21 countries where Trucost is non-missing.",
    "inference_method": "OLS with firm and year fixed effects, controls measured at start of year, standard errors clustered at firm. Association only — the authors state the evidence is mainly descriptive and make no causal claim. IMPORTANT: the metric decomposition is by SUBJECT MATTER, not by metric grammar; the Column (1) vs Column (2) contrast isolates topic specificity. Carbon is the only domain in the paper tested against a physical outcome; all other categories are tested against ratings and financial performance only, so this is a one-domain result.",
    "observation_window": "2011-2020 — entirely pre-generative. Transfers to AI incentive metrics as an analogue, never as evidence about them."
  }
}'