Stafford Brief — 2026-09-01 (Tuesday)
The scan had an unusually strong day — Gallup's twelve-year span-of-control series, EMA's abandoned-pilot residue, one escalation and one principled refusal to escalate. Read research-evidence/briefs/2026-09-01.md; I am not repeating any of it. This brief does one thing. Yesterday I logged a counterargument that I called the base's cheapest open item and said would resolve in one direction or the other. It resolved today, in one day, and it went the other way from what the abstract advertised: the strongest standing objection to Momentum Mirage is dead, killed by the paper that raised it.
📖 TODAY'S READ
Executive Compensation Tied to ESG Performance: International Evidence — Cohen, Kadach, Ormazabal & Reichelstein, Journal of Accounting Research 61(3): 805–853, June 2023. Full text obtained from the Universität Mannheim repository (madoc.bib.uni-mannheim.de/63834/1/ESGPay.pdf) after Wiley, SSRN, Taylor & Francis and the ECGI PDF all failed yesterday.
Yesterday's brief said the entire soft-metric challenge turned on one unread paragraph. It turned out to be one unread table, and it says the opposite of the abstract.
The abstract reports that ESG pay adoption is "accompanied by improvements in ESG performance." That single phrase aggregates three results against three different dependent variables. Table 8 — 21,715 firm-years, 4,395 firms, 21 countries, 2011–2020, firm and year fixed effects, standard errors clustered at firm — regresses ∆CO2, Scope 1 emissions in tons, measured by Trucost. Not a rating. A physical quantity counted by a third party.
Column (1), the indicator for incorporating any ESG metric: −0.07, t = −0.85. A null. Column (2), decomposed by metric type: the carbon-specific metric, −0.77, t = −2.88, significant at 1% — and it is the only significant coefficient in the column.
Same firms. Same contract vehicle. Same specification. Same dependent variable. The only thing that varies is whether the incentive names a quantity somebody counts. The soft metric did nothing to emissions; the metric that named a result cut them.
Then Table 9 supplies the other half: against ESG ratings, the generic metric is +0.233 (10%) on Sustainalytics and +1.004 (5%) on KLD. The generic metric moved the scores and did not move the tons.
That is Momentum Mirage, measured — the appearance metric moves while the substance metric does not — and it is the first empirical support this base has ever had for an inference it has been making on grammar alone.
Four disciplines, and they are not optional. (1) The load-bearing comparison is the narrow one: ESG Pay versus Carbon emissions, both against ∆CO2. Nine of the ten categories are being tested against an outcome most were never meant to move — a diversity metric failing to cut emissions is a mismatched dependent variable, not a failed incentive, and anyone stating the wide version will be correctly taken apart. (2) The authors call their evidence "mainly descriptive, which cautions against making strong causal claims" and say plainly that "the results in Tables 8-10 are not statistically strong." (3) The window is 2011–2020, entirely pre-generative — an analogue for AI metrics, not evidence about them. (4) The authors do not endorse this reading. They read Tables 8–9 as a "strengthened pledge" story and reject window-dressing. The decomposition is theirs; the contrast drawn from it is ours. Say so, or a referee will.
One correction to yesterday's file, since Brandon may already have the citation: it is 61(3), June 2023, not 61(4).
🥊 CHALLENGE
The soft-metric objection — logged OPEN yesterday, marked DEFEATED today. Full resolution block in knowledge/thesis-challenges.md; yesterday's entry is preserved verbatim beneath it, because a ledger that edits its own losing entries is worthless.
Yesterday I wrote that the honest position was that the base had tagged a breakpoint from metric grammar without ever testing whether metric grammar predicts anything. That is no longer the honest position. It has been tested, on 4,395 firms in 21 countries against a physical outcome, and grammar predicted it.
But the win produces a better objection, and I am logging it OPEN in the same commit — this is the part Brandon should actually spend time on.
The measurability confound. Carbon worked because CO2 is countable in tons by Trucost — an independent, standardized, cross-firm measurement that existed before anyone wrote the metric. Look at what failed: culture, employee satisfaction, compliance, governance, "other." Every null category is a domain with no Trucost. So the distinction the framework calls activity-defined versus result-defined may be substantially confounded with whether the domain admits third-party measurement at all.
And in 2026, AI has no Trucost. There is no independent, standardized, cross-firm measure of an AI transformation's result. If no company can yet write a defensible AI outcome metric, then Qorvo's "exploration and deployment of AI tools" is evidence of an attribution problem, not an organizational failure — and the framework is diagnosing a measurement limit as a breakpoint. That is a category error, and it is precisely the error this base has spent a fortnight accusing other people of.
The cheapest test, and it is inside a PDF already on disk. Safety and security is the counterexample to press. Injury and incident rates are externally reported in many jurisdictions — so if the confound were the whole story, safety should have behaved like carbon. It did not: −0.05, t = −0.34, a null. If Cohen et al.'s Appendix B shows those safety metrics worded as programme activity rather than as injury-rate targets, then grammar is operating independently of measurability inside their own data, and this challenge weakens sharply on evidence already downloaded. Queued for tomorrow's exploratory slot. Today's is spent, and I am not spending it twice.
⚙️ FOUR FORCES CONCLUSION
Commitment, and for once it is priced rather than asserted.
The framework defines Commitment as decision-makers acting consistently with stated direction under pressure. Table 10 is the pressure. The carbon metric — the one that actually moved emissions — carries the largest negative stock return in the sample (−0.079, t = −2.66, 1%) and a negative ROA coefficient (−0.015, t = −1.89). The generic metric costs nothing and delivers nothing measurable.
Read those two tables together and you get the most useful sentence available from this paper: the metric that worked is the metric that hurt.
This is where I part company with the obvious framing, and Brandon should too. A board writing "win the AI opportunity" instead of a result-defined metric is not necessarily drafting badly. Activity metrics are popular because they are free — no measurable cost to the executives who accept them, no measurable movement in the outcome. That converts the diagnosis from an incompetence story into a revealed-preference story: the organization chose the metric that could not fail because the alternative had a price attached and somebody read the price. That is a harder claim, a more defensible one, and considerably more interesting than "compensation committees write vague goals."
💭 OPEN QUESTION
If a result-defined metric is expensive to accept, who in the organization is supposed to bear that cost — and what makes them?
Brandon's Paper 3 mechanism 3 asks that a meaningful share of executive compensation be tied to specific multi-year transformation outcomes. Mechanism 3 has never had evidence attached to it. It has some now, and the evidence contains a sting the mechanism does not address: the executives who accepted the specific metric took a measurable hit. The prescription asks people to volunteer for a worse expected payout.
So mechanism 3 is incomplete as written. It names what a good incentive looks like and is silent on what makes anyone accept one. Cohen et al. point at the answer without developing it — their Tables 5–7 find ESG Pay adoption associated with institutional investor engagement, voting support and increased fund holdings. The metric that costs the executive something got written because an outside party with standing made it costly not to.
The question for Brandon, and it is a live gap in the published series: in an AI transformation, who plays the institutional investor? There is no equivalent constituency demanding result-defined AI metrics — no ISS recommendation, no Big Three engagement campaign, no disclosure regime. If mechanism 3 requires external pressure to be adoptable, and no external pressure exists for AI, then mechanism 3 is currently unimplementable and the paper does not say so. That is worth a paragraph in the next revision, and it is the kind of gap that is much better found by his own analyst than by a reader.
🧱 ASSET SUGGESTION
Asset #4 — The Proxy-Statement Instrument — UNBLOCKED. It was blocked on exactly one step: read Cohen et al.'s outcome variable. Done. Full update in knowledge/asset-suggestions.md; three changes.
The precedent transfers, and better than hoped. The outcome is physical, not rating-based, so the headline result ports. More importantly, the activity-versus-result coding scheme this asset proposed is no longer a construct a referee can call arbitrary — it is the published decomposition of a top-three accounting journal, and it produced that paper's only significant outcome coefficient. Cite Table 8 Column (2) as the template.
The design should be three-way, not two-way. Cohen et al. run three dependent variables and the finding is that they disagree. Port that: code the AI metric at T, measure at T+2 against (a) a substantive AI outcome, (b) commercial "AI maturity" scoring, (c) financial performance. If the activity metrics move the scores and not the substance, that is Momentum Mirage measured on AI in public documents. A two-way design misses the very thing that makes the ESG paper citable.
Measure the cost, not just the outcome. Table 10 is half the story and it is the non-obvious half.
The design risk, stated now rather than discovered at analysis: AI has no Trucost. The T+2 outcome definition is this asset's central problem, not a detail — fix it before any coding starts, per the 2026-08-30 discipline, and be willing to conclude the honest study is smaller than the ambitious one. Asset #4 and the new measurability challenge now share a critical path: whoever defines the AI outcome variable resolves both. Of the four assets in the ledger, this now has the fewest open dependencies.
🎓 SPARRING NOTE
Growth edge #3 — from intervention design to intervention evidence. Yesterday I set an exercise: write down what would show mechanism 3 working and what would show it failing, before looking at what exists.
Do that exercise before reading further into this brief's evidence, because today makes it easy to skip and skipping it wastes the only clean instance you will get. Mechanism 3 is the first of Paper 3's six mechanisms to meet real evidence, and the evidence did three things at once: it confirmed the mechanism's operative word (specific), attached a price to it that the mechanism does not mention, and exposed a precondition — external pressure — that the paper never states.
The sharper version of yesterday's question: your falsification criterion for mechanism 3 almost certainly said something like "executives paid on transformation outcomes produce more transformation outcomes." Cohen et al. satisfies that and it is still not enough, because it also shows the mechanism is costly and pressure-dependent. A mechanism that works only when an outside party forces its adoption is a different claim than a mechanism that works — and Paper 3 currently makes the second claim while the evidence supports the first.
So the exercise grows one line, and it is the line that matters: for each of the six mechanisms, write the falsification criterion and the adoption condition — what has to be true for anyone to accept it. Six mechanisms, six criteria, six adoption conditions. That is a publishable appendix, the most direct available answer to the unfalsifiability challenge, and — on today's evidence — it is where the series is actually thin. Your mechanisms are better specified than your account of who would ever adopt them.
SOURCES & THREADS
Threads: none due, and none chased. Nearest next-check is the middle-manager thread on 2026-09-09. The scan advanced three threads ahead of schedule today (agent-governance criterion (c) now met, Regime Fork still unmet, middle-manager sharpened) and I have nothing to add to any of them that its brief did not already log. Nothing opened, nothing closed. No new thread proposed — today's find resolves a question rather than opening a story, and the "any disclosed AI metric that names a result" watch already lives on the Where does the saved time go? thread.
Exploratory query (logged to knowledge/query-log.md): what is the outcome variable in Cohen et al.? — routed around yesterday's four HTTP 403s by hunting institutional-repository copies rather than publisher pages, then extracted locally with PyMuPDF when the fetch summarizer choked on the binary. Yield: one challenge defeated, one successor challenge opened, one KB file rewritten and re-tagged, one citation corrected, one asset unblocked.
Target selection, stated because it departed from the standing rule. The slot is specified against the top evidence gap in the paper case files. I spent it on a challenge-ledger item instead. The reason: an unresolved counterargument aimed at a mechanism Brandon has already published outranks an unevidenced claim in a paper still in draft. Paper 4's Claim 5 can be softened before publication; Paper 3 mechanism 3 is in print and was, as of yesterday, refutable by a JAR paper. Claim 5 also got a null from the scan's own slot today, so a second pass at it would have been the more redundant choice as well.
Source candidates: none proposed, and that is deliberate. Yesterday logged Journal of Accounting Research at 1 signal. Today is the same paper read properly, not a second find. Reading one source twice is not two signals, and quietly incrementing the count would corrupt the promotion bar the ledger exists to enforce. JAR stays at 1.
Judgment files updated: thesis-challenges.md (soft-metric objection → DEFEATED with the resolving evidence; measurability confound → OPEN), asset-suggestions.md (#4 unblocked, design revised three ways), query-log.md, knowledge-base/cohen-kadach-ormazabal-reichelstein-esg-pay-international-2023.md (rewritten on full text).
Routing
- strategy — the one thing to act on. Yesterday's routing said the compensation-committee argument was not safe to publish until Cohen et al.'s outcome variable was read, because anyone writing it could be refuted by a 2023 JAR paper. That block is lifted, and it lifts into a stronger position than the one it blocked. The argument is no longer "the market inverted my prescription and here is why that is bad" — it is "the market inverted my prescription, and the one time this experiment was run, the inverted version produced nothing while the version I specified worked." Write it with the four disciplines above attached, especially the pre-generative window and the narrow-comparison rule. The three sentences that carry it: −0.07 and not significant for any ESG metric; −0.77 at the 1% level for the metric that named the result; and the metric that worked is the one that cost its executives money.
- strategy — a gap in Paper 3 that today opened and I would not have found without it. Mechanism 3 specifies what a good incentive looks like and is silent on what makes anyone accept a costly one. Cohen et al.'s Tables 5–7 point at institutional-investor pressure as the adoption mechanism; AI has no equivalent constituency. Mechanism 3 may be currently unimplementable for AI, and the paper does not say so. Worth a paragraph in the next revision — see the Open Question.
- governance — Asset #4 is unblocked and is now the most tractable item in the ledger. The gating decision is Brandon's and it is narrow: define the T+2 AI outcome variable. That single definition also resolves the new measurability challenge. Everything else about the design is settled and has a published template.
- personal — an ops note that is also a lesson about this system. Yesterday four publisher routes returned 403 and the finding was filed as unverifiable. Today the same paper was fully readable from a university repository in one search, and the abstract turned out to have been misleading about its own result. Two things follow. The pipeline should treat "publisher 403" as "try institutional repositories, arXiv, RePEc, and the authors' own pages" rather than as a stopping condition — this is the second time in three days that a wall turned out to be one hop deep. And the harder lesson: an abstract that aggregates across dependent variables can invert its own paper's finding. This base staged an entry yesterday on the strength of an abstract and got the sign wrong. The rule that should follow: no source is load-bearing in a draft on an abstract alone. That is hard rule 10's re-verification requirement, and it just paid for itself in one day.
STAGE-CANDIDATE — supersedes the 2026-08-31 payload, which must NOT be run
⚠️ Yesterday's brief staged this entry with published_date: null, forces: [], a COUNTERARGUMENT rationale and three UNVERIFIED instrument fields, and said explicitly: "If the full text is obtained first, re-run with the real instrument fields rather than patching a duplicate." The full text was obtained. Run this payload; discard yesterday's. If yesterday's has already been staged, re-run this one — the duplicate path patches missing fields, but check that the breakpoint rationale is replaced and not merely appended, because the two rationales assert opposite conclusions.
published_date stays null: the issue is June 2023 but no exact publication day could be verified from a primary. Forces are now tagged, because the paper was read.
python3 build/add-research-json.py --json '{
"title": "Executive Compensation Tied to ESG Performance: International Evidence",
"publisher": "Journal of Accounting Research",
"url": "https://onlinelibrary.wiley.com/doi/10.1111/1475-679X.12481",
"published_date": null,
"credibility": "high",
"breakpoints": [
{"name": "momentum_mirage", "rationale": "Holding firms, contract vehicle, specification and dependent variable constant and varying only whether the incentive names a measurable result, an indicator for incorporating any ESG metric is a null on Scope 1 emissions (-0.07, t=-0.85) while the carbon-specific metric cuts them (-0.77, t=-2.88, p<0.01) — and the generic metric is meanwhile associated with rising commercial ESG ratings (+0.233 Sustainalytics, +1.004 KLD), i.e. the appearance measure moves while the substance measure does not."},
{"name": "strategic_disconnection", "rationale": "Eight of the ten metric categories firms write into executive pay are aggregate or activity-shaped (corporate culture, employee satisfaction and development, other) and none is associated with movement in the one physical outcome the paper measures — vague intent producing the appearance of alignment in the most binding statement of intent a public company files."}
],
"forces": [
{"name": "Commitment", "rationale": "The carbon metric that moved emissions is also the metric carrying the sample largest negative stock return (-0.079, t=-2.66) and a negative ROA coefficient (-0.015, t=-1.89): a result-defined commitment cost the firms that made it, while the generic commitment cost nothing and delivered nothing measurable."},
{"name": "Purpose", "rationale": "Clarity of intent is the only variable distinguishing Table 8 Column (1) from the single significant coefficient in Column (2) — same firms, same vehicle, same fixed effects, differing in whether the stated objective names a quantity a third party counts."}
],
"key_findings": [
"Table 8 (N=21,715 firm-years, firm and year FE, SEs clustered at firm): ESG Pay indicator for any ESG metric on change in Scope 1 CO2 is -0.07 (t=-0.85), not significant; the carbon-specific metric is -0.77 (t=-2.88), significant at 1% and the only significant coefficient among ten metric categories",
"Table 9: against change in ESG rating, the generic ESG Pay indicator is +0.233 (10%) on Sustainalytics and +1.004 (5%) on KLD but a precise -0.001 (t=-0.17) on Refinitiv — vendor-divergent, and the authors cite Berg et al. 2022 on rating disagreement",
"Table 10: the carbon-specific metric carries the largest negative financial coefficients in the paper — annual stock return -0.079 (t=-2.66, 1%) and change in ROA -0.015 (t=-1.89) — so the metric that moved the outcome is the one that cost its executives",
"Adoption reached 1,198 firms, 31% of sample firms, by 2020; ESG Pay is most common in oil and petroleum, utilities and automakers, and less frequent in the U.S. than in Europe, Australia and Canada",
"Authors caveat explicitly that the evidence is mainly descriptive, cautioning against strong causal claims, and that the results in Tables 8-10 are not statistically strong"
],
"sample_size": 22603,
"instrument": {
"data_source": "ISS Executive Compensation Analytics (contract terms), Trucost (Scope 1 CO2 emissions), Refinitiv / Sustainalytics / MSCI-KLD (ESG ratings), Datastream-WorldScope (accounting and market), FactSet/LionShares (institutional ownership). Selection: 53,565 ECA firm-years to 35,076 (6,262 firms) with Datastream and FactSet coverage, to 22,603 firm-years and 4,395 firms in 21 countries where Trucost is non-missing.",
"inference_method": "OLS with firm and year fixed effects, controls measured at start of year, standard errors clustered at firm; the metric taxonomy is hand-collected from disclosed compensation contracts. Association only — the authors state the evidence is mainly descriptive and make no causal claim.",
"observation_window": "2011-2020 — entirely pre-generative. Transfers to AI incentive metrics as an analogue, never as evidence about them."
}
}'
PROPOSED-UPDATE — knowledge/emerging-patterns.md, "Paid for Deployment" pattern
⚠️ This supersedes the 2026-08-31 PROPOSED-UPDATE for the same pattern. I checked today's scan commit (f63cd74): yesterday's block was not applied, so nothing needs unwinding — but do not apply it now. It states that the pattern's Momentum Mirage mapping carries a live counterargument, and that counterargument was defeated today. Apply this block instead.
Append to the pattern entry. Recurrence count unchanged at 2 — this adds no member. It converts a threatened mapping into a supported one.
Update 2026-09-01 (Stafford, delta layer) — the Momentum Mirage mapping has been tested and it holds. This replaces the counterargument flagged on 2026-08-31, which was raised and defeated within one day on a full-text read of the same paper. This pattern reads activity-shaped AI metrics as Momentum Mirage on the grounds that a goal with no stated result condition pays identically for a transformation and for a mirage. Cohen, Kadach, Ormazabal & Reichelstein, Journal of Accounting Research 61(3):805–853, June 2023 — the previous non-financial metric to enter executive pay — test that inference and confirm it. Their outcome is ∆CO2, Scope 1 emissions in tons from Trucost, a physical third-party measure, not a rating. Table 8 (N = 21,715 firm-years, 4,395 firms, 21 countries, 2011–2020, firm and year FE, SEs clustered at firm): the indicator for incorporating any ESG metric is −0.07 (t = −0.85), a null; the carbon-specific metric — the one that names a measurable result — is −0.77 (t = −2.88, p < 0.01), the only significant coefficient among ten metric categories. Table 9 completes the shape: the generic metric is associated with rising commercial ESG ratings (+0.233 Sustainalytics at 10%, +1.004 KLD at 5%; a precise −0.001 on Refinitiv). The appearance measure moves; the substance measure does not. Table 10 adds the finding this pattern should carry as its most useful line: the carbon metric that worked also carries the sample's largest negative stock return (−0.079, t = −2.66) and a negative ROA coefficient (−0.015, t = −1.89) — result-defined metrics are expensive to accept, which is a candidate explanation for why boards write activity metrics that is considerably stronger than "boards draft badly." Five disciplines, all mandatory in any use. (1) The load-bearing comparison is only ESG Pay versus Carbon emissions, both against ∆CO2 — nine categories are tested against an outcome they were never meant to move, and the wide version of this claim is indefensible. (2) The authors call the evidence "mainly descriptive, which cautions against making strong causal claims" and "not statistically strong." (3) Window 2011–2020, entirely pre-generative — an analogue for AI metrics, never evidence about them. (4) The authors do not endorse this reading; they read Tables 8–9 as a strengthened-pledge story and reject window-dressing. The decomposition is theirs, the contrast is ours. (5) Do not run a window-dressing story on Refinitiv — it is the self-reported-disclosure-based score and it is precisely the one that did not move. Successor objection now tracked as the measurability confound in
stafford-research/knowledge/thesis-challenges.md: carbon may have worked because CO2 has Trucost, and every null category is a domain with no equivalent — AI included. File:stafford-research/knowledge-base/cohen-kadach-ormazabal-reichelstein-esg-pay-international-2023.md
PROPOSED-UPDATE — knowledge/source-candidates.md
Amendment to yesterday's block, not a new one. Yesterday proposed Journal of Accounting Research and Journal of Corporate Law Studies at 1 signal each; that block still stands and should be applied as written. Do not increment JAR to 2 on the strength of today's work — today is the same paper read in full, not a second find. One source read twice is one signal. Add to the JAR entry only:
Note 2026-09-01: full text later obtained via the Universität Mannheim institutional repository after four publisher 403s; the paper's abstract was found to aggregate across three dependent variables in a way that inverts the reading of its own Table 8. Signal count unchanged at 1.
PROPOSED-UPDATE — knowledge/field-map.md
Yesterday's Dell'Erba & Gomtsyan block stands and should be applied as written. One addition, because today changes what to watch them for:
Addendum 2026-09-01: their normative critique — that aggregate ESG measures "fail to highlight specific areas that require immediate improvements" — now has an empirical result behind it that they do not cite: Cohen et al. (JAR 61(3), 2023) Table 8, where the aggregate metric is a null on emissions and the specific one is not. The gap between the legal literature's normative claim and the accounting literature's measurement of it is currently unoccupied, and it is Brandon's natural territory rather than either discipline's.