Stafford

← 2026-08-30 · all briefs · 2026-09-01 →

Stafford Brief — 2026-08-31 (Monday)

The scan had its best day in a fortnight and I am not going to repeat it — read research-evidence/briefs/2026-08-31.md for the Equilar disclosures, the Hugging Face incident, the escalated coverage-gap pattern and the four-week window. This brief does one thing: it takes the scan's lead find to the literature that has already run this experiment once, and comes back with a counterargument the scan had no route to. The headline is that today's best find is weaker than it looked at noon, and the reason is a peer-reviewed paper that says soft metrics work.


📖 TODAY'S READ

Executive Compensation Tied to ESG Performance: International Evidence — Cohen, Kadach, Ormazabal & Reichelstein, Journal of Accounting Research 61(4): 805–853 (2023) (Wiley; working paper ECGI N° 825/2022 / CEPR DP17267, March 2023)

The scan's Equilar entry reads activity-shaped AI metrics in executive pay — Qorvo's "exploration and deployment of AI tools" at 20% of an LTIP, Juniper's "win the AI opportunity" at 10% — as Momentum Mirage written into a compensation contract, on the reasoning that a metric with no stated result condition pays identically for a transformation and for a mirage. That reasoning has been tested once, on the previous non-financial metric to enter executive pay, and it lost. Cohen et al. document ESG metrics spreading through executive compensation internationally and report that "the adoption of ESG variables in managerial performance measures is accompanied by improvements in ESG performance."

This matters because ESG-in-pay is the soft metric par excellence. Aggregate, discretionary, frequently unquantified — Dell'Erba & Gomtsyan build an entire legal critique on exactly that softness. And the outcome still moved. If softness does not prevent the outcome from moving, then reading Qorvo's language as Momentum Mirage is a claim about metric grammar that the only available evidence does not support.

What Brandon should take from it, in one sentence: before writing a word about the compensation committee, know that there is a top-three accounting journal saying the mechanism works even when the metric is vague — and that the entire counterargument turns on one thing nobody in this base has yet read.

The discipline that has to travel with all of the above. The authors' verb is "accompanied by" — association, and their abstract claims nothing more. Wiley, SSRN and Taylor & Francis all returned HTTP 403; the ECGI PDF is font-subset encoded and would not extract. Only the abstract was read. No sample size, country count, window or identification strategy is cited here, because none could be verified. A "3% in 2010 to over 30% in 2021" adoption series is circulating in search summaries with no locatable primary — do not use it. File: knowledge-base/cohen-kadach-ormazabal-reichelstein-esg-pay-international-2023.md.


🔭 THINKER TO WATCH

Marco Dell'Erba (University of Zurich) and Suren Gomtsyan (LSE Law)"Regulatory and Investor Demands to Use ESG Performance Metrics in Executive Compensation: Right Instrument, Wrong Method," Journal of Corporate Law Studies (2024); authors' summary on the Harvard Law School Forum, 2024-12-02.

They argue that "aggregate ESG measures must be avoided because they fail to highlight specific areas that require immediate improvements" and that hitting the short-term goal "does not necessarily translate into better overall financial performance or more responsible corporate behaviour in the long-term." Their prescription: "a standard approach to the use of ESG metrics in pay must be avoided."

Why they are worth a slot rather than a footnote: that is Brandon's Momentum Mirage argument, reached independently, in corporate law scholarship, about the previous metric, two years early. This is a discipline that has been arguing about incentive-metric design quality — not incentive levels, not agency theory generally, but whether the words in the plan name a result — while organizational design has been arguing about culture and change management. Nobody in field-map.md occupies it.

How Brandon's position differs, and it is a real difference rather than a courtesy one. Dell'Erba and Gomtsyan treat a badly written metric as a drafting failure to be fixed by better-tailored metrics. Brandon treats it as a diagnostic symptom — the metric is vague because the strategy behind it is vague, which is Strategic Disconnection, and the remedy is upstream of the compensation committee. Their fix is better drafting; his is that better drafting of an incoherent intent produces a better-drafted incoherent intent. Their evidence is on his side and their theory of the problem is not, which is the most useful shape an adjacent scholar can have.


🏢 CASE IN THE WILD

Microsoft 2014 and Qorvo 2026 — the same instrument, pointed opposite ways, and Brandon already published one half of it.

Paper 3, Designed to Stall, mechanism 3 is "Compensation tied to multi-year transformation outcomes": "A meaningful share of executive compensation, not a token slice, is tied to specific multi-year transformation outcomes." The paper's working example is Microsoft restructuring executive long-term incentive grants around multi-year cloud adoption outcomes.

Qorvo weights "exploration and deployment of AI tools to enhance organizational productivity" at 20% of a long-term incentive plan. Same vehicle. Meaningful share, not a token slice — 20% clears Brandon's own bar. What is missing is the only thing Paper 3 actually asked for: the outcome.

The mechanism, named precisely: this is not a market that ignored Paper 3's prescription. It is a market that adopted the form of the prescription and inverted the content — which is a considerably more interesting failure than neglect, and it is Momentum Mirage operating on the remedy rather than on the disease.

The open question about what happens next, and it is not rhetorical. Paper 3 named Microsoft and Berkshire as exceptions because outcome-tied compensation was rare. AI metrics are now spreading through incentive plans fast. If the vehicle Brandon prescribed becomes common while carrying activity metrics, does his mechanism 3 get credited or discredited when the results come in? He has a citation risk here that runs in both directions and no control over which way it resolves.


💭 OPEN QUESTION

Is Momentum Mirage a claim about how a goal is worded, or a claim about what happens when it is missed?

The base has been tagging Momentum Mirage off metric grammar — no threshold, no success condition, therefore appearance-over-substance. Cohen et al. is the first evidence that grammar may not be where the action is. If soft ESG metrics moved ESG outcomes, the operative variable was something other than whether the sentence named a number: board attention, disclosure obligation, investor pressure, or simply that someone senior now had to talk about it quarterly.

That would not weaken the framework. It would relocate it — from the wording of the metric to the accountability architecture around it, which is where the Four Forces layer already lives and where Commitment specifically lives. It would also make the diagnosis harder to perform from a proxy statement, which is a real cost, since reading the wording is cheap and auditing the architecture is not.

The version of this Brandon cannot answer today and should want to: what would a result-defined AI metric even look like in 2026? "Win the AI opportunity" is mockable, but AI's business results are genuinely two to four years out and genuinely hard to attribute. If no company can write a defensible outcome metric yet, then activity metrics are not evidence of a mirage — they are evidence of an attribution problem, and the framework is diagnosing a measurement limit as an organizational failure. That is the strongest form of today's objection and it does not depend on Cohen et al. at all.


🥊 CHALLENGE

The soft-metric objection — logged OPEN today in knowledge/thesis-challenges.md.

At full strength: the framework infers Momentum Mirage from the form of a metric, and that inference has never been tested. It has now been tested once, on ESG, by four authors at San Diego State, IESE, IESE/CEPR/ECGI and Mannheim/Stanford, in the Journal of Accounting Research, and the outcome moved. Worse for the base: this is the same evidence shape it would cite approvingly if the sign went the other way.

What it would take for the objection to win: Cohen et al.'s outcome variable turning out to be substantive and externally measured — verified emissions, injury rates, audited diversity data — plus any AI-specific replication linking activity-defined metrics to subsequent conversion. That would force Momentum Mirage to be restated as a claim about accountability architecture rather than about wording.

What would defeat it, and this is the live possibility: if ESG "performance" is measured by ratings that are themselves largely disclosure- and activity-scored, then the finding partly reduces to paying executives to score better on a disclosure index makes them score better on a disclosure index — which leaves Momentum Mirage untouched and, read properly, is an instance of it.

So the entire challenge turns on one unread paragraph. Get the JAR article and read its outcome construction. It is the cheapest open item in the challenge ledger and it resolves in one direction or the other. Until then the honest position is that the base tagged a breakpoint from metric grammar without ever having tested whether metric grammar predicts anything — and I am not softening the "Paid for Deployment" pattern on an abstract, because at two sources with no prevalence claim it is not over-extended.

No Four Forces section today. Only abstracts were readable, so Forces are an honest null on my find; the scan's brief carries the day's Forces reading.


🧱 ASSET SUGGESTION

Asset #4 — The Proxy-Statement Instrument. Full spec in knowledge/asset-suggestions.md.

The pitch, in two sentences: asset 3b — the prospective study that would answer the unfalsifiability challenge — is stalled because it is priced at ~30 sites plus a partner with access to live deployments, on the assumption that the ex-ante organizational variable has to be collected. For the incentive layer it is already filed: code every disclosed AI metric in the S&P 500 as activity-defined versus result-defined at time T, measure AI outcomes at T+2 — no access required, universe rather than convenience sample, and Cohen et al. is the published template proving the design clears a top journal.

Form: instrument and dataset first, essay after. Feeds: the Equilar entry and the "Paid for Deployment" pattern (which explicitly says it is watching for "any instrument reporting the share of a defined universe that discloses an AI metric" — this builds it rather than waits for it); the unfalsifiability challenge; Paper 3 mechanism 3.

The bound, stated so it does not get lost: a disclosed incentive metric is not a Five Breakpoints readiness score. It measures one breakpoint at one altitude and says nothing about process friction or capability below the compensation committee. Pitching it as the framework's readiness instrument is exactly the over-claim the 2026-08-30 note warned against. Blocked on the same cheap step as the challenge: read Cohen et al.'s outcome variable first.


🎓 SPARRING NOTE

Growth edge #3 — from intervention design to intervention evidence.

Paper 3 gives six mechanisms. Today, for the first time, one of them met evidence: mechanism 3 has a published near-analogue (ESG-linked pay) with a peer-reviewed result attached, and a live natural experiment running (AI metrics entering LTIPs now).

The exercise, and it should take twenty minutes on paper, not a search: for mechanism 3, write down the single measurement that would show it working, and the single measurement that would show it failing — before looking at what exists. Then check them against Cohen et al. The question worth sitting with is whether your falsification criterion for mechanism 3 is one you could have satisfied from public filings all along, or whether you would have needed access no consultant gets. If it is the former, edges #2 and #3 are the same problem and the readiness instrument has been cheaper than assumed since Paper 3 was published. Run the same exercise on mechanism 1 next; six mechanisms, six falsification criteria, is a publishable appendix and the most direct available answer to the unfalsifiability challenge.


SOURCES & THREADS

Threads: no thread's next-check date has arrived — nearest is the middle-manager thread on 2026-09-09 — and the scan's brief confirms it checked none on schedule today, so there was no due thread for me to cover. Two threads were advanced ahead of schedule by the scan's own finds; I have nothing to add to either that the scan did not already log. Nothing opened, nothing closed.

Exploratory query (logged to knowledge/query-log.md): has anyone measured whether disclosed non-financial executive incentive metrics predict the outcome they name? — the ESG-in-pay literature as the precedent case for the scan's brand-new "Paid for Deployment" pattern. Deliberately outside the standing categories and outside both query logs. Yield: one KB file, one new thesis challenge, one new asset case — and the yield is a counterargument rather than support, which is the outcome this slot is supposed to be capable of producing on a non-Wednesday and rarely does.

Source candidates: two proposed below, both from the same query, both in a territory (corporate-law scholarship on incentive-metric design) that field-map.md does not currently reach.

Judgment files updated: thesis-challenges.md (new OPEN challenge), asset-suggestions.md (new asset #4), brandon-development.md — not updated; the sparring note above is the first ever issued and the session log should record Brandon's response, not my prompt.


Routing

STAGE-CANDIDATE — for the next research-evidence scan run

Ready to run from the research-evidence repo root. Forces are an honest null (abstract only); no sample_size field, because none could be verified.

python3 build/add-research-json.py --json '{
  "title": "Executive Compensation Tied to ESG Performance: International Evidence",
  "publisher": "Journal of Accounting Research",
  "url": "https://onlinelibrary.wiley.com/doi/10.1111/1475-679X.12481",
  "published_date": null,
  "credibility": "high",
  "breakpoints": [
    {"name": "momentum_mirage", "rationale": "COUNTERARGUMENT: the authors report that adoption of ESG variables in managerial performance measures is accompanied by improvements in ESG performance, which cuts against inferring Momentum Mirage from a metric that states no result condition — the reasoning this base applied to the Equilar AI disclosures on the same day."}
  ],
  "forces": [],
  "key_findings": [
    "Adoption of ESG variables in managerial performance measures is accompanied by improvements in ESG performance and meaningful changes in executive compensation (authors: accompanied by, i.e. association, not causation)",
    "The practice varies at country, industry and firm level in ways the authors read as consistent with efficient incentive contracting",
    "Reliance on ESG metrics is associated with institutional investor engagement, voting and trading",
    "Abstract-only read: outcome variable, sample size, country count and observation window all unverified — Wiley, SSRN and T&F returned HTTP 403"
  ],
  "instrument": {
    "data_source": "UNVERIFIED — not disclosed in the abstract; full text inaccessible (HTTP 403 at Wiley, SSRN, Taylor & Francis; ECGI PDF font-subset encoded and unextractable)",
    "inference_method": "UNVERIFIED — identification strategy not stated in the abstract; the authors word the key result as association",
    "observation_window": "UNVERIFIED — no window stated in the abstract; working-paper version cover-dated March 2023, so the panel necessarily predates it and is entirely pre-generative for AI purposes"
  }
}'

Staging caution for whoever runs it: this entry is deliberately thin because the paper could not be read. If the full text is obtained first, re-run with the real instrument fields rather than patching a duplicate — and the outcome variable is the field that matters.

PROPOSED-UPDATE — knowledge/emerging-patterns.md, "Paid for Deployment" pattern

Append to the pattern entry. Recurrence count unchanged at 2 — this adds no member, it adds a challenge to the mapping.

Update 2026-08-31 (Stafford, delta layer) — the pattern's Momentum Mirage mapping now carries a named counterargument, and it should be stated inside the pattern rather than discovered by a reviewer. This pattern reads activity-shaped AI metrics as Momentum Mirage on the grounds that a goal with no stated result condition pays identically for a transformation and for a mirage. That inference has been tested once, on the previous non-financial metric to enter executive pay, and it lost. Cohen, Kadach, Ormazabal & Reichelstein, Journal of Accounting Research 61(4):805–853 (2023), report that "the adoption of ESG variables in managerial performance measures is accompanied by improvements in ESG performance." ESG-in-pay is the soft metric par excellence — aggregate, discretionary, frequently unquantified — and the outcome moved anyway. Three disciplines. (1) The authors' verb is "accompanied by": association, and nothing in the abstract claims more. (2) Only the abstract was read (Wiley/SSRN/T&F all HTTP 403), so no sample size, window or identification strategy may be cited for it. (3) The objection turns entirely on the outcome variable, which is unread: if ESG performance is rating-based, and those ratings are themselves disclosure- and activity-scored, the finding partly reduces to paying executives to score better on a disclosure index — which would leave this mapping untouched and would be an instance of Momentum Mirage rather than a refutation. Do not weaken the pattern on this; do carry it. Logged as the soft-metric objection in stafford-research/knowledge/thesis-challenges.md. Action that resolves it: obtain the JAR article and read the outcome construction. File: stafford-research/knowledge-base/cohen-kadach-ormazabal-reichelstein-esg-pay-international-2023.md

PROPOSED-UPDATE — knowledge/source-candidates.md

Two candidates, both first seen 2026-08-31, both at 1 signal, both surfaced by Stafford's exploratory query on non-financial incentive metrics.

Journal of Corporate Law Studies (1 signal) — surfaced via Dell'Erba & Gomtsyan, "Regulatory and Investor Demands to Use ESG Performance Metrics in Executive Compensation: Right Instrument, Wrong Method" (2024). Logged for the genre: peer-reviewed corporate-law scholarship on incentive-metric design quality — whether the words in a compensation plan name a result — which is Brandon's Momentum Mirage territory approached from a discipline no source in sources.md covers. Promotion bar: a second find in this venue bearing on incentive-metric design or AI governance obligations, not a second ESG paper.

Journal of Accounting Research (1 signal) — surfaced via Cohen, Kadach, Ormazabal & Reichelstein (2023). Logged because it is the only venue that has so far produced a test of whether a disclosed non-financial incentive metric moves the outcome it names, which is the measurement this base most lacks. Promotion bar: a second find measuring a disclosed organizational or incentive variable against a subsequent outcome. Caution on both: these are high-credibility academic venues, so the temptation will be to promote on one find — hold probation, as the rules require.

PROPOSED-UPDATE — knowledge/field-map.md

Marco Dell'Erba (University of Zurich) and Suren Gomtsyan (LSE Law), added 2026-08-31 via Stafford's exploratory slot. Corporate-law scholars on incentive-metric design; "Right Instrument, Wrong Method," Journal of Corporate Law Studies (2024). Relevance: they reach Brandon's Momentum Mirage conclusion about compensation metrics independently, from law rather than organizational design — "aggregate ESG measures must be avoided because they fail to highlight specific areas that require immediate improvements." Differentiation: they treat the vague metric as a drafting failure fixable by better-tailored metrics; Brandon treats it as a symptom of upstream strategic vagueness, so better drafting of an incoherent intent yields a better-drafted incoherent intent. Their evidence supports him; their theory of the problem does not. Watch for: any work by either on AI-specific performance metrics in pay, which would put them directly on this base's territory.