Stafford

← 2026-09-02 · all briefs · 2026-09-04 →

Stafford Brief — 2026-09-03

Delta layer on today's scan (1 staged, 10 skipped; flagship: the Alibaba agentic field experiment). Today's scan brief is strong and self-contained — read it first. This brief adds one thing it could not: the base was missing the entire canonical AI-field-experiment literature, and the most important missing paper comes back the other way.


📖 TODAY'S READ

Brynjolfsson, Li & Raymond, Generative AI at Work — NBER WP 31161 (2023) / QJE 140(2):889–942 (2025). Stafford's own find; the delta on the scan's flagship.

The scan called yesterday's Alibaba paper "the first randomized entry in the case file." That was true — and the reason it was true is the story. Across 492 KB files the base holds not one of the canonical AI field experiments: not this one, not Dell'Acqua et al.'s Jagged Frontier, not Peng et al.'s Copilot RCT, not Noy & Zhang in Science. The base's flagship "first randomized test of the thesis's mechanism" was first only because the most-cited AI RCT in existence — whose co-author (Brynjolfsson) is already in the base via the Stanford Enterprise AI Playbook — had never been ingested.

It matters because it comes back the other way, in the same function. BLR staggered a generative-AI assistive chat tool across 5,172 customer-support agents and measured, off platform logs: +14% resolutions/hour (+34% for novices, ~0 for experts), improved customer NPS, lower attrition. No role redesign. That is the exact falsifier the base's own 09-02 falsification sweep hunted — "measured AI returns without role change" — declared not to exist. It exists, it is famous, and the scan's own sweep missed it because it searched 2026 SEO phrasings for a 2023 paper.

This is not a contradiction of Alibaba; it is the axis the scan asked for. Alibaba is agentic-with-handoff and cut the touched-work rating −0.412; BLR is assistive-in-the-loop and raised NPS. Same function, opposite sign, differing on one variable: autonomy. The base can now draw that contrast for the first time — and it should hold both, not one.


🔭 THINKER TO WATCH

Danielle Li (MIT Sloan) — labor economist, BLR co-author. The scan named Lauren Lu (operations management) as the discipline running the experiments; Li is the other one, and the pairing is the point. Lu runs experiments inside live service queues (ops); Li runs them on administrative personnel data (labor economics). Both disciplines are entirely unsampled by this base, whose experiment coverage until today was zero, and both are running Brandon's exact question — what AI does to the shape of work — with instruments his practitioner and consulting sources cannot match.

Brandon's differentiation: Li reads AI as skill compression — the tool moves tacit expertise down the distribution, lifting novices. That is a Capability mechanism and it is one altitude below the framework: it explains who gains, not whether the organization can hold the gain once the routine work is gone and the residual is the hard part (which is precisely where yesterday's Alibaba result found the human doing less). Li measures the augmentation; Brandon names what the augmentation leaves unaddressed.


🏢 CASE IN THE WILD

A Fortune 500 business-process-software support floor — the assistive mirror of yesterday's Taobao floor. Same work (customer chat resolution), same measurement genre (platform logs, a customer-satisfaction score), and the opposite outcome, because the design was opposite. The Taobao floor let an agent handle chats and hand the failures back to a human mid-stream; this floor let the AI suggest while the human stayed on every chat start to finish. The first degraded the customer rating on the work the AI touched; the second raised NPS and retained newer workers.

The breakpoint mechanism is the handoff, not the tool. In BLR the human never receives a half-failed interaction — they receive a suggestion and remain the decision-maker, so there is no mid-failure handoff to a disengaged supervisor. In Alibaba the human inherits the AI's emotional failures and does measurably less on them. The open question this leaves — and it is Brandon's, because neither paper asks it: is the assistive/agentic line the real design boundary, or is it a proxy for who owns the outcome? BLR's agent owns the whole chat; Alibaba's supervisor owns only the exceptions the AI could not close. Autonomy and accountability move together in these two cases and nothing here separates them.


⚙️ FOUR FORCES CONCLUSION

Capability — and the two experiments together show it is not one thing. The scan concluded Capability from Alibaba: oversight capability is conditional, and nobody measures the condition. BLR sharpens that into a distinction. In BLR, Capability behaves as a substitutable stock — the tool fills the novice's missing skill directly, which is why the gain is 34% at the bottom and ~0 at the top. In Alibaba, Capability behaves as a conditional — the trained, authorized supervisor's capability to recover a failure depends on the failure type and on what they are still measured on. Assistive AI adds capability where it is missing; agentic AI exposes whether the surrounding capability is real when the work gets handed back. Same Force, two mechanisms, and the paper-2 argument needs both: the augmentation case (Capability supplied) and the oversight case (Capability assumed and found conditional).


💭 OPEN QUESTION

Does "AI as a knowledge-dissemination layer" count as the organizational redesign the framework requires — or is it exactly the tool-on-unchanged-conditions the framework says fails? BLR's mechanism is that the AI codified top agents' tacit knowledge and moved it to novices. That is either (a) the tool doing the org's capability work for it — in which case BLR confirms the thesis (readiness, supplied by the tool, is what produced the gain) — or (b) a genuine counterexample, tech deployed on an unchanged role producing durable measured gains. The framework cannot have it both ways silently. Which reading is right turns on whether Brandon can state Claim 1 so that a knowledge-flow change counts as a design change — and that is the same grounding (Garicano's knowledge hierarchy) that growth edge #1 has been asking for. The definitional question and the empirical counterexample are the same problem.


🥊 CHALLENGE

Tool-deployment returns without redesign — the BLR counterexample. Logged CONTESTED today (knowledge/thesis-challenges.md). At full strength: the framework's prerequisite claim says deploying AI onto an unchanged role yields appearance, not substance; BLR is the largest, cleanest, most-cited case of the opposite — externally-measured quality and productivity gains with no role redesign, in the same function as the base's flagship. What keeps it CONTESTED rather than won: (1) BLR's mechanism is a capability/knowledge intervention — the tool did organizational work the firm had not — which is inside the Four Forces; (2) it is assistive, human-in-the-loop throughout, and pairs with Alibaba to isolate autonomy rather than to refute anything; (3) it carries its own Momentum-Mirage tail (top workers follow AI even when it is worse). What it would take to win: a redesign-free deployment with durable external gains whose mechanism is demonstrably not capability the org failed to supply itself. Do not cite the redesign-prerequisite claim in public as if BLR did not exist — pair them, and win the definitional point in the Open Question above before leaning on the claim.


🧱 ASSET SUGGESTION

"The Human-in-the-Loop Control That Isn't" — essay + a three-question audit card for the practice (asset #5, logged today). The base now holds both arms of the autonomy axis in one function: assistive (BLR, quality up) and agentic-with-handoff (Alibaba, quality down). The consulting-usable object is the audit the scan's 09-03 strategy routing already drafted — which failure classes must the human recover, what evidence that they can, what is the review measured on that isn't throughput — and the intellectual spine is that "human in the loop" collapses two designs that behave oppositely and calls both a control. Not time-priced (the scan's own read: the material is fourteen weeks old and unread), so this is argument value. Feeds: both experiment KB files, the verification-cost thread, the governance-survey prevalence. Ships independent of the essay if Brandon wants the audit in the practice this week.


🔍 SOURCES & THREADS

Exploratory query (logged, knowledge/query-log.md): targeted the 09-03 top gap (Claim 5, "a second experiment where role design is the treatment") by first checking whether the base lacked the field-experiment canon rather than searching the world — the scan's "first randomized entry" framing implied it did. Yield: 1 KB file + 1 STAGE-CANDIDATE (BLR) and the base's biggest structural gap named. Lesson worth keeping: before spending the slot on a live search, check whether the gap is a coverage hole (the paper exists and was never ingested) rather than a world hole. Today it was coverage — the most-cited AI RCT in existence, by an author already in the base.

Threads: none due (nearest 2026-09-09, middle-manager). No thread opened or closed by Stafford today.

Candidates / field map: two PROPOSED-UPDATE blocks below — the AI-field-experiment canon as a field-map cluster, and the remaining three canonical experiments as ingest targets. No promotions/expiries by Stafford.

Judgment files updated (this repo): thesis-challenges.md (1 new CONTESTED challenge; 2 dated updates to OPEN challenges — unfalsifiability weaker on its "autopsy" limb, measurability-confound gains one adjacent limiting data point); asset-suggestions.md (asset #5 opened); brandon-development.md (sparring note not issued — weekly cap of two already spent on 08-31/09-01; BLR→Garicano hook parked for the first /coach); query-log.md (today's query).


Routing

STAGE-CANDIDATE — Brynjolfsson, Li & Raymond, Generative AI at Work (ready for the next scan run or Brandon to stage in research-evidence)

python3 build/add-research-json.py --json '{
  "title": "Generative AI at Work",
  "publisher": "NBER Working Paper 31161 / Quarterly Journal of Economics 140(2):889-942",
  "url": "https://www.nber.org/papers/w31161",
  "published_date": "2023-04-23",
  "credibility": "high",
  "breakpoints": [
    {"name": "technology_illusion", "rationale": "COUNTER-EVIDENCE: an assistive GenAI tool deployed onto the existing agent role with no workflow redesign raised resolutions/hour 14% and improved externally-measured customer NPS - a substantive, quality-positive result the Technology-Illusion reading must absorb as a capability intervention (the tool codified and disseminated top-agent tacit knowledge) rather than dismiss."},
    {"name": "momentum_mirage", "rationale": "The authors report top workers increasingly adhere to AI recommendations even when those recommendations are worse than their own unaided judgment - a measured instance of a best-practice signal displacing expert substance, inside an otherwise pro-tool result."}
  ],
  "forces": [
    {"name": "Capability", "rationale": "The effect structure is a Capability story: +34% for novices, ~0 for experts, because the tool substitutes for missing capability where it is lowest - the AI functions as a real-time knowledge-transfer layer, not autonomous throughput."}
  ],
  "key_findings": [
    "Access to the assistive tool raised productivity (issues resolved per hour) 14% on average, 34% for novice/low-skilled workers, minimal for experienced workers",
    "AI assistance improved customer sentiment (NPS) and reduced attrition, the retention gain concentrated among newer workers",
    "Mechanism: the model disseminates the tacit best practices of the most able workers, compressing the experience curve",
    "Cautionary tail: top workers increasingly follow AI suggestions even when the suggestions are worse than their own"
  ],
  "sample_size": 5172,
  "instrument": {
    "data_source": "Administrative platform logs from a Fortune 500 business-process-software firm customer-support operation (agents mostly in the Philippines)",
    "inference_method": "Staggered team-by-team rollout exploited as difference-in-differences, with an August 2020 RCT-analysis subsample",
    "observation_window": "2020-2021 (GPT-family assistive tool; ~5% of workers had access Oct 2020, ~70% by Jan 2021) - pre-ChatGPT vintage, argumentative value not a 2026-frontier claim"
  }
}'

PROPOSED-UPDATE — research-evidence knowledge/paper-2-case-file.md (append under the 2026-09-03 evidence section)

Structural-gap note, 2026-09-03 (via Stafford). The 09-03 entry called the Alibaba paper "the first entry in the case file where an AI deployment was randomly assigned." Correct — but the reason is a coverage hole, not a scarcity of experiments: the base holds none of the canonical AI field-experiment literature. Brynjolfsson, Li & Raymond (QJE 2025; NBER 31161) is now filled locally (stafford-research/knowledge-base/brynjolfsson-li-raymond-generative-ai-at-work-2023-04.md; STAGE-CANDIDATE ready) and is the assistive comparator to Alibaba in the same function: +14% productivity (+34% novices), improved NPS, no role redesign. The pair isolates autonomy — assistive-in-the-loop improved quality; agentic-with-handoff degraded it. Highest-value next acquisitions to close the gap: Dell'Acqua et al., Navigating the Jagged Frontier (BCG consultants, 2023); Peng et al., GitHub Copilot RCT (2023); Noy & Zhang, Science (2023). Ingesting the class lets the case file argue Claim 3/Claim 5 from the whole evidence base rather than one experiment — and lets Claim 4's Momentum-Mirage mapping cite BLR's "top workers follow worse AI" tail.

PROPOSED-UPDATE — research-evidence knowledge/field-map.md (new cluster) and knowledge/source-candidates.md

New field-map cluster — "AI field-experiment economics" (labor economics + operations management). The base has sampled economics, consulting and computer science and never the experimental wings. Anchor names: Danielle Li (MIT Sloan) and Lindsey Raymond (BLR); Lauren Xiaoyuan Lu (Tuck) already logged by the 09-03 scan as an OM candidate. Add to source-candidates.md at 1 signal each: Fabrizio Dell'Acqua / Ethan Mollick (Jagged Frontier), Shakked Noy & Whitney Zhang (MIT, Science 2023). First-seen 2026-09-03, surfaced by the BLR ingest. Promotion bar: a second on-thesis experimental find in a different function. This is the third structural source-gap the base has named in a fortnight (after the résumé-panel and the OM gaps) and it should be closed deliberately.

Routing by audience