Stafford
Waiting on you — 7 open items · see the decision queue

Latest brief · 2026-09-06 · previous (2026-09-05) · all briefs

Stafford Brief — 2026-09-06

Sunday. The scan's brief stands on its own today — Gallup's 43,262-response flat line and the Applause rollback rate are both correctly read there and are not restated here. This brief does one thing: it reports that the gap the scan named on 09-04 as this base's highest-value acquisition was already sitting in the base on 08-27, and that the entry holding it carries three errors, one of which has already propagated into the shared thread file.


⚠️ OPS FLAG — a two-day hole in the judgment layer, and today's finding sits inside it

Stafford did not run on 09-04 or 09-05. briefs/ goes 08-29 → 09-03 and resumes today. The scan ran all three days and its output is intact; what is missing is the review layer over it.

That is not a bookkeeping note. The single most consequential thing the scan did in those three days was the 09-04 Friday synthesis — two patterns escalated, the Claim 5 gap restated, and a new story thread opened. Nobody checked it. The miss reported below is a 09-04 miss, and it survived two days precisely because the layer whose job is to catch it was not running.


📖 TODAY'S READ

Rotenstein LS, Holmgren A, Thombley R, … Adler-Milstein J, Mishuris RG. "Changes in Clinician Time Expenditure and Visit Quantity With Adoption of Artificial Intelligence–Powered Scribes: A Multisite Study." JAMA. 2026;335(16):1408–1417.

Not a new find. It has been in the base since 2026-08-27, and it answers the question the scan opened a thread about on 09-04.

On 09-04 the scan escalated the Migrating Bottleneck pattern on NBER WP 35275, named Claim 5's highest-value acquisition as "a measurement of the attenuation where downstream steps are owned by someone other than the producer," and opened the thread "Does the writing-to-shipping attenuation replicate outside software?" — next check 2026-12-04. That thread's criterion asks for "any measurement of an AI productivity gain traced through two or more successive stages of an organizational production chain in a non-software function" and then names five candidate functions. The first one it names is clinical documentation.

The base ingested a clinical-documentation instrument eight days earlier and filed it against a different thread. Its title is Clinician Time Expenditure and Visit Quantity — the two stages, in the title.

What it measures. 8,581 ambulatory clinicians at five US academic health systems, 1,809 adopters against 6,772 non-adopters, scribes introduced June 2023 – August 2025, outcomes EHR-system-derived rather than self-reported. Stage one, the producer's own step, per 8 scheduled patient hours: documentation time −16.0 minutes (95% CI, 13.7–18.3), total EHR time −13.4 (9.1–17.7). Stage two, the organizational output: +0.49 weekly visits (0.17–0.81). And the measure the deployments were justified by — EHR time outside work hours — did not change significantly.

Why it is a better instrument for this gap than 35275 is, on the one axis the gap is about. 35275's stated limitation is that its population is GitHub developers who often own the whole chain, so "the instrument cannot separate an organizational bottleneck from individual capacity." In ambulatory medicine that confound cannot arise. A physician who saves sixteen minutes of documentation cannot convert it into a patient, because they do not own the schedule. The template, the slot length and the panel are held by someone else. The producer and the downstream step are separated by construction, not by luck — which is exactly what the 09-04 gap statement asked for.

The two instruments are complementary, and each has precisely what the other lacks. 35275 traces a multi-stage chain with a generational gradient and a matched event-study design, and cannot say the weak link is organizational. Rotenstein has the ownership separation and administrative outcomes, and is an opt-in observational cohort measuring two endpoints rather than a chain. Neither closes the gap alone; stated as a pair they very nearly do, and that sentence exists in neither the case file nor the thread.


🔭 THINKER TO WATCH

A. Jay Holmgren and Julia Adler-Milstein — second and penultimate authors, and the reason this paper is worth more than its abstract.

The scan featured Harter today; this is a different territory and does not overlap it. Holmgren and Adler-Milstein lead the empirical literature on EHR-derived measurement of clinical work — the discipline of taking audit-log and system-derived time data seriously as an instrument, which is why this paper reports minutes with confidence intervals where the rest of the AI-productivity market reports percentages from surveys. This base has thirty-odd instruments on AI and organizational outcomes and, before 08-27, essentially none from a research community that had already spent a decade building administrative measures of the work itself.

Brandon's differentiation, and it is the standing one. They measure burden — how much time the EHR takes from a clinician, and whether a tool gives some back. It is a clinician-welfare frame, and on that frame this paper's honest headline is disappointment: the off-hours number, the one closest to burnout, did not move. Brandon's frame asks the question they do not: given that the time was demonstrably freed, what in the organization decided where it went? Their own paper answers it in a sentence they treat as a study limitation and he should treat as the finding — quoted in the next section.

Affiliations not verified in this session — taken from the author list only. Worth ten seconds before either name goes in a draft.


🏢 CASE IN THE WILD

Five academic health systems bought capacity and, by their own account, never connected it to the schedule.

Verbatim from the paper: "none of the organizations in the study required that clinicians book additional patients to qualify for AI scribe use."

Read that as an operating fact rather than as a methods note. Five large institutions deployed a tool whose business case is capacity, across 1,809 clinicians over twenty-six months, and not one of them attached a single condition linking the tool to the outcome that justified it. The scribe was granted as a benefit. The visit template was left alone. Nobody owned the join.

That is Strategic Disconnection with a control group — the stated purpose (capacity, burden relief) and the operational reality (an unchanged booking template) never met, and the paper documents the non-meeting in one clause.

And the second verbatim line is the one that should stop anyone about to cite the revenue figure. The authors' own account of where the marginal revenue came from: "this marginal revenue may have derived from increased visits booked in available or nontemplated time or changes in the composition of level-4 and level-5 E/M visits coded facilitated by better documentation."

Part of the measured gain may be better-documented billing rather than more care delivered. The note improved, so the visit coded higher. That is Claim 4 in its purest form — the appearance of output, produced by the documentation layer, in a peer-reviewed paper, named by the authors themselves. The $167 per clinician per month figure sitting in this base's KB entry is secondary-sourced and is partly this. It should not be cited as productivity by anyone, including Brandon.

Open question on the case: which of the five systems has since changed a booking template, and would any of them be able to tell you?


⚙️ FOUR FORCES CONCLUSION

Commitment — and this is a cleaner instance of it than the base usually gets, because the failure is documented rather than inferred.

The base's ordinary Commitment evidence is a gap between what executives say and what surveys later report. Here the gap is a procedural fact recorded inside the study: the organizations stated a capacity rationale, funded it, ran it for twenty-six months, and imposed no condition connecting the tool to the schedule. Commitment is defined in this system as decision-makers acting consistently with stated direction under pressure. This is what its absence looks like when someone happens to write it down.

Capability is implicated too and points somewhere unexpected — see the Challenge, where the dose-response says the capability was not the constraint.


💭 OPEN QUESTION

If the conversion rate from freed time to output does not degrade with dose — and on this instrument it does not — then "organizational readiness" is not a conversion efficiency at all. It is a threshold. So what is the threshold, and is it a property of the organization or of the individual?

The numbers force this question and they were not in the base until today. Against the all-adopter estimates, the heavy-user subgroup (≥50% of visits) shows total EHR time −21.3 minutes (13.9–28.7), 1.59×; documentation −27.3 (23.1–31.6), 1.71×; and visits +1.0 weekly (0.5–1.6), 2.04×.

The output gain scales at least as steeply as the input saving. The framework's instinct — that organizations leak the gain in conversion — predicts the opposite: more time saved, proportionally less arriving. That is not what this shows. The conversion held, and the population-level result is small for a different reason entirely: only 21% of eligible clinicians adopted, and only about a third of those used it on half their visits. The shortfall is breadth, not leakage.

Which is a materially different diagnosis and a more actionable one. It also has a natural mechanism — fixed costs. Below some usage intensity you save minutes that never aggregate into a bookable slot; above it, they do. Returns that are threshold-shaped rather than linear are precisely what "readiness" ought to mean operationally, and this base has never had a number for it.

The discipline that keeps this honest, and it is severe. These are separate subgroup estimates on a self-selected subgroup, not an interaction test, in a cohort where adoption was opt-in at four of five sites. Clinicians who use a tool on half their visits are plausibly different from those who do not, in ways that also predict seeing more patients. The point estimates are consistent with threshold-shaped returns and the design cannot distinguish that from selection. Anyone citing the 2.04× as a dose-response is overreading it, and that includes this brief — which is why the claim above is stated as a question.


🥊 CHALLENGE

The Brynjolfsson–Li–Raymond counterexample gets a second, independent instrument — and it is a better fit for the objection than Brynjolfsson is. Claim 5's wording will not survive publication in its current form.

The CONTESTED entry logged on 09-03 sets out what it would take for the objection to win: "a redesign-free deployment producing durable, externally-measured gains where the mechanism is demonstrably not a capability/knowledge intervention the org failed to do itself — genuinely 'tool alone,' with the organizational conditions left untouched."

Rotenstein meets that description more cleanly than Brynjolfsson does, on all four clauses.

  1. The mechanism is not knowledge dissemination. The framework's escape route from BLR is that the tool was the intervention — it codified top agents' tacit practice and moved it to novices, which is Capability and therefore inside the Four Forces. An ambient scribe does no such thing. It transcribes an encounter and drafts a note. It substitutes for a documentation task; it disseminates nobody's expertise. The escape route is closed here.
  2. The organizational conditions were verifiably untouched"none of the organizations in the study required that clinicians book additional patients." Not inferred. Stated.
  3. Externally measured — EHR-system-derived, not self-report.
  4. Durable — twenty-six months, five institutions, 6,772-clinician comparison group.

So the objection strengthens, and the honest response is to change the claim rather than defend it. Claim 5 asserts that the prerequisite is organizational readiness rather than coordination technology. On this evidence, deploying a tool onto an entirely unchanged workflow produced a real, externally measured, statistically significant organizational output gain. Small — half a visit a week — but the strong form of "prerequisite" predicts approximately zero, and it is not zero.

The defensible restatement, and it is stronger than what it replaces because it is falsifiable: without redesign you get a real but small gain, captured by the minority who changed their own practice, while the outcome the investment was justified by does not move at all. That is a claim about magnitude and distribution, not about a precondition. Every number in this paper supports it: +0.49 population, +1.0 among heavy users, 21% adoption, and off-hours time flat — the thing they bought it for, unmoved.

And it converges with today's scan brief from a completely different instrument. The scan reached for the same restatement off Applause's rollback rate — "not 'AI without redesign produces no gain' but 'AI without redesign produces a gain too small to be worth running.'" Two instruments, one day apart, independently pushing Claim 5 from a prerequisite claim to a threshold claim. The two are not convergent evidence — different constructs, different populations, and one of them will not disclose its denominator. They are convergent on the wording, which is the thing that needs fixing.

Brandon: this is Paper 3's 70%-figure move, applied to your own unpublished claim, before a referee does it for you. The brand asset in asset-suggestions.md #1 is the person who checks the numbers everyone repeats. Checking your own is the same asset and costs nothing while the paper is still a draft.

Ledger updated: knowledge/thesis-challenges.md, existing CONTESTED entry strengthened rather than a new challenge opened — the objection is the same objection, and inflating the ledger with a second copy of it would be exactly the double-counting this base disciplines elsewhere.


🔍 SOURCES & THREADS

Exploratory query — the slot was spent on a coverage check rather than a live search, and that was the right call for the second time in four days. Target: Claim 5's named highest-value acquisition from 09-04. Method, applying the 09-03 lesson verbatim — check whether the gap is a coverage hole in the base before assuming it is a hole in the world — grep across all 498 KB files for the clinical-documentation, radiology and EHR literature, then verify the one hit against primary. Yield: no new source, one thread criterion met from stock, three corrections to an existing entry, and one challenge materially strengthened. Logged to knowledge/query-log.md.

The lesson repeats and should now be treated as a standing procedure rather than an anecdote. On 09-03 the exploratory slot found that the most-cited AI field experiment in existence had never been ingested. Today it found the opposite failure: an instrument that was ingested, correctly, and then not connected to the two places that needed it eight days later. Both are bootstrap failures. Both were caught by the same move. Proposed standing rule, in Routing: when a thread is opened or a case-file gap is restated, grep the base against the new criterion before setting a next-check date. The 09-04 thread would have been opened with its first criterion already met.

Verification performed (primary, this session). JAMA full record and the PMC copy of the same article read directly. Confirmed: full citation JAMA. 2026;335(16):1408–1417; visit outcome is weekly visits while time outcomes are minutes per 8 scheduled patient hours (different denominators — no conversion ratio may be computed between them); heavy-user subgroup figures 21.3 / 27.3 / +1.0 all present in the paper with confidence intervals; both verbatim quotations above. Mass General Brigham's newsroom URL now 301s to a site root and the release could not be re-read — the $167/month and 32% figures remain secondary-sourced and unverified.

Threads: none due today, and none forced. Nearest is the middle-manager thread on 2026-09-09, which the scan advanced ahead of schedule today. One thread advanced from stock rather than by search: "Does the writing-to-shipping attenuation replicate outside software?" — its first criterion is met on the evidence already held, and the proposed log entry is in Routing. No threads opened or closed by Stafford today.

Candidates: one proposed — Holmgren / Adler-Milstein and the EHR-derived-measurement literature as a monitored territory, not a single outlet. Rationale in Routing. No promotions, no expiries proposed.

Judgment files updated: thesis-challenges.md (CONTESTED entry strengthened, restatement recorded), query-log.md. asset-suggestions.md deliberately unchanged — the Claim 5 restatement is a paper revision, not a standalone asset, and opening asset #6 for it would inflate the ledger to record a decision Brandon can make in an afternoon.

SPARRING NOTE withheld, third consecutive day, same reason. Two notes remain outstanding this week (08-31, 09-01) with no recorded response and the cap is two. Today's threshold-versus-efficiency question is the best hook growth edge #3 has had — parked in brandon-development.md, not issued.


Routing

PROPOSED-UPDATE — knowledge-base/rotenstein-jama-ai-scribes-clinician-time-multisite-2026-04.md (research-evidence)

Three corrections, one addition. Anchor: the Summary section's third paragraph and the Relevance section.

  1. Strike "16 minutes a day." The Relevance section reads "16 minutes a day of documentation time saved converts to 0.49 additional weekly visits." The time outcomes are minutes per 8 scheduled patient hours; the visit outcome is weekly visits. Different denominators. It is not "a day," and the two figures cannot be divided into a conversion ratio without weekly scheduled hours, which the paper does not supply. This error has already propagated into knowledge/story-threads.md, "Where does the saved time go?", log entry 2026-08-27 — correct it in both places.
  2. Correct the heavy-user attribution and magnitudes. The entry states the 2×/3× figures are "from secondary coverage only and are not on the JAMA page." They are in the paper, with confidence intervals: total EHR time −21.3 (95% CI, 13.9–28.7); documentation −27.3 (23.1–31.6). Against the all-adopter estimates those are 1.59× and 1.71×, not "twice" and "three times." If the secondary coverage's 2×/3× was a heavy-versus-light comparison rather than heavy-versus-average, the entry must say so; as written it reads as heavy-versus-average and is wrong on that reading.
  3. Add the missing number, which is the most decision-relevant one in the paper for this base: heavy users delivered +1.0 additional weekly visits (95% CI, 0.5–1.6) against +0.49 for all adopters — 2.04×, i.e. the output gain scales at least as steeply as the input saving. Absent from the entry and from every thread log.
  4. Add both verbatim lines and complete the citation. "none of the organizations in the study required that clinicians book additional patients to qualify for AI scribe use" and "…changes in the composition of level-4 and level-5 E/M visits coded facilitated by better documentation." Citation: JAMA. 2026;335(16):1408–1417; first authors Rotenstein LS, Holmgren A, Thombley R (Mishuris RG is last author, not second). Attach a caution to the $167/clinician/month figure: it is secondary-sourced and, on the authors' own account, may be partly E/M coding composition rather than delivered care.

PROPOSED-UPDATE — knowledge/story-threads.md, "Does the writing-to-shipping attenuation replicate outside software?" (research-evidence)

Log entry, 2026-09-06 — the thread's first criterion was met before the thread was opened, from a source already in the base. The criterion asks for a measurement traced through two or more successive stages of an organizational production chain in a non-software function and names clinical documentation first. Rotenstein et al., JAMA 2026;335(16):1408–1417, ingested 2026-08-27, is that measurement: 8,581 ambulatory clinicians, 1,809 adopters against 6,772 non-adopters, five academic health systems, June 2023–August 2025, EHR-system-derived outcomes. Stage one — documentation time −16.0 min/8 scheduled patient hours (13.7–18.3); stage two — +0.49 weekly visits (0.17–0.81); and off-hours EHR time unchanged.

What it does to the thread's central question, which is why the thread stays open. The thread exists to separate an organizational bottleneck from individual capacity. On that axis this instrument is stronger than 35275 and the thread should say so: the clinician does not own the schedule, so producer and downstream step are separated by construction rather than by chance. But it measures two endpoints, not a traced chain, and it is an opt-in observational cohort rather than a matched event study. The two instruments are complementary — 35275 has the chain and the generational gradient and cannot locate the weak link; Rotenstein locates the ownership boundary and has no chain.

And the finding cuts against the thread's founding expectation, which must be recorded rather than smoothed. Heavy users show 1.59× the EHR-time saving, 1.71× the documentation saving and 2.04× the visit gain. The conversion does not attenuate with dose. The population-level shortfall is produced by breadth — 21% adoption, ~32% of adopters at ≥50% intensity — not by leakage in conversion. Subgroup estimates on a self-selected subgroup; consistent with threshold-shaped returns, not distinguishable from selection. Next check: bring forward from 2026-12-04 to 2026-10-06 and narrow the criteria to the two things now missing: an interaction test or any design that separates dose from selection, and a non-software instrument in a function without a scheduled, countable output unit.

PROPOSED-UPDATE — knowledge/paper-2-case-file.md, gap ranking (research-evidence)

Claim 5's gap statement of 09-04 should be amended rather than restated. It reads that the highest-value acquisition is a measurement of the attenuation where downstream steps are owned by someone other than the producer. The file already holds one and has since 08-27. The remaining gap is narrower and should be written as such: an instrument that separates dose from selection, in a function without a scheduled, countable output unit. Ambulatory medicine has a booking template — the closest thing to a best case for conversion, which the thread's own 08-27 entry already noted. Claim 5's evidence should be read down accordingly, not up.

PROPOSED-UPDATE — knowledge/emerging-patterns.md, Migrating Bottleneck (research-evidence)

Do not recruit Rotenstein as a fifth member, and record why. It measures two endpoints rather than a chain, and the pattern's value rests on instrumented multi-stage measurement. What it earns is a boundary note attached to the pattern: on the one instrument in this base where the producer demonstrably does not own the downstream step, the gain did not attenuate with dose. The pattern claims the bottleneck migrates; this is the first evidence in the file that where the handoff is clean and slot-shaped, it can also be crossed. Recurrence count unchanged at 4.

PROPOSED-UPDATE — knowledge/source-candidates.md and knowledge/field-map.md (research-evidence)

Candidate (territory, not outlet): EHR-derived measurement of clinical work — Holmgren, Adler-Milstein, Melnick, Tai-Seale and the JAMA/JAMIA/Annals channel they publish in. First seen 2026-09-06 via Rotenstein et al. Signal count 1. Rationale: this base measures AI-and-organization almost entirely in surveys and résumé panels; this community has a decade of administrative time-measurement and now points it at AI deployments. Promotion test: a second instrument from this territory measuring an AI deployment against an organizational output. Affiliations unverified in this session — confirm before either name enters the field map proper.

strategy

Claim 5's wording needs to change before Paper 2 goes out, and the change makes it stronger. Two instruments, from different functions, now show a real organizational output gain from deployment onto an entirely unchanged workflow. The prerequisite framing predicts roughly zero and gets +0.49 weekly visits with a confidence interval that excludes zero. Restate as magnitude and distribution: without redesign the gain is real, small, concentrated in the minority who changed their own practice, and absent on the outcome the investment was justified by. Falsifiable, defensible, and consistent with everything in the base — including the two findings that currently sit uncomfortably against the strong form.

A client question that costs nothing, as the counterpart to the scan's spend-ratio question. "When you deployed it, did you change anything downstream — a template, a quota, a target, a queue?" Five academic health systems, 1,809 clinicians, twenty-six months, and the answer on record is no. The value is the same as the spend-ratio question: most organizations cannot name a single downstream change, and the inability to name one is the finding.

governance

A citation hazard worth naming before it reaches a draft. The AI-scribe revenue figure circulating in secondary coverage may partly reflect E/M coding-level composition — better documentation producing higher-coded visits — rather than more care delivered. That distinction is the difference between a productivity claim and a billing claim, and in healthcare the second one has regulators attached to it. Do not cite the $167/month.

personal

The discipline that worked today was not analytical, it was procedural: grep before you search. Twice in four days the exploratory slot returned more by interrogating the base than by querying the world — once finding a canonical paper absent, once finding a held paper unconnected. The system's failure mode has shifted. At 498 files the risk is no longer that the base is thin; it is that the base knows things the briefs do not. That argues for spending the exploratory slot on coverage checks more often than the current cadence assumes, and it argues harder for the standing rule proposed above.

And the uncomfortable half, stated because it is the same lesson the scan wrote in its own personal note today. The finding that strengthens the counterargument against Claim 5 is one this base already owned and had read as support for Claim 5 — the 08-27 entry files it under Technology Illusion and Momentum Mirage and reads the small conversion as the organization failing to capture the gain. The dose-response says otherwise, and it was in the paper the whole time. Applying the bar when the source agrees with you is exactly what the scan said last night. This is what it costs when you don't.