Before You Scale Clinical AI: Did the Pilot Remove Work or Move It?

Illustrative image of a folder being handed between two team members at a clinic table.

If the first output is ready sooner, the expansion decision still depends on what happened afterward: who prepared the input, checked the result, corrected it, approved it and handled the follow-up? A clinical AI pilot may reduce total staff work, move work between roles, or add effort that a clinic accepts for a better result. The distinction depends on the pilot’s original purpose. If the goal was capacity, an earlier draft alone does not show that capacity was gained. If the goal was a clinic-defined improvement in the approved output, extra review time may be an acceptable tradeoff—but it should be visible and deliberate.

For a clinician-founder, medical director or operations lead, the useful question is: What changed across one complete, comparable episode, and did that change meet the objective we set? This is a way to read the clinic’s own observations after a bounded pilot. It does not assume that a particular AI system saves time or improves clinical care.

What do clinical AI trials tell us about time and performance?

Published experiments show why a good output and a time saving should be measured separately. In a 2024 randomized diagnostic-reasoning study in JAMA Network Open, 50 physicians worked through up to six clinical vignettes with either access to GPT-4 and conventional resources or access to conventional resources alone. LLM availability did not significantly improve the study’s primary structured-reflection score. The authors also did not confirm a difference in time spent per case. This was a physician vignette task, not an end-to-end clinic workflow.

A different 2025 randomized study in Nature Medicine (accessible author manuscript) assigned 92 physicians to GPT-4 plus conventional resources or conventional resources alone for five simulated management cases. Physicians with LLM access scored higher on that study’s expert-developed management rubric and spent more time per case. The authors called for validation in real clinical practice; the rubric’s validity outside the study was not established. A 17 February 2025 publisher correction adds omitted equal-contribution and joint-supervision authorship footnotes; it does not describe a change to the study results. These were different tasks and endpoints, so their findings should not be treated as a direct comparison of diagnostic and management workflows.

Neither experiment followed a clinic’s coordinator, clinician and support staff through preparation, review, approval and communication. Neither measured this article’s central outcome: whether work disappeared across the complete episode or was redistributed. The trials give a reason to ask separate questions about time and assessed output. The clinic must gather the episode-level evidence for its own expansion decision.

Which episode should you compare?

Choose a unit of work that starts and ends at points the team can recognize in both routes. For example, if the pilot concerns a periodic review, the unit might begin when the required information is available for preparation and end after the clinician-approved plan is communicated and the agreed follow-up is assigned. That is a possible comparison boundary, not a claim about any clinic’s workflow or a prescribed clinical protocol.

Before interpreting the results, write down the baseline route and the assisted route, their entry and exit points, and every role involved. Compare episodes with a similar mix of ordinary and difficult cases. Record how many were attempted and completed, how onboarding or a new workflow affected early attempts, and which cases were incomplete, escalated or required an exception. Keep the work spent on those cases in the comparison. A result from the first unfamiliar use should be distinguishable from a result after the team has used the process repeatedly.

Keep the clinic’s original objective in view. Was the pilot meant to release scarce clinician time, reduce total active staff effort, shorten elapsed turnaround, make an approved output more consistent, or address more than one of these? These goals can diverge. The DECIDE-AI reporting guideline asks early live evaluation reports to describe the clinical workflow, when AI was used and who reached the final supported decision; it also asks for usability and user learning-curve reporting. It is reporting guidance, not a validation of the comparison method proposed here. Qualified clinic professionals must define what constitutes an acceptable clinical output and how exceptions are handled.

Where did the work go?

Follow work by person and step, including attempted episodes that were incomplete or escalated. Include preparation and reconciliation of source information, AI-assisted drafting or synthesis where applicable, clinician checking, correction, approval, member communication and follow-up. Record active minutes for each role in the baseline and assisted routes. Then record elapsed waiting between handoffs separately: a case can move faster on the calendar while using more staff time, or use less active effort while waiting longer for approval.

Table 1. Three dimensions for comparing a complete pilot episode.
Dimension to compare What to record in both routes What the comparison can reveal
Active work by role Minutes spent by each person on preparation, draft work, checking, corrections, approval and follow-up Whether effort fell overall, shifted to a scarce role, or created new rework
Elapsed time Time from the agreed start to the agreed finish, including handoff waits Whether the episode reached completion sooner, independent of staff minutes
Output and exceptions The clinic’s own approved-output criteria, corrections, unresolved cases and exception handling Whether a change in time accompanied a result the clinic actually accepts

Table: Three distinct readings of one complete pilot episode. Populate every cell with observed baseline and assisted episodes; the table contains no benchmark or trial-derived threshold.

Do not combine these dimensions into one “efficiency” figure. If an AI-generated first draft takes less time but a clinician spends more time checking and rewriting it, the draft-time gain may be offset elsewhere. If total work is unchanged but appropriately assigned tasks move from a clinician to a role with available capacity, that may still serve the clinic’s stated objective; qualified clinic professionals must define clinician approval and clinical responsibility. The reverse shift can create a bottleneck even when total minutes fall. These are conditional interpretations, not findings from the cited trials or a claim about a specific product.

There is another defensible outcome: the assisted route takes longer, yet the clinic judges the final, clinician-approved result to be better on a criterion it selected in advance. Leadership may choose that tradeoff, provided the extra effort and exceptions are sustainable. The published management-vignette trial illustrates that higher rubric scores and longer case time can coexist in an experiment; it does not prove that a real clinic’s extra time buys better patient outcomes. If the pilot inputs later become part of a financial case, treat the work comparison as an input to that analysis, then use a separate clinical intelligence ROI framework with locally verified assumptions. A published formula is not evidence of this clinic’s savings.

Expand, redesign or stop?

Read the result against the purpose agreed before the pilot. Expand the bounded workflow only when comparable episodes meet that purpose, the people who own each step can sustain it, and exceptions remain manageable under the clinic’s separate clinical and governance review. An observed time saving alone does not clear those reviews.

Redesign and recheck when the overall result is promising but a specific handoff or role bears the new burden. A shorter first draft followed by repeated clinician correction, for example, points to a review or input-quality problem to investigate. Change the affected step, then compare another set of similar complete episodes. Do not treat the redesigned route as already proven by the original pilot.

Withhold expansion or stop when the objective is not met, a critical exception remains unresolved, or no one can reliably own the added work. “Stop” can mean stopping that particular AI-assisted route; it does not answer whether a different workflow or product could perform differently. Record the decision and its reason so an attractive isolated output does not quietly replace the clinic’s original goal.

A short decision review can stay focused: What was the objective? Were baseline and assisted episodes comparable? What happened to active work by role, elapsed time and approved output? Which exceptions remained, and who owns the next step? These are management questions for interpreting observed work, not a universal scorecard or clinical safety threshold.

What should a live demonstration show next?

If a supplier is part of the next step, ask it to demonstrate the same workflow boundary and review handoff you used in the pilot comparison, with synthetic or otherwise non-identifiable examples. Ask what the current product demonstrably supports, where staff still prepare, decide or correct, and which steps depend on an integration or a local process. This keeps the conversation anchored to the work leadership needs to evaluate.

HolistiCare’s clinician use-case overview is a starting point for a product conversation. In a demonstration, test the relevant handoff against the product as it exists today. The pilot comparison above does not establish a HolistiCare time saving, clinical improvement or passed evaluation.

What do you think?

What to read next