When our small behavioral-health practice first considered AI for administrative work, the proposed workflow looked clean on paper. We wanted faster intake follow-up, cleaner internal summaries, and less time spent rewriting routine messages. The business case seemed straightforward. What we actually got was a hard stop from legal and privacy review, followed by a narrower design that took longer to implement but could be defended. That sequence is the subject of this post.
This is not a story about technology failing. It is a story about a workflow that asked the organization to carry more risk than it could reasonably accept, and about the redesign that followed. The experience sits at the center of how I now think about AI workflow approval in small practices.
The Workflow We Wanted
The original idea was simple. Staff would paste limited referral and intake details into an AI tool to generate follow-up language, appointment reminders, and internal status notes. The tool would return drafted text. A clinician or coordinator would review it, make light edits, and send. We expected measurable time savings on routine outreach.
We mapped the steps:
Coordinator collects basic referral information already present in the EHR.
Selected fields are copied into the AI interface.
AI produces a draft message or summary.
Human reviews and sends or files the result.
On the surface, the design included human review. In practice, the review step was thin. There was no clear rule about which fields could never leave the practice environment, no separate account structure, and no documented exception log. The assumption was that “someone will look at it” would be enough.
It was not.
Why Legal Rejected the First Version
Legal and privacy review raised three concrete objections. First, the data boundary was too porous. Even limited referral details can contain identifiers and clinical context. Second, the human-review step lacked enforceable ownership and documentation. Third, the practice had no clear process for recording near-misses or deciding when to stop the pilot.
The review did not say AI was forbidden. It said this particular workflow asked the organization to accept risk it could not yet manage. That distinction matters. The tool itself was not the problem. The operating design around the tool was the problem.
We had treated a vendor feature set as if it were a completed risk assessment. It was not. That’s a judgment call, not a tool question.


The Redesign We Built Instead
We rebuilt the workflow around four constraints that the first version had ignored.
Data minimization first. Only non-clinical, non-identifying administrative language could enter the AI interface. Referral names, dates of birth, diagnosis codes, and free-text clinical notes stayed inside the EHR. Drafts were generated from structured templates the coordinator already controlled.
Separate accounts and access controls. No shared logins. Each user who touched the tool had an individual account with the minimum permissions required. When someone left the practice, access was revoked as part of the standard offboarding checklist.
Explicit human-review ownership. Every draft carried a named reviewer. The review was not a checkbox. It required the reviewer to confirm that no restricted data had entered the prompt and that the output did not introduce new clinical language. That confirmation was logged.
Exception and near-miss recording. We created a simple one-page log. Any time the AI produced something unexpected, or a staff member felt the boundary had been approached, the event was written down. We reviewed the log weekly during the pilot.
The resulting workflow was slower than the original proposal. Time savings were real but modest. What we gained was a process the practice could explain and defend.
What the Pilot Actually Showed
Over six weeks we ran the narrowed workflow on a limited set of administrative tasks. Three findings stood out.
First, staff questions surfaced earlier and more usefully than resistance. When people asked “Can I paste this field?” the answer was almost always no, and the question itself became evidence that the boundary was being respected.
Second, the exception log captured several near-misses that would have been invisible under the first design. One involved a coordinator who almost included a free-text note. Another involved an output that drifted into clinical phrasing. Both were stopped before they left the practice.
Third, the narrower design forced clearer ownership. When something needed explanation later, we had a named reviewer and a dated log entry rather than a vague recollection that “someone looked at it.”
These outcomes did not prove the workflow was risk-free. They proved it was reviewable. That is the standard a small practice can reasonably meet.
Practical Takeaways for Other Practices
If you are evaluating an AI workflow for administrative tasks, start with the rejection criteria rather than the feature list. Ask:
What data is allowed to leave the practice environment, and what is never allowed?
Who owns the human-review step, and how is that ownership recorded?
How will exceptions and near-misses be captured while the pilot is still small?
What happens to accounts and access when a staff member leaves?
The answers do not need to be complete on day one. They need to be explicit enough that the practice can defend the decision later.
We wanted a faster workflow. Legal rejected the first version because the operating design around the tool was incomplete. What we built instead was narrower, slower to launch, and far easier to explain. That trade-off is the one most small practices will face. The tool does not sign the note. You do.