When the note-writing outlasts the visit: applied AI for behavioral-health teams

Documentation has quietly become the job in community mental health. Separate the rules from the judgment and human-in-the-loop AI gives clinicians their hours back — without touching the record's integrity.

Eli Wood headshot

Eli Wood

July 22, 2026 3 min read
Two facing chairs, an open clinical notebook and an hourglass, evoking a behavioral-health clinician getting documentation time back.

The problem

Ask a behavioral-health clinician what they'd do with an extra hour a day and none of them says "see fewer clients." They say "finish my notes before 9pm." Documentation in community mental health has quietly become the job — intake assessments, progress notes, treatment plans, and the encounter data that both reimbursement and certification depend on. A clinician can spend nearly as long writing about a session as they spent in it.

That was tolerable when budgets were flat. It isn't now. An across-the-board Medicaid rate cut has landed on organizations already at break-even, and the first capacity to go is administrative — the people who used to buffer clinicians from paperwork. So the documentation lands harder on the clinicians who remain, right as there are fewer of them. Serve the same community with less staff, or shrink. That's the squeeze in one sentence.

The insight

The instinct is to treat this as a technology-adoption problem: buy an AI scribe, bolt it on, move on. That misreads where the value and the risk actually sit.

Behavioral-health documentation isn't one task — it's two different kinds of work wearing the same clothes. There's the rules part: which fields a note must contain, which service maps to which code, which quality measures must be reported. And there's the judgment part: what actually happened in the room, what the clinician noticed, what the clinical reasoning was. The mistake most tools make is letting a model do both at once — which is exactly how you get a fast note that's subtly wrong in a way an auditor, or a court, will find.

Separate the two and the problem gets tractable. Let a deterministic rules layer own the structure — required fields, coding, completeness checks — where being consistent matters more than being clever. Let the model do the first-draft narrative from the session, and then put the clinician in the loop as the editor, not the author. The clinician's scarce judgment is spent verifying and correcting, not typing from scratch. That's the whole move: start where the work is expensive and repetitive, keep a human in the loop wherever being wrong is costly, and earn trust by making every AI-drafted line easy to check against the source.

The orgs that win here won't be the ones that adopt the flashiest scribe. They'll be the ones that redesign the documentation workflow around that human-in-the-loop split — so a 45-minute intake write-up becomes an 8-minute review, and the clinical record actually gets cleaner because the rules layer never forgets a field.

The path — what to do in two days

You don't need a year-long transformation to find out whether this works. In two focused days you can:

  • Pick one workflow, not all of them. Intake assessments and crisis/urgent-care documentation are usually where the load concentrates and where the ROI shows up in weeks. Choose the one that's hurting most.
  • Map the rules vs. judgment split for that workflow — what must be structurally correct (fields, codes, measures) vs. what needs clinical narrative. This is where the human-in-the-loop boundary gets drawn.
  • Prototype the review, not the magic. Draft the note, then design the clinician's 8-minute verification step — what they see, what's pre-filled, what they can override. If clinicians don't trust the review, nothing else matters.
  • Instrument the time saved and the audit quality on a handful of real (de-identified) cases, so you can prove capacity recovered without any drop in record integrity.

That's a Foundation Sprint. It's deliberately small, deliberately concrete, and deliberately about a workflow your clinicians already do every day — because the goal isn't to "do AI," it's to give clinical hours back to the people the community can't afford to lose.

If this is your situation, spend two days with us. We call it a Foundation Sprint.

About the author

Eli Wood headshot
Eli Wood

CEO, Black Flag Design

Eli Wood leads Black Flag Design, a creative technology company focused on shipping ambitious digital products, AI systems, and design-forward software with a direct point of view on how technology changes work.

Black Flag Journal

Field notes from building AI-native software in the real world.

Get the Journal

Related stories

More from the journal

Pen-and-ink sketch of a small clockwork robot working at a tool-covered workbench late at night while a human sleeps peacefully on a couch in the background, a wall clock reading 2:00 above
ai April 24, 2026 13 min read

The Agent Stays Up Late, Not Me

Every senior engineer knows the right way to set up a codebase. None of them do it. Here’s the four-stage framework we use — The Ratchet — to take a vibe-coded project all the way to a thing you’d trust in production, and the punchline about why this only just became worth doing.

Most teams have always known they should be running tests, type-checking, security audits, accessibility checks, dead-code analysis, prose linting, and a coverage floor. Most teams run two of those. Here’s why that math has finally inverted, and the four-stage framework we use to ratchet a vibe-coded project to a hardened one.

Keith Pattison, Founder of Black Flag Design

Keith Pattison

AI Builder @ Black Flag Design

Read
What a Year of Claude Code Trails Tells You About Your Team editorial image
claude code April 20, 2026 5 min read

What a Year of Claude Code Trails Tells You About Your Team

Claude Code leaves evidence — sessions, commits, PRs, review notes. Read it like a logbook and you'll find what devs actually need to know before they go deeper.

After a year of shipping with Claude Code across real client work, the signal isn't in any single session — it's in the trails. Here's what those trails told us about where Claude Code shines, where it drifts, and the habits devs should build before they lean in harder.

Eli Wood headshot

Eli Wood

CEO, Black Flag Design

Read
The Black Flag Playbook: Six Principles for Shipping with AI editorial image
playbook April 20, 2026 6 min read

The Black Flag Playbook: Six Principles for Shipping with AI

Battle-tested principles for teams building real software with AI-generated code. Human judgment, tight scope, and weekly evidence — the disciplines that keep AI-built systems reliable.

The six rules we use to ship production software with AI. Small scope, weekly demos, human-led oversight, and continuous improvement — drawn from six months of real client engagements.

Keith Pattison, Founder of Black Flag Design

Keith Pattison

AI Builder @ Black Flag Design

Read
Get the Journal