Skip to content

Latest commit

 

History

1 Commit

Folders and files

Repository files navigation

chart-note

A clinician speaks a chart note. A speech-to-speech model drafts it into a synthetic dental chart. Nothing commits until the clinician says so. Corrections are logged, so over time the system can show it edits less.

How a note moves from voice to record

Built in one recorded session. Every patient and every chart entry in this repo is invented.

Why it is shaped this way

Four things came out of a conversation with a practicing dentist who is building AI tooling for his own office. This build answers each one.

Voice is one layer, not the whole stack. The model handles listening and speaking. It does not own the record. It calls three narrow tools (draft_chart_note, correct_chart_note, commit_chart_note) and the chart lives in ordinary code that the practice controls.

The human is in the loop where the consequence is. A spoken note lands as a draft, in amber. The model reads it back. The clinician corrects by voice or confirms by voice. Only an explicit confirmation commits, and the commit tool refuses without it (tests/chart.test.ts). The 90% the model gets right saves the typing; the 10% the clinician fixes is a logged correction, and the count of corrections per committed note is on screen so you can watch it fall.

Data stays where the practice can see it. The chart is in-process. The only thing that leaves the machine is the audio stream to the model. In production the same design points at a self-hosted speech model so PHI never leaves infrastructure the practice controls; this demo uses Amazon Nova 2 Sonic on Bedrock because that is where the model runs today, and the README says so rather than pretending otherwise.

A workflow with a place to land. Dictate, draft, confirm, commit. That is the first thing a practice would want, and it is small enough to make trustable before anything bigger.

Run it

Chrome only (the official Nova Sonic samples are Chrome-optimized and so is this).

npm install
cp .env.example .env        # IAM access-key pair + AWS_REGION=us-east-1
npm test                    # the gate, offline
npm run smoke               # one text turn through the live model, prints latency and usage
npm start                   # http://localhost:3033, allow the mic, Start session, speak

The page walks you through it in three steps, with the line to say on screen at each step and a "type it instead" button if you have no mic:

  1. Say the note. "Avery, fourteen MO composite, A2, two percent lido with epi one carp, rubber dam, tolerated well, recall six months."
  2. Fix one thing. "Make that fifteen."
  3. Sign it. "Sign it."

The step you are on comes from the server (GET /api/state, stage socket event), derived from what actually happened in the chart, never from the page. A plumbing strip (mic, model, tool, gate, chart) lights up as each hop fires, with the model's reply time next to it.

npm run smoke:onboarding   # the three lines through the live model, asserts draft -> correct -> commit

The note fields follow what a dental progress note carries (tooth, surfaces in M-O-D order, procedure, shade, anesthetic with concentration and carpule count, isolation, occlusion, tolerance, post-op instructions given, follow-up), from published charting guidance. Not verified by a clinician; a practice would adjust them.

What we learned wiring it

  • The Bedrock long-term API key does not work for this model. InvokeModelWithBidirectionalStream returns AccessDeniedException: This operation does not support API Keys. You need an IAM access-key pair (SigV4). The SDK also prefers AWS_BEARER_TOKEN_BEDROCK whenever it is set, so the client removes it from the environment when a key pair is present.
  • Nova 2 Sonic only generates while an audio content is open and receiving frames. A typed turn with no audio flowing gets silence forever. With no mic active the server streams 16 kHz PCM16 silence (1024 samples every 64 ms), the same trick the official text-mode sample uses.
  • Measured on the smoke test: stream open in ~300 ms, first reply text at ~1.5 s after the turn, about 1,000 tokens for a full draft-and-read-back turn. At the published rates that is well under a cent per note.

Honest limits

  • Synthetic chart, three invented patients, no EHR integration. This proves the loop, not a product.
  • One model, English only. Voice is tiffany by default (SONIC_VOICE env; US English has only matthew and tiffany). Delivery is shaped by the system prompt more than the voice: Nova 2 over-uses listed phrases, so the prompt uses one example instead. Tooth numbering is universal (1 to 32); procedures are plain words, not billing codes.
  • The gate is ours; Bedrock Guardrails do not apply to Nova 2 Sonic.
  • Turn latency is what the model gives; there is no endpointing tuning here yet.

Layout

src/
  sonic.ts     the bidirectional stream: session, prompt, audio, tool loop, silence
  tools.ts     the three tools and the system prompt
  chart.ts     the synthetic chart and the confirm gate
  server.ts    Express + socket.io bridge to the browser
  smoke.ts     one live text turn, prints latency and usage
public/        the one-screen client: mic, transcript, chart, corrections
tests/         gate tests, offline
docs/          blueprint

Event protocol adapted from aws-samples/amazon-nova-samples (speech-to-speech, Nova 2 Sonic).

Hosting

railway.json runs it as one service. Set the IAM pair, AWS_REGION, and DEMO_KEY as service variables; the link then needs ?k=<DEMO_KEY> once and a cookie keeps it. /api/health stays open for the platform check. WebSockets are exempt from Railway's idle timeout, which is why it is there.

License

MIT. Every patient, note, and name in the chart is invented.

About

A clinician speaks a chart note; Nova 2 Sonic drafts it into a synthetic chart; nothing commits until the clinician confirms.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages