A clinician speaks a chart note. A speech-to-speech model drafts it into a synthetic dental chart. Nothing commits until the clinician says so. Corrections are logged, so over time the system can show it edits less.
Built in one recorded session. Every patient and every chart entry in this repo is invented.
Four things came out of a conversation with a practicing dentist who is building AI tooling for his own office. This build answers each one.
Voice is one layer, not the whole stack. The model handles listening and speaking. It does
not own the record. It calls three narrow tools (draft_chart_note, correct_chart_note,
commit_chart_note) and the chart lives in ordinary code that the practice controls.
The human is in the loop where the consequence is. A spoken note lands as a draft, in amber.
The model reads it back. The clinician corrects by voice or confirms by voice. Only an explicit
confirmation commits, and the commit tool refuses without it (tests/chart.test.ts). The 90% the
model gets right saves the typing; the 10% the clinician fixes is a logged correction, and the
count of corrections per committed note is on screen so you can watch it fall.
Data stays where the practice can see it. The chart is in-process. The only thing that leaves the machine is the audio stream to the model. In production the same design points at a self-hosted speech model so PHI never leaves infrastructure the practice controls; this demo uses Amazon Nova 2 Sonic on Bedrock because that is where the model runs today, and the README says so rather than pretending otherwise.
A workflow with a place to land. Dictate, draft, confirm, commit. That is the first thing a practice would want, and it is small enough to make trustable before anything bigger.
Chrome only (the official Nova Sonic samples are Chrome-optimized and so is this).
npm install
cp .env.example .env # IAM access-key pair + AWS_REGION=us-east-1
npm test # the gate, offline
npm run smoke # one text turn through the live model, prints latency and usage
npm start # http://localhost:3033, allow the mic, Start session, speakThe page walks you through it in three steps, with the line to say on screen at each step and a "type it instead" button if you have no mic:
- Say the note. "Avery, fourteen MO composite, A2, two percent lido with epi one carp, rubber dam, tolerated well, recall six months."
- Fix one thing. "Make that fifteen."
- Sign it. "Sign it."
The step you are on comes from the server (GET /api/state, stage socket event), derived from
what actually happened in the chart, never from the page. A plumbing strip (mic, model, tool,
gate, chart) lights up as each hop fires, with the model's reply time next to it.
npm run smoke:onboarding # the three lines through the live model, asserts draft -> correct -> commitThe note fields follow what a dental progress note carries (tooth, surfaces in M-O-D order, procedure, shade, anesthetic with concentration and carpule count, isolation, occlusion, tolerance, post-op instructions given, follow-up), from published charting guidance. Not verified by a clinician; a practice would adjust them.
- The Bedrock long-term API key does not work for this model.
InvokeModelWithBidirectionalStreamreturnsAccessDeniedException: This operation does not support API Keys. You need an IAM access-key pair (SigV4). The SDK also prefersAWS_BEARER_TOKEN_BEDROCKwhenever it is set, so the client removes it from the environment when a key pair is present. - Nova 2 Sonic only generates while an audio content is open and receiving frames. A typed turn with no audio flowing gets silence forever. With no mic active the server streams 16 kHz PCM16 silence (1024 samples every 64 ms), the same trick the official text-mode sample uses.
- Measured on the smoke test: stream open in ~300 ms, first reply text at ~1.5 s after the turn, about 1,000 tokens for a full draft-and-read-back turn. At the published rates that is well under a cent per note.
- Synthetic chart, three invented patients, no EHR integration. This proves the loop, not a product.
- One model, English only. Voice is
tiffanyby default (SONIC_VOICEenv; US English has onlymatthewandtiffany). Delivery is shaped by the system prompt more than the voice: Nova 2 over-uses listed phrases, so the prompt uses one example instead. Tooth numbering is universal (1 to 32); procedures are plain words, not billing codes. - The gate is ours; Bedrock Guardrails do not apply to Nova 2 Sonic.
- Turn latency is what the model gives; there is no endpointing tuning here yet.
src/
sonic.ts the bidirectional stream: session, prompt, audio, tool loop, silence
tools.ts the three tools and the system prompt
chart.ts the synthetic chart and the confirm gate
server.ts Express + socket.io bridge to the browser
smoke.ts one live text turn, prints latency and usage
public/ the one-screen client: mic, transcript, chart, corrections
tests/ gate tests, offline
docs/ blueprint
Event protocol adapted from aws-samples/amazon-nova-samples (speech-to-speech, Nova 2 Sonic).
railway.json runs it as one service. Set the IAM pair, AWS_REGION, and DEMO_KEY as service
variables; the link then needs ?k=<DEMO_KEY> once and a cookie keeps it. /api/health stays open
for the platform check. WebSockets are exempt from Railway's idle timeout, which is why it is there.
MIT. Every patient, note, and name in the chart is invented.
