The full example from Build a voice agent for the browser with JavaScript, plus the relay the guide describes, so you can run it.
The page captures microphone audio, streams it to Dialog-RSN-1 through a same-origin relay, and renders the reply as it arrives. It uses plain browser APIs, with no framework and no build step.
You need Node 18 or later and a Dialog-RSN-1 API key.
npm install
cp .env.example .env # add your DIALOGUE_API_KEY
set -a && . ./.env && set +a
npm start # http://localhost:8787Open the page, click Start talking, and allow the microphone.
| File | What it is |
|---|---|
public/index.html |
The guide's page, unchanged |
public/app.js |
The guide's script. One change: it connects to the relay on the page's own origin instead of your-app.example.com |
public/pcm-worklet.js |
The guide's AudioWorkletProcessor, unchanged |
relay.js |
Serves public/ and relays /realtime-relay to Dialog-RSN-1, adding your key server-side |
A browser's native WebSocket can't set headers, and Dialog-RSN-1 only accepts the key in the
X-API-KEY header. Even if it could, a key in page code is readable by anyone who opens dev tools.
The relay holds the key and the browser never sees it:
Browser -> ws://localhost:8787/realtime-relay (no key, same-origin)
Relay -> wss://api.us.poly.ai/v1/realtime (holds the real key)
If you write your own relay in Node, set a User-Agent on the upstream connection. The API
rejects a handshake without one with a bare 403, and the ws package sends none by default.
Apache 2.0. See LICENSE.