Ship small interactive tools on a website. A tool is one function and one JSON file. Toolbench builds the form, runs it, and renders the result.
Try it. The live bench, no clone required.
// tools/reverse/index.ts
import type { Tool } from "@toolbench/sdk";
export default {
run({ text }) {
return { kind: "fields", fields: [{ label: "reversed", value: [...text].reverse().join("") }] };
},
} satisfies Tool<{ text: string }>;<tool-host tool="reverse" mode="card"></tool-host>The tool has no idea it is on the web. The website needs no framework.
- Live demo
- What it looks like
- Why
- Install
- Quick start
- Display modes
- How a tool runs
- Result shapes
- Worker mode
- Theming
- Size and cost
- Browser and runtime support
- Non-goals
- Repository layout
- Documentation
The same bench this repository builds is hosted at eknowledger.github.io/toolbench. Four pages, because the display modes are four different situations and one page would let a bug in one hide behind another:
| Page | What it shows |
|---|---|
| Gallery | The example tools as cards, each loading nothing until it is clicked. Below them, a highlighted code result and the two lifecycle states, deprecated and retired. |
| Full page | One tool on its own route, with its help text and its sample inputs. |
| In an article | Embed mode: two tools inside prose, inheriting the page's type. |
| When a tool fails | A tool that spins forever, throws, or ignores cancellation, and a tool attribute naming an id that does not exist. |
The badly behaved fixtures are on their own page, and the card grid filters them out along
with anything not live. A first visitor to the demo should not meet a tool whose job is to
time out, or be offered a retired one as though it were current.
That is a GitHub project Pages site, so Vite base is /toolbench/. Local
pnpm bench and pnpm bench:build keep base at /, which is why
http://localhost:5180 still works. Set BENCH_BASE=/toolbench/ to preview the
deployed layout (the preview is then at http://localhost:4173/toolbench/).
A maintainer enables the site once: Settings → Pages → Source → GitHub Actions.
The workflow is .github/workflows/pages.yml. Until
that toggle is flipped, the try-it URL 404s. Flipping it is the remaining click.
Every screenshot below is the test bench in bench/, which is the same runtime a host installs. Nothing
is mocked up.
On the left the card is static markup: a result computed at build time by seed(), with not one byte
of the tool's code downloaded. On the right, the same card after one click.
![]() |
![]() |
| Closed: markup only, 0 bytes of tool code | Opened: one chunk fetched, form generated from the manifest |
samples in the manifest becomes a row under the form. A click fills the inputs and stops there, because
Run is the trigger and a sample is an input change like any other: the answer already on screen dims to
say it belongs to the inputs that produced it, rather than being quietly replaced. Below, one click has
loaded ten measurements and switched the definition, and the figures underneath are still the previous
run's.
The one worth shipping is the example whose input the tool rejects. It is the one a reader will never type, because nobody sits down to invent a broken input, and it is how a tool shows its failure path on purpose rather than at the least convenient moment.
![]() |
| One click fills the form. The stale result stays put, so the old answer is still there to compare against |
A tool returns one of a closed set of shapes and the runtime draws it. Below is bytes, for wire
formats: offsets, hex, a printable gutter, and named ranges. The highlight on the four-byte character
spans the row break, which is what a table cannot do and why the kind exists.
The same tool in both themes. A tool inherits the page it is on; theming is a dozen CSS custom properties and nothing else.
![]() |
![]() |
A chart is an image, and a screen reader gets nothing from an image. So every series result renders the
same data as a table inside a <details>, automatically. The tool author does nothing to get this.
![]() |
![]() |
Interactive widgets on a website are usually one-offs: a script per widget, wired into one page by hand. The second widget repeats the form handling, the error states, the loading behaviour and the styling. By the fourth they have drifted apart, and none of them has tests.
Toolbench turns the shape of a tool into a contract. In exchange:
- Fixtures are the test suite. A tool ships known-answer cases that run in Node, with no browser and no build step. A broken tool fails your build instead of someone's afternoon.
- One interface for every tool, accessible by default: real labels, errors tied to the input that caused them, a status region, keyboard operation, reduced-motion support.
- Nothing loads until it is needed. A card downloads no tool code until someone opens it.
- Old tools keep working. The contract is versioned, changes are additive only, and compatibility is enforced by tests rather than intent. See docs/versioning.md.
pnpm add @toolbench/sdk @toolbench/runtime@toolbench/sdk is what tool authors import. It has no dependencies and does not touch the DOM.
@toolbench/runtime is what the website imports. It has no dependencies either, and no framework.
Three steps. The result is a working tool page.
1. Write the tool. Four files in a directory:
tools/reverse/
tool.json identity, inputs, result shapes
index.ts the function
cases.json known-answer fixtures
README.md help text
sdk is 1 on purpose, because a v1 manifest still being valid is the compatibility promise made
concrete: the current contract is 3, and declaring 1 opts out of everything added since (bytes in v2,
samples in v3). See docs/versioning.md.
2. Register it. Once per site:
import { defineToolHost, RegistrySource } from "@toolbench/runtime";
import manifest from "./tools/reverse/tool.json";
defineToolHost({
source: new RegistrySource({
reverse: { manifest, load: () => import("./tools/reverse/index.ts") },
}),
pageUrl: (id) => `/tools/${id}/`,
// Optional. Return a Node, never a string: the runtime will not assign innerHTML.
// highlight: (source, lang) => yourHighlighter(source, lang),
});With a bundler that supports directory globs, the source is a loop over the glob instead of a literal
map. bench/src/registry.ts is the version this repository uses, and it
picks up new tools with no code change.
3. Put it on a page.
<tool-host tool="reverse" mode="page"></tool-host>4. Test it. cases.json holds inputs and expected results:
[
{
"name": "reverses a string",
"input": { "text": "abc" },
"expect": { "kind": "fields", "fields": [{ "label": "reversed", "value": "cba" }] }
}
]node --test tools/cases.test.tsOn your own site that suite is three lines, since the directory walk is exported:
import { describe, it } from "node:test";
import { checkToolDirectory } from "@toolbench/sdk/fixtures";
checkToolDirectory(new URL("../src/tools/", import.meta.url), { describe, it });It validates every manifest, checks each tool's id matches its directory, refuses an empty
cases.json, verifies no case expects a kind the manifest failed to declare, runs every case, and runs
every sample the manifest declares. Wire it into CI and a change that breaks a tool fails your build
instead of someone's afternoon.
Full walkthrough: docs/authoring-a-tool.md.
One declaration drives three presentations. The mode also decides when the tool's code is fetched, which is the difference between a page that stays fast and one that does not.
mode |
Shows | Fetches the tool's code |
|---|---|---|
card |
Name, blurb, the inputs marked primary, the first result part |
When the reader opens it. Until then the card is static markup. |
page |
Every input, the full result, the tool's links | When it scrolls within two viewports |
embed |
Same as page without the title, sized for the middle of an article |
When it scrolls within two viewports |
A card shows one result part by default, which is too few for a tool whose answer is a number and a
curve: the number alone hides how steep the curve is, and the curve alone is a set of unnamed lines. Give
the card more room with parts:
<tool-host tool="queue-explorer" mode="card" parts="2"></tool-host>An attribute rather than a manifest key, deliberately. How much room a card has is a property of the page it sits on, not of the tool: the same tool is a one-part card in a sidebar and a two-part card leading a section, and a manifest cannot know which. Whatever does not fit is counted, not hidden, and a compact chart keeps its legend so its lines stay identifiable.
parts works in any mode, not only on a card. An embedded tool with no parts still shows everything,
because that is what it has always done and a cap nobody asked for would truncate existing pages on
upgrade.
In prose, "the first part, then a line pointing somewhere else" is the wrong offer: the reader is already
where they want to be. more="expand" turns that line into a button that reveals the rest in place.
<tool-host tool="percentiles" mode="embed" parts="1" more="expand"></tool-host>more |
The notice under a truncated result |
|---|---|
link (default) |
+2 more results on the full tool, as plain text. Right for a card, whose job is to send the reader to the page |
expand |
Show 2 more results, as a button. Pressing it reveals them here and changes the label to Hide 2 results |
Expanding does not run the tool again: the result is redrawn from the answer already in hand, so it
costs nothing even for a worker-mode tool, where a re-run would mean a round trip. It is a real <button> with aria-expanded, keyboard
operable, keeping focus across the toggle, and it appears only when something is actually hidden. Whether
it is open survives a re-run, because it is the reader saying what they want to see rather than a property
of one answer.
Opening it reveals the whole result, not only the extra parts. Inside a card, where fields are capped too, the disclosure is the single notice: the reader is not told about "2 more fields" they have no control over, and opening it shows those as well.
A tool used as the opening figure of a page about something else should lead with its answer, with the controls underneath for a reader who wants to check it:
<tool-host tool="percentiles" mode="embed" layout="answer-first"></tool-host>Or the page can own the controls entirely. controls="none" draws the result and nothing else (no form,
no Run, no status line) and the page drives it with the values setter and run():
<tool-host tool="percentiles" mode="embed" controls="none" id="figure"></tool-host>
<script type="module">
const figure = document.querySelector("#figure");
mySlider.addEventListener("input", () => { figure.values = { values: mySlider.value }; figure.run(); });
</script>With controls="none" the host does not mark a result stale when values change either: whoever owns the
control that changed it owns saying so. Neither attribute changes a card.
A card is a facade: markup with no behaviour and no download. If your site renders HTML ahead of time, give the card a seed: the tool's result for its default inputs, computed during your build. The card then shows a real result with no JavaScript at all.
Compute it with seed() rather than writing it by hand. A hand-written seed is correct on the day it is
typed and drifts silently afterwards, showing a confidently wrong answer to every reader who does not
press Run. That is not hypothetical: the seed in this repository's own bench was hardcoded and had drifted
by the time seed() replaced it.
// In your build: a Vite plugin, an Astro integration, a script that writes HTML.
import { seed, serialiseSeed } from "@toolbench/sdk";
const output = await seed({ manifest, tool });
// A card renders only the first part of a group, so there is no point shipping the rest.
const forCard = output.kind === "group" ? output.parts[0] : output;
html = `<tool-host tool="percentiles" mode="card">
<script type="application/json" data-toolbench-seed>${serialiseSeed(forCard)}</script>
</tool-host>`;serialiseSeed, not JSON.stringify. An HTML parser ends a <script> at the first </script in
its text however the JSON is quoted, so a tool that echoes any part of its input could otherwise truncate
your page. bench/vite.config.ts is a working 30-line version of the above.
seed() throws if a tool crashes on its own defaults or rejects them, because both are authoring bugs
and a card seeded with an error is worse than an unseeded one. Wrap it if you would rather have a missing
seed than a failed build.
The reader decides. Nothing runs on its own.
- Opening a card, or scrolling a tool into view, loads its code and shows the form. It does not run the tool.
- Run runs it. So does Enter in a single-line input, or Ctrl/Cmd + Enter in a textarea.
- Changing an input marks the current result stale: it dims, and the status line says so. The previous result stays on screen because it is still the last true one, and keeping it is what lets you compare before and after.
- A tool that is genuinely instant can set
"autoRun": trueand update as the reader types. It is off by default, and the validator rejects it for worker-mode tools.
The progress bar appears only if a run is still going after 400 ms. A bar that flashes for 20 ms draws the eye to report that nothing happened.
A tool can ship its own examples. samples is a list of labelled inputs, which the runtime draws as
a row of buttons under the form:
"samples": [
{ "label": "Definitions disagree", "input": { "values": "3 4 4 5 6 7 9 14 40 260", "method": "linear" } },
{ "label": "Not a number", "input": { "values": "12, 14, 15, 18ms, 21, 24, 31, 44" } }
]Clicking one fills the form, which is an input change like any other: the result goes stale and waits for
Run, or re-runs if the tool set autoRun. Each input is partial, so a sample changes the one thing it is
about and leaves the rest as the reader left it. The sample most worth shipping is the one the tool
rejects, because nobody types a broken input by hand, and it is the only way a reader sees the failure path
before meeting it for real. Cards show no samples: there is room there for one input and a Run button.
Added in contract version 3, and
docs/authoring-a-tool.md covers choosing them.
A host can prefill the form without putting those examples in the tool, and without reaching into
the shadow root. values fills inputs the way typing does. It does not run the tool. run() is
opt-in:
const host = document.querySelector("tool-host");
host.values = { packet: "80 e0 12 34 …", view: "all" }; // partial; does not run
host.run(); // optionalUnspecified inputs keep their current value. Out-of-range numbers are clamped, an invalid select
falls back to that input's default, and a string longer than maxLength is cut: a host is not more
trustworthy than a reader. An existing result goes stale, the same as typing. This works before the
form has opened, so a card can be prefilled. A shareable ?in=… deep link is then a host feature,
with no further contract change.
A tool returns one of a closed set of shapes, so the runtime can draw anything a tool produces and a tool cannot invent something nobody can render.
| Kind | Use for |
|---|---|
fields |
Label and value pairs, optionally grouped under headings |
table |
Columns and rows, with alignment and per-cell emphasis |
series |
A chart with axes, legend and annotations. Ships the same numbers as a table for readers who cannot see it |
text |
Prose or preformatted output |
code |
Source, with a language tag. The runtime renders preformatted text. A host that already has a highlighter passes highlight?: (source, lang) => Node to defineToolHost |
bytes |
Raw bytes as a reader of a wire format wants them: offsets, hex, a printable gutter, and named ranges that can wrap a row |
group |
Several of the above in one result. A decoder that returns fields and a table is the common case |
error |
The input was wrong. Naming the input marks that control invalid and attaches the message to it |
Returning error means the input was bad and the tool worked. Throwing means the tool has a bug. The
two render differently on purpose.
A tool whose running time depends on its input sets "thread": "worker". That is the only mode where a
timeout can stop a run, because on the main thread there is nothing to terminate.
The host then needs one file, and it has to live in the host project because only the host's bundler can resolve the host's tools:
// tool.worker.ts
import { createToolWorker } from "@toolbench/runtime/worker";
import { registry } from "./registry.ts";
createToolWorker({ load: (id) => registry[id].load() });defineToolHost({
source,
workerFactory: () => new Worker(new URL("./tool.worker.ts", import.meta.url), { type: "module" }),
});Two things to know before you spend an afternoon on them:
- Vite defaults
worker.formatto"iife", which cannot code-split. A worker that imports tools builds in development and fails the production build. Setworker: { format: "es" }. - If no
workerFactoryis configured, a worker-mode tool runs on the main thread and logs a warning. A slow tool is better than a missing one.
The runtime renders inside a shadow root. Your CSS cannot reach in and break a tool, and the tool's CSS cannot leak out and break your page. Theming is therefore a list of custom properties, and that list is the whole styling API:
tool-host {
--tb-accent: #0b6b5f;
--tb-bg: #ffffff;
--tb-surface: #f4fbf9;
--tb-border: #b9dbd4;
--tb-radius: 2px;
--tb-font: Georgia, serif;
--tb-mono: "SF Mono", monospace;
}--tb-accent is a fill colour. Accent-coloured text (links, the card's hint, annotation labels) reads
--tb-accent-text, and text on an accent fill (the Run button's label) reads --tb-accent-ink. They
default to the accent and to the surface colour. Set them when the accent is bright, since a bright
amber that works as a button cannot be small text on white, and white text on it is unreadable:
tool-host { --tb-accent: #f59e0b; --tb-accent-text: #a46805; --tb-accent-ink: #000; }Defaults use light-dark(), so a tool follows the page's colour scheme before anyone configures
anything. The full list is at the top of
packages/runtime/src/styles.ts.
Measured on the built bench with gzip, not estimated:
| Transfer | |
|---|---|
| Runtime plus the bench's page wiring, once per page that uses a tool | 23.4 KB |
| Worker entry, only for pages with a worker-mode tool | 4.1 KB |
percentiles tool chunk |
1.3 KB |
queue-explorer tool chunk |
1.2 KB |
regex-explainer tool chunk |
2.8 KB |
histogram tool chunk |
0.8 KB |
| A page with no tool on it | 0 bytes |
| A card nobody opens | 0 bytes of tool code |
Reproduce with pnpm bench:build && node scripts/size-check.mjs. CI runs the same check against a
budget, so these numbers cannot rot.
In the browser the runtime needs custom elements, shadow DOM, IntersectionObserver, module workers
and CSS container queries: Chrome 105+, Firefox 114+, Safari 16.4+.
CI runs the bench suite against Playwright's Chromium, Firefox and WebKit, on every pull request. That
turns "it works in three engines" from a claim into something checked. Note what it does not establish: those
are the engines Playwright ships today, not the version floors above, so the floors remain the oldest engines
the code is written against rather than the oldest ones tested. Locally, after pnpm bench:build:
pnpm exec playwright install firefox # or webkit, or chromium
PLAYWRIGHT_BROWSER=firefox pnpm test:benchUnset, pnpm test:bench still opens system Chrome (CHROME_CHANNEL, default chrome). CI sets
CHROME_CHANNEL=chromium so the runner uses the browser Playwright just installed.
No runtime gap turned up on those three engines. One test-harness difference did: Playwright's page
request listener reports a worker's module import() as script on Chromium, xhr on WebKit, and
not at all on Firefox. The suite reads the worker's own performance timeline for that allowlist, which
is the observation that holds on every engine.
Two newer features are used with plain fallbacks in front of them, so a browser without either gets
light colours rather than a broken layout: light-dark() for automatic dark mode (Chrome 123+,
Firefox 120+, Safari 17.5+) and color-mix() for one focus ring. Anything older than the baseline
above shows the static facade and its seeded result, which is one reason seeding is worth doing.
Development and tests: Node 24 or newer. Toolbench runs TypeScript directly, which is how fixtures execute with no build step. Any bundler for the host site; the examples use Vite 7.
- Not a code playground. Readers do not write code here. That needs a compiler in the browser, which is megabytes.
- Not a notebook. No dataflow between tools, no execution order.
- Not a sandbox. A tool is code you wrote and bundled. Worker mode is a stability boundary that lets a slow run be stopped; it is not a security boundary, and the docs never pretend otherwise. Code you did not write needs a different design.
- The contract is pure functions only. No file input, no network access, at contract version 4. Both are planned as additive versions, which is the case the versioning policy exists to handle.
| Path | What it is |
|---|---|
packages/sdk |
The contract: types, manifest validation, version migration, fixture runner. No dependencies, no DOM. |
packages/runtime |
<tool-host>: the element, the form, the renderers, the runner, the worker protocol. |
tools/ |
Three example tools (percentiles, queue-explorer, utf8-bytes), with fixtures and help. |
bench/ |
The test bench, also the live demo: three display modes, both threading modes, host theming, and a tool that fails on purpose (on its own page). |
docs/ |
Architecture, authoring, versioning. |
pnpm install
pnpm check # typecheck, unit tests, every tool's fixtures
pnpm new-tool <id> # scaffold tools/<id>/ (four files, already green)
pnpm bench # the bench on http://localhost:5180
pnpm test:bench # the same bench, in Chromium, Firefox, or WebKit
pnpm coverage # Node suite coverage: uncovered list, no threshold
# live: https://eknowledger.github.io/toolbench/| Document | For |
|---|---|
| docs/architecture.md | How it works, where the boundaries are, how to change it, how to contribute |
| docs/authoring-a-tool.md | Writing, testing and shipping a tool |
| docs/versioning.md | The compatibility policy and how to raise the contract version |
| docs/releasing.md | How versions are decided and how a release reaches npm |
| CONTRIBUTING.md | Setup, the loop, and which opinions are load bearing |
| SECURITY.md | The threat model, and what is in scope to report |
| ROADMAP.md | What is planned, in what order, and why that order |
MIT.






