DocxBench is a standalone local benchmark for evidence about reading, writing, and editing an existing Word document (.docx). It keeps invocation and scoring separate, records semantic and preservation results for every attempted run, and retains run-relative raw artifacts. It mirrors PPTBench's architecture and run.json schema version 2 (same top-level keys and row shapes), so one importer pattern covers both.
DocxBench measures one thing: given one frozen synthetic document, does an adapter's output still carry each declared WordprocessingML feature (round trip), and does it apply four exact targeted edits without disturbing anything else (mutation)?
It does not claim:
- broad Word fidelity, rendering fidelity, or layout equivalence (nothing is rendered or laid out);
- that a pass on one small synthetic fixture generalizes to real-world documents;
- that byte equality matters (byte comparison is a recorded, unscored observation);
- library-only speed or memory (timings cover the whole child process, including interpreter, JVM, .NET,
soffice, or container start).
The fixture is project-authored and synthetic, not Office-authored or enterprise ground truth (not_ground_truth: true).
Requires Python 3.11+ and uv.
uv sync --extra adapters --extra dev
uv run pytest
uv run docxbench benchmark --output /tmp/docxbench-run --adapter python-docx --iterations 1
uv run docxbench report /tmp/docxbench-run/run.json--output must name a new or empty non-symlink directory; DocxBench never overwrites an existing run. Without --adapter, every scored adapter runs. --lane selects a subset of lanes; --iterations, --warmups, --timeout (per child, seconds), and --overall-timeout (whole run) control repetition and budgets.
docxbench report <run.json> prints one line per result row (copied from run.json, never rescored), and writes an HTML dashboard and an SVG outcome heatmap next to the run.
The wheel ships one fixture as an immutable package resource, src/docxbench/data/fixtures/mixed-doc.docx. It is built by docxbench.generator, a stdlib-only OOXML writer: zipfile plus hand-built XML, members in sorted order, stored entries with fixed 1980-01-01 timestamps and mode 0644, and pinned core-property timestamps. The bytes depend only on that module, so regeneration is byte-reproducible.
mixed-doc contains:
- heading styles 1 to 3 and body paragraphs with run formatting (bold, italic, underline, color, font size);
- a table captioned
DocxBenchTablewith a horizontal and a vertical merge and cell shading; - a three-level numbered list and a two-level bulleted list (
numbering.xml); - an inline PNG image;
- a header and footer, the footer with a
PAGEfield; - a footnote and a comment;
- a tracked insertion and a tracked deletion;
- a section break into a landscape section;
- a TOC field;
- an internal hyperlink to the bookmark
DocxBenchTargetand an external hyperlink; - a plain-text content control (
w:sdt); - custom document properties (
lpwstr,bool,i4); - an opaque custom XML part (
customXml/item1.xmlwith its properties part) that must be preserved.
Every run verifies the fixture's SHA-256 against the frozen manifest (src/docxbench/data/manifest.json) and copies the exact packaged bytes under the run root. Runs never regenerate fixtures. Maintainers making an intentional corpus change run uv run python -c 'from docxbench.fixtures import regenerate_frozen_data; regenerate_frozen_data()' and commit the fixture and manifest together.
A read/write round trip with no edits. A package-readable precondition, then one declared semantic check per authored feature:
heading-styles, run-formatting, table-name-and-text, table-merged-cells, table-cell-shading, numbered-list-multilevel, bulleted-list, inline-image, header-footer-content, page-number-field, footnotes, comments, tracked-insertion, tracked-deletion, section-break-landscape, toc-field, internal-hyperlink-and-bookmark, external-hyperlink, content-control, custom-document-properties, opaque-custom-xml-part.
Feature extractors resolve relationships, styles, and numbering by meaning rather than by part name or relationship id. For example, list checks compare each item's rendered label (2., a., a bullet glyph) and level; header and footer checks compare what renders on first, odd, and even pages of each section; custom properties compare name, type, and value. An equivalent re-serialization by a different engine passes, while lost or changed content fails.
Four exact targeted edits, located structurally:
| Check | Locator | Old text | New text |
|---|---|---|---|
bookmark-paragraph-updated |
paragraph holding bookmark DocxBenchTarget |
Bookmarked paragraph original text. |
UPDATED-BOOKMARK-PARAGRAPH |
table-cell-1-1-updated |
grid cell (1,1) of DocxBenchTable |
R1C1 |
UPDATED-CELL-A |
table-cell-3-2-updated |
grid cell (3,2) of DocxBenchTable |
R3C2 |
UPDATED-CELL-B |
numbered-list-item-2-updated |
second item of the first numbered list | Numbered item two |
UPDATED-LIST-ITEM |
only-declared-text-nodes-changedcompares the whole-package semantic model of the expected result (input plus exactly those four text replacements) against the output, so edits to the wrong paragraph and unrelated collateral fail.- Every
feature-matrixcheck is re-run as a-preservedcheck against the expected edited document. - A byte-only observation compares raw bytes of every untouched part.
Signatures are namespace-resolved and attribute-order-insensitive. The equivalences below are applied because the schema or packaging model defines them, uniformly for every adapter:
- Part names, relationship ids, ZIP member order, and XML prefixes carry no meaning. The whole-package model walks the relationship graph from the package root and identifies each reached part by its content type, canonical content, and outgoing relationships (type plus target), so a renamed part, a renumbered
r:id, or media stored under a new name compares equal. Prefixed values (xsi:type, markup-compatibility prefix lists) are resolved to namespace URIs. Parts no relationship reaches state nothing and are not modeled. Theopaque-custom-xml-partcheck finds the opaque part by its content, whatever it is named. - Internal identifiers that only pair references with definitions (numbering
numId/abstractNumId; comment, note, bookmark, and revisionw:id; drawing-object and paragraph ids) compare by order of first use, not by value. - Save-time metadata is dropped:
dcterms:modified,cp:lastModifiedBy,cp:revision, the extended-properties application identity, flags, and statistics,w:rsidsandw:rsid*attributes,w14:textId, listw:nsid/w:tmpl, and custom-propertypid. Children of core, extended, and custom property containers are unordered, and an empty property element states nothing; a property part with nothing left equals its absence. xml:spaceon a text node without leading or trailing whitespace is inert; attributes typedxsd:booleanorST_OnOffcompare by value (1andtrueare the same); and an attribute written with its schema default (ST_OnOfftrueon a toggle element,falseelsewhere,w:vMergecontinue,w:pgSzportrait) equals its absence.
Other numeric and string values are never normalized.
Every check declares category (semantic, feature, preservation, or byte-only) and scored. Byte-only observations (raw-untouched-part-equality, scored: false) are information only; a byte difference alone is never called corruption. Scoring first re-verifies that the input is the exact frozen artifact (fixture-source-frozen). A lane succeeds only when every measured iteration passes every scored check.
Result outcomes are success, failure, unsupported (the library has no public API for the lane; the row carries the reason and is never run), unavailable (the adapter or its runtime is not installed; the row carries recovery text), and timeout.
Every attempt runs in a new child process. run.json adapters[] records each adapter's library, exact version, and runtime identity.
python-docx1.2.0: public API in a child Python process.docxtpl0.20.2:DocxTemplate.render({})thensavefor the round trip.template-mutationisunsupported: docxtpl replaces only Jinja tags authored into a template and has no API that addresses existing literal content by bookmark, cell, or list position.wolfdocx-wheel: the installedwolfdocxdistribution, driven through its documentedwolfdocx author report-targetsandwolfdocx author reportCLI with discovered target bindings.feature-matrixisunsupported: wolfdocx has no open/save round-trip operation (its authoring operations require at least one bound edit). Build provenance is recorded with--wolfdocx-artifact,--wolfdocx-source-revision, and--wolfdocx-build-mode.apache-poi: Apache POI XWPF (org.apache.poi:poi-ooxml5.5.1) on an Eclipse Temurin 21 JRE, helper inadapters/apache-poi/, imagedocxbench/apache-poi:5.5.1.docx4j:org.docx4j:docx4j-JAXB-ReferenceImpl17.2.1 on an Eclipse Temurin 21 JRE, helper inadapters/docx4j/, imagedocxbench/docx4j:17.2.1.open-xml-sdk:DocumentFormat.OpenXml3.5.1 on the .NET 9 runtime, helper inadapters/open-xml-sdk/, imagedocxbench/open-xml-sdk:3.5.1.libreoffice: local headless LibreOffice Writer.adapters/libreoffice/docxbench_uno.pyruns insidesofficeas a user-profile Python-UNO script in a fresh, attempt-owned profile. It loads through the Word import filter, edits through the UNO text API, and stores with theMS Word 2007 XMLexport filter.pandoc:pandoc --from=docx --to=docx --track-changes=all --reference-doc=<input>, a round trip through pandoc's document AST.template-mutationapplies the edits with the Lua filteradapters/pandoc/docxbench-mutation.lua, which rewrites the inline text of the bookmarked paragraph, the two table cells, and the list item in the AST. pandoc then writes a new package, so anything its AST does not model is lost; the preservation checks measure that.
The Java and .NET helpers read the input document on stdin and write the output package on stdout. Each attempt runs docker run --rm -i --network none --pull never <image> <lane>, so no host directory is mounted. Build the images once per Docker engine:
docker build -t docxbench/apache-poi:5.5.1 adapters/apache-poi
docker build -t docxbench/docx4j:17.2.1 adapters/docx4j
docker build -t docxbench/open-xml-sdk:3.5.1 adapters/open-xml-sdkContainer adapters use Docker's current context, or the context named by DOCXBENCH_DOCKER_CONTEXT (for example a remote Linux engine reached over SSH). Their identity records the library and version labels, the runtime label, the engine platform, and the image ID. Helper programs are found under DOCXBENCH_ADAPTER_HOME, or the checkout's adapters/ directory when DocxBench runs from a source checkout. A missing helper, runtime, or image produces unavailable rows with recovery text.
Timing comparability: elapsed time covers the whole child. For the container adapters it includes Docker CLI and container start, transport to the engine, and JVM or .NET start; their peak RSS measures only the local Docker client. For libreoffice it includes soffice start and profile creation. Do not read these numbers as library-only processing speed or memory.
Listed, never scored (neither can open an existing .docx and write one back):
docx(npm): generation-only.mammoth: converts.docxto HTML or Markdown only.
uv sync --extra adapters
uv pip install <wolfdocx wheel>
docker --context <engine> build -t docxbench/apache-poi:5.5.1 adapters/apache-poi
docker --context <engine> build -t docxbench/docx4j:17.2.1 adapters/docx4j
docker --context <engine> build -t docxbench/open-xml-sdk:3.5.1 adapters/open-xml-sdk
DOCXBENCH_DOCKER_CONTEXT=<engine> uv run docxbench benchmark \
--output evidence/<date>/run --iterations 3 \
--wolfdocx-artifact <wolfdocx wheel> \
--wolfdocx-source-revision <commit> --wolfdocx-build-mode source-build-wheel
uv run docxbench report evidence/<date>/run/run.json- Raw run report: evidence/2026-10-02/run-fair/run.json
- Report dashboard: evidence/2026-10-02/run-fair/report/index.html
- Outcome heatmap: evidence/2026-10-02/run-fair/report/heatmap.svg
run-fair was scored by the name-independent whole-package model described above. The earlier run in evidence/2026-10-02/run/, scored by a whole-package comparison keyed by part name and relationship id, is retained unchanged; every row's outcome and failed-check set is identical in both runs.
| Adapter | feature-matrix |
template-mutation |
|---|---|---|
python-docx 1.2.0 |
success | success |
docxtpl 0.20.2 |
success | unsupported |
wolfdocx-wheel 0.1.0 |
unsupported | failure |
apache-poi 5.5.1 |
success | success |
docx4j 17.2.1 |
success | success |
libreoffice 26.8.0.3 |
failure | failure |
pandoc 3.11 |
failure | failure |
open-xml-sdk 3.5.1 |
success | success |
Failed scored checks, as recorded in run.json:
wolfdocx-wheeltemplate-mutation:wolfdocx author reportrefused the request withtext_structure_unsupported, so no output was written. Its text edits accept only paragraphs whose direct children are runs, and the bookmarked target paragraph also holdsw:bookmarkStart/w:bookmarkEnd. In a separate manual probe that is not part of the retained run, the same request without the bookmark edit (the two cells and the list item) was published; the lane requires all four edits.libreoffice:table-name-and-text(the table'sw:tblCaptionname andw:tblDescriptionare dropped; cell text survives) andcustom-document-properties(thei4propertyDocxBenchRevisionis written back asr8); intemplate-mutation, the same two as-preservedchecks plusonly-declared-text-nodes-changed. All four edits were applied exactly.pandoc: infeature-matrix,run-formatting,table-name-and-text,table-cell-shading,bulleted-list,header-footer-content,page-number-field,comments,section-break-landscape,toc-field,content-control,custom-document-properties, andopaque-custom-xml-part; intemplate-mutation, the same features as-preservedchecks plusonly-declared-text-nodes-changed. All four edits were applied exactly.
The wolfdocx candidate was a wheel built out-of-tree from source revision 788278b0e947a10fcb400247b777a3e71b3d36a8 (receipt source-build-wheel). Container adapters ran on a remote linux/amd64 Docker engine; the Python adapters (CPython 3.12.12), LibreOffice, and pandoc ran locally on macOS arm64. Each row ran three measured iterations, no warmups, with a 120 second child timeout. Timings include process and runtime start and were taken with OS cache state uncontrolled. Rows that are unsupported were not run.
MIT; see LICENSE. See PROVENANCE.md for dependency licenses.