Skip to content
SynthGLPublic

About

Benchmark for libraries that read, write and edit .docx files: feature, template-mutation and preservation checks

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

DocxBench

DocxBench is a standalone local benchmark for evidence about reading, writing, and editing an existing Word document (.docx). It keeps invocation and scoring separate, records semantic and preservation results for every attempted run, and retains run-relative raw artifacts. It mirrors PPTBench's architecture and run.json schema version 2 (same top-level keys and row shapes), so one importer pattern covers both.

What it measures, and what it does not claim

DocxBench measures one thing: given one frozen synthetic document, does an adapter's output still carry each declared WordprocessingML feature (round trip), and does it apply four exact targeted edits without disturbing anything else (mutation)?

It does not claim:

  • broad Word fidelity, rendering fidelity, or layout equivalence (nothing is rendered or laid out);
  • that a pass on one small synthetic fixture generalizes to real-world documents;
  • that byte equality matters (byte comparison is a recorded, unscored observation);
  • library-only speed or memory (timings cover the whole child process, including interpreter, JVM, .NET, soffice, or container start).

The fixture is project-authored and synthetic, not Office-authored or enterprise ground truth (not_ground_truth: true).

Quick start

Requires Python 3.11+ and uv.

uv sync --extra adapters --extra dev
uv run pytest
uv run docxbench benchmark --output /tmp/docxbench-run --adapter python-docx --iterations 1
uv run docxbench report /tmp/docxbench-run/run.json

--output must name a new or empty non-symlink directory; DocxBench never overwrites an existing run. Without --adapter, every scored adapter runs. --lane selects a subset of lanes; --iterations, --warmups, --timeout (per child, seconds), and --overall-timeout (whole run) control repetition and budgets.

docxbench report <run.json> prints one line per result row (copied from run.json, never rescored), and writes an HTML dashboard and an SVG outcome heatmap next to the run.

Fixture

The wheel ships one fixture as an immutable package resource, src/docxbench/data/fixtures/mixed-doc.docx. It is built by docxbench.generator, a stdlib-only OOXML writer: zipfile plus hand-built XML, members in sorted order, stored entries with fixed 1980-01-01 timestamps and mode 0644, and pinned core-property timestamps. The bytes depend only on that module, so regeneration is byte-reproducible.

mixed-doc contains:

  • heading styles 1 to 3 and body paragraphs with run formatting (bold, italic, underline, color, font size);
  • a table captioned DocxBenchTable with a horizontal and a vertical merge and cell shading;
  • a three-level numbered list and a two-level bulleted list (numbering.xml);
  • an inline PNG image;
  • a header and footer, the footer with a PAGE field;
  • a footnote and a comment;
  • a tracked insertion and a tracked deletion;
  • a section break into a landscape section;
  • a TOC field;
  • an internal hyperlink to the bookmark DocxBenchTarget and an external hyperlink;
  • a plain-text content control (w:sdt);
  • custom document properties (lpwstr, bool, i4);
  • an opaque custom XML part (customXml/item1.xml with its properties part) that must be preserved.

Every run verifies the fixture's SHA-256 against the frozen manifest (src/docxbench/data/manifest.json) and copies the exact packaged bytes under the run root. Runs never regenerate fixtures. Maintainers making an intentional corpus change run uv run python -c 'from docxbench.fixtures import regenerate_frozen_data; regenerate_frozen_data()' and commit the fixture and manifest together.

Lanes and scoring

feature-matrix

A read/write round trip with no edits. A package-readable precondition, then one declared semantic check per authored feature:

heading-styles, run-formatting, table-name-and-text, table-merged-cells, table-cell-shading, numbered-list-multilevel, bulleted-list, inline-image, header-footer-content, page-number-field, footnotes, comments, tracked-insertion, tracked-deletion, section-break-landscape, toc-field, internal-hyperlink-and-bookmark, external-hyperlink, content-control, custom-document-properties, opaque-custom-xml-part.

Feature extractors resolve relationships, styles, and numbering by meaning rather than by part name or relationship id. For example, list checks compare each item's rendered label (2., a., a bullet glyph) and level; header and footer checks compare what renders on first, odd, and even pages of each section; custom properties compare name, type, and value. An equivalent re-serialization by a different engine passes, while lost or changed content fails.

template-mutation

Four exact targeted edits, located structurally:

Check Locator Old text New text
bookmark-paragraph-updated paragraph holding bookmark DocxBenchTarget Bookmarked paragraph original text. UPDATED-BOOKMARK-PARAGRAPH
table-cell-1-1-updated grid cell (1,1) of DocxBenchTable R1C1 UPDATED-CELL-A
table-cell-3-2-updated grid cell (3,2) of DocxBenchTable R3C2 UPDATED-CELL-B
numbered-list-item-2-updated second item of the first numbered list Numbered item two UPDATED-LIST-ITEM
  • only-declared-text-nodes-changed compares the whole-package semantic model of the expected result (input plus exactly those four text replacements) against the output, so edits to the wrong paragraph and unrelated collateral fail.
  • Every feature-matrix check is re-run as a -preserved check against the expected edited document.
  • A byte-only observation compares raw bytes of every untouched part.

Semantic signatures

Signatures are namespace-resolved and attribute-order-insensitive. The equivalences below are applied because the schema or packaging model defines them, uniformly for every adapter:

  • Part names, relationship ids, ZIP member order, and XML prefixes carry no meaning. The whole-package model walks the relationship graph from the package root and identifies each reached part by its content type, canonical content, and outgoing relationships (type plus target), so a renamed part, a renumbered r:id, or media stored under a new name compares equal. Prefixed values (xsi:type, markup-compatibility prefix lists) are resolved to namespace URIs. Parts no relationship reaches state nothing and are not modeled. The opaque-custom-xml-part check finds the opaque part by its content, whatever it is named.
  • Internal identifiers that only pair references with definitions (numbering numId/abstractNumId; comment, note, bookmark, and revision w:id; drawing-object and paragraph ids) compare by order of first use, not by value.
  • Save-time metadata is dropped: dcterms:modified, cp:lastModifiedBy, cp:revision, the extended-properties application identity, flags, and statistics, w:rsids and w:rsid* attributes, w14:textId, list w:nsid/w:tmpl, and custom-property pid. Children of core, extended, and custom property containers are unordered, and an empty property element states nothing; a property part with nothing left equals its absence.
  • xml:space on a text node without leading or trailing whitespace is inert; attributes typed xsd:boolean or ST_OnOff compare by value (1 and true are the same); and an attribute written with its schema default (ST_OnOff true on a toggle element, false elsewhere, w:vMerge continue, w:pgSz portrait) equals its absence.

Other numeric and string values are never normalized.

Check outcomes

Every check declares category (semantic, feature, preservation, or byte-only) and scored. Byte-only observations (raw-untouched-part-equality, scored: false) are information only; a byte difference alone is never called corruption. Scoring first re-verifies that the input is the exact frozen artifact (fixture-source-frozen). A lane succeeds only when every measured iteration passes every scored check.

Result outcomes are success, failure, unsupported (the library has no public API for the lane; the row carries the reason and is never run), unavailable (the adapter or its runtime is not installed; the row carries recovery text), and timeout.

Adapters

Every attempt runs in a new child process. run.json adapters[] records each adapter's library, exact version, and runtime identity.

  • python-docx 1.2.0: public API in a child Python process.
  • docxtpl 0.20.2: DocxTemplate.render({}) then save for the round trip. template-mutation is unsupported: docxtpl replaces only Jinja tags authored into a template and has no API that addresses existing literal content by bookmark, cell, or list position.
  • wolfdocx-wheel: the installed wolfdocx distribution, driven through its documented wolfdocx author report-targets and wolfdocx author report CLI with discovered target bindings. feature-matrix is unsupported: wolfdocx has no open/save round-trip operation (its authoring operations require at least one bound edit). Build provenance is recorded with --wolfdocx-artifact, --wolfdocx-source-revision, and --wolfdocx-build-mode.
  • apache-poi: Apache POI XWPF (org.apache.poi:poi-ooxml 5.5.1) on an Eclipse Temurin 21 JRE, helper in adapters/apache-poi/, image docxbench/apache-poi:5.5.1.
  • docx4j: org.docx4j:docx4j-JAXB-ReferenceImpl 17.2.1 on an Eclipse Temurin 21 JRE, helper in adapters/docx4j/, image docxbench/docx4j:17.2.1.
  • open-xml-sdk: DocumentFormat.OpenXml 3.5.1 on the .NET 9 runtime, helper in adapters/open-xml-sdk/, image docxbench/open-xml-sdk:3.5.1.
  • libreoffice: local headless LibreOffice Writer. adapters/libreoffice/docxbench_uno.py runs inside soffice as a user-profile Python-UNO script in a fresh, attempt-owned profile. It loads through the Word import filter, edits through the UNO text API, and stores with the MS Word 2007 XML export filter.
  • pandoc: pandoc --from=docx --to=docx --track-changes=all --reference-doc=<input>, a round trip through pandoc's document AST. template-mutation applies the edits with the Lua filter adapters/pandoc/docxbench-mutation.lua, which rewrites the inline text of the bookmarked paragraph, the two table cells, and the list item in the AST. pandoc then writes a new package, so anything its AST does not model is lost; the preservation checks measure that.

The Java and .NET helpers read the input document on stdin and write the output package on stdout. Each attempt runs docker run --rm -i --network none --pull never <image> <lane>, so no host directory is mounted. Build the images once per Docker engine:

docker build -t docxbench/apache-poi:5.5.1 adapters/apache-poi
docker build -t docxbench/docx4j:17.2.1 adapters/docx4j
docker build -t docxbench/open-xml-sdk:3.5.1 adapters/open-xml-sdk

Container adapters use Docker's current context, or the context named by DOCXBENCH_DOCKER_CONTEXT (for example a remote Linux engine reached over SSH). Their identity records the library and version labels, the runtime label, the engine platform, and the image ID. Helper programs are found under DOCXBENCH_ADAPTER_HOME, or the checkout's adapters/ directory when DocxBench runs from a source checkout. A missing helper, runtime, or image produces unavailable rows with recovery text.

Timing comparability: elapsed time covers the whole child. For the container adapters it includes Docker CLI and container start, transport to the engine, and JVM or .NET start; their peak RSS measures only the local Docker client. For libreoffice it includes soffice start and profile creation. Do not read these numbers as library-only processing speed or memory.

Listed, never scored (neither can open an existing .docx and write one back):

  • docx (npm): generation-only.
  • mammoth: converts .docx to HTML or Markdown only.

Reproduce the retained evidence run

uv sync --extra adapters
uv pip install <wolfdocx wheel>
docker --context <engine> build -t docxbench/apache-poi:5.5.1 adapters/apache-poi
docker --context <engine> build -t docxbench/docx4j:17.2.1 adapters/docx4j
docker --context <engine> build -t docxbench/open-xml-sdk:3.5.1 adapters/open-xml-sdk
DOCXBENCH_DOCKER_CONTEXT=<engine> uv run docxbench benchmark \
  --output evidence/<date>/run --iterations 3 \
  --wolfdocx-artifact <wolfdocx wheel> \
  --wolfdocx-source-revision <commit> --wolfdocx-build-mode source-build-wheel
uv run docxbench report evidence/<date>/run/run.json

Observed run evidence (2026-10-02)

run-fair was scored by the name-independent whole-package model described above. The earlier run in evidence/2026-10-02/run/, scored by a whole-package comparison keyed by part name and relationship id, is retained unchanged; every row's outcome and failed-check set is identical in both runs.

Adapter feature-matrix template-mutation
python-docx 1.2.0 success success
docxtpl 0.20.2 success unsupported
wolfdocx-wheel 0.1.0 unsupported failure
apache-poi 5.5.1 success success
docx4j 17.2.1 success success
libreoffice 26.8.0.3 failure failure
pandoc 3.11 failure failure
open-xml-sdk 3.5.1 success success

Failed scored checks, as recorded in run.json:

  • wolfdocx-wheel template-mutation: wolfdocx author report refused the request with text_structure_unsupported, so no output was written. Its text edits accept only paragraphs whose direct children are runs, and the bookmarked target paragraph also holds w:bookmarkStart/w:bookmarkEnd. In a separate manual probe that is not part of the retained run, the same request without the bookmark edit (the two cells and the list item) was published; the lane requires all four edits.
  • libreoffice: table-name-and-text (the table's w:tblCaption name and w:tblDescription are dropped; cell text survives) and custom-document-properties (the i4 property DocxBenchRevision is written back as r8); in template-mutation, the same two as -preserved checks plus only-declared-text-nodes-changed. All four edits were applied exactly.
  • pandoc: in feature-matrix, run-formatting, table-name-and-text, table-cell-shading, bulleted-list, header-footer-content, page-number-field, comments, section-break-landscape, toc-field, content-control, custom-document-properties, and opaque-custom-xml-part; in template-mutation, the same features as -preserved checks plus only-declared-text-nodes-changed. All four edits were applied exactly.

The wolfdocx candidate was a wheel built out-of-tree from source revision 788278b0e947a10fcb400247b777a3e71b3d36a8 (receipt source-build-wheel). Container adapters ran on a remote linux/amd64 Docker engine; the Python adapters (CPython 3.12.12), LibreOffice, and pandoc ran locally on macOS arm64. Each row ran three measured iterations, no warmups, with a 120 second child timeout. Timings include process and runtime start and were taken with OS cache state uncontrolled. Rows that are unsupported were not run.

License

MIT; see LICENSE. See PROVENANCE.md for dependency licenses.

About

Benchmark for libraries that read, write and edit .docx files: feature, template-mutation and preservation checks

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages