Skip to content

Run route B over the OCR markdown of a PDF set - #43

Merged
peterbjohnson merged 7 commits into
mainfrom
wb/t48
Sep 23, 2026
Merged

peterbjohnson merged 7 commits into
mainfrom
wb/t48

Conversation

@peterbjohnson

@peterbjohnson peterbjohnson commented Sep 23, 2026 •

Copy link
Copy Markdown
Member

Every ticket writes its tests before its code, and its pull request reports a run over a target set or a corpus folder with the number of flagged fields. See docs/plan.md. Write every run's output to a file under out/ before reading it.

In the corpus sweep, route B did not run on any PDF set: the sweep hands the PDF itself to
pandoc, which reads no PDF, so the MECH60014 PDF sheets report route B failed: Unknown input format pdf and the UCL_MechEng folder no filter: pandoc reads no sheet of this set. For a PDF, route B reads the Mathpix markdown that route A reads: the filter is
written from that markdown's tree and pandoc runs it over that markdown. The ticket is done
when the UCL_MechEng sweep runs both routes on both sheets and reports a comparison.


Workbench ticket t48.

Live sweep over UCL_MechEng with both routes (from the implementer's run, run outputs kept out of the commit)

Sheet Questions Parts Fields Agreed Adjudicated Flagged Not verbatim Tokens
Worksheet_1.pdf 12 13 63 16 26 3 1 81,259
Worksheet_2.pdf 5 10 40 6 0 4 0 10,351

@peterbjohnson
peterbjohnson merged commit c91108b into main Sep 23, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant