paper-extracting is the active cohort extraction pipeline in this repository. It extracts structured cohort metadata from one or more PDF papers, runs a local llama.cpp-based inference setup on a Slurm GPU node, fetches live EMX2 schema and ontology inputs, optionally uses a shared OCR model for weak PDFs, and writes a cohort workbook plus detailed run artifacts.
This README explains the complete workflow:
- Sync the repository to the cluster.
- Sync one or more PDF files to the cluster.
- Build and download the cluster runtime with
setup_cluster_runtime.sbatch. - Start a Slurm extraction job with
run_cluster_cohort.sbatch. - Wait for the workbook and run artifacts.
- Sync the workbook back to your local machine.
If you already have cluster access and just want the shortest working route, use this sequence.
These examples assume:
- local repo path:
/Users/p.jansma/Documents/GitHub/paper-extracting/ - cluster repo path:
/groups/umcg-gcc/tmp02/users/umcg-pjansma/Repositories/paper-extracting/ - SSH host alias:
tunnel+nibbler
Adjust those paths if your own setup differs.
rsync -avhP \
/Users/p.jansma/Documents/GitHub/paper-extracting/ \
tunnel+nibbler:/groups/umcg-gcc/tmp02/users/umcg-pjansma/Repositories/paper-extracting/rsync -avhP \
/Users/p.jansma/Documents/GitHub/paper-extracting/data/oncolifes.pdf \
tunnel+nibbler:/groups/umcg-gcc/tmp02/users/umcg-pjansma/Repositories/paper-extracting/data/If your PDF lives somewhere else locally, sync that file instead.
ssh tunnel+nibbler
cd /groups/umcg-gcc/tmp02/users/umcg-pjansma/Repositories/paper-extractingsbatch --export=ALL,WORKDIR=/groups/umcg-gcc/tmp02/users/umcg-pjansma,BUILD_LLAMA=1 setup_cluster_runtime.sbatchThis is the standard and preferred setup route for this project. It runs setup_cluster_runtime.sh on a GPU node and prepares a git-ignored runtime inside this repo.
By default it creates:
./.runtime/llama.cpp./.runtime/GGUF/gemma-4-31B-it-Q4_K_M.gguf./.venv.cluster/
watch -n 10 "squeue -u $USER"You can inspect the setup job logs with:
ls -1 setup_cluster_runtime_*.out setup_cluster_runtime_*.err | tailsbatch --export=ALL run_cluster_cohort.sbatch -p all --pdfs data/oncolifes.pdf -o oncolifes_cohort.xlsxThis is the normal batch route for one PDF. The workbook will be written in the repo root on the cluster as:
oncolifes_cohort.xlsxThe detailed run artifacts are written under:
logs/runs/<run_id>/watch -n 10 "squeue -u $USER"The Slurm wrapper logs normally appear as:
cohort_extract-<jobid>.out
cohort_extract-<jobid>.errrsync -avhP \
tunnel+nibbler:/groups/umcg-gcc/tmp02/users/umcg-pjansma/Repositories/paper-extracting/oncolifes_cohort.xlsx \
/Users/p.jansma/Documents/cluster_data/ls -1dt /groups/umcg-gcc/tmp02/users/umcg-pjansma/Repositories/paper-extracting/logs/runs/* | headInside the newest run directory, the first files to inspect are usually:
status.jsonlpipeline_issues.jsonconfig.runtime.tomlprompts.runtime.toml
src/run_cluster_cohort.sh is the supported runtime wrapper around the cohort extraction pipeline. In a normal production run it does the following:
- Creates a run directory under
logs/runs/<run_id>/. - Copies
config.cohort.tomlto a run-specificconfig.runtime.toml. - Builds the run-specific prompt file
prompts.runtime.toml. - Fetches the latest EMX2 ontology and model CSV files from
molgenis/molgenis-emx2. - Builds a live dynamic EMX2 runtime registry for the
UMCGCohortsStagingprofile. - Optionally schema-syncs the prompt file against the live EMX2 schema.
- Starts two local
llama-serverinstances and a local TCP load balancer. - Optionally starts the OCR
llama-serverand prefetches OCR text for weak PDFs. - Runs
src/main_cohort.pyto extract the selected PDF files into one cohort workbook. - Writes the workbook, run logs, prompt artifacts, OCR artifacts, and issue files.
- Optionally syncs the workbook elsewhere with
rsync.
The recommended way to start this project is:
- Sync the repository to the cluster.
- Sync your PDF file or files to the cluster.
- Log in to the cluster.
- Move into the repository directory.
- Submit
setup_cluster_runtime.sbatch. - Wait until setup finishes.
- Submit
run_cluster_cohort.sbatch.
ssh tunnel+nibbler
cd /groups/umcg-gcc/tmp02/users/umcg-pjansma/Repositories/paper-extractingThis is the normal route for this project. Run:
sbatch --export=ALL,WORKDIR=/groups/umcg-gcc/tmp02/users/umcg-pjansma,BUILD_LLAMA=1 setup_cluster_runtime.sbatchThis wrapper job:
- allocates a GPU node through Slurm
- loads the required modules
- runs
setup_cluster_runtime.sh - clones
ggml-org/llama.cpp - checks out the pinned working commit used by this project
- optionally builds
llama.cpp - downloads and verifies the default Gemma GGUF used by this repo
- creates or updates
./.venv.cluster/ - installs the project plus
pypdfium2,pillow, andxlsxwriter
For most users, this is the correct way to do setup.
It does not download the shared OCR model used by default in src/run_cluster_cohort.sh. The OCR defaults still point to the existing shared cluster paths:
/groups/umcg-gcc/tmp02/users/umcg-pjansma/Models/GGUF/GLM-OCR/GLM-OCR-Q8_0.gguf/groups/umcg-gcc/tmp02/users/umcg-pjansma/Models/GGUF/GLM-OCR/mmproj-GLM-OCR-Q8_0.gguf
If those files are unavailable on your cluster, either:
- set
OCR_VLM_ENABLE=0, or - override
OCR_VLM_MODEL_PATH,OCR_VLM_MMPROJ_PATH, andOCR_VLM_LLAMA_BIN
If you want to build llama.cpp manually, do not do that on the login node. Start an interactive compute session first:
srun --cpus-per-task=4 --mem=32G --nodes=1 --gres=gpu:a40:2 --time=04:00:00 --pty bash -iOnce you are inside the interactive compute shell, load the modules and run setup:
module purge
module load GCCcore/11.3.0
module load CMake/3.23.1-GCCcore-11.3.0
module load CUDA/12.2.0
module load Python/3.10.4-GCCcore-11.3.0
WORKDIR=/groups/umcg-gcc/tmp02/users/umcg-pjansma \
BUILD_LLAMA=1 \
bash setup_cluster_runtime.shThis manual route does the same core setup work, but setup_cluster_runtime.sbatch remains the preferred route.
If llama.cpp is already built and you only want to refresh the venv and model file, you can run:
WORKDIR=/groups/umcg-gcc/tmp02/users/umcg-pjansma \
BUILD_LLAMA=0 \
bash setup_cluster_runtime.shIf you want a separate test runtime instead of ./.runtime/ and ./.venv.cluster/, use:
SETUP_SUFFIX=_testThat creates:
./.runtime_test/./.venv.cluster_test/
Example:
SETUP_SUFFIX=_test BUILD_LLAMA=1 bash setup_cluster_runtime.shThe active cohort application now lives directly in the repository root and src/.
Important files:
setup_cluster_runtime.sh: repo-local runtime bootstrap forllama.cpp, Gemma, and the cluster venvsetup_cluster_runtime.sbatch: preferred Slurm wrapper for runtime bootstraprun_cluster_cohort.sbatch: preferred Slurm extraction entrypointsrc/run_cluster_cohort.sh: cluster runtime wrapper that starts the servers and runs the extractionsrc/main_cohort.py: main PDF-to-workbook extraction programconfig.cohort.toml: base runtime configprompts/prompts_cohort.toml: baseline prompt setsrc/emx2_dynamic_runtime.py: live EMX2 schema and ontology runtime buildersrc/cohort_prompt_schema_updater.py: prompt schema sync logic
Before you start, make sure you have all of the following.
The current scripts assume a GPU cluster and a Slurm scheduler.
The setup batch script currently requests:
1node2A40 GPUs4CPUs32Gmemory04:00:00walltime
The extraction batch script currently requests:
1node1task2A40 GPUs4CPUs32Gmemory01:00:00walltime
If your cluster uses different GPU names, partitions, or time limits, edit:
setup_cluster_runtime.sbatchrun_cluster_cohort.sbatch
You need local shell access to:
- sync this repo to the cluster
- sync your PDF files to the cluster
- sync the resulting workbook back to your machine
By default the runner expects the shared GLM-OCR paths shown above. If those are missing, either disable OCR or point to your own OCR model files.
The pipeline extracts structured cohort information from PDF papers. You must place the selected PDFs on the cluster before starting a run.
The current defaults are defined in setup_cluster_runtime.sh and src/run_cluster_cohort.sh.
Default binary:
./.runtime/llama.cpp/build/bin/llama-serverDefault path:
./.runtime/GGUF/gemma-4-31B-it-Q4_K_M.ggufThe same model is started on both GPUs by default.
Preferred default:
./.venv.cluster/Fallback if present:
./.venv/Default OCR model:
/groups/umcg-gcc/tmp02/users/umcg-pjansma/Models/GGUF/GLM-OCR/GLM-OCR-Q8_0.ggufDefault OCR mmproj:
/groups/umcg-gcc/tmp02/users/umcg-pjansma/Models/GGUF/GLM-OCR/mmproj-GLM-OCR-Q8_0.ggufDefault OCR llama-server:
/groups/umcg-gcc/tmp02/users/umcg-pjansma/Repositories/llama.cpp-glmtest/build/bin/llama-server- load balancer:
18000 - GPU server 0:
18080 - GPU server 1:
18081 - OCR server:
18090
- repo:
molgenis/molgenis-emx2 - ref:
main - profile:
UMCGCohortsStaging
Some runtime status messages still say qwen_starting and qwen_ready. That is only a leftover status label. The current default MODEL_PATH is Gemma, not Qwen.
From your local machine, sync the repository to the cluster.
Example:
rsync -avhP \
/Users/p.jansma/Documents/GitHub/paper-extracting/ \
tunnel+nibbler:/groups/umcg-gcc/tmp02/users/umcg-pjansma/Repositories/paper-extracting/Then log in and move into the project directory:
ssh tunnel+nibbler
cd /groups/umcg-gcc/tmp02/users/umcg-pjansma/Repositories/paper-extractingYou must also place the PDF files on the cluster before running.
Example for one file:
rsync -avhP \
/Users/p.jansma/Documents/GitHub/paper-extracting/data/oncolifes.pdf \
tunnel+nibbler:/groups/umcg-gcc/tmp02/users/umcg-pjansma/Repositories/paper-extracting/data/Example for multiple files:
rsync -avhP \
/Users/p.jansma/Documents/GitHub/paper-extracting/data/ \
tunnel+nibbler:/groups/umcg-gcc/tmp02/users/umcg-pjansma/Repositories/paper-extracting/data/The preferred route is:
sbatch --export=ALL,WORKDIR=/groups/umcg-gcc/tmp02/users/umcg-pjansma,BUILD_LLAMA=1 setup_cluster_runtime.sbatchWait for the job:
watch -n 10 "squeue -u $USER"Inspect setup logs:
ls -1 setup_cluster_runtime_*.out setup_cluster_runtime_*.err | tailIf setup completes successfully, you should now have:
./.runtime/llama.cpp/build/bin/llama-server./.runtime/GGUF/gemma-4-31B-it-Q4_K_M.gguf./.venv.cluster/bin/python
The main production route is the Slurm wrapper:
sbatch --export=ALL run_cluster_cohort.sbatch -p all --pdfs data/oncolifes.pdf -o oncolifes_cohort.xlsxIf you call run_cluster_cohort.sbatch without extra arguments, it uses a safe default:
sbatch run_cluster_cohort.sbatchThat default run becomes roughly:
bash src/run_cluster_cohort.sh -p all --pdfs data/oncolifes.pdf -o cohort_<jobid>.xlsxExample:
sbatch --export=ALL run_cluster_cohort.sbatch \
-p all \
--pdfs data/oncolifes.pdf data/concrete.pdf \
--paper-names oncolifes concrete \
-o multi_cohort.xlsxExample:
sbatch --export=ALL run_cluster_cohort.sbatch \
-p A C E \
--pdfs data/oncolifes.pdf \
-o oncolifes_partial.xlsxExample:
sbatch --export=ALL run_cluster_cohort.sbatch \
--ocr \
-p all \
--pdfs data/oncolifes.pdf \
-o oncolifes_ocr.xlsxExample:
sbatch --export=ALL run_cluster_cohort.sbatch \
--ocr-force-use \
-p all \
--pdfs data/oncolifes.pdf \
-o oncolifes_ocr_forced.xlsxExample:
sbatch --export=ALL run_cluster_cohort.sbatch \
--ocr-dump \
-p all \
--pdfs data/oncolifes.pdf \
-o oncolifes_ocr_dump.xlsxIf the shared OCR model paths are unavailable or you do not want OCR:
sbatch --export=ALL,OCR_VLM_ENABLE=0 run_cluster_cohort.sbatch \
-p all \
--pdfs data/oncolifes.pdf \
-o oncolifes_no_ocr.xlsxsqueue -u $USERwatch -n 10 "squeue -u $USER"watch -n 10 'squeue -u $USER -o "%.18i %.20j %.8T %.10M %.10L %.6D %R"'sacct -j <jobid> --format=JobID,JobName%30,State,ExitCode,Elapsedtail -f cohort_extract-<jobid>.out cohort_extract-<jobid>.errThe extraction wrapper creates a run directory under:
logs/runs/<run_id>/The first files to inspect there are usually:
status.jsonlpipeline_issues.jsonconfig.runtime.tomlprompts.runtime.toml
If prompt schema sync ran, you may also see:
prompt_schema_sync.report.jsonprompt_schema_sync.compare.mdprompt_schema_sync.prompt.diffprompt_schema_sync.before_after.md
If OCR prefetch or OCR dumps ran, you may also see:
ocr_prefetch/text_compare/logs/ocr_prefetch_failed.txt
The safest default is to pull the workbook manually from your local machine after the job finishes.
Example:
rsync -avhP \
tunnel+nibbler:/groups/umcg-gcc/tmp02/users/umcg-pjansma/Repositories/paper-extracting/oncolifes_cohort.xlsx \
/Users/p.jansma/Documents/cluster_data/If you wrote to a different output filename, pull that file instead.
The runner supports automatic sync through:
LOCAL_RSYNC_DESTLOCAL_RSYNC_HOSTSYNC_OUTPUT_ENABLESYNC_REQUIRED
Use that only if your cluster node can actually reach the target receiver. If that connectivity is uncertain, the manual pull route above is more reliable.
If you are already inside an interactive compute session and want to run without the Slurm wrapper script, use:
bash src/run_cluster_cohort.sh -p all --pdfs data/oncolifes.pdf -o oncolifes_cohort.xlsxThat is useful for debugging, but run_cluster_cohort.sbatch is the normal production route.
The main output is the Excel workbook you pass with -o.
Example:
oncolifes_cohort.xlsxThe run directory always contains:
logs/runs/<run_id>/config.runtime.tomllogs/runs/<run_id>/status.jsonllogs/runs/<run_id>/pipeline_issues.json
The run directory usually also contains:
logs/runs/<run_id>/prompts.runtime.tomllogs/runs/<run_id>/logs/
Additional output files may be written next to the workbook:
<output>.dynamic_prompts.toml<output>.dynamic_emx2_registry.json<output>.dynamic_prompt_constraints.json<output>.issues.json
RUNTIME_ROOT=/path/to/custom_runtime bash setup_cluster_runtime.shVENV_DIR=/path/to/custom_venv bash setup_cluster_runtime.shLLAMA_BIN=/path/to/llama-server \
bash src/run_cluster_cohort.sh -p all --pdfs data/oncolifes.pdf -o out.xlsxMODEL_PATH=/path/to/model.gguf \
bash src/run_cluster_cohort.sh -p all --pdfs data/oncolifes.pdf -o out.xlsxMOLGENIS_EMX2_REPO="PieterJansma/molgenis-emx2" \
MOLGENIS_EMX2_REF="main" \
sbatch --export=ALL run_cluster_cohort.sbatch -p all --pdfs data/oncolifes.pdf -o fork_test.xlsxCOHORT_PROMPT_SCHEMA_SYNC=0 \
COHORT_DYNAMIC_PROMPTS=1 \
sbatch --export=ALL run_cluster_cohort.sbatch -p all --pdfs data/oncolifes.pdf -o dynamic_only.xlsx- Use
setup_cluster_runtime.sbatchas the normal setup route. - Do not build
llama.cppon the login node. - After code-only changes, syncing the repo is usually enough because the venv is installed in editable mode.
- After dependency changes in
pyproject.toml, rerunsetup_cluster_runtime.shorsetup_cluster_runtime.sbatch. - If
llama-serveris missing, rerun setup withBUILD_LLAMA=1. - If the Gemma file is missing, rerun setup and inspect the setup logs.
- If OCR fails because the shared GLM-OCR files are unavailable, set
OCR_VLM_ENABLE=0or point to your own OCR model files. - If the workbook sync fails, the workbook still remains on the cluster in the project directory unless you wrote it somewhere else explicitly.
llama.cpp repository:https://github.com/ggml-org/llama.cppGemma 4 31B IT GGUF:https://huggingface.co/ggml-org/gemma-4-31B-it-GGUFmolgenis-emx2:https://github.com/molgenis/molgenis-emx2