PolitiBench collects PolitiFact articles, prepares standalone claims, extracts article evidence, and evaluates claim feasibility and verdict prediction. This guide covers the PolitiFact workflow.
Run all commands below from the repository root. Relative data paths resolve from the repository root even when a script is launched from another directory.
PolitiFact website
→ PolitiFact_RawScrape.py Download article HTML into a Parquet dataset
→ PolitiFact_ScrapeData.py Extract claims, metadata, analysis, and verdicts
→ Claim_Augmentation.py Prefix metadata, then decontextualize propositions
→ Evidence_Extraction.py Extract and format evidence
→ Verdict evaluation Predict labels and compute accuracy and macro F1
Augmented claims → Feasibility.py → Assess whether each claim can be fact-checked
Feasibility is an independent evaluation. OpenAI and Qwen verdict evaluation are alternative backends; neither needs to run before the other.
| Directory | Contents |
|---|---|
| Code/Data_Collection | Download article HTML and parse it into structured datasets. |
| Code/Data_Component | Combined claim augmentation and evidence extraction. |
| Code/Evaluation | Feasibility assessment and OpenAI/Qwen verdict evaluation. |
| Prompt | Model instructions, few-shot examples, and message builders. |
Data/ |
Default raw HTML and parsed article datasets. |
Data/Components/ |
Default augmented claims and extracted evidence. |
Results/Evaluation/ |
Default evaluation results and checkpoints. |
The scripts create output directories when needed. Prompt files and shared helper modules are imported by the scripts; they do not need to be run separately.
Use Python 3.10 or newer. A virtual environment keeps project dependencies separate:
python3 -m venv .venv
source .venv/bin/activate
python3 -m pip install requests \
-r Code/Data_Component/requirements.txt \
-r Code/Evaluation/requirements.txtFor local Qwen verdict evaluation or feasibility assessment, also install:
python3 -m pip install -r Code/Evaluation/requirements-local.txtAugmentation, evidence extraction, and OpenAI verdict evaluation require OPENAI_API_KEY in the environment or a .env file at the repository root:
OPENAI_API_KEY=your_api_key_hereLocal model loading also reads HF_TOKEN if needed. Environment values take precedence over .env. Model loading and API client initialization happen when inference is needed, not when displaying --help.
The configured model defaults are gpt-5.6-luna for API stages, Qwen/Qwen3-4B for local verdict evaluation, and Qwen/Qwen3-8B for feasibility. Use --model to select another compatible model available to you. Local models require sufficient memory and download uncached weights when loaded.
Run each step after the preceding step succeeds. These commands use matching default file paths.
python3 Code/Data_Collection/PolitiFact_RawScrape.py --limit 1Output: Data/PolitiFact_HTML.parquet.
--limit counts listing pages, not articles. The dataset stores each article's HTML in an html column alongside its URL and download metadata; it does not create separate HTML files. The scraper's default limit is 200 pages, and -1 removes the page limit. An optional --time-stamp YYYYMMDD sets the earliest article date.
python3 Code/Data_Collection/PolitiFact_ScrapeData.pyOutput: Data/PolitiFact_DATA.parquet.
This extracts claims, verdicts, speaker and statement metadata, summaries, sources, and article analysis. analysis_text and ruling_text are separate. The empty Leakage_Control_PolitiFact.py placeholder is not part of the runnable pipeline.
python3 Code/Data_Component/Claim_Augmentation.py --data-size 20Outputs: Data/Components/Claims_Decontextualized.parquet and .json.
The combined script first builds augmented_claim_cat using available metadata, then calls the model to resolve contextual references in the proposition. It retains the original claim and adds decontextualized_claim, status/error fields, and token counts. A separate metadata-generation run is not required.
Omit --data-size to process all rows. Use --start-row to select a zero-based starting row. For local prefixing only:
python3 Code/Data_Component/Claim_Augmentation.py --prefix-onlyThat mode requires no API call and writes Claims_Metadata.parquet and .json; it does not produce the decontextualized input required by the next stage.
python3 Code/Data_Component/Evidence_Extraction.py --k 2Outputs: Data/Components/Evidence.parquet, .json, and Evidence.run.json.
This reads the decontextualized claims and analysis_text. It adds structured Extracted_Evidence and a Formatted_Evidence column with full support, partial support, contradiction, and context sections. --k accepts 2, 3, or 4 facts per extraction category. Missing or blank claims/article text are excluded.
Choose the OpenAI backend:
python3 Code/Evaluation/verdict_evaluation.py \
--claim-col decontextualized_claimOutput: Results/Evaluation/Results_Formatted_Evidence_GPT.json.
Or use local Qwen:
python3 Code/Evaluation/verdict_evaluation_qwen.py \
--claim-col decontextualized_claimOutput: Results/Evaluation/Results_Formatted_Evidence_QWEN.json.
Both read Data/Components/Evidence.parquet and use Formatted_Evidence by default. They report accuracy and macro F1 over PolitiFact's six verdict labels. Gold labels are used for scoring, not passed to the model. Use --evidence-cols analysis_text Formatted_Evidence to compare evidence representations on the same valid rows.
The default claim column is the original claim. The explicit --claim-col decontextualized_claim above evaluates the output of augmentation instead.
python3 Code/Evaluation/Feasibility.py \
--text-column decontextualized_claim \
--output-name Feasibility_DecontextualizedOutput: Results/Evaluation/Feasibility_Decontextualized.parquet, plus its .run.json manifest.
This reads the augmented claims directly and rates whether each is sufficiently specified for fact-checking: 0 means not assessable, 1 means ambiguous or missing context, and 2 means sufficiently clear. It does not decide truth or perform web searches. Use --text-column claim with another output name to assess the original claims.
All processing stages accept --output-dir; parsing, component, and evaluation stages accept --input-path. Changing one stage's output directory does not automatically change the next stage's input.
| Stage | Naming and resume behavior |
|---|---|
| Raw collection | --output-name specifies the full raw filename. Reusing a path overwrites the file; runs do not append. |
| Article parsing | --output-path specifies a complete output path and overrides --output-dir. |
| Augmentation | --output-name is a stem for both Parquet and JSON. Saves every 25 rows; no automatic resume. Use a new name and --start-row for continuation. |
| Evidence extraction | --output-name is a stem. Add --resume with the same input, model, and k; keep the JSON checkpoint and run manifest together. |
| Verdict evaluation | --output-prefix names results. Matching checkpoints resume automatically; --overwrite starts a fresh run. |
| Feasibility | --output-name names the default Parquet output. Add --resume with matching input/settings, or use --output-path for another format. |
Use fresh output names when changing prompts or comparing experiments. Empty crawls and parser runs with no successfully parsed records now fail explicitly rather than reporting success without a usable dataset.
Run the PolitiFact-only end-to-end integration test:
python3 -m unittest discover \
-s Code/Data_Component/tests \
-k test_collection_to_all_evaluations_for_politifact -vRun the evaluation regression tests:
python3 -m unittest discover -s Code/Evaluation/tests -vThe integration test runs the real script entry points and writes/reloads every intermediate file, with HTTP and model responses mocked. It checks schema compatibility, paths, and handoffs without paid API calls or model downloads.