Skip to content

Repository files navigation

ProcureSight — Which government contracts get changed the most after they're signed?

When a U.S. government agency hires a company, the contract often gets changed after it's signed — to add work, extend the deadline, or adjust the money. These changes are called modifications, and they're a normal part of doing business. But some contracts get changed a lot in their first year, and that's worth spotting early so the people who manage contracts can plan their time and attention.

This project takes 713,260 real government contract records from the Department of Justice and turns them into a simple answer to one question: on the day a contract is signed, can we tell which ones are headed for heavy changes? It turns out we can, well enough to be useful.


The short version

  • We looked at 297,098 Department of Justice contracts worth $29.6 billion, awarded between 2018 and 2023.
  • About 1 in 12 (8.6%) got changed heavily — three or more times — in their first year.
  • Using only what's known the day a contract is signed, a model can rank contracts by how likely they are to be changed heavily. When it flags its top 10% riskiest contracts, about 41% of them really do get changed heavily — versus about 8% if you just picked at random. That's a 5× improvement over guessing.
  • The single biggest clue is the type of contract. "Time-and-materials" contracts (where you pay for hours and supplies as you go) get changed heavily about twice as often as "fixed-price" contracts (where the price is locked in up front).

Important: getting changed a lot is not the same as doing something wrong. Changes are a routine way to manage contracts. This tool points to contracts likely to change, not contracts doing anything improper.

The question, plainly

On the day a contract is awarded — before any changes have happened — which contracts are most likely to be changed three or more times in the next 12 months?

The data

We used USAspending.gov, the U.S. government's public record of federal spending (it's free and open to anyone). We pulled Department of Justice contracts signed between fiscal years 2018 and 2023, plus one extra year of records (2024) so that even the newest contracts have a full 12 months of history to look at.

Where it comes from USAspending.gov (public data)
Who it covers Department of Justice
Contracts studied 297,098
Total value $29.6 billion
Individual records processed 713,260 (each award plus every change to it)

Every step is reproducible: we saved the exact data request, a fingerprint (checksum) of each file, and the record counts, and we kept the original files untouched.

How we did it (in plain terms)

  1. Define "heavy change." We grouped every record by its contract, found the original award, and counted how many changes happened in the next 12 months. Three or more = a "heavily changed" contract. We only kept contracts old enough to have a full year of history.
  2. Only use what's known at signing. The model is allowed to see the things you'd know the day the contract is signed — the type of contract, how much it's worth, who won it, what it's for, whether it was competed, and where the work happens. It is not allowed to peek at anything that happened afterward. That's what makes the prediction honest.
  3. Test it on years it never saw. We taught the model on 2018–2021 contracts, tuned it on 2022, and then tested it on 2023 contracts it had never seen — the same way you'd find out if it actually works on next year's contracts.
  4. Try two approaches. A simple model and a more flexible one, and keep whichever does better at picking out the heavily-changed contracts.

What we found

  • Changes are common; heavy changes are not. Almost half of contracts (46%) get at least one change in the first year, but only 8.6% get changed heavily. That makes "heavy change" a meaningful thing to flag.
  • Changes happen fast. When a contract does get changed, the first change comes after about 71 days on average — so a check-in around the three-month mark would catch most of them.
  • Contract type matters most. "Time-and-materials" contracts get changed heavily 15% of the time — nearly double the 8% rate for "fixed-price" contracts. How you structure the deal up front is a strong hint about what comes later.
  • The spending is spread out. No handful of companies dominates — the 25 largest vendors account for only about 27% of the money.
  • Where the money goes. The FBI is the biggest spender ($7.1B). The largest categories are IT services ($4.2B) and inmate healthcare ($2.0B) — which fits the DOJ's law-enforcement and prison responsibilities.

Does the prediction actually work?

We measured it on the 2023 contracts the model had never seen. The clearest way to say it:

If you sorted all contracts by the model's risk score and looked only at the riskiest 10%, about 41 out of every 100 would truly go on to be heavily changed — compared with about 8 out of 100 if you didn't use the model at all.

That's a 5× improvement, which means a contract manager could focus their attention on a short, high-yield list instead of reviewing everything equally.

Technical scorecard (for the curious)

Held-out 2023 test year:

Model ROC-AUC PR-AUC Top-10% precision
Logistic regression (baseline) 0.749 0.240 0.272
Random forest (chosen) 0.815 0.390 0.410
Random guessing (base rate) — — 0.077

The winner was chosen using the tuning year only, so the test-year scores are an honest estimate of real-world performance. All runs are tracked in MLflow.

What it means

  • Heavy contract changes are somewhat predictable from day-one information — enough to help agencies decide where to focus limited oversight, instead of treating every contract the same.
  • The type of contract is the most useful early signal — a lever that's decided up front.
  • A check-in about three months after signing would catch most first changes.

Honest limitations

  • This is public, self-reported government data; a few fields can be missing.
  • "Three or more changes" is a reasonable cutoff, not the only possible one.
  • We count how often a contract changed, not how big the changes were.
  • These patterns are specific to the Department of Justice and shouldn't be assumed to hold for other agencies without re-checking.

How it's built (for a technical reader)

Real contract data flows through a standard analytics pipeline and ends in a dashboard anyone can click through:

USAspending.gov API  ->  AWS S3 (raw storage)  ->  Databricks (clean + organize)
   ->  DuckDB (fast query layer)  ->  Streamlit dashboard  +  plain-English AI summary
  • Extraction pulls the data reproducibly (saved request, checksums, manifest).
  • Databricks (with Spark) turns raw records into a clean, contract-level table in three layers (raw → cleaned → analysis-ready), stored back in S3.
  • DuckDB is the free, open-source database that powers the dashboard — used here in place of a paid cloud data warehouse. It reads the finished data straight from S3.
  • The model (Python / scikit-learn, tracked with MLflow) runs against that finished data and writes its risk scores back to S3.
  • Streamlit is the four-page dashboard; a controlled AI step writes a short summary using only the calculated numbers, with guardrails against over-claiming.
Folder What's in it
extraction/ Downloads the data and records checksums / a manifest
databricks/ The clean-up and organize steps (Spark notebooks)
modeling/ Trains the model and saves its risk scores
warehouse/ Builds the DuckDB database the dashboard reads
app/ The Streamlit dashboard (four pages)
llm/ The controlled plain-English summary
docs/ Findings, data dictionary, demo script, results
tests/ Automated checks for every stage

Reproducing it

python3.12 -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
cp .env.example .env          # set your S3 bucket + region; AWS keys live in ~/.aws

python extraction/request_download.py   # 1. download the data
python extraction/poll_download.py
python extraction/extract_files.py
python extraction/manifest.py           #    checksums + upload to S3
#   2. run databricks/01..04 on the workspace to clean and organize
python modeling/train_models.py         # 3. train the model, save scores
python warehouse/build_duckdb.py        # 4. build the dashboard's database
streamlit run app/Home.py               # 5. open the dashboard
pytest -q                               # 6. run the checks

The dashboard's numbers are checked against the cleaned data before use, so they always match. A short walkthrough is in docs/demo_script.md.

Data: USAspending.gov (public domain), Department of Justice contract transactions.

About

Predicting which U.S. federal contracts get changed the most after signing — 713k real DOJ contract records through S3, Databricks, DuckDB, and a Streamlit dashboard.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages