I have a background in procurement and logistics and am working towards a junior data engineering role. I’m developing my skills through practical projects with Python, SQL and open data, focusing on reproducible pipelines, data validation and tested software.
Based in The Hague, I’m looking for an opportunity to contribute my domain experience, tackle technical problems and grow alongside experienced engineers. I’m particularly interested in work that supports public services and informed decision-making.
dataset-prober — Released v0.1.0 (working towards v0.2.0)
A Python command-line tool for inspecting open-data resources and optionally loading approved data into DuckDB. The current release supports CSV and CBS/OData inspection, with explicit approval before database loading and protection against overwriting existing tables.
The project demonstrates Python packaging, data validation, transactional loading, automated testing and continuous integration. (Project is ongoing.)
Python · DuckDB · pytest · Ruff · GitHub Actions
The aim of this project is to study and build a production-oriented pipeline focused on ingestion, validation, transformation and SQL analysis. In this case study the main focus will be analysing and processing Dutch seaport cargo volumes from the official CBS StatLine Dataportal. The current status is planning and initial data exploration.
Foundational tools compound in value, so I work through them deliberately rather than picking up fragments as needed.
Working through the classic Unix text-processing tools: ed, sed, awk, grep.
Learning the C programming language through a structured series of small, hands-on projects including PowerShell equivalents.
- Project work — Python, DuckDB, pytest, Ruff, GitHub Actions, Git
- Current learning focus — SQL, relational modelling and data pipelines
- Development environment — Linux, zsh, Miniconda
- Additional experience — R, pandas, Jupyter, RStudio and Quarto
Reproducibility is a first-class requirement, not an afterthought. Public data belongs to the public — and so do the analyses built from it. Provenance matters: where data came from, and when, is part of the result.

