Skip to content
View GirelliG-it's full-sized avatar
💻
Focused
💻
Focused

Block or report GirelliG-it

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
GirelliG-it/README.md

Giovanni Girelli

I have a background in procurement and logistics and am working towards a junior data engineering role. I’m developing my skills through practical projects with Python, SQL and open data, focusing on reproducible pipelines, data validation and tested software.

Based in The Hague, I’m looking for an opportunity to contribute my domain experience, tackle technical problems and grow alongside experienced engineers. I’m particularly interested in work that supports public services and informed decision-making.


Data engineering projects

dataset-prober — Released v0.1.0 (working towards v0.2.0)

A Python command-line tool for inspecting open-data resources and optionally loading approved data into DuckDB. The current release supports CSV and CBS/OData inspection, with explicit approval before database loading and protection against overwriting existing tables.

The project demonstrates Python packaging, data validation, transactional loading, automated testing and continuous integration. (Project is ongoing.)

Python · DuckDB · pytest · Ruff · GitHub Actions

github-copilot-sql-lab

The aim of this project is to study and build a production-oriented pipeline focused on ingestion, validation, transformation and SQL analysis. In this case study the main focus will be analysing and processing Dutch seaport cargo volumes from the official CBS StatLine Dataportal. The current status is planning and initial data exploration.


Learning

Foundational tools compound in value, so I work through them deliberately rather than picking up fragments as needed.

unix-power-tools

Working through the classic Unix text-processing tools: ed, sed, awk, grep.

c-projects

Learning the C programming language through a structured series of small, hands-on projects including PowerShell equivalents.


Tools I use and am learning

  • Project work — Python, DuckDB, pytest, Ruff, GitHub Actions, Git
  • Current learning focus — SQL, relational modelling and data pipelines
  • Development environment — Linux, zsh, Miniconda
  • Additional experience — R, pandas, Jupyter, RStudio and Quarto

Principles

Reproducibility is a first-class requirement, not an afterthought. Public data belongs to the public — and so do the analyses built from it. Provenance matters: where data came from, and when, is part of the result.


Links

Popular repositories Loading

  1. dataset-prober dataset-prober Public

    Safety-first Python CLI for discovering open-data resources, deterministically classifying inspected content, and explicitly loading verified datasets into DuckDB. v0.1.0 has been released. Working…

    Python

  2. GirelliG-it GirelliG-it Public

    README

  3. c-projects c-projects Public

    This repository documents my progress learning C through a structured series of small, hands-on projects.

    C

  4. unix-power-tools unix-power-tools Public

    Working through the classic Unix text-processing tools: ed, sed, awk, and grep.

    sed

  5. github-copilot-sql-lab github-copilot-sql-lab Public

    The aim of this project is to study and build a production-oriented pipeline focused on ingestion, validation, transformation and SQL analysis and the ability of GitHub Copilot to analyze, interpre…

    Python