Skip to content
View lucalullo's full-sized avatar

Block or report lucalullo

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
lucalullo/README.md

Luca Lullo - Data & AI Engineer

Independent Data & AI Engineer focused on data engineering, advanced data cleaning, applied machine learning, and the development of AI systems.

I build reproducible data pipelines, high-quality datasets, predictive models, AI agents, and language models from scratch, with a strong emphasis on transparency, experimentation, and practical implementation.

My work spans:

  • data integration, cleaning, validation, and quality control;
  • exploratory analysis and reproducible research;
  • feature engineering and predictive machine learning;
  • explainable AI and model interpretation;
  • AI agents with routing, planning, tools, memory, and human escalation;
  • agentic AutoML and experiment-driven ML systems;
  • language models and Transformers implemented from first principles;
  • analysis of public, socioeconomic, institutional, and climate data.

I primarily use Python to transform complex data into reliable, understandable, and actionable systems and insights.


🏆 Kaggle Expert

Kaggle

Expert in Datasets and Notebooks

  • Datasets: Top 30
  • Notebooks: Top 511
  • Multiple Kaggle medals across datasets and notebooks

🛠️ Tech Stack

Python Pandas Scikit-learn TensorFlow PyTorch XGBoost LightGBM SQL Git


📂 Featured Projects

Project Area Highlights
CodAdapt Machine Learning / Tabular Data Experimental ML algorithm based on adaptive coded memory and shared multi-resolution encoding for classification and regression
Neuro Tabular Deep Learning / Tabular Data Neural-network framework for tabular learning with preprocessing, training engine, and scikit-learn-compatible API
BurOCRazia OCR / Data Engineering Local and deterministic system for recognizing Italian administrative documents with OCR, traceable SQLite catalog, and offline processing
Building Agentic AutoML Agentic AI / AutoML Experiment-driven AutoML system evolving from a simple baseline to a senior ML agent
Building AI Agent AI Agent from Scratch Routing, planning, parsing, tool execution, memory, and agent skills
Building LLM LLM from Scratch Progressive implementation from statistical language modeling to a decoder-only Transformer
Customer Support Agent AI Agent / Human-in-the-Loop Customer-support agent with classification, automation, and human escalation
Home Credit Default Risk Machine Learning / Credit Risk Feature engineering, XGBoost, Optuna, and SHAP interpretability
Global Emissions & Temperature - 1950-2024 Climate / Time Series 75 years of CO₂, greenhouse-gas, and global temperature analysis

🔎 Areas of Interest

  • Data engineering and data quality
  • Advanced data cleaning
  • Applied machine learning
  • Explainable AI
  • Agentic AI and AI agents
  • AutoML and experiment automation
  • Natural Language Processing
  • Language models and Transformers
  • AI systems from scratch
  • Public and socioeconomic data
  • Climate and time-series analysis
  • Reproducible research

📬 Connect

LinkedIn Kaggle GitHub

Pinned Loading

  1. codadapt codadapt Public

    Experimental tabular machine-learning algorithm based on adaptive coded memory and shared multi-resolution encoding.

    Python

  2. neuro-tabular neuro-tabular Public

    Experimental neural network library for tabular data with a scikit-learn-style API and automatic preprocessing.

    Python

  3. burocrazia burocrazia Public

    Applicazione locale e deterministica per riconoscere con prudenza documenti amministrativi italiani, con OCR, catalogo tracciabile e funzionamento offline.

    Python

  4. building-agentic-automl building-agentic-automl Public

    Building an Agentic AutoML system from scratch, step by step, from a simple baseline to an experiment-driven senior ML agent.

    Jupyter Notebook 4

  5. building-llm building-llm Public

    A step-by-step educational journey from a character-level statistical language model to a small decoder-only Transformer.

    Jupyter Notebook 10 1

  6. home-credit-default-risk home-credit-default-risk Public

    Machine learning project to predict credit default risk with feature engineering, XGBoost and SHAP interpretability.

    Jupyter Notebook 8