Skip to content
View soulipaco's full-sized avatar

Block or report soulipaco

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
soulipaco/README.md
Business analyst turned data & AI engineer

The analyst half decides what is worth measuring. The engineer half builds the thing that measures it — governed Databricks systems whose numbers are published and enforced, including the ones that came out badly.

Technical portfolio → · LinkedIn


Currently

Testing platform claims before building on them. Can one Databricks SQL function replace a tuned forecasting pipeline? databricks-forecast-planning-cockpit answers that under a protocol frozen before the test year was downloaded — and publishes the part of the answer that came out badly. Released as v1.0.0.


Selected work

Structure-aware, multilingual PII reduction for Databricks · released v0.1.0 · Python · Presidio + spaCy

Reduces PII inside free-text columns an operator names. A ticket id survives, a timestamp and a speaker label survive; the name and the email do not. English, German and Greek. Every published number is a regression gate — 56 of them, across three corpora and both provider chains — so no figure in the documentation can move without a test failing.

Executed driver-path parity on a real Azure Databricks workspace; the service hosted as a Databricks App and driven over HTTPS
Not executed — and it says so the distributed mapInPandas path is shipped and has never run; ADDRESS is in the taxonomy and nothing detects it; Greek PERSON recall is published as 0.500 rather than rounded up
Not claimed it is not an estate scanner, and it promises no compliance outcome and no guaranteed anonymization

Inspect · what was actually executed · 36 decision records · the measured baseline · providers and their limits

Frozen-protocol forecasting benchmark · released v1.0.0 · Databricks ai_forecast · Prophet · Delta · Asset Bundles · AI/BI

ai_forecast v2 against a tuned Prophet pipeline and a weekly seasonal baseline on 21 NYC 311 daily-demand series, 28 days ahead from 12 monthly origins across 2025. The SQL function won on accuracy — median series error 17.1% against 19.5%, best on 17 of 21 series — with no training code and no tuning. It also ran 14.3% low, below actual demand at every origin, so under the pass rule written before 2025 was downloaded it does not qualify as a replacement.

Executed all 756 forecast cells scored with zero failures, v2 as fresh inference on a Databricks trial workspace; predictions stored in Delta and served by a code-managed AI/BI dashboard; 85 tests in CI
Not like-for-like v2 ran on Databricks while Prophet and the baseline ran locally, so runtimes are reported but not compared as a speed test
Not claimed a retrospective backtest on one public dataset — no universal winner, and no staffing, capacity or savings result

Inspect · benchmark protocol · evidence index · what v2 needed to run

Released Databricks accelerator · v1.0.0 · analytics engineering · forecasting · AI/BI · Genie

Answers an operating question rather than a modelling one: when is a new-hire cohort becoming production-ready, and what evidence supports the decision? Four governed source tables become learning curves, volume-aware diagnostics, forecasts, process-control views, an AI/BI dashboard, a Genie space and an optional evidence-grounded action workflow. Release quality gates run in CI.

Inspect · validation record · architecture

Released reference · v0.1.0 · governed retrieval · retrieval evaluation

Built around one testable claim: fixed-window retrieval loses document structure and evidence relationships. It preserves the hierarchy and expands exact one-hop CFR references as separately citable evidence, against a date-pinned public eCFR corpus with a committed evaluation set and live Databricks evidence.

Inspect · evaluation design · retrieval design


Also in the portfolio

prophet-forecasting-mlops A compact, reproducible batch-forecasting reference. Forecasting behaviour stays in testable Python; Databricks-specific code is confined to delivery, tracking and persistence. A seeded synthetic source makes the contracts reviewable without private data — and the recorded run counts are execution and contract checks, not accuracy claims.
databricks-genie-deployment-kit Semantic analytics managed as code: room configuration, semantic metadata, SQL examples, benchmark questions, deployment scripts and operating playbooks as reviewable assets, with a public-data Olist example. Durable repository-native dashboard evidence is still pending, because the published dashboard is not anonymously accessible.
speechanalytics-databricks-pipeline A 16-stage contract-first speech-analytics design with per-call failure isolation and guards against raw transcript text reaching analytical outputs. No recorded successful Databricks pipeline execution, so it stays labelled a prototype.

How I try to make the work checkable

  • Gates before claims. A published number is enforced by a regression test, not restated from a notebook run nobody can repeat.
  • Evaluation kept out of the pipeline. Ground truth comes from a generation manifest, so it is derived rather than reverse-engineered after the fact.
  • Reproducibility as the default. Seeded synthetic data, pinned upstream revisions, locked environments, public-safe fixtures, CI on every push.
  • Limitations written down. Each repository states what has not been executed and what it does not claim. In someone else's work, that is the list I read first.

Earlier work

Three learning-stage repositories stay public because they show how the work developed, not because they stand beside the systems above — an early Spark ML notebook comparing loan-default classification approaches, and two absenteeism studies covering feature engineering and workforce modelling and model selection across departments. They have notebook outputs and written findings; they do not have the reproducible environments, tests, CI, deployment boundaries or public validation data the newer work does.


Technical portfolio — the full map, with case studies and where to inspect each proof · LinkedIn

Pinned Loading

  1. pii-reduction pii-reduction Public

    Structure-aware, multilingual PII reduction for Databricks. Every published number is a regression gate. Not an estate scanner — it reduces PII in columns you name.

    Python

  2. contact-center-new-hire-intelligence contact-center-new-hire-intelligence Public

    Databricks accelerator for new-hire ramp analytics, learning curves, forecasting, AI/BI, Genie, and evidence-grounded action planning.

    Python

  3. structure-aware-rag-databricks structure-aware-rag-databricks Public

    Structure-aware RAG on Databricks with governed Delta data, exact CFR relationship expansion, and reproducible public-data evaluation.

    Python

  4. prophet-forecasting-mlops prophet-forecasting-mlops Public

    Reproducible batch forecasting with Prophet, Optuna, MLflow, Delta, and Databricks Asset Bundles.

    Python

  5. technical-portfolio technical-portfolio Public

    Selected projects and open-source work across data, analytics, Databricks, ML systems, and applied AI.

  6. databricks-forecast-planning-cockpit databricks-forecast-planning-cockpit Public

    Frozen-protocol benchmark: Databricks ai_forecast v2 vs tuned Prophet vs seasonal naive on NYC 311 daily requests, with Delta, Asset Bundles and an AI/BI dashboard

    Python