Skip to content

About

AutoMLOps Studio is an "end-to-end" educational platform designed to simplify the Machine Learning lifecycle. Developed by a student, for students, the project provides an intuitive interface to explore everything from data ingestion to production model monitoring.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

Β 

History

217 Commits

Folders and files

NameName
Last commit message
Last commit date
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 

Repository files navigation

πŸš€ AutoMLOps Studio

Comprehensive Automated Machine Learning & MLOps Platform

Version Python 3.13 License: MIT Live Demo Docker MLflow Streamlit

AutoMLOps Studio is an end-to-end educational and practical platform designed to simplify the Machine Learning lifecycle. Developed by a student, for students, the project provides an intuitive interface to explore everything from data ingestion to production model monitoring, applying the best MLOps principles at every stage.

πŸ”— Live Demo: the platform is deployed on Streamlit Cloud.

πŸ“– For the full in-depth documentation, see docs/DOCUMENTATION.md.


🎯 Objective & Problem Statement

Learning MLOps often requires dealing with complex infrastructures before even understanding the core concepts. This project solves that by centralizing:

  • Unified Workflow: A clear journey from data upload to deployment across multiple domains (Tabular, Vision, Reinforcement Learning).
  • Visual Experimentation: Visualize the impact of hyperparameters and architectures in real-time.
  • Production Concepts: Learn about Data Drift, Model Serving, and Performance Monitoring without the need to configure complex servers.
  • Autonomy: Train models with automatic tracking of parameters, metrics, models, and dependencies using MLflow.

🧠 Educational Concepts Built-In

Since AutoMLOps Studio is built for learning, it natively enforces and exposes industry-standard MLOps practices:

1. The Tri-Split Rule (Train / Validation / Test)

One of the most confusing concepts for beginners is why data needs to be split multiple times. AutoMLOps enforces a strict, professional evaluation pipeline:

  • Train Split: The "textbook" the model uses to learn patterns.
  • Validation Split: The "practice exam". During AutoML, the system tests thousands of hyperparameter combinations and evaluates them here. Since the model is tuned based on this score, the result is artificially optimistic.
  • Test Split (Global Holdout): The "final exam". This data is isolated entirely at the beginning of the pipeline. The model only sees it once, at the very end, providing an unbiased, real-world performance metric free of Data Leakage.

2. Preventing Data Leakage

The platform strictly handles preprocessing (like scaling, imputation, or SMOTE) inside Scikit-Learn pipelines. This ensures that transformations are fitted only on training data and applied safely to validation/test sets, preventing future information from leaking into the training phase.

3. Model Telemetry & Data Drift

Models degrade over time. By serving the model via the FastAPI endpoint, all incoming predictions and inputs are logged to a SQLite telemetry store. The system uses these logs to simulate and detect Data Drift, showing students what happens when real-world distributions shift away from the original training data.


🌟 Key Features

1. πŸ–₯️ A Unified GUI with 8 Sections

The entire interface is a single Streamlit application (app.py) organized into 8 navigable sections: Overview, Data, AutoML, Reinforcement Learning, Experiments, Registry & Deploy, Monitoring, and What-If Simulator.

2. πŸ€– Multi-Domain Machine Learning

  • Tabular Data: An Optuna-driven AutoML engine covering Classification, Regression (including NLP regression via fine-tuned Transformers, Multi-Output / Multivariate Regression, and Quantile Regression with a calibrated prediction interval), Forecasting (including LSTM/TCN deep models), Forecast Classification, Survival Analysis, Uplift Modeling, Clustering, Time-Series Clustering, Anomaly Detection (11 detectors), Density Estimation, Ranking, Multi-Label & Multi-Task, Association Rules, and Dimensionality Reduction β€” with presets, early stopping, and time budgets.

    Survival Analysis expects two target columns (duration, event flag) and Uplift Modeling expects two as well (treatment, outcome); both are selected in the wizard as a multi-column target. Dimensionality Reduction is fitted against a label so that PCA, TruncatedSVD, LDA, NCA and PLS all stay trainable.

  • Customizable Preprocessing: Choose the numeric scaler (Automatic/Standard, None, MinMax, Robust, MaxAbs, Quantile, Power/Yeo-Johnson), the imputation strategy (median, mean, most frequent, constant), the categorical encoding mode (automatic, One-Hot only, Ordinal only) with a configurable cardinality threshold, and optional Winsorization (quantile clipping) of outliers. Scaling can also be overridden per model in the Model Selection step. For text columns, the NLP Text Preprocessing panel lets you choose the vectorizer β€” TF-IDF, Bag-of-Words (counts), Binary Bag-of-Words, Feature Hashing, Contextual Embeddings (Sentence-Transformers) or raw-text passthrough β€” plus cleaning mode (standard / god mode / none), n-gram range, vocabulary size, stop-word removal with language selection, and sublinear TF scaling.
  • Computer Vision: Train models for Image Classification, Image Multi-Label, Semantic Segmentation (DeepLabV3) and Image Anomaly Detection, with selectable backbones (ResNet, MobileNetV2, EfficientNet, DenseNet, VGG). Object Detection (Faster R-CNN) and Pose Estimation (Keypoint R-CNN) are still constructible in CVAutoMLTrainer, but their training path is not implemented and now raises NotImplementedError instead of reporting a fake zero-loss run, so the Vision Studio no longer offers them.
  • Reinforcement Learning: A complete module for training agents with PPO, DQN, A2C, SAC, and TD3 (stable-baselines3) on 9 Gymnasium environments (CartPole, LunarLander, BipedalWalker, …), plus offline RL via d3rlpy, custom Gymnasium environment upload, training wrappers, Optuna hyperparameter tuning, and live reward visualization.

3. πŸ§ͺ Experiments & MLOps Integration

  • MLflow Tracking: Every experiment run is automatically tracked β€” configurations, hyperparameters, metrics, and model artifacts (including rl_config.json/rl_config.yaml for RL runs). Optional remote tracking via DagsHub.
  • Job Manager: Background training jobs run in separate subprocesses with pause/resume/cancel control, so heavy training never locks the UI.
  • Data Lake: A versioned, filesystem-backed catalog for datasets and RL agent trajectories (states, actions, rewards, terminal signals).
  • SHAP Explainability: Model explanations generated via SHAP for trained pipelines.
  • Whitebox Notebook Generation: The winning AutoML pipeline is automatically exported as a reproducible Jupyter notebook (logged to MLflow as an artifact).
  • 5 Pillars of ML: Every model receives an anatomical profile across the 5 Pillars (see below).

4. πŸ“‘ Serving, Monitoring & Deployment

  • FastAPI Serving: Production-ready API (api.py) for real-time inference, protected by an API key (x-api-key header), with liveness/readiness health endpoints.
  • Live Telemetry: Every /predict call is logged to a SQLite telemetry database for drift and performance analysis.
  • Drift Monitoring: Statistical drift detection (KS test for numeric features, chi-squared for categorical) plus Deepchecks data-integrity reports.
  • What-If Simulator: Interactively test models against hypothetical inputs.
  • Deployment Options: Push models to the Hugging Face Hub, or export a standalone API bundle (FastAPI app + requirements + Dockerfile) as a zip.

πŸ“‹ Supported Task Types & Business Objectives

AutoMLOps Studio adopts the formal taxonomy from the Machine Learning Mind Map, distinguishing between:

  1. Genuine Task Types (Output Structure): Strictly defined by the mathematical format of the generated target data (e.g., discrete class label, continuous scalar value, time-to-event pair ((T, E)), counterfactual vector).
  2. Business Objectives / Applications (Cross-Paradigm): Operational business needs that can be addressed through multiple statistical paradigms and different underlying models.

Practical Example - Anomaly Detection (anomaly_detection): It is a Business Objective, not a rigid Task Type. The platform supports solving it through multiple mathematical avenues across tabular, time-series, and NLP data (11 built-in detectors, plus supervised classification when labels exist):

  • Spatial Isolation: IsolationForest (random feature splits).
  • Local Density: LocalOutlierFactor ((k)-NN density comparison).
  • Gaussian Statistical Envelope: EllipticEnvelope (Mahalanobis Distance).
  • Support Boundary: OneClassSVM (Support hyperplane in Hilbert space).
  • Statistical Tests: zscore_detector (classic Z-Score) and modified_zscore (robust Median/MAD).
  • Correlation-Aware Distance: mahalanobis (empirical or robust MCD covariance).
  • Histogram Density: hbos (Histogram-Based Outlier Score).
  • Distance-Based: knn_outlier (average distance to (k) nearest neighbors).
  • Subspace Reconstruction: pca_residual (PCA reconstruction error).
  • Temporal Residuals: rolling_residual (rolling-median baseline + MAD) for time series.
  • Supervised Classification: Via classification when historical anomaly labels exist ((y \in {0, 1})); when labels are provided to the detectors, semi-supervised F1 scoring is used during optimization.
Modality Type / Objective Classification Brief Description Main Metrics
Tabular classification Task Type Predict a discrete class label (Binary or Multiclass). Also covers Time-Series Classification when Sequential data is enabled (chronological splits + TimeSeriesSplit validation). accuracy, f1, precision, recall, roc_auc
Tabular regression Task Type Predict a continuous numeric target (Gaussian, Poisson, Gamma GLMs). Supports NLP regression via fine-tuned Transformers (BERT/DistilBERT regression heads) on text columns. r2, rmse, mae, poisson_deviance, gamma_deviance
Tabular multi_regression Task Type Multi-Output (Multivariate) Regression: predict several continuous targets simultaneously β€” the full regression catalog wrapped with MultiOutputRegressor, with averaged and per-output metrics. r2, rmse, mae, mape
Tabular survival_analysis Task Type Predict time-to-event with right-censoring ((T, E)). c_index (Concordance Index)
Tabular uplift_modeling Task Type Estimate Individual Treatment Effect (ITE / Causal Inference) via S/T-Learners. qini_score / AUUC
Tabular forecast Objective Predict future values from historical temporal data (Lags, Rolling, PyTorch LSTM/TCN). r2, rmse, mae, mape
Tabular forecast_classification Task Type Predict the future class/state of a sequence (e.g. up/down, regime change) using lag features derived from the categorical target and chronological validation. accuracy, f1, precision, recall, roc_auc
Tabular anomaly_detection Objective Detect outliers or rare abnormal patterns with 11 detectors: IsolationForest, LOF, EllipticEnvelope, OneClassSVM, Z-Score, Modified Z-Score (MAD), Mahalanobis, HBOS, KNN distance, PCA residual, and Rolling Residual (time series). decision_score, f1
Tabular clustering Task Type Group samples by similarity without labels. silhouette
Tabular ts_clustering Task Type Time-Series Clustering: segments a series into sliding windows (configurable size/step), extracts summary features (mean, std, min, max, median, skew, trend) and clusters temporal regimes with the standard clustering algorithms. silhouette
Tabular density_estimation Task Type Model the probability density of the data via Kernel Density Estimation, Gaussian Mixture models, or Histograms. Works on tabular, time-series, and vectorized text (latent SVD projection for high-dimensional NLP features). log_likelihood
Tabular ranking Task Type Score items for ordered relevance. ndcg
Tabular multi_label Task Type Predict multiple labels per row (multi-target). f1_micro, subset_accuracy
Tabular multi_task Task Type Predict multiple disparate classification targets concurrently. f1_micro, subset_accuracy
Tabular association_rules Task Type Discover co-occurrence rules via a custom pairwise rule miner (support / confidence / lift). rule_score, lift
Tabular dimensionality_reduction Task Type Reduce feature space (PCA, TruncatedSVD, LDA, NCA, PLS). explained variance / reconstruction
Computer Vision image_classification Task Type Assign one class to each image. val_acc, val_loss
Computer Vision image_multi_label Task Type Assign multiple labels to each image. val_acc, val_loss
Computer Vision image_segmentation Task Type Pixel-wise semantic segmentation. val_score, val_loss
Computer Vision object_detection Task Type Detect objects and bounding boxes. Benchmark metrics
Computer Vision pose_estimation Task Type Estimate keypoints/body joints. Keypoint accuracy
Reinforcement Learning rl_agent Task Type Train an agent to maximize reward via PPO, DQN, A2C, SAC, or TD3. episode_reward, mean_reward

πŸ›οΈ Architecture: The 5 Pillars of ML

Every model trained in AutoMLOps Studio undergoes an anatomical profile analysis based on the 5 Pillars of ML:

  1. Pillar 1 (Structure / Skeleton): White-box (Linear/Tree/GLM) vs Black-box (Ensemble/Neural Nets).
  2. Pillar 2 (Signal Source): Supervised (SL), GLM (Poisson/Gamma), Censored (Survival), Counterfactual (Uplift), Reward (RL), Self-Supervised.
  3. Pillar 3 (Criterion / Loss & Assumed Distribution): Mathematical loss derived from the distribution family (Gaussian, Bernoulli, Poisson, Gamma, Cox Likelihood, Qini).
  4. Pillar 4 (Regularization): Explicit (L1, L2, ElasticNet, tree max depth) vs Implicit (SGD optimizer bias).
  5. Pillar 5 (Optimizer Engine): L-BFGS, Adam, SGD, Optuna TPE, Tree Splitter.

πŸ†• What's New (Recent)

  • Forecast Classification (forecast_classification): Predict future classes/states of sequences (e.g. up/down movement, regime labels) β€” the full classification model catalog with lag features derived from the categorical target, chronological holdouts, and TimeSeriesSplit validation.
  • Time-Series Classification: Enabling Sequential data on a classification task now automatically applies chronological splits and temporal cross-validation.
  • TS Clustering (ts_clustering): Window-based time-series clustering β€” configurable window size/step, 7 summary features per series (mean, std, min, max, median, skew, trend), clustered with K-Means, DBSCAN, GMM, and friends.
  • Density Estimation (density_estimation): KDE, Gaussian Mixture, and Histogram density models optimized by held-out log-likelihood β€” for tabular data, time series, and NLP (TF-IDF features are projected to a compact latent space automatically).
  • Extended Anomaly Detection: 11 detectors now available (IsolationForest, LOF, EllipticEnvelope, OneClassSVM, Z-Score, Modified Z-Score/MAD, Mahalanobis (incl. robust MCD), HBOS, KNN distance, PCA residual, and Rolling Residual for time series), with semi-supervised F1 optimization when labels exist.
  • NLP Regression: Regression targets on text data via fine-tuned Transformer heads (BERT / DistilBERT -reg variants).
  • Multi-Output Regression (multi_regression): Predict multiple continuous targets at once β€” every regressor in the catalog (Ridge, Random Forest, XGBoost, LightGBM, SVR, KNN, MLP, …) is automatically wrapped with sklearn's MultiOutputRegressor; reports averaged r2/rmse/mae/mape plus a per-output RΒ² breakdown.
  • Preprocessing Customization: User-selectable scaling (Auto/Standard, None, MinMax, Robust, MaxAbs, Quantile, Power), imputation strategy, categorical encoding mode, one-hot cardinality threshold, and Winsorization β€” plus per-model scaling overrides in the Model Selection step.
  • NLP Preprocessing Customization: Selectable text vectorization β€” TF-IDF, Bag-of-Words (raw counts), Binary Bag-of-Words, Feature Hashing (fixed-width, for huge vocabularies), Contextual Embeddings (Sentence-Transformers), or raw-text passthrough for Transformer models β€” with cleaning modes (standard, god mode, none), n-gram ranges (1,1 β†’ 2,2), vocabulary size control, multilingual stop-word removal, and sublinear TF scaling.
  • Temporal & Text Characteristics: Tabular datasets support "Contains Temporal Data" (automatic chronological validation splits plus lag/rolling-window features) and "Contains Text / NLP Data" (vectorization of text columns with the chosen strategy).
  • Forecast Task Type: A dedicated forecasting engine (including LSTM and TCN models implemented in pure PyTorch) integrated across all frameworks.
  • Multi-Task Classification: Predict multiple target columns concurrently; the interface automatically orchestrates separate training runs per target when needed.
  • Semi-Supervised Learning: Self-Training classification for targets containing unlabeled samples (-1 or NaN), dynamically wrapping base classifiers in a SelfTrainingClassifier.

πŸ“‚ Project Structure

automlops-studio/
β”œβ”€β”€ app.py                  # Entire Streamlit GUI (design system, 8 sections, wizard pipeline)
β”œβ”€β”€ api.py                  # FastAPI model-serving API (API-key protected, SQLite telemetry)
β”œβ”€β”€ automl_engine.py        # Compatibility facade re-exporting engines for api.py / tests
β”œβ”€β”€ debug_manager.py        # Manual debug script exercising the Job Manager
β”œβ”€β”€ electron-main.js        # Electron desktop wrapper (spawns Streamlit, embeds it)
β”œβ”€β”€ electron-preload.js     # Electron preload (exposes desktop API via contextBridge)
β”œβ”€β”€ src/
β”‚   β”œβ”€β”€ core/               # Data processor, orchestrator, data lake, drift,
β”‚   β”‚                       #   API-bundle exporter, whitebox notebook generator
β”‚   β”œβ”€β”€ engines/            # ML engines: classical AutoML, computer vision,
β”‚   β”‚                       #   reinforcement learning, stability analysis,
β”‚   β”‚                       #   PyTorch LSTM/TCN forecasters
β”‚   β”œβ”€β”€ tracking/           # Job Manager (subprocess workers), MLflow tracking, telemetry
β”‚   β”œβ”€β”€ deploy/             # Hugging Face Hub deployment helpers
β”‚   └── utils/              # SHAP explainers, 5-Pillars profiles, model cards
β”œβ”€β”€ tests/                  # pytest suite (run with `pytest -q tests/`)
β”œβ”€β”€ data_lake/              # Versioned datasets, RL trajectories, telemetry DB
β”œβ”€β”€ mlruns/                 # MLflow artifacts (metadata in sqlite:///mlflow.db)
β”œβ”€β”€ models/                 # Saved pipelines / RL agents
β”œβ”€β”€ .streamlit/config.toml  # Streamlit server config (headless, port 8501)
β”œβ”€β”€ Dockerfile              # Python 3.13-slim image serving the app on port 7860
β”œβ”€β”€ docker-compose.yml      # 3-service stack: api, dashboard, mlflow
└── requirements.txt        # Fully pinned dependency environment

Note: the GUI lives entirely in app.py (a single-file Streamlit app).


βš™οΈ Setup

The platform targets Python 3.13 (the Dockerfile and CI both use 3.13; the included devcontainer uses a 3.11 image). The dependency set is heavy (PyTorch and friends β€” several GB), so installation may take a while.

  1. Install the dependencies:
pip install -r requirements.txt
  1. Create your environment file (mandatory):
copy .env.example .env    # Windows (use `cp` on Linux/macOS)

Then edit .env. Variables actually used by the code:

Variable Required Description
API_SECRET_KEY Yes Secret key for the serving API. api.py refuses to start without it; clients send it in the x-api-key header.
MLFLOW_TRACKING_URI No Defaults to sqlite:///mlflow.db (local SQLite backend with artifacts in mlruns/).
MLFLOW_TRACKING_USERNAME No Username for remote MLflow tracking (e.g. DagsHub).
MLFLOW_TRACKING_PASSWORD No Password/token for remote MLflow tracking.

Note: LOG_LEVEL, MODEL_REGISTRY_PATH, and DATA_LAKE_PATH were removed from .env.example because they are not consumed by the code (see docs/DOCUMENTATION.md Β§3.3).

🐳 Dev Container (GitHub Codespaces friendly)

The repository ships a .devcontainer/devcontainer.json based on the Python 3.11 dev container image. It installs the project dependencies automatically on build and launches the Streamlit GUI on port 8501 when the container attaches, so opening the repo in GitHub Codespaces (or VS Code Dev Containers) yields a running app with no local setup.


πŸš€ How to Run

🐍 Local Streamlit Dashboard

python -m streamlit run app.py

Opens the GUI at http://localhost:8501.

πŸ“‘ Serving API

python api.py
# or
uvicorn api:app --host 0.0.0.0 --port 8000

Serves the newest pipeline in models/ at http://localhost:8000 (/, /health/live, /health/ready, /predict). Requires API_SECRET_KEY in .env.

πŸ“Š MLflow UI (optional)

Because the tracking backend is SQLite, pass the backend store explicitly:

mlflow ui --backend-store-uri sqlite:///mlflow.db --default-artifact-root mlruns --port 5000

🐳 Docker Compose (full stack)

docker compose up --build

Spins up three services (all read .env, so complete the setup step first):

  • Dashboard (Streamlit): http://localhost:8501
  • Serving API (FastAPI): http://localhost:8000
  • MLflow UI: http://localhost:5000

πŸ€— Hugging Face Spaces

The project Dockerfile follows the HF Spaces convention: the container serves the Streamlit app on port 7860, which is how the live Space is deployed. Pushing this repo to a Space (Docker SDK) is all that is required.

πŸ–₯️ Electron Desktop App

An Electron wrapper bundles the app as a native desktop application:

npm install
npm start        # runs the app inside Electron
npm run dist     # builds installers via electron-builder (NSIS / AppImage / DMG)

Installers for Windows, Linux, and macOS are produced automatically by CI (see below).


πŸ§ͺ Testing

The project ships with a pytest suite covering the engines, tracking, data lake, API, and drift logic:

pytest -q tests/

πŸ” CI/CD

Three GitHub Actions workflows keep the project healthy:

  • CI (ci.yml): on every push/PR β€” installs dependencies on Python 3.13 and runs pytest -q tests/.
  • Build Desktop App (build-electron.yml): on pushes to main/master (and manually) β€” builds Electron installers on Windows, macOS, and Ubuntu and uploads them as artifacts.
  • Release (release.yml): on a v* tag push (and manually) β€” builds the same three installers and publishes them as a GitHub Release with notes taken from git log. The installers bundle the Electron shell and app source only; they launch python -m streamlit run app.py, so Python with requirements.txt installed must already be available (a venv/ next to the app, or system Python).

πŸ“œ License

This project is released under the MIT License β€” see the LICENSE file for details.

Copyright (c) 2026 Pedro Morato Lahoz.


Developed by Pedro Morato Lahoz.

About

AutoMLOps Studio is an "end-to-end" educational platform designed to simplify the Machine Learning lifecycle. Developed by a student, for students, the project provides an intuitive interface to explore everything from data ingestion to production model monitoring.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages