AutoMLOps Studio is an end-to-end educational and practical platform designed to simplify the Machine Learning lifecycle. Developed by a student, for students, the project provides an intuitive interface to explore everything from data ingestion to production model monitoring, applying the best MLOps principles at every stage.
π Live Demo: the platform is deployed on Streamlit Cloud.
π For the full in-depth documentation, see
docs/DOCUMENTATION.md.
Learning MLOps often requires dealing with complex infrastructures before even understanding the core concepts. This project solves that by centralizing:
- Unified Workflow: A clear journey from data upload to deployment across multiple domains (Tabular, Vision, Reinforcement Learning).
- Visual Experimentation: Visualize the impact of hyperparameters and architectures in real-time.
- Production Concepts: Learn about Data Drift, Model Serving, and Performance Monitoring without the need to configure complex servers.
- Autonomy: Train models with automatic tracking of parameters, metrics, models, and dependencies using MLflow.
Since AutoMLOps Studio is built for learning, it natively enforces and exposes industry-standard MLOps practices:
One of the most confusing concepts for beginners is why data needs to be split multiple times. AutoMLOps enforces a strict, professional evaluation pipeline:
- Train Split: The "textbook" the model uses to learn patterns.
- Validation Split: The "practice exam". During AutoML, the system tests thousands of hyperparameter combinations and evaluates them here. Since the model is tuned based on this score, the result is artificially optimistic.
- Test Split (Global Holdout): The "final exam". This data is isolated entirely at the beginning of the pipeline. The model only sees it once, at the very end, providing an unbiased, real-world performance metric free of Data Leakage.
The platform strictly handles preprocessing (like scaling, imputation, or SMOTE) inside Scikit-Learn pipelines. This ensures that transformations are fitted only on training data and applied safely to validation/test sets, preventing future information from leaking into the training phase.
Models degrade over time. By serving the model via the FastAPI endpoint, all incoming predictions and inputs are logged to a SQLite telemetry store. The system uses these logs to simulate and detect Data Drift, showing students what happens when real-world distributions shift away from the original training data.
The entire interface is a single Streamlit application (app.py) organized into 8 navigable sections:
Overview, Data, AutoML, Reinforcement Learning, Experiments, Registry & Deploy, Monitoring, and What-If Simulator.
- Tabular Data: An Optuna-driven AutoML engine covering Classification, Regression (including NLP regression via fine-tuned Transformers, Multi-Output / Multivariate Regression, and Quantile Regression with a calibrated prediction interval), Forecasting (including LSTM/TCN deep models), Forecast Classification, Survival Analysis, Uplift Modeling, Clustering, Time-Series Clustering, Anomaly Detection (11 detectors), Density Estimation, Ranking, Multi-Label & Multi-Task, Association Rules, and Dimensionality Reduction β with presets, early stopping, and time budgets.
Survival Analysis expects two target columns (duration, event flag) and Uplift Modeling expects two as well (treatment, outcome); both are selected in the wizard as a multi-column target. Dimensionality Reduction is fitted against a label so that PCA, TruncatedSVD, LDA, NCA and PLS all stay trainable.
- Customizable Preprocessing: Choose the numeric scaler (Automatic/Standard, None, MinMax, Robust, MaxAbs, Quantile, Power/Yeo-Johnson), the imputation strategy (median, mean, most frequent, constant), the categorical encoding mode (automatic, One-Hot only, Ordinal only) with a configurable cardinality threshold, and optional Winsorization (quantile clipping) of outliers. Scaling can also be overridden per model in the Model Selection step. For text columns, the NLP Text Preprocessing panel lets you choose the vectorizer β TF-IDF, Bag-of-Words (counts), Binary Bag-of-Words, Feature Hashing, Contextual Embeddings (Sentence-Transformers) or raw-text passthrough β plus cleaning mode (standard / god mode / none), n-gram range, vocabulary size, stop-word removal with language selection, and sublinear TF scaling.
- Computer Vision: Train models for Image Classification, Image Multi-Label, Semantic Segmentation (DeepLabV3) and Image Anomaly Detection, with selectable backbones (ResNet, MobileNetV2, EfficientNet, DenseNet, VGG). Object Detection (Faster R-CNN) and Pose Estimation (Keypoint R-CNN) are still constructible in
CVAutoMLTrainer, but their training path is not implemented and now raisesNotImplementedErrorinstead of reporting a fake zero-loss run, so the Vision Studio no longer offers them. - Reinforcement Learning: A complete module for training agents with PPO, DQN, A2C, SAC, and TD3 (stable-baselines3) on 9 Gymnasium environments (CartPole, LunarLander, BipedalWalker, β¦), plus offline RL via d3rlpy, custom Gymnasium environment upload, training wrappers, Optuna hyperparameter tuning, and live reward visualization.
- MLflow Tracking: Every experiment run is automatically tracked β configurations, hyperparameters, metrics, and model artifacts (including
rl_config.json/rl_config.yamlfor RL runs). Optional remote tracking via DagsHub. - Job Manager: Background training jobs run in separate subprocesses with pause/resume/cancel control, so heavy training never locks the UI.
- Data Lake: A versioned, filesystem-backed catalog for datasets and RL agent trajectories (states, actions, rewards, terminal signals).
- SHAP Explainability: Model explanations generated via SHAP for trained pipelines.
- Whitebox Notebook Generation: The winning AutoML pipeline is automatically exported as a reproducible Jupyter notebook (logged to MLflow as an artifact).
- 5 Pillars of ML: Every model receives an anatomical profile across the 5 Pillars (see below).
- FastAPI Serving: Production-ready API (
api.py) for real-time inference, protected by an API key (x-api-keyheader), with liveness/readiness health endpoints. - Live Telemetry: Every
/predictcall is logged to a SQLite telemetry database for drift and performance analysis. - Drift Monitoring: Statistical drift detection (KS test for numeric features, chi-squared for categorical) plus Deepchecks data-integrity reports.
- What-If Simulator: Interactively test models against hypothetical inputs.
- Deployment Options: Push models to the Hugging Face Hub, or export a standalone API bundle (FastAPI app + requirements + Dockerfile) as a zip.
AutoMLOps Studio adopts the formal taxonomy from the Machine Learning Mind Map, distinguishing between:
- Genuine Task Types (Output Structure): Strictly defined by the mathematical format of the generated target data (e.g., discrete class label, continuous scalar value, time-to-event pair ((T, E)), counterfactual vector).
- Business Objectives / Applications (Cross-Paradigm): Operational business needs that can be addressed through multiple statistical paradigms and different underlying models.
Practical Example - Anomaly Detection (
anomaly_detection): It is a Business Objective, not a rigid Task Type. The platform supports solving it through multiple mathematical avenues across tabular, time-series, and NLP data (11 built-in detectors, plus supervised classification when labels exist):
- Spatial Isolation:
IsolationForest(random feature splits).- Local Density:
LocalOutlierFactor((k)-NN density comparison).- Gaussian Statistical Envelope:
EllipticEnvelope(Mahalanobis Distance).- Support Boundary:
OneClassSVM(Support hyperplane in Hilbert space).- Statistical Tests:
zscore_detector(classic Z-Score) andmodified_zscore(robust Median/MAD).- Correlation-Aware Distance:
mahalanobis(empirical or robust MCD covariance).- Histogram Density:
hbos(Histogram-Based Outlier Score).- Distance-Based:
knn_outlier(average distance to (k) nearest neighbors).- Subspace Reconstruction:
pca_residual(PCA reconstruction error).- Temporal Residuals:
rolling_residual(rolling-median baseline + MAD) for time series.- Supervised Classification: Via
classificationwhen historical anomaly labels exist ((y \in {0, 1})); when labels are provided to the detectors, semi-supervised F1 scoring is used during optimization.
| Modality | Type / Objective | Classification | Brief Description | Main Metrics |
|---|---|---|---|---|
| Tabular | classification |
Task Type | Predict a discrete class label (Binary or Multiclass). Also covers Time-Series Classification when Sequential data is enabled (chronological splits + TimeSeriesSplit validation). | accuracy, f1, precision, recall, roc_auc |
| Tabular | regression |
Task Type | Predict a continuous numeric target (Gaussian, Poisson, Gamma GLMs). Supports NLP regression via fine-tuned Transformers (BERT/DistilBERT regression heads) on text columns. | r2, rmse, mae, poisson_deviance, gamma_deviance |
| Tabular | multi_regression |
Task Type | Multi-Output (Multivariate) Regression: predict several continuous targets simultaneously β the full regression catalog wrapped with MultiOutputRegressor, with averaged and per-output metrics. |
r2, rmse, mae, mape |
| Tabular | survival_analysis |
Task Type | Predict time-to-event with right-censoring ((T, E)). | c_index (Concordance Index) |
| Tabular | uplift_modeling |
Task Type | Estimate Individual Treatment Effect (ITE / Causal Inference) via S/T-Learners. | qini_score / AUUC |
| Tabular | forecast |
Objective | Predict future values from historical temporal data (Lags, Rolling, PyTorch LSTM/TCN). | r2, rmse, mae, mape |
| Tabular | forecast_classification |
Task Type | Predict the future class/state of a sequence (e.g. up/down, regime change) using lag features derived from the categorical target and chronological validation. | accuracy, f1, precision, recall, roc_auc |
| Tabular | anomaly_detection |
Objective | Detect outliers or rare abnormal patterns with 11 detectors: IsolationForest, LOF, EllipticEnvelope, OneClassSVM, Z-Score, Modified Z-Score (MAD), Mahalanobis, HBOS, KNN distance, PCA residual, and Rolling Residual (time series). | decision_score, f1 |
| Tabular | clustering |
Task Type | Group samples by similarity without labels. | silhouette |
| Tabular | ts_clustering |
Task Type | Time-Series Clustering: segments a series into sliding windows (configurable size/step), extracts summary features (mean, std, min, max, median, skew, trend) and clusters temporal regimes with the standard clustering algorithms. | silhouette |
| Tabular | density_estimation |
Task Type | Model the probability density of the data via Kernel Density Estimation, Gaussian Mixture models, or Histograms. Works on tabular, time-series, and vectorized text (latent SVD projection for high-dimensional NLP features). | log_likelihood |
| Tabular | ranking |
Task Type | Score items for ordered relevance. | ndcg |
| Tabular | multi_label |
Task Type | Predict multiple labels per row (multi-target). | f1_micro, subset_accuracy |
| Tabular | multi_task |
Task Type | Predict multiple disparate classification targets concurrently. | f1_micro, subset_accuracy |
| Tabular | association_rules |
Task Type | Discover co-occurrence rules via a custom pairwise rule miner (support / confidence / lift). | rule_score, lift |
| Tabular | dimensionality_reduction |
Task Type | Reduce feature space (PCA, TruncatedSVD, LDA, NCA, PLS). | explained variance / reconstruction |
| Computer Vision | image_classification |
Task Type | Assign one class to each image. | val_acc, val_loss |
| Computer Vision | image_multi_label |
Task Type | Assign multiple labels to each image. | val_acc, val_loss |
| Computer Vision | image_segmentation |
Task Type | Pixel-wise semantic segmentation. | val_score, val_loss |
| Computer Vision | object_detection |
Task Type | Detect objects and bounding boxes. | Benchmark metrics |
| Computer Vision | pose_estimation |
Task Type | Estimate keypoints/body joints. | Keypoint accuracy |
| Reinforcement Learning | rl_agent |
Task Type | Train an agent to maximize reward via PPO, DQN, A2C, SAC, or TD3. | episode_reward, mean_reward |
Every model trained in AutoMLOps Studio undergoes an anatomical profile analysis based on the 5 Pillars of ML:
- Pillar 1 (Structure / Skeleton): White-box (Linear/Tree/GLM) vs Black-box (Ensemble/Neural Nets).
- Pillar 2 (Signal Source): Supervised (SL), GLM (Poisson/Gamma), Censored (Survival), Counterfactual (Uplift), Reward (RL), Self-Supervised.
- Pillar 3 (Criterion / Loss & Assumed Distribution): Mathematical loss derived from the distribution family (Gaussian, Bernoulli, Poisson, Gamma, Cox Likelihood, Qini).
- Pillar 4 (Regularization): Explicit (L1, L2, ElasticNet, tree max depth) vs Implicit (SGD optimizer bias).
- Pillar 5 (Optimizer Engine): L-BFGS, Adam, SGD, Optuna TPE, Tree Splitter.
- Forecast Classification (
forecast_classification): Predict future classes/states of sequences (e.g. up/down movement, regime labels) β the full classification model catalog with lag features derived from the categorical target, chronological holdouts, and TimeSeriesSplit validation. - Time-Series Classification: Enabling Sequential data on a
classificationtask now automatically applies chronological splits and temporal cross-validation. - TS Clustering (
ts_clustering): Window-based time-series clustering β configurable window size/step, 7 summary features per series (mean, std, min, max, median, skew, trend), clustered with K-Means, DBSCAN, GMM, and friends. - Density Estimation (
density_estimation): KDE, Gaussian Mixture, and Histogram density models optimized by held-out log-likelihood β for tabular data, time series, and NLP (TF-IDF features are projected to a compact latent space automatically). - Extended Anomaly Detection: 11 detectors now available (IsolationForest, LOF, EllipticEnvelope, OneClassSVM, Z-Score, Modified Z-Score/MAD, Mahalanobis (incl. robust MCD), HBOS, KNN distance, PCA residual, and Rolling Residual for time series), with semi-supervised F1 optimization when labels exist.
- NLP Regression: Regression targets on text data via fine-tuned Transformer heads (BERT / DistilBERT
-regvariants). - Multi-Output Regression (
multi_regression): Predict multiple continuous targets at once β every regressor in the catalog (Ridge, Random Forest, XGBoost, LightGBM, SVR, KNN, MLP, β¦) is automatically wrapped with sklearn'sMultiOutputRegressor; reports averagedr2/rmse/mae/mapeplus a per-output RΒ² breakdown. - Preprocessing Customization: User-selectable scaling (Auto/Standard, None, MinMax, Robust, MaxAbs, Quantile, Power), imputation strategy, categorical encoding mode, one-hot cardinality threshold, and Winsorization β plus per-model scaling overrides in the Model Selection step.
- NLP Preprocessing Customization: Selectable text vectorization β TF-IDF, Bag-of-Words (raw counts), Binary Bag-of-Words, Feature Hashing (fixed-width, for huge vocabularies), Contextual Embeddings (Sentence-Transformers), or raw-text passthrough for Transformer models β with cleaning modes (standard, god mode, none), n-gram ranges (1,1 β 2,2), vocabulary size control, multilingual stop-word removal, and sublinear TF scaling.
- Temporal & Text Characteristics: Tabular datasets support "Contains Temporal Data" (automatic chronological validation splits plus lag/rolling-window features) and "Contains Text / NLP Data" (vectorization of text columns with the chosen strategy).
- Forecast Task Type: A dedicated forecasting engine (including LSTM and TCN models implemented in pure PyTorch) integrated across all frameworks.
- Multi-Task Classification: Predict multiple target columns concurrently; the interface automatically orchestrates separate training runs per target when needed.
- Semi-Supervised Learning: Self-Training classification for targets containing unlabeled samples (
-1orNaN), dynamically wrapping base classifiers in aSelfTrainingClassifier.
automlops-studio/
βββ app.py # Entire Streamlit GUI (design system, 8 sections, wizard pipeline)
βββ api.py # FastAPI model-serving API (API-key protected, SQLite telemetry)
βββ automl_engine.py # Compatibility facade re-exporting engines for api.py / tests
βββ debug_manager.py # Manual debug script exercising the Job Manager
βββ electron-main.js # Electron desktop wrapper (spawns Streamlit, embeds it)
βββ electron-preload.js # Electron preload (exposes desktop API via contextBridge)
βββ src/
β βββ core/ # Data processor, orchestrator, data lake, drift,
β β # API-bundle exporter, whitebox notebook generator
β βββ engines/ # ML engines: classical AutoML, computer vision,
β β # reinforcement learning, stability analysis,
β β # PyTorch LSTM/TCN forecasters
β βββ tracking/ # Job Manager (subprocess workers), MLflow tracking, telemetry
β βββ deploy/ # Hugging Face Hub deployment helpers
β βββ utils/ # SHAP explainers, 5-Pillars profiles, model cards
βββ tests/ # pytest suite (run with `pytest -q tests/`)
βββ data_lake/ # Versioned datasets, RL trajectories, telemetry DB
βββ mlruns/ # MLflow artifacts (metadata in sqlite:///mlflow.db)
βββ models/ # Saved pipelines / RL agents
βββ .streamlit/config.toml # Streamlit server config (headless, port 8501)
βββ Dockerfile # Python 3.13-slim image serving the app on port 7860
βββ docker-compose.yml # 3-service stack: api, dashboard, mlflow
βββ requirements.txt # Fully pinned dependency environment
Note: the GUI lives entirely in
app.py(a single-file Streamlit app).
The platform targets Python 3.13 (the Dockerfile and CI both use 3.13; the included devcontainer uses a 3.11 image). The dependency set is heavy (PyTorch and friends β several GB), so installation may take a while.
- Install the dependencies:
pip install -r requirements.txt- Create your environment file (mandatory):
copy .env.example .env # Windows (use `cp` on Linux/macOS)Then edit .env. Variables actually used by the code:
| Variable | Required | Description |
|---|---|---|
API_SECRET_KEY |
Yes | Secret key for the serving API. api.py refuses to start without it; clients send it in the x-api-key header. |
MLFLOW_TRACKING_URI |
No | Defaults to sqlite:///mlflow.db (local SQLite backend with artifacts in mlruns/). |
MLFLOW_TRACKING_USERNAME |
No | Username for remote MLflow tracking (e.g. DagsHub). |
MLFLOW_TRACKING_PASSWORD |
No | Password/token for remote MLflow tracking. |
Note:
LOG_LEVEL,MODEL_REGISTRY_PATH, andDATA_LAKE_PATHwere removed from.env.examplebecause they are not consumed by the code (see docs/DOCUMENTATION.md Β§3.3).
The repository ships a .devcontainer/devcontainer.json based on the Python 3.11 dev container image. It installs the project dependencies automatically on build and launches the Streamlit GUI on port 8501 when the container attaches, so opening the repo in GitHub Codespaces (or VS Code Dev Containers) yields a running app with no local setup.
python -m streamlit run app.pyOpens the GUI at http://localhost:8501.
python api.py
# or
uvicorn api:app --host 0.0.0.0 --port 8000Serves the newest pipeline in models/ at http://localhost:8000 (/, /health/live, /health/ready, /predict). Requires API_SECRET_KEY in .env.
Because the tracking backend is SQLite, pass the backend store explicitly:
mlflow ui --backend-store-uri sqlite:///mlflow.db --default-artifact-root mlruns --port 5000docker compose up --buildSpins up three services (all read .env, so complete the setup step first):
- Dashboard (Streamlit):
http://localhost:8501 - Serving API (FastAPI):
http://localhost:8000 - MLflow UI:
http://localhost:5000
The project Dockerfile follows the HF Spaces convention: the container serves the Streamlit app on port 7860, which is how the live Space is deployed. Pushing this repo to a Space (Docker SDK) is all that is required.
An Electron wrapper bundles the app as a native desktop application:
npm install
npm start # runs the app inside Electron
npm run dist # builds installers via electron-builder (NSIS / AppImage / DMG)Installers for Windows, Linux, and macOS are produced automatically by CI (see below).
The project ships with a pytest suite covering the engines, tracking, data lake, API, and drift logic:
pytest -q tests/Three GitHub Actions workflows keep the project healthy:
- CI (
ci.yml): on every push/PR β installs dependencies on Python 3.13 and runspytest -q tests/. - Build Desktop App (
build-electron.yml): on pushes tomain/master(and manually) β builds Electron installers on Windows, macOS, and Ubuntu and uploads them as artifacts. - Release (
release.yml): on av*tag push (and manually) β builds the same three installers and publishes them as a GitHub Release with notes taken fromgit log. The installers bundle the Electron shell and app source only; they launchpython -m streamlit run app.py, so Python withrequirements.txtinstalled must already be available (avenv/next to the app, or system Python).
This project is released under the MIT License β see the LICENSE file for details.
Copyright (c) 2026 Pedro Morato Lahoz.
Developed by Pedro Morato Lahoz.