An interactive, desktop-native Machine Learning workbench built in Python with Tkinter, Scikit-Learn, and Matplotlib. Designed for exploratory data analysis, interactive feature selection, automated preprocessing, and multi-model benchmarking across classical regression and classification algorithms.
The studio removes boilerplate code by providing a stateful graphical interface for dataset inspection, feature matrix partitioning, train/test splitting, model training, residual scatter plotting, dynamic confusion matrix generation, and full classification reporting.
-
Reactive State Machine UI: Step-gated button states ensure rigorous pipeline ordering (
File Ingest$\to$ Feature Selection$\to$ Data Splitting$\to$ Model Training$\to$ Metric Inspection). - Interactive Dual-Axis Table Viewer: Custom virtualized canvas rendering tabular CSV data with horizontal and vertical scrollbars, zebra striping, and auto-sizing.
-
Dynamic Feature Matrix Partitioning:
- Independent variable (
$X$ ) multi-selection dialog withSelect All,Select None, and column checkbox toggles. - Dependent target variable (
$y$ ) radio-selection modal enforcing strict mutual exclusion ($y \notin X$ ).
- Independent variable (
-
Integrated Machine Learning Algorithms:
- Ordinary Least Squares (OLS) Linear Regression: Intercept & coefficient inspector, Matplotlib residual scatter plots, MAE/MSE/RMSE metrics.
-
Logistic Regression (
liblinear): Multiclass and binary classification, label-encoded target restoration, dynamic confusion matrix, classification reports. - Decision Tree Classifier (CART): Non-linear hierarchical feature splitting, Gini impurity optimization.
- Random Forest Classifier: Ensemble bagging over 100 decorrelated decision trees.
-
Embedded Analytical Visualizations: Native Matplotlib integration (
FigureCanvasTkAgg) rendering interactive scatter plots and dynamically generated metrics grids. -
Bundled Sample Datasets: Includes standard benchmark datasets (
samples/iris.csvfor classification andsamples/housing.csvfor regression) for instant out-of-the-box experimentation.
flowchart TD
A["CSV Dataset Ingest<br/>(e.g., 'samples/iris.csv')"] --> B["Data Validation & In-Memory Loading<br/>(Pandas DataFrame)"]
B --> C["Table Viewer Window<br/>(Scrollable Virtualized Canvas)"]
B --> D["Feature Partitioning Dialogs"]
D --> D1["SelectionX Modal<br/>Multi-select Independent Features (X)"]
D --> D2["SelectionY Modal<br/>Single-select Target Variable (y)"]
D1 & D2 --> E["Train / Test Splitting<br/>• Configurable Test Ratio α ∈ (0, 1)<br/>• Deterministic Random Seed<br/>• Automatic LabelEncoder for Categorical y"]
E --> F{"Select ML Algorithm"}
subgraph Regression_Flow ["Regression Subsystem"]
F --> G["Linear Regression (OLS)"]
G --> G1["Coefficients Modal<br/>β₀ Intercept & βᵢ Weights"]
G --> G2["Scatter Plot (Matplotlib)<br/>y_test vs. y_pred"]
G --> G3["Regression Error Modal<br/>MAE, MSE, RMSE"]
end
subgraph Classification_Flow ["Classification Subsystem"]
F --> H["Logistic Regression<br/>(liblinear solver)"]
F --> I["Decision Tree Classifier<br/>(Gini Impurity)"]
F --> J["Random Forest Classifier<br/>(100 Estimators)"]
H & I & J --> K1["Dynamic Confusion Matrix<br/>Predicted vs. Actual Table"]
H & I & J --> K2["Classification Report<br/>Precision, Recall, F1, Support"]
H & I & J --> K3["Classification Error Metrics<br/>MAE, MSE, RMSE"]
end
Given feature matrix
Predictions
For classification with liblinear solver, which optimizes the regularized negative log-likelihood:
$$\min{\mathbf{w}} \frac{1}{2} \mathbf{w}^T \mathbf{w} + C \sum_{i=1}^n \log\left(1 + e^{-y_i \mathbf{w}^T \mathbf{x}_i}\right)$$
Recursive binary partitioning splits feature space at node
Ensemble bagging constructs
The studio automatically calculates class-wise and macro/weighted aggregate metrics:
ml-algorithm-studio/
├── GUI.py # Main studio application & GUI state controller
│ ├── MachineLearning # Primary window, state gating, and model orchestrator
│ ├── Table # Virtualized dual-scrollbar tabular dataset viewer
│ ├── SelectionX # Checkbox-driven independent feature modal
│ ├── SelectionY # Radiobutton-driven target variable modal
│ ├── ConfusionMatrix # Styled multi-class confusion matrix grid window
│ ├── Errors # Residual error metrics inspector (MAE, MSE, RMSE)
│ ├── ClassificationReport # Formatted precision, recall, F1, and support table
│ ├── Scatter # Matplotlib canvas rendering y_test vs. y_pred
│ └── Coefficients # Intercept and per-feature coefficient inspector
├── samples/ # Bundled validation datasets
│ ├── iris.csv # Multi-class classification (Setosa, Versicolor, Virginica)
│ └── housing.csv # Multivariate regression (Area, Bedrooms, Bathrooms, Price)
├── requirements.txt # Project dependencies (pandas, scikit-learn, matplotlib, numpy)
├── .gitignore # Python cache & IDE exclusion rules
├── py.ico # Application window icon
└── LICENSE # MIT License
- Python 3.8 or higher.
- Tkinter installed (standard with standard Python distributions on Windows/macOS; on Linux:
sudo apt-get install python3-tk).
git clone https://github.com/FortunateSpy5/ml-algorithm-studio.git
cd ml-algorithm-studio
pip install -r requirements.txtpython GUI.py- Load File: In the
File Nameentry, entersamples/iris.csvand click Select. - Preview Data: Click Show to open the scrollable table view displaying all rows and column headers (
sepal_length,sepal_width,petal_length,petal_width,species). - Partition Features:
- Click X: Click Select All, uncheck
species, and click Confirm. - Click y: Select
speciesand click Confirm.
- Click X: Click Select All, uncheck
- Split Dataset: Set
Test Sizeto0.20, setRandom Stateto42, and click Split. - Train & Evaluate:
- Under Random Forest, click Predict.
- Click Confusion Matrix to view the actual vs. predicted distribution matrix.
- Click Classification Report to view class-level Precision, Recall, and F1 scores.
- Load File: In the
File Nameentry, entersamples/housing.csvand click Select. - Partition Features:
- Click X: Select
area,bedrooms,bathrooms,stories, andparking, then click Confirm. - Click y: Select
priceand click Confirm.
- Click X: Select
- Split Dataset: Set
Test Sizeto0.25, setRandom Stateto101, and click Split. - Train & Evaluate:
- Under Linear Regression, click Predict.
- Click Coefficients to inspect the learned baseline bias and feature weights.
- Click Scatter Plot to open the interactive Matplotlib window comparing actual test prices vs. model predictions.
- Click Error to review MAE, MSE, and RMSE.
This project is licensed under the MIT License - see the LICENSE file for details.