A machine learning project based on Kaggle's House Prices: Advanced Regression Techniques competition.
The goal of this project was not simply to maximize a Kaggle leaderboard score, but to build a complete regression workflow and understand why different modelling approaches perform differently.
The project started with a simple Linear Regression model, progressed to Random Forests and hyperparameter tuning, and eventually moved toward Ridge Regression with a logarithmically transformed target variable.
The final Ridge Regression model achieved a Kaggle score of:
0.12773
This was a substantial improvement over the Random Forest baseline:
| Model | Kaggle Score |
|---|---|
| Random Forest | 0.14463 |
| Ridge Regression + log(SalePrice) | 0.12773 |
Because lower is better for the competition's metric, the final model represents an improvement of approximately 11.7% over the Random Forest.
The project uses the Kaggle House Prices: Advanced Regression Techniques dataset.
The training dataset contains:
- 1,460 houses
- 81 columns
- Numerical and categorical variables
SalePriceas the target variable
The features describe many aspects of each property, including:
- Overall quality
- Living area
- Basement size
- Garage characteristics
- Number of bathrooms
- Year built
- Neighborhood
- Exterior materials
- Property condition
- Sale conditions
The dataset contains a mixture of numerical and categorical information, making preprocessing an important part of the project.
The project was structured into separate components rather than keeping the entire workflow in one script.
House Prices - Advanced Regression Techniques/
│
├── data/
│ ├── raw/
│ │ ├── train.csv
│ │ ├── test.csv
│ │ └── data_description.txt
│ │
│ └── submissions/
│ └── submission.csv
│
├── src/
│ ├── analysis.py
│ ├── feature_engineering.py
│ ├── preprocess.py
│ ├── train.py
│ ├── train-ridge.py
│ └── evaluate.py
│
├── Experiment.md
├── README.md
└── requirements.txt
The separation of responsibilities made it easier to experiment without repeatedly rewriting the entire project.
The first stage was understanding the dataset.
I inspected:
- Number of observations and columns
- Numerical and categorical variables
- Feature distributions
- Correlations between numerical variables
- Relationships between features and
SalePrice
This revealed that the dataset contains considerable redundancy.
For example:
| Feature pair | Correlation |
|---|---|
| Garage Area / Garage Cars | 0.88 |
| Gr Liv Area / TotalSF | 0.87 |
| TotalBsmtSF / TotalSF | 0.83 |
| YearBuilt / GarageYrBlt | 0.83 |
| Gr Liv Area / TotRmsAbvGrd | 0.83 |
| TotalBsmtSF / 1stFlrSF | 0.82 |
| 1stFlrSF / TotalSF | 0.80 |
These relationships became particularly important later when experimenting with Ridge Regression.
The dataset contains both numerical and categorical variables.
A ColumnTransformer was used to apply different preprocessing pipelines to each type.
Missing numerical values were handled using median imputation.
Categorical missing values were filled using the most frequent category and then transformed using one-hot encoding.
This preprocessing was placed inside a Scikit-learn Pipeline, ensuring that preprocessing was learned from the training data and consistently applied to validation and test data.
This also prevented data leakage during cross-validation.
Several additional features were created based on the meaning of the original variables.
HouseAge = YrSold - YearBuilt
This represents how old the property was when it was sold.
YearsSinceRemodel = YrSold - YearRemodAdd
This represents how recently the property was remodelled.
GarageAge = YrSold - GarageYrBlt
TotalSF = TotalBsmtSF + 1stFlrSF + 2ndFlrSF
A combined measure was created from the different porch-area variables.
Half bathrooms were given half the weight of full bathrooms.
TotalBathrooms =
FullBath
+ 0.5 × HalfBath
+ BsmtFullBath
+ 0.5 × BsmtHalfBath
Polynomial features were also investigated, including squared versions of variables such as:
OverallQualTotalSFHouseAgeGarageCarsTotalBathrooms
Feature engineering did not immediately improve the Random Forest model, but it became useful for subsequent experimentation. It makes sense that the polynomials didn't affect the Random Forest model much, since it's questions are dividing, rather than taking account the exact differences. Instead of asking: (Square Foot < 500) It'll just ask (Square Foot < 22.36^2). which will give the same division of properties.
The first model was a basic Linear Regression model.
The results were:
| Metric | Result |
|---|---|
| MAE | £20,466 |
| RMSE | £31,295 |
| R² | 0.872 |
| Within 5% | 29.5% |
| Within 10% | 53.8% |
| Within 20% | 83.9% |
Compare with average Sale Price of all properties: $180,921 A difference of $20,466 shows there's still clear room for improvement, but there are definitely useful patterns found and used already.
This provided a useful baseline.
The model was reasonably capable of predicting house prices, but there was considerable room for improvement.
The next step was to test a more powerful non-linear model.
A RandomForestRegressor was used because house prices are unlikely to follow simple linear relationships.
The initial Random Forest substantially improved the results:
| Metric | Result |
|---|---|
| MAE | £17,395 |
| RMSE | £28,396 |
| R² | 0.895 |
| Within 5% | 41.8% |
| Within 10% | 68.2% |
| Within 20% | 86.3% |
This was a significant improvement over Linear Regression.
The result made intuitive sense: Random Forests can model complex interactions and non-linear relationships without requiring those relationships to be explicitly specified.
The engineered features were then introduced to the Random Forest model.
The result was slightly worse:
| Metric | Original Random Forest | With Feature Engineering |
|---|---|---|
| MAE | £17,395 | £17,908 |
| RMSE | £28,396 | £29,798 |
| R² | 0.895 | 0.884 |
| Within 20% | 86.3% | 86.6% |
The engineered model was therefore not considered an improvement.
However, this experiment demonstrated an important principle:
A feature that makes intuitive sense does not necessarily improve a machine learning model.
The Random Forest may already have been able to extract much of the information contained in these variables from the original features.
Instead of manually selecting Random Forest parameters, GridSearchCV was introduced.
The parameters investigated included:
max_depthn_estimatorsmin_samples_leaf
The first grid search produced:
max_depth = 30
min_samples_leaf = 2
n_estimators = 100
with a best cross-validation MAE of approximately:
£17,324
A second, more focused grid search was then performed around the promising parameters.
The second search found:
max_depth = 30
min_samples_leaf = 1
n_estimators = 150
with a CV MAE of approximately:
£17,270
The corresponding validation results were:
| Metric | Result |
|---|---|
| MAE | £17,617 |
| RMSE | £29,278 |
| R² | 0.888 |
| Within 5% | 43.8% |
| Within 10% | 68.5% |
| Within 20% | 86.3% |
The improvement from tuning was relatively small.
This suggested that further Random Forest tuning was unlikely to produce a dramatic improvement.
At this point, I began investigating the structure of the dataset rather than simply tuning the Random Forest.
Several pairs of variables had very high correlations.
For example:
GarageArea ↔ GarageCars 0.88
GrLivArea ↔ TotalSF 0.87
TotalBsmtSF ↔ TotalSF 0.83
YearBuilt ↔ GarageYrBlt 0.83
This led to an investigation into multicollinearity.
Multicollinearity is particularly relevant to linear models because highly correlated predictors can make it difficult for the model to determine which variable should receive the explanatory weight.
This was much less concerning for the Random Forest because tree-based models do not rely on the same coefficient-based representation.
I then investigated Ridge Regression.
Ridge Regression extends linear regression by applying L2 regularization, which penalizes excessively large coefficients.
This makes it particularly useful when many predictors contain overlapping information.
The first Ridge results were considerably weaker than the Random Forest:
| Metric | Result |
|---|---|
| MAE | £34,970 |
| RMSE | £52,516 |
| R² | 0.640 |
| Within 5% | 20.2% |
| Within 10% | 33.9% |
| Within 20% | 58.9% |
Removing highly correlated variables made the result slightly worse.
This was an important finding:
Ridge Regression was already handling the redundant information through regularization, so manually removing correlated variables did not improve the model.
The numerical features were then examined for skewness.
Several variables were strongly right-skewed.
Examples included:
MiscVal
PoolArea
LotArea
3SsnPorch
TotalSF
LowQualFinSF
MasVnrArea
TotalPorchSF
GrLivArea
A common approach for heavily right-skewed positive variables is a log1p transformation.
The reasoning is that a variable with a small number of extremely large values can be compressed by a logarithmic transformation, producing a distribution that is often more suitable for linear models.
This investigation became particularly important because the target variable itself, SalePrice, is strongly right-skewed.
Room for improvement: Some skewed variables like PoolArea may even include many 0's, meaning it doesn't even have a pool. In these cases it may have been better to create a binary value, like "HasPool". In this project, however, I kept it simple and put all "skewed variables" into the same category.
The most important improvement came from transforming the target variable.
Instead of training Ridge directly on:
SalePrice
the model was trained on:
log1p(SalePrice)
Predictions were then converted back to normal prices using:
expm1(prediction)
Conceptually:
SalePrice
↓
log1p()
↓
Ridge Regression
↓
predicted log-price
↓
expm1()
↓
predicted SalePrice
This was especially appropriate for the competition because its evaluation metric is based on the logarithm of the predicted and actual sale prices.
The resulting model was substantially more promising.
Cross-validation produced:
Mean CV RMSE: 0.1323
Standard deviation: 0.0166
The final Kaggle submission achieved:
This was significantly better than the Random Forest score of:
0.14463
| Approach | Kaggle Score |
|---|---|
| Linear Regression | — |
| Random Forest | 0.14463 |
| Ridge Regression + log target | 0.12773 |
Ridge Regression + log target: Mean Absolute Error: $16,783 Root Mean Squared Error: $27,421 R² Score: 0.902 Average SalePrice: $180,921 Predictions within 5%: 38.7% Predictions within 10%: 67.5% Predictions within 20%: 92.1%
Random Forest: Mean Absolute Error: $17,617 Root Mean Squared Error: $29,278 R² Score: 0.888 Average SalePrice: $180,921 Predictions within 5%: 43.8% Predictions within 10%: 68.5% Predictions within 20%: 86.3%
The final Ridge model therefore improved the Kaggle score by approximately 11.7% relative to the Random Forest.
The most important lesson was that increasing model complexity was not necessarily the best way to improve performance.
The largest improvement came from understanding the distribution of the target variable and selecting a modelling approach that matched the structure of the problem.
This project taught me considerably more than simply how to train a regression model.
- How to work with datasets containing many columns
- How to investigate relationships between variables
- How to use correlation coefficients to identify redundant information
- How to analyze skewness
- How transformations can change the usefulness of a variable
- How to build a preprocessing pipeline
- How to use
ColumnTransformer - How to combine preprocessing and models with Scikit-learn
Pipeline - How Random Forests handle non-linear relationships
- How Ridge Regression uses regularization
- Why multicollinearity matters for linear models
- Why tree-based models and linear models respond differently to correlated features
- How target transformations can dramatically affect model performance
- Mean Absolute Error
- Root Mean Squared Error
- R²
- Cross-validation
- Grid search
- Hyperparameter tuning
- Comparing validation performance with Kaggle performance
- Understanding the difference between improving a leaderboard score and improving the underlying modelling approach
The project taught me to approach machine learning experimentally:
Observe → form a hypothesis → implement it → evaluate it → understand the result → iterate.
Rather than assuming that a more complicated model must be better, I learned that understanding the data and the evaluation metric can be more valuable than simply adding complexity.
The project could be extended in several directions.
Potential next steps include:
- More systematic feature selection
- More carefully designed transformations of skewed predictors
- Lasso and Elastic Net regression
- XGBoost
- LightGBM
- Model ensembles
- More systematic cross-validation
- Residual analysis
- Feature importance and coefficient analysis
- More advanced handling of missing values
- Interaction features
- More rigorous comparison of different target transformations
These would be useful future experiments, but the current project already demonstrates a complete machine learning workflow from raw data through analysis, preprocessing, modelling, experimentation, cross-validation, and final Kaggle submission.