Skip to content

Latest commit

 

History

4 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

House Prices — Advanced Regression Techniques

A machine learning project based on Kaggle's House Prices: Advanced Regression Techniques competition.

The goal of this project was not simply to maximize a Kaggle leaderboard score, but to build a complete regression workflow and understand why different modelling approaches perform differently.

The project started with a simple Linear Regression model, progressed to Random Forests and hyperparameter tuning, and eventually moved toward Ridge Regression with a logarithmically transformed target variable.

Results

The final Ridge Regression model achieved a Kaggle score of:

0.12773

This was a substantial improvement over the Random Forest baseline:

Model Kaggle Score
Random Forest 0.14463
Ridge Regression + log(SalePrice) 0.12773

Because lower is better for the competition's metric, the final model represents an improvement of approximately 11.7% over the Random Forest.

1. Dataset

The project uses the Kaggle House Prices: Advanced Regression Techniques dataset.

The training dataset contains:

  • 1,460 houses
  • 81 columns
  • Numerical and categorical variables
  • SalePrice as the target variable

The features describe many aspects of each property, including:

  • Overall quality
  • Living area
  • Basement size
  • Garage characteristics
  • Number of bathrooms
  • Year built
  • Neighborhood
  • Exterior materials
  • Property condition
  • Sale conditions

The dataset contains a mixture of numerical and categorical information, making preprocessing an important part of the project.

2. Project Structure

The project was structured into separate components rather than keeping the entire workflow in one script.

House Prices - Advanced Regression Techniques/
│
├── data/
│   ├── raw/
│   │   ├── train.csv
│   │   ├── test.csv
│   │   └── data_description.txt
│   │
│   └── submissions/
│       └── submission.csv
│
├── src/
│   ├── analysis.py
│   ├── feature_engineering.py
│   ├── preprocess.py
│   ├── train.py
│   ├── train-ridge.py
│   └── evaluate.py
│
├── Experiment.md
├── README.md
└── requirements.txt

The separation of responsibilities made it easier to experiment without repeatedly rewriting the entire project.

3. Initial Data Analysis

The first stage was understanding the dataset.

I inspected:

  • Number of observations and columns
  • Numerical and categorical variables
  • Feature distributions
  • Correlations between numerical variables
  • Relationships between features and SalePrice

This revealed that the dataset contains considerable redundancy.

For example:

Feature pair Correlation
Garage Area / Garage Cars 0.88
Gr Liv Area / TotalSF 0.87
TotalBsmtSF / TotalSF 0.83
YearBuilt / GarageYrBlt 0.83
Gr Liv Area / TotRmsAbvGrd 0.83
TotalBsmtSF / 1stFlrSF 0.82
1stFlrSF / TotalSF 0.80

These relationships became particularly important later when experimenting with Ridge Regression.

4. Preprocessing

The dataset contains both numerical and categorical variables.

A ColumnTransformer was used to apply different preprocessing pipelines to each type.

Numerical features

Missing numerical values were handled using median imputation.

Categorical features

Categorical missing values were filled using the most frequent category and then transformed using one-hot encoding.

This preprocessing was placed inside a Scikit-learn Pipeline, ensuring that preprocessing was learned from the training data and consistently applied to validation and test data.

This also prevented data leakage during cross-validation.

5. Feature Engineering

Several additional features were created based on the meaning of the original variables.

House age

HouseAge = YrSold - YearBuilt

This represents how old the property was when it was sold.

Years since remodel

YearsSinceRemodel = YrSold - YearRemodAdd

This represents how recently the property was remodelled.

Garage age

GarageAge = YrSold - GarageYrBlt

Total square footage

TotalSF = TotalBsmtSF + 1stFlrSF + 2ndFlrSF

Total porch area

A combined measure was created from the different porch-area variables.

Total bathrooms

Half bathrooms were given half the weight of full bathrooms.

TotalBathrooms =
    FullBath
    + 0.5 × HalfBath
    + BsmtFullBath
    + 0.5 × BsmtHalfBath

Polynomial features were also investigated, including squared versions of variables such as:

  • OverallQual
  • TotalSF
  • HouseAge
  • GarageCars
  • TotalBathrooms

Feature engineering did not immediately improve the Random Forest model, but it became useful for subsequent experimentation. It makes sense that the polynomials didn't affect the Random Forest model much, since it's questions are dividing, rather than taking account the exact differences. Instead of asking: (Square Foot < 500) It'll just ask (Square Foot < 22.36^2). which will give the same division of properties.

6. Linear Regression Baseline

The first model was a basic Linear Regression model.

The results were:

Metric Result
MAE £20,466
RMSE £31,295
R² 0.872
Within 5% 29.5%
Within 10% 53.8%
Within 20% 83.9%

Compare with average Sale Price of all properties: $180,921 A difference of $20,466 shows there's still clear room for improvement, but there are definitely useful patterns found and used already.

This provided a useful baseline.

The model was reasonably capable of predicting house prices, but there was considerable room for improvement.

7. Random Forest

The next step was to test a more powerful non-linear model.

A RandomForestRegressor was used because house prices are unlikely to follow simple linear relationships.

The initial Random Forest substantially improved the results:

Metric Result
MAE £17,395
RMSE £28,396
R² 0.895
Within 5% 41.8%
Within 10% 68.2%
Within 20% 86.3%

This was a significant improvement over Linear Regression.

The result made intuitive sense: Random Forests can model complex interactions and non-linear relationships without requiring those relationships to be explicitly specified.

8. Feature Engineering Experiment

The engineered features were then introduced to the Random Forest model.

The result was slightly worse:

Metric Original Random Forest With Feature Engineering
MAE £17,395 £17,908
RMSE £28,396 £29,798
R² 0.895 0.884
Within 20% 86.3% 86.6%

The engineered model was therefore not considered an improvement.

However, this experiment demonstrated an important principle:

A feature that makes intuitive sense does not necessarily improve a machine learning model.

The Random Forest may already have been able to extract much of the information contained in these variables from the original features.

9. Grid Search and Cross-Validation

Instead of manually selecting Random Forest parameters, GridSearchCV was introduced.

The parameters investigated included:

  • max_depth
  • n_estimators
  • min_samples_leaf

The first grid search produced:

max_depth = 30
min_samples_leaf = 2
n_estimators = 100

with a best cross-validation MAE of approximately:

£17,324

A second, more focused grid search was then performed around the promising parameters.

The second search found:

max_depth = 30
min_samples_leaf = 1
n_estimators = 150

with a CV MAE of approximately:

£17,270

The corresponding validation results were:

Metric Result
MAE £17,617
RMSE £29,278
R² 0.888
Within 5% 43.8%
Within 10% 68.5%
Within 20% 86.3%

The improvement from tuning was relatively small.

This suggested that further Random Forest tuning was unlikely to produce a dramatic improvement.

10. Investigating Correlation

At this point, I began investigating the structure of the dataset rather than simply tuning the Random Forest.

Several pairs of variables had very high correlations.

For example:

GarageArea ↔ GarageCars      0.88
GrLivArea ↔ TotalSF          0.87
TotalBsmtSF ↔ TotalSF        0.83
YearBuilt ↔ GarageYrBlt      0.83

This led to an investigation into multicollinearity.

Multicollinearity is particularly relevant to linear models because highly correlated predictors can make it difficult for the model to determine which variable should receive the explanatory weight.

This was much less concerning for the Random Forest because tree-based models do not rely on the same coefficient-based representation.

11. Ridge Regression

I then investigated Ridge Regression.

Ridge Regression extends linear regression by applying L2 regularization, which penalizes excessively large coefficients.

This makes it particularly useful when many predictors contain overlapping information.

The first Ridge results were considerably weaker than the Random Forest:

Metric Result
MAE £34,970
RMSE £52,516
R² 0.640
Within 5% 20.2%
Within 10% 33.9%
Within 20% 58.9%

Removing highly correlated variables made the result slightly worse.

This was an important finding:

Ridge Regression was already handling the redundant information through regularization, so manually removing correlated variables did not improve the model.

12. Investigating Skewness

The numerical features were then examined for skewness.

Several variables were strongly right-skewed.

Examples included:

MiscVal
PoolArea
LotArea
3SsnPorch
TotalSF
LowQualFinSF
MasVnrArea
TotalPorchSF
GrLivArea

A common approach for heavily right-skewed positive variables is a log1p transformation.

The reasoning is that a variable with a small number of extremely large values can be compressed by a logarithmic transformation, producing a distribution that is often more suitable for linear models.

This investigation became particularly important because the target variable itself, SalePrice, is strongly right-skewed.

Room for improvement: Some skewed variables like PoolArea may even include many 0's, meaning it doesn't even have a pool. In these cases it may have been better to create a binary value, like "HasPool". In this project, however, I kept it simple and put all "skewed variables" into the same category.

13. Log-Transforming SalePrice

The most important improvement came from transforming the target variable.

Instead of training Ridge directly on:

SalePrice

the model was trained on:

log1p(SalePrice)

Predictions were then converted back to normal prices using:

expm1(prediction)

Conceptually:

SalePrice
    ↓
log1p()
    ↓
Ridge Regression
    ↓
predicted log-price
    ↓
expm1()
    ↓
predicted SalePrice

This was especially appropriate for the competition because its evaluation metric is based on the logarithm of the predicted and actual sale prices.

The resulting model was substantially more promising.

Cross-validation produced:

Mean CV RMSE: 0.1323
Standard deviation: 0.0166

The final Kaggle submission achieved:

0.12773

This was significantly better than the Random Forest score of:

0.14463

14. Final Comparison

Approach Kaggle Score
Linear Regression —
Random Forest 0.14463
Ridge Regression + log target 0.12773

Ridge Regression + log target: Mean Absolute Error: $16,783 Root Mean Squared Error: $27,421 R² Score: 0.902 Average SalePrice: $180,921 Predictions within 5%: 38.7% Predictions within 10%: 67.5% Predictions within 20%: 92.1%

Random Forest: Mean Absolute Error: $17,617 Root Mean Squared Error: $29,278 R² Score: 0.888 Average SalePrice: $180,921 Predictions within 5%: 43.8% Predictions within 10%: 68.5% Predictions within 20%: 86.3%

The final Ridge model therefore improved the Kaggle score by approximately 11.7% relative to the Random Forest.

The most important lesson was that increasing model complexity was not necessarily the best way to improve performance.

The largest improvement came from understanding the distribution of the target variable and selecting a modelling approach that matched the structure of the problem.

15. What I Learned

This project taught me considerably more than simply how to train a regression model.

Data analysis

  • How to work with datasets containing many columns
  • How to investigate relationships between variables
  • How to use correlation coefficients to identify redundant information
  • How to analyze skewness
  • How transformations can change the usefulness of a variable

Machine learning

  • How to build a preprocessing pipeline
  • How to use ColumnTransformer
  • How to combine preprocessing and models with Scikit-learn Pipeline
  • How Random Forests handle non-linear relationships
  • How Ridge Regression uses regularization
  • Why multicollinearity matters for linear models
  • Why tree-based models and linear models respond differently to correlated features
  • How target transformations can dramatically affect model performance

Model evaluation

  • Mean Absolute Error
  • Root Mean Squared Error
  • R²
  • Cross-validation
  • Grid search
  • Hyperparameter tuning
  • Comparing validation performance with Kaggle performance
  • Understanding the difference between improving a leaderboard score and improving the underlying modelling approach

Most importantly

The project taught me to approach machine learning experimentally:

Observe → form a hypothesis → implement it → evaluate it → understand the result → iterate.

Rather than assuming that a more complicated model must be better, I learned that understanding the data and the evaluation metric can be more valuable than simply adding complexity.

16. Future Improvements

The project could be extended in several directions.

Potential next steps include:

  • More systematic feature selection
  • More carefully designed transformations of skewed predictors
  • Lasso and Elastic Net regression
  • XGBoost
  • LightGBM
  • Model ensembles
  • More systematic cross-validation
  • Residual analysis
  • Feature importance and coefficient analysis
  • More advanced handling of missing values
  • Interaction features
  • More rigorous comparison of different target transformations

These would be useful future experiments, but the current project already demonstrates a complete machine learning workflow from raw data through analysis, preprocessing, modelling, experimentation, cross-validation, and final Kaggle submission.

Releases

Packages

Contributors

Languages