Skip to content

Latest commit

 

History

11 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Stellar Classification — Kaggle Playground Series S6E6

Predicts stellar class (GALAXY / QSO / STAR) from photometric and spectroscopic features, evaluated on balanced accuracy (macro recall).

Final result: 0.9683 OOF balanced accuracy (0.968 on the public leaderboard), up from a 0.9640 single-model LightGBM baseline.

The key insight

Sky position (alpha/delta — right ascension and declination) turned out to carry a strong, non-obvious signal: an ablation test showed that removing it entirely cost ~1.4 points of balanced accuracy. This dataset has class-dense regions of sky (a synthetic-data artifact of how the underlying survey fields were likely generated), and most of this project's improvement over a naive baseline comes from properly exploiting that:

  1. Converting RA/Dec to cartesian unit-sphere coordinates (sin/cos transform) so models correctly see the 0°/360° wraparound in right ascension, which raw-degree splits mishandle.
  2. Computing leakage-free K-nearest-neighbor class-density features — for each point, what fraction of its k nearest sky-neighbors belong to each class (k = 25, 100, 500). A single coordinate split can't efficiently approximate a dense local region; this feature hands the model that information directly.

Project structure

.
├── README.md                  
├── requirements.txt
├── run_pipeline.py             <- orchestrates the full pipeline end-to-end
├── index.html             
├── data/
│   ├── train.csv
│   ├── test.csv
│   └── sample_submission.csv
├── src/
│   ├── config.py                <- all paths, hyperparameters, CV settings
│   ├── features.py              <- feature engineering + KNNSpatialFeaturizer
│   ├── checkpointing.py          <- per-fold checkpoint/resume utility
│   ├── pseudo_labeling.py        <- pseudo-label selection + pool augmentation
│   ├── train_lightgbm.py         <- base model 1
│   ├── train_xgboost.py          <- base model 2
│   ├── train_catboost.py         <- base model 3 (most diverse, best single model)
│   ├── generate_pseudo_labels.py <- stage-2 pseudo-label generation
│   └── ensemble.py               <- final blend + submission.csv
└── outputs/                     <- OOF/test prediction arrays + submission.csv

Setup

pip install -r requirements.txt

Place train.csv and test.csv in data/

Running

Run everything end to end:

python run_pipeline.py

This takes roughly 2-3 hours on a modest machine (each of the 4 model-training stages does 5-fold CV with per-fold KNN feature construction over 500K-700K rows). Every fold is checkpointed to checkpoints/<model>_<tag>/ as it completes — if the run is interrupted, just re-run it; finished folds load from disk instantly instead of retraining.

To run stages individually:

cd src

# Stage 1: base models on original data
python train_lightgbm.py --tag v2
python train_xgboost.py --tag v2

# Stage 2: pseudo-labels from the stage-1 blend
python generate_pseudo_labels.py \
    --inputs ../outputs/lightgbm_v2_oof_test.npz ../outputs/xgboost_v2_oof_test.npz

# Stage 3: retrain on original + pseudo-labeled data
python train_lightgbm.py --pseudo-labels ../outputs/pseudo_labels.npz --tag v3
python train_catboost.py --pseudo-labels ../outputs/pseudo_labels.npz --tag v3

# Stage 4: final blend
python ensemble.py

Why these four models, and why this blend

  • LightGBM is the workhorse — fast, native categorical support, used to validate every feature engineering idea before scaling up.
  • XGBoost, trained on identical features/folds, agreed with LightGBM on 99.4% of predictions — i.e. almost no ensemble value on its own (confirmed: a stacked meta-learner on top of these two actually hurt OOF score, a classic sign of two models with nothing complementary to offer each other).
  • CatBoost (ordered boosting + symmetric trees, structurally different from gradient-boosted decision trees in both) agreed with LightGBM on only 97.9% of predictions — genuine diversity. It turned out to be both the strongest single model and the model that improved the blend the most.
  • Pseudo-labeling: ~194K test rows at ≥0.99 prediction confidence (verified ~99.7% accurate via an OOF proxy check) were added back into training, a ~33% increase in training data, for a modest but real accuracy gain.

The final submission blends CatBoost + LightGBM (both trained on the pseudo-labeled pool), equal weight. A weight-optimization search over all four models scored marginally higher on the full OOF set but lower on a held-out half-split — a sign of overfitting the weight search to OOF noise — so the simpler, more robust equal-weight blend of the two strongest/most-different models was kept instead.

About

Ensemble of CatBoost + LightGBM (5-fold CV, class-weighted, pseudo-labeled with ~194K high-confidence test rows) using engineered color indices, redshift transforms, and leak-free KNN spatial-density features on cartesian sky coordinates; CV balanced accuracy 0.96833.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Contributors

Languages