Predicts stellar class (GALAXY / QSO / STAR) from photometric and
spectroscopic features, evaluated on balanced accuracy (macro recall).
Final result: 0.9683 OOF balanced accuracy (0.968 on the public leaderboard), up from a 0.9640 single-model LightGBM baseline.
Sky position (alpha/delta — right ascension and declination) turned
out to carry a strong, non-obvious signal: an ablation test showed that
removing it entirely cost ~1.4 points of balanced accuracy. This
dataset has class-dense regions of sky (a synthetic-data artifact of
how the underlying survey fields were likely generated), and most of
this project's improvement over a naive baseline comes from properly
exploiting that:
- Converting RA/Dec to cartesian unit-sphere coordinates
(
sin/costransform) so models correctly see the 0°/360° wraparound in right ascension, which raw-degree splits mishandle. - Computing leakage-free K-nearest-neighbor class-density features — for each point, what fraction of its k nearest sky-neighbors belong to each class (k = 25, 100, 500). A single coordinate split can't efficiently approximate a dense local region; this feature hands the model that information directly.
.
├── README.md
├── requirements.txt
├── run_pipeline.py <- orchestrates the full pipeline end-to-end
├── index.html
├── data/
│ ├── train.csv
│ ├── test.csv
│ └── sample_submission.csv
├── src/
│ ├── config.py <- all paths, hyperparameters, CV settings
│ ├── features.py <- feature engineering + KNNSpatialFeaturizer
│ ├── checkpointing.py <- per-fold checkpoint/resume utility
│ ├── pseudo_labeling.py <- pseudo-label selection + pool augmentation
│ ├── train_lightgbm.py <- base model 1
│ ├── train_xgboost.py <- base model 2
│ ├── train_catboost.py <- base model 3 (most diverse, best single model)
│ ├── generate_pseudo_labels.py <- stage-2 pseudo-label generation
│ └── ensemble.py <- final blend + submission.csv
└── outputs/ <- OOF/test prediction arrays + submission.csv
pip install -r requirements.txtPlace train.csv and test.csv in data/
Run everything end to end:
python run_pipeline.pyThis takes roughly 2-3 hours on a modest machine (each of the 4
model-training stages does 5-fold CV with per-fold KNN feature
construction over 500K-700K rows). Every fold is checkpointed to
checkpoints/<model>_<tag>/ as it completes — if the run is
interrupted, just re-run it; finished folds load from disk instantly
instead of retraining.
To run stages individually:
cd src
# Stage 1: base models on original data
python train_lightgbm.py --tag v2
python train_xgboost.py --tag v2
# Stage 2: pseudo-labels from the stage-1 blend
python generate_pseudo_labels.py \
--inputs ../outputs/lightgbm_v2_oof_test.npz ../outputs/xgboost_v2_oof_test.npz
# Stage 3: retrain on original + pseudo-labeled data
python train_lightgbm.py --pseudo-labels ../outputs/pseudo_labels.npz --tag v3
python train_catboost.py --pseudo-labels ../outputs/pseudo_labels.npz --tag v3
# Stage 4: final blend
python ensemble.py- LightGBM is the workhorse — fast, native categorical support, used to validate every feature engineering idea before scaling up.
- XGBoost, trained on identical features/folds, agreed with LightGBM on 99.4% of predictions — i.e. almost no ensemble value on its own (confirmed: a stacked meta-learner on top of these two actually hurt OOF score, a classic sign of two models with nothing complementary to offer each other).
- CatBoost (ordered boosting + symmetric trees, structurally different from gradient-boosted decision trees in both) agreed with LightGBM on only 97.9% of predictions — genuine diversity. It turned out to be both the strongest single model and the model that improved the blend the most.
- Pseudo-labeling: ~194K test rows at ≥0.99 prediction confidence (verified ~99.7% accurate via an OOF proxy check) were added back into training, a ~33% increase in training data, for a modest but real accuracy gain.
The final submission blends CatBoost + LightGBM (both trained on the pseudo-labeled pool), equal weight. A weight-optimization search over all four models scored marginally higher on the full OOF set but lower on a held-out half-split — a sign of overfitting the weight search to OOF noise — so the simpler, more robust equal-weight blend of the two strongest/most-different models was kept instead.