A gaze-target estimation project built on the GazeFollow dataset.
python -m venv .venv
.venv\Scripts\activate # Windows
pip install -r requirements.txtpython -m src.train
python -m src.evaluate| Run | Model | Test mean distance | Test min distance | Test MSE |
|---|---|---|---|---|
| Constant guess (image center) | none | 0.317 | n/a | n/a |
| BASE01 | 4-layer CNN, global avg pooling | 0.253 | n/a | 0.0404 |
| E1A | ResNet18 scene + head mask, heatmap output | 0.211 | 0.137 | 0.0401 |
| E1B | E1A + head-crop pathway | 0.187 | 0.114 | 0.0311 |
- Single seed (42) per run. Seeds 1 and 2 for E1B are running; add "mean ± std over 3 seeds" here once they finish.
- Distance is computed per annotation row against normalized (x, y) in [0, 1]. This is not the standard GazeFollow protocol (which compares against the average of annotators), so these numbers are not comparable to published results.
- E1A has a similar MSE to the baseline but a lower mean distance, because the heatmap model sometimes lands on a wrong region, which gives larger errors.
- Training data is only 3,824 images (the HF copy contains just the GazeFollow test split, which we split by image into train/val/test).
Setup:
pip install -r requirements.txt
python download_gazefollow.py
python src/utils/prepare_gazefollow.pyRun (about 1.5 min/epoch on a Colab T4 GPU, about 45 min for 30 epochs):
python -m src.train_spatial --run-id E1A --variant scene_mask --epochs 30
python -m src.train_spatial --run-id E1B --variant two_pathway --epochs 30Every run appends a row to experiments/log.csv. Checkpoints are saved to outputs/<run-id>/best_model.pth and are not committed to Git.
src/data/gazefollow_cached.py— GPU-resident data pipeline (one-time cache, GPU augmentation)src/models/gaze_spatial.py—scene_maskandtwo_pathwaymodels with soft-argmax heatmap outputsrc/train_spatial.py— training, evaluation, and logging