A bio-inspired CV experiment — honestly reported.
Find salient image regions, simulate a Physarum polycephalum network between them, feed it to a CNN as an extra channel. See if biology helps. It doesn't, at this scale.
Results · Why It Didn't Work · Pipeline · Reproduce · Retrospective
Important
Headline result: null. Adding a slime-mold-inspired connectivity channel to a ResNet-20 on CIFAR-10 changes test accuracy by +0.10% / -0.04% / +0.12% across three variants — all within single-seed noise (~±0.3–0.5%). At 32×32 with input-channel stacking, the prior does not help. This repo documents the pipeline, the result, and what I'd change next.
| Mode | Best | Final | Δ vs RGB |
|---|---|---|---|
| RGB baseline | 0.8985 | 0.8985 | +0.00% |
| RGB + Dijkstra slime | 0.8995 | 0.8995 | +0.10% |
| RGB + Tero PDE slime | 0.8981 | 0.8981 | −0.04% |
| RGB + Dijkstra + PDE | 0.8997 | 0.8997 | +0.12% |
In 2010, Tero et al. showed that Physarum polycephalum — a slime mold with no nervous system — grows a transport network between food sources that closely matches the Tokyo rail network. Local rules: edges carrying high flow get reinforced, idle edges decay. The result is an efficient, redundant graph.
The analogy this project tested:
human attention → scans salient regions, ignores background
slime mold → builds efficient network between food sources
↓
hypothesis: saliency peaks are "cities", slime network between them
encodes spatial structure a CNN could exploit
The hypothesis was clean. The result was null. The rest of this README walks through why.
flowchart LR
A["<b>Input</b><br/>32×32×3"] --> B["<b>Spectral Residual</b><br/>Saliency<br/><sub>Hou & Zhang 2007</sub>"]
B --> C["<b>Top-5 Peaks</b><br/>Local maxima<br/>as 'Tokyo cities'"]
C --> D{"Slime<br/>method"}
D -->|fast| E["<b>Dijkstra</b><br/>weighted shortest paths<br/>between seeds"]
D -->|faithful| F["<b>Tero PDE</b><br/>pressure Laplacian +<br/>conductivity update"]
E --> G["<b>Slime map</b><br/>32×32×1"]
F --> G
G --> H["<b>Stack</b><br/>RGB + slime<br/>→ 32×32×4 or ×5"]
H --> I["<b>ResNet-20</b><br/>30 epochs, cosine LR"]
I --> J["Class<br/>prediction"]
style A fill:#1e293b,stroke:#475569,color:#f8fafc
style B fill:#1e293b,stroke:#475569,color:#f8fafc
style C fill:#1e293b,stroke:#475569,color:#f8fafc
style D fill:#fbbf24,stroke:#f59e0b,color:#0a0e1a
style E fill:#1e293b,stroke:#fbbf24,color:#f8fafc
style F fill:#1e293b,stroke:#fbbf24,color:#f8fafc
style G fill:#1e293b,stroke:#fbbf24,color:#f8fafc
style H fill:#1e293b,stroke:#475569,color:#f8fafc
style I fill:#1e293b,stroke:#475569,color:#f8fafc
style J fill:#1e293b,stroke:#475569,color:#f8fafc
The two slime implementations:
-
slime_mold_fast— saliency-weighted shortest paths between seeds viascipy.sparse.csgraph.dijkstrawith a reinforcement pass. ~8 ms/image. The pragmatic version. -
slime_mold_tero_pde— faithful Tero et al. 2010 simulation: solves the pressure Laplacian via Kirchhoff's law, then applies the nonlinear conductivity update$f(Q) = Q^\gamma / (1 + Q^\gamma)$ with$\gamma > 1$ to encourage redundant network structure. ~30–50 ms/image. The biologically motivated version.
All four configurations converge to indistinguishable test accuracy. Curves overlap past epoch 20.
Per-class deltas show no consistent pattern — if slime added signal, we'd expect coherent improvements (e.g., all rigid objects, or all animals). Instead we see noise.
| Class | Dijkstra | PDE | Both |
|---|---|---|---|
| airplane | +0.30% | −0.60% | −0.60% |
| automobile | −0.20% | −0.40% | +0.00% |
| bird | +0.00% | −1.00% | +1.90% |
| cat | −0.20% | +0.50% | −0.80% |
| deer | −0.60% | −1.00% | −0.80% |
| dog | +0.80% | −0.30% | +1.00% |
| frog | −0.80% | +0.60% | −0.50% |
| horse | +1.30% | +0.10% | +0.60% |
| ship | −0.10% | −0.10% | −0.80% |
| truck | +0.50% | +1.80% | +1.20% |
The mild bumps on truck and horse are suggestive but single-seed — almost certainly noise without 3+ seed replication.
Note
The full post-mortem lives in docs/retrospective.md. Short version:
- No topology to discover at 32×32. With 5 seeds spaced ~10 px apart, the "network" is 4–5 short line segments. Tokyo had room for Steiner-point-like junctions and redundant loops; CIFAR doesn't.
- The slime map is a deterministic function of RGB. A sufficiently expressive CNN can learn an equivalent feature internally. Hand-crafted priors mostly help when the network can't otherwise learn them — not the case here.
- Input-channel stacking is a weak mechanism. The slime map sits passively next to RGB. A principled version would gate intermediate feature maps multiplicatively — actual attention, not just an extra color.
- Single-seed runs hide nothing here. The deltas are smaller than typical seed variance. Multi-seed replication would almost certainly average them to zero.
If anyone (including future-me) wants to revive this idea, the version with a real chance:
- Higher resolution where attention demonstrably matters — CUB-200 birds, Stanford Cars, or aerial / medical imagery. Fine-grained tasks have headroom for spatial priors.
- Slime as multiplicative attention on intermediate features, not an input channel.
- Real ablations — vs. raw saliency map, vs. Gaussian blobs at seeds, vs. minimum spanning tree, vs. Voronoi. These tell you whether slime specifically matters, or whether any structured prior would do.
- Differentiable end-to-end seed selection — learn where to look rather than hand-tuning spectral residual + top-K.
- 3+ seeds per configuration, always.
slime-mold-vision/
├── assets/
│ └── banner.{svg,png}
├── notebooks/
│ └── slime_mold_vision.ipynb # Interactive walkthrough
├── src/
│ ├── saliency.py # Spectral residual + peak finding
│ ├── slime.py # Dijkstra & Tero PDE simulations
│ ├── dataset.py # CIFAR10MultiSlime wrapper
│ ├── model.py # ResNet-20 with variable input channels
│ └── train.py # Training loop + per-class eval
├── scripts/
│ ├── precompute_slime.py # Generates ./slime_cache/*.pt
│ └── run_experiment.py # End-to-end 4-way comparison
├── results/
│ ├── training_log.txt
│ ├── final_results.md
│ ├── test_accuracy.png
│ ├── per_class_accuracy.png
│ └── slime_examples.png
└── docs/
├── method.md # Math + references
└── retrospective.md # Full post-mortem
Tip
Single GPU runtime: ~1 hour 5 min on a Tesla T4 for all 4 configurations × 30 epochs.
# 1. Install dependencies
pip install -r requirements.txt
# 2. Precompute slime maps (~10 min for Dijkstra, ~30–45 min for PDE)
# Saves ~469 MB of cached .pt tensors to ./slime_cache/
python scripts/precompute_slime.py
# 3. Train all 4 configurations and write results/
python scripts/run_experiment.py --epochs 30 --batch-size 512Or open notebooks/slime_mold_vision.ipynb for the narrated version with inline visualizations.
- A. Tero et al. "Rules for Biologically Inspired Adaptive Network Design." Science 327(5964):439–442, 2010. [DOI]
- T. Nakagaki, H. Yamada, Á. Tóth. "Maze-solving by an amoeboid organism." Nature 407:470, 2000. [DOI]
- X. Hou, L. Zhang. "Saliency Detection: A Spectral Residual Approach." CVPR, 2007. [PDF]
- K. He et al. "Deep Residual Learning for Image Recognition." CVPR, 2016. [arXiv]
MIT — use freely, attribute kindly.