Artifact-based PyTorch LLM training system for synthetic data generation, GPT/MoE experiments, distributed training, evaluation, and reproducible exports.
LLM-Light is a research codebase for building and validating small language-model training systems end to end. It covers dataset ingestion, synthetic Python data generation, tokenizer training, packed autoregressive shards, dense and mixture-of-experts GPT models, distributed training, checkpoint resume, evaluation, inference, TensorBoard observability, and compact artifact export.
The current models are intentionally budget-sized research artifacts. They are useful for architecture and training-system experiments, not general-purpose assistants or production code models.
The main validation target is TinyPython: a synthetic task-to-function Python
dataset generated by this project and published as
BertilBraun/TinyPython.
The latest documented work is an aggregated architecture sweep over dense GPT,
classic MoE GPT, and modern MoE GPT variants.
The aggregated sweep contains 42 selected runs grouped into 29 experiment families. Repeated experiments are reported as mean +/- population standard deviation so equivalent reruns can be compared directly.
Highlights from the current sweep:
- Best final pass rate:
python_modern_moe_deep10_vocab2000_aux020at 88.14%. - Dense 10M active-parameter control:
python_modern_dense_active10m_vocab2000at 86.86%. - Best sub-2M-active-parameter model:
python_modern_moe_vocab2000_aux020at 69.37%. - Vocab 2000 improved the small-model direction compared with vocab 4000.
- Modern GPT blocks substantially outperformed the earlier dense and MoE baselines.
- Top-k=2 routing improved balance but did not justify the extra active compute in the current 4-expert setup.
- The deep 10M-active MoE was the strongest model, but throughput was much lower than the dense control, so it is not automatically the best default tradeoff.
Full results and plots are documented in docs/PYTHON_MODEL_SWEEP_RESULTS.md. The experiment rationale is documented in docs/PYTHON_MODEL_SWEEP_THREE_RATIONALE.md.
- docs/TRAINING.md: setup, tests, smoke runs, Python MoE reproduction commands, generation, evaluation, and bundle export.
- docs/ARCHITECTURE.md: pipeline stages, artifacts, configuration surface, datasets, models, evaluation, distributed behavior, and implementation limits.
- docs/PYTHON_MODEL_SWEEP_RESULTS.md: aggregated Python architecture sweep results and plots.
- docs/PYTHON_MODEL_SWEEP_THREE_RATIONALE.md: rationale for the focused third sweep.
- docs/RESULTS.md: earlier validated experiment summaries.
- docs/ORCHESTRATION.md: artifact-store execution, local subprocess jobs, and async evaluation architecture.
- docs/EXTENDING.md: guides for adding evaluators, dataset sources, and model architectures.
- docs/PROJECT_PLAN.md: future work and open project gaps.
- docs/CODING_STANDARDS.md: local coding and testing standards.
The ordered pipeline supports:
raw_dataset
processed_dataset
tokenizer
packed_dataset
pretraining
post_training
evaluation
Implemented capabilities include:
- Inline, local text, and Hugging Face dataset ingestion.
- Synthetic TinyPython corpus generation with local vLLM teacher models, semantic task seeds, parsing, Python validation, resumable JSONL output, and invalid-sample tracking.
- Ordered preprocessing with Unicode normalization, line-ending normalization, length filters, exact deduplication, split assignment, and Python function extraction.
- Parallel preprocessing over workers and split-sharded text artifacts.
- Character tokenizer, Python byte-level BPE, and Rust-backed byte-level BPE.
- Packed fixed-length autoregressive datasets backed by shard files and indexes.
- Dense GPT, modern dense GPT, classic MoE GPT, and modern top-k MoE GPT models.
- Causal language-model pretraining with checkpoint resume.
- Single-process training and distributed data-parallel training through
torchrun. - Full checkpoints and rank-local sharded checkpoints for distributed runs.
- Greedy and sampled generation through naive or KV-cache inference engines.
- Exact reproduction, perplexity, fixed-prompt generation, and Python completion evaluation with parse, execution, and check-pass metrics.
- Training-time evaluation, throughput metrics, structured artifacts, and TensorBoard traces.
- Direct preference optimization utilities and generated Python DPO data flow.
- Compact run-bundle export for moving completed experiments without committing large run directories.
Fully sharded data parallel training and model-parallel variants are still future scale-out work, not validated project results.
Install dependencies with uv:
uv sync --extra devRun tests:
uv run python -m pytestRun the one-sentence smoke pipeline:
uv run python -m llm_lite.scripts.run_plan \
--config configs/verify_one_sentence.yamlGenerate from a completed run:
uv run python -m llm_lite.scripts.generate \
--config configs/verify_one_sentence.yaml \
--prompt "" \
--maximum-new-tokens 20For a fresh GPU instance, use the training helper to generate and run the TinyPython model sweep pilot:
bash scripts/train.shTo run the pilot and then the full sweep, reusing completed pilot artifacts:
SWEEP_MODE=pilot_then_full bash scripts/train.shSee docs/TRAINING.md for full commands.
Run directories under runs/ contain datasets, tokenizer files, packed shards,
checkpoints, metrics, TensorBoard logs, and evaluation reports. Large artifacts
and checkpoints should stay out of source control.
Completed runs can be exported into compact bundles:
uv run python -m llm_lite.scripts.run_plan \
--config configs/python_moe_full.yaml \
--from export \
--to exportThe aggregated sweep release bundle uses experiment names rather than artifact
hashes, with suffixes such as _1 and _2 for repeated runs. The release should
include:
runs/: selected run artifacts grouped by experiment name.release_manifest.json: mapping from release names back to source artifacts.docs/PYTHON_MODEL_SWEEP_RESULTS.md: aggregated report.docs/images/python_model_sweep/: generated comparison plots.
The next high-signal architecture experiment is router balancing for MoE. The current auxiliary-loss approach improved some runs but still left severe expert dominance in others. A useful follow-up is to compare the existing auxiliary loss against DeepSeek-style auxiliary-loss-free routing, where an expert-wise routing bias is updated from observed load outside backpropagation.
Other open directions:
- Test router-bias balancing on the fast sub-2M-active modern MoE before scaling deeper models.
- Revisit larger MoE only if the routing and throughput tradeoff improves.
- Investigate whether shared experts or a larger expert count help specialization more than top-k=2 with only four experts.
- Validate FSDP or model-parallel training before treating larger-scale runs as a supported project path.