Skip to content
BertilBraunPublic

About

Artifact-based PyTorch LLM training system for synthetic data generation, GPT/MoE experiments, distributed training, evaluation, and reproducible exports.

Resources

Stars

1 star

Watchers

0 watching

Forks

Latest commit

 

History

123 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

LLM-Light

Artifact-based PyTorch LLM training system for synthetic data generation, GPT/MoE experiments, distributed training, evaluation, and reproducible exports.

LLM-Light is a research codebase for building and validating small language-model training systems end to end. It covers dataset ingestion, synthetic Python data generation, tokenizer training, packed autoregressive shards, dense and mixture-of-experts GPT models, distributed training, checkpoint resume, evaluation, inference, TensorBoard observability, and compact artifact export.

The current models are intentionally budget-sized research artifacts. They are useful for architecture and training-system experiments, not general-purpose assistants or production code models.

Current Results

The main validation target is TinyPython: a synthetic task-to-function Python dataset generated by this project and published as BertilBraun/TinyPython. The latest documented work is an aggregated architecture sweep over dense GPT, classic MoE GPT, and modern MoE GPT variants.

Python model sweep checkpoint pass rate

The aggregated sweep contains 42 selected runs grouped into 29 experiment families. Repeated experiments are reported as mean +/- population standard deviation so equivalent reruns can be compared directly.

Highlights from the current sweep:

  • Best final pass rate: python_modern_moe_deep10_vocab2000_aux020 at 88.14%.
  • Dense 10M active-parameter control: python_modern_dense_active10m_vocab2000 at 86.86%.
  • Best sub-2M-active-parameter model: python_modern_moe_vocab2000_aux020 at 69.37%.
  • Vocab 2000 improved the small-model direction compared with vocab 4000.
  • Modern GPT blocks substantially outperformed the earlier dense and MoE baselines.
  • Top-k=2 routing improved balance but did not justify the extra active compute in the current 4-expert setup.
  • The deep 10M-active MoE was the strongest model, but throughput was much lower than the dense control, so it is not automatically the best default tradeoff.

Full results and plots are documented in docs/PYTHON_MODEL_SWEEP_RESULTS.md. The experiment rationale is documented in docs/PYTHON_MODEL_SWEEP_THREE_RATIONALE.md.

Documentation Map

Implemented Surface

The ordered pipeline supports:

raw_dataset
processed_dataset
tokenizer
packed_dataset
pretraining
post_training
evaluation

Implemented capabilities include:

  • Inline, local text, and Hugging Face dataset ingestion.
  • Synthetic TinyPython corpus generation with local vLLM teacher models, semantic task seeds, parsing, Python validation, resumable JSONL output, and invalid-sample tracking.
  • Ordered preprocessing with Unicode normalization, line-ending normalization, length filters, exact deduplication, split assignment, and Python function extraction.
  • Parallel preprocessing over workers and split-sharded text artifacts.
  • Character tokenizer, Python byte-level BPE, and Rust-backed byte-level BPE.
  • Packed fixed-length autoregressive datasets backed by shard files and indexes.
  • Dense GPT, modern dense GPT, classic MoE GPT, and modern top-k MoE GPT models.
  • Causal language-model pretraining with checkpoint resume.
  • Single-process training and distributed data-parallel training through torchrun.
  • Full checkpoints and rank-local sharded checkpoints for distributed runs.
  • Greedy and sampled generation through naive or KV-cache inference engines.
  • Exact reproduction, perplexity, fixed-prompt generation, and Python completion evaluation with parse, execution, and check-pass metrics.
  • Training-time evaluation, throughput metrics, structured artifacts, and TensorBoard traces.
  • Direct preference optimization utilities and generated Python DPO data flow.
  • Compact run-bundle export for moving completed experiments without committing large run directories.

Fully sharded data parallel training and model-parallel variants are still future scale-out work, not validated project results.

Quick Start

Install dependencies with uv:

uv sync --extra dev

Run tests:

uv run python -m pytest

Run the one-sentence smoke pipeline:

uv run python -m llm_lite.scripts.run_plan \
  --config configs/verify_one_sentence.yaml

Generate from a completed run:

uv run python -m llm_lite.scripts.generate \
  --config configs/verify_one_sentence.yaml \
  --prompt "" \
  --maximum-new-tokens 20

For a fresh GPU instance, use the training helper to generate and run the TinyPython model sweep pilot:

bash scripts/train.sh

To run the pilot and then the full sweep, reusing completed pilot artifacts:

SWEEP_MODE=pilot_then_full bash scripts/train.sh

See docs/TRAINING.md for full commands.

Artifacts

Run directories under runs/ contain datasets, tokenizer files, packed shards, checkpoints, metrics, TensorBoard logs, and evaluation reports. Large artifacts and checkpoints should stay out of source control.

Completed runs can be exported into compact bundles:

uv run python -m llm_lite.scripts.run_plan \
  --config configs/python_moe_full.yaml \
  --from export \
  --to export

The aggregated sweep release bundle uses experiment names rather than artifact hashes, with suffixes such as _1 and _2 for repeated runs. The release should include:

  • runs/: selected run artifacts grouped by experiment name.
  • release_manifest.json: mapping from release names back to source artifacts.
  • docs/PYTHON_MODEL_SWEEP_RESULTS.md: aggregated report.
  • docs/images/python_model_sweep/: generated comparison plots.

Future Work

The next high-signal architecture experiment is router balancing for MoE. The current auxiliary-loss approach improved some runs but still left severe expert dominance in others. A useful follow-up is to compare the existing auxiliary loss against DeepSeek-style auxiliary-loss-free routing, where an expert-wise routing bias is updated from observed load outside backpropagation.

Other open directions:

  • Test router-bias balancing on the fast sub-2M-active modern MoE before scaling deeper models.
  • Revisit larger MoE only if the routing and throughput tradeoff improves.
  • Investigate whether shared experts or a larger expert count help specialization more than top-k=2 with only four experts.
  • Validate FSDP or model-parallel training before treating larger-scale runs as a supported project path.

About

Artifact-based PyTorch LLM training system for synthetic data generation, GPT/MoE experiments, distributed training, evaluation, and reproducible exports.

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Contributors

Languages