Open-source AI Accelerator (GPU/TPU) Hardware-Software Co-Design Toolkit & Roofline Memory Tile Visualizer
-
Updated
Aug 21, 2026 - Python
Open-source AI Accelerator (GPU/TPU) Hardware-Software Co-Design Toolkit & Roofline Memory Tile Visualizer
Final Year Research & Engineering Project — A hardware-aware compression pipeline that profiles layer-by-layer CPU execution metrics. It balances model perplexity against physical latency, eliminates microarchitectural overhead, and compiles optimized mixed-precision graphs into deployable .gguf format.
Reproducible study of how tensor layout, reduction length, and pointer alignment drive vendor GEMM dispatch and latency cliffs.
To associate your repository with the deep-learning-compilers topic, visit your repo's landing page and select "manage topics."