You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Mixed-precision quantization scheme (16/8/4bit mixed quantization) for the Wan2.2-Animate-14B model. Compresses the original 35GB base model to 17GB, balancing inference performance and model size.
Final Year Research & Engineering Project — A hardware-aware compression pipeline that profiles layer-by-layer CPU execution metrics. It balances model perplexity against physical latency, eliminates microarchitectural overhead, and compiles optimized mixed-precision graphs into deployable .gguf format.