ML model optimization product to accelerate inference.
-
Updated
Jun 2, 2025 - Python
ML model optimization product to accelerate inference.
Benchmark llama.cpp on your own GPU/CPU boxes and compare against everyone else's — sweep runner, an sha256-keyed index of ~4M GGUF files, and a shared results database you query in plain language. Self-host it, or plug a machine into llamatoaster.com.
A simple tensorflow C++ REST API server
Inference Time Performance stats for various backbone networks.
Benchmarking the impact of compiler-level graph optimizations on NPU inference performance and memory wall bottlenecks.
Check the fastText's inference performance for OOV.
Optimising train, inference and throughput of expensive ML models
Linear and Multiple Regression with data manipulation using SQL and R functions.
To associate your repository with the inference-performance topic, visit your repo's landing page and select "manage topics."