139.7M-parameter model trained from scratch for multilingual text-to-SQL and MCP-style tool calling. English, Hindi and Hinglish.
-
Updated
Sep 13, 2026 - Python
139.7M-parameter model trained from scratch for multilingual text-to-SQL and MCP-style tool calling. English, Hindi and Hinglish.
SageGPT: A ~7.2M parameter Sanskrit-only decoder-only Transformer SLM trained from scratch on a 139 MB tokenized Sanskrit corpus containing 72.8M SentencePiece model-token IDs, derived from a 105.2M-character purified Sanskrit text corpus using NVIDIA DGX Spark. Architecture: 6 layers, 8 attention heads, 256 embedding dim, 1024 context, 8K vocab.
SageGPT 7.25M param SLM trained from scratch on 56.89M Sanskrit tokens using Apple MLX. 4-layer decoder-only Transformer with 8K vocabulary for Apple Silicon inference.
To associate your repository with the trained-from-scratch topic, visit your repo's landing page and select "manage topics."