Skip to content
Navigation Menu
Sign in
Appearance settings
Platform
AI CODE CREATION
GitHub Copilot
Write better code with AI
GitHub Copilot app
Direct agents from issue to merge
MCP Registry
Integrate external tools
DEVELOPER WORKFLOWS
Actions
Automate any workflow
Codespaces
Instant dev environments
Issues
Plan and track work
Code Review
Manage code changes
Code Quality
Enforce quality at merge
APPLICATION SECURITY
GitHub Advanced Security
Find and fix vulnerabilities
Code security
Secure your code as you build
Secret protection
Stop leaks before they start
EXPLORE
Why GitHub
Documentation
Blog
Changelog
Marketplace
View all features
Solutions
BY COMPANY SIZE
Enterprises
Small and medium teams
Startups
Nonprofits
BY USE CASE
App Modernization
DevSecOps
DevOps
CI/CD
View all use cases
BY INDUSTRY
Healthcare
Financial services
Manufacturing
Government
View all industries
View all solutions
Resources
EXPLORE BY TOPIC
AI
Software Development
DevOps
Security
View all topics
EXPLORE BY TYPE
Customer stories
Events & webinars
Ebooks & reports
Business insights
GitHub Skills
SUPPORT & SERVICES
Documentation
Customer support
Community forum
Trust center
Partners
View all resources
Open Source
COMMUNITY
GitHub Sponsors
Fund open source developers
PROGRAMS
Security Lab
Maintainer Community
GitHub Stars
Archive Program
REPOSITORIES
Topics
Trending
Collections
Enterprise
ENTERPRISE SOLUTIONS
Enterprise platform
AI-powered developer platform
AVAILABLE ADD-ONS
GitHub Advanced Security
Enterprise-grade security features
Copilot for Business
Enterprise-grade AI features
Premium Support
Enterprise-grade 24/7 support
Pricing
Search
/
Sign in
Sign up
Appearance settings
You signed in with another tab or window.
Reload
to refresh your session.
You signed out in another tab or window.
Reload
to refresh your session.
You switched accounts on another tab or window.
Reload
to refresh your session.
Dismiss alert
{{ message }}
NVIDIA
/
Personal-AI-Router
Public
Notifications
You must be signed in to change notification settings
Fork
254
Star
1.5k
Code
Issues
45
Pull requests
44
Discussions
Actions
Projects
Security and quality
0
Insights
Additional navigation options
Code
Issues
Pull requests
Discussions
Actions
Projects
Security and quality
Insights
Add managed llama.cpp engine with Windows CUDA support
- #116
#116
Open
nv-pgoode
wants to merge 13 commits into
NVIDIA:develop
NVIDIA/Personal-AI-Router:develop
from
nv-pgoode:codex/llamacpp-core-minimal
nv-pgoode/Personal-AI-Router:codex/llamacpp-core-minimal
Copy head branch name to clipboard
Conversation
Commits
13
(13)
Checks
Files changed
Open
Add managed llama.cpp engine with Windows CUDA support
#116
nv-pgoode
wants to merge 13 commits into
NVIDIA:develop
NVIDIA/Personal-AI-Router:develop
from
nv-pgoode:codex/llamacpp-core-minimal
nv-pgoode/Personal-AI-Router:codex/llamacpp-core-minimal
Copy head branch name to clipboard
Commits
Commits on Sep 22, 2026
Add managed llama.cpp as a third engine across the services
Show description for a48d357
3 people
committed
a48d357
View commit details
Copy full SHA for a48d357
Browse repository at this point
Walk ancestors in the Unix owned-path guard
Show description for 85fb8b4
pgoode41
and
claude
committed
85fb8b4
View commit details
Copy full SHA for 85fb8b4
Browse repository at this point
Add llama.cpp to the desktop
Show description for ffc7d5d
3 people
committed
ffc7d5d
View commit details
Copy full SHA for ffc7d5d
Browse repository at this point
Document llama.cpp and teach the tooling about it
Show description for 7a178ae
3 people
committed
7a178ae
View commit details
Copy full SHA for 7a178ae
Browse repository at this point
Focus workload display and repair engine event consumers
Show description for 56e837d
pgoode41
committed
56e837d
View commit details
Copy full SHA for 56e837d
Browse repository at this point
Defer explicit workload cancellation from the llama.cpp integration
Show description for dd958ba
pgoode41
committed
dd958ba
View commit details
Copy full SHA for dd958ba
Browse repository at this point
Include embedded engine manifests in the service build fingerprint
Show description for 66eb109
pgoode41
and
claude
committed
66eb109
View commit details
Copy full SHA for 66eb109
Browse repository at this point
Repair llama.cpp integration findings from the whole-engine assessment
Show description for 21873ca
pgoode41
and
claude
committed
21873ca
View commit details
Copy full SHA for 21873ca
Browse repository at this point
Route llama.cpp like the other engines and load models on demand
Show description for 931647a
pgoode41
and
claude
committed
931647a
View commit details
Copy full SHA for 931647a
Browse repository at this point
Finish the routing wording in the desktop docs
Show description for 172c8bd
pgoode41
and
claude
committed
172c8bd
View commit details
Copy full SHA for 172c8bd
Browse repository at this point
Preserve live engine status across delayed startup snapshots
Show description for f305c3b
pgoode41
committed
f305c3b
View commit details
Copy full SHA for f305c3b
Browse repository at this point
Install the official CUDA llama.cpp build on Windows x64 NVIDIA GPUs
Show description for 26271a0
pgoode41
committed
26271a0
View commit details
Copy full SHA for 26271a0
Browse repository at this point
Commits on Sep 23, 2026
test: respect Unix cache environment key case
Show description for 81624e8
pgoode41
committed
81624e8
View commit details
Copy full SHA for 81624e8
Browse repository at this point
You can’t perform that action at this time.