A lightweight backend router for Gemini API traffic. It exposes OpenAI-, DeepSeek-, and Ollama-compatible HTTP surfaces while routing requests across a pool of Gemini API keys with automatic fallback, local quota tracking, and resilient error handling - plus direct routes to a local Ollama for embeddings and vision.
The optional Codex provider uses your authenticated account for text inference:
gpt-5.5, gpt-5.6-luna, gpt-5.6-sol, gpt-6-astra, with model-specific thinking.
Enable it per app; on exhausted quota it can fall back to that app's authorized Gemini chain.
The dashboard supports two isolated saved logins and manual account selection,
with private aliases, public quota bars and per-account measured request tokens. No MCP,
personal-chat wake, browser controller or coding tools. See Codex.
It is designed to squeeze the maximum useful throughput out of free-tier Gemini accounts: pool many accounts, always try the strongest model first, and gracefully scale down a fallback chain that ends on very high-quota models - so most traffic is served at zero API cost.
GemRouter makes a fleet of free Gemini accounts behave like one large, reliable, OpenAI-compatible endpoint - without a paid plan and without external billing/monitoring services.
- Multiplied free quota. Daily quota is per account. Pool N free accounts and the usable
daily capacity multiplies by N. Example with 7 accounts:
gemini-2.5-flash~140 req/day,gemini-3.1-flash-lite~3,500/day,gemma-4-31b-it/gemma-4-26b-a4b-it~9,000/day each → ~21,000+ text requests/day at $0. - Strongest-first, graceful degradation. Each request tries the best model first and only falls back down an ordered chain (flagship flash → flash-lite → gemma) when a model is exhausted/overloaded. You always get the best model you can serve for free, and you keep serving traffic instead of failing when the top models run out.
- No wasted quota on transient failures. 429s roll back the local counter so a rejected request doesn't burn quota; an escalating backoff (1 min → 5 min → daily) parks genuinely depleted model+account pairs; a model-wide 503 "high demand" is parked for 30 s instead of hammering every key. Net effect: fewer burnt requests, higher effective free throughput.
- No empty/failed billable retries. An empty completion is never returned as success - it is retried with a larger budget, then falls to the next model. Clients don't pay (in tokens or quota) for blank answers, and don't need their own retry glue.
- Local embeddings & vision = $0 and offloaded.
bge-m3(embeddings) and a vision model run on your own Ollama box via dedicated/v1/embeddingsand/v1/visionendpoints - no per-token embedding/vision API bill, and they run as separate flows that never touch the Gemini quota or queue. - No external cost to operate. Quota is tracked entirely in-process (a local JSON ledger);
there is no Cloud Monitoring, Service Usage, or
gclouddependency to pay for or maintain. - Drop-in OpenAI compatibility. Point any OpenAI SDK at GemRouter's base URL and keep your code unchanged - you swap a paid endpoint for a free-tier pool with one config line.
- Operate without redeploys. Add/remove accounts, set priorities, reorder the routed model
set, and mint client keys live from the admin UI - no
.envedits, no restarts, no downtime.
- Multi-account routing across many Gemini API keys (AI Studio free-tier and Tier 1), with capacity-aware rotation (equal-priority accounts with more headroom are preferred).
- Local quota ledger - RPM / TPM / RPD tracked per quota-group + model; RPD resets at midnight America/Los_Angeles (Google's real boundary). No Google Cloud calls.
- Resilient fallback - ordered model chain, strongest→weakest; escalating 429 backoff; 503 high-demand handling; empty-completion retry/fallback; cross-backend fallback to Ollama.
- Local Ollama integration - dedicated
/v1/embeddings(bge-m3) and/v1/vision(e.g. minicpm) routes, direct to a local server, with no Gemini fallback and their own daily counters. - Compatibility surfaces - OpenAI, DeepSeek, and Ollama-compatible endpoints from one backend.
- Operator admin UI at
/admin- live quota, account manager (add/remove/priority/enable, per-account model discovery), routed-model editor, app/key management (with custom key prefixes), interaction telemetry filterable by app, and an outbound-proxy manager. - Optional Codex inference - account login, explicit app permission, thinking, quota-only Gemini fallback, buffered SSE and measured token accounting.
| Endpoint | Auth | Purpose |
|---|---|---|
POST /v1/chat/completions, /chat/completions |
client key | Gemini, NVIDIA or an authorized Codex model |
POST /v1/responses |
client key | Responses-compatible text inference |
POST /v1/embeddings, /embeddings |
client key | Embeddings via local Ollama (bge-m3) |
POST /v1/vision, /vision |
client key | Vision via local Ollama (separate flow, no queue) |
GET /health, GET /dashboard/summary |
none | Health and guest-safe quota/stats |
GET /admin, POST /admin/* |
admin | Operator dashboard and management APIs |
- Node
24.18.0(see.nvmrc) pnpm10.26.1
pnpm install --frozen-lockfile
cp .env.example .env
# edit .env - set at minimum GEMROUTER_ADMIN_TOKEN, GEMROUTER_BOOTSTRAP_API_KEY, GEMROUTER_GEMINI_API_KEYS
pnpm buildpnpm start
# or for production
./scripts/start-gemrouter.shDefault port: 4024. Health check:
curl -fsS http://127.0.0.1:4024/healthcurl -sS http://127.0.0.1:4024/v1/chat/completions \
-H "Authorization: Bearer $GEMROUTER_BOOTSTRAP_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gemini-2.5-flash",
"messages": [{ "role": "user", "content": "Reply only with OK." }]
}'Use it from the OpenAI SDK by pointing base_url at http://<host>:4024/v1.
| Guide | Contents |
|---|---|
| Setup | Installation, configuration, first run |
| Configuration reference | All environment variables |
| Routing and quota | Multi-key routing, fallback, cooldowns, local Ollama, endpoints |
| Architecture and workflow | Startup, model discovery, app policy, routing, quota lifecycle, production checklist |
| Operations | Deployment, systemd, security, live admin management, troubleshooting |
| Codex provider | Login, models/thinking, quota, fallback, usage and verification |
Third-party attribution for design references is recorded in THIRD_PARTY_NOTICES.md.
Account metadata/keys live in data/gemini-api-accounts.json (gitignored; see
docs/gemini-api-accounts.example.json for the format).
Licensing scope and preserved third-party permissions are documented in LICENSING.md. The 0xfunboy Non-Commercial License covers eligible original material only.
The failed widget-wake microtest and preceding browser-controller experiment are
archived in archive/widget-wake-probe-2026-09-16 (12af0c3), including code,
tests and chronological evidence, without private profiles or credentials.
See the archive pointer.
The active checkout uses Codex as an inference provider. Read-only account tools:
pnpm --silent codex:account status|login|usage|models.
See Codex account and usage for dashboard configuration,
measurement limits and the distinction between automated tests and real reads.