ZYVEN is a local-first, animated desktop AI companion built for Windows using Python and PyQt6. Designed as a singular, cohesive desktop presence rather than a detached chatbot window or cloud frontend, ZYVEN combines local LLM reasoning (via Ollama and Qwen 2.5), deterministic high-speed reflex actions, and an animated avatar into a single, responsive application.
ZYVEN is the assistant. All subsystems belong directly to ZYVEN and communicate through a centralized coordinator (ZyvenController) using PyQt6 signals and slots.
ZYVEN Companion (PyQt6 UI)
│
[ User Text / Voice ]
│
▼
ZyvenController
│
ReflexRouter
┌──────┴──────────────────────────┐
[ High Confidence Match ] [ General Chat / Reasoning ]
│ │
▼ ▼
ToolDispatcher OllamaWorker (QThread)
(Application Allowlist / (Qwen 2.5 7B Q4_K_M)
Browser / System Control) │
│ ▼
│ Validated JSON Proposal
│ │
▼ ▼
State Transition ◄──────────────────────┘
(IDLE / THINKING /
ACTING / ERROR)
│
▼
Avatar / Dialogue Bubble
│
[ Optional: Piper TTS ]
- Single Process Event Loop: The PyQt6 Qt event loop owns the application. Expensive operations (LLM inference, speech recognition, audio playback, tool actions) run exclusively on background
QThreadworkers. The UI never freezes. - VRAM Discipline: Built specifically for 6 GB VRAM GPUs (NVIDIA GeForce RTX 4050 Laptop). Context size (
num_ctx = 2048), generation limits (num_predict = 128), and keep-alive intervals (keep_alive = "2m") are strictly bounded. - Strict Security Boundary: The LLM is never treated as trusted code. LLM output is strictly parsed as JSON proposals. Tool execution is handled solely by a
ToolDispatcherwith an explicit application allowlist.subprocesscalls never useshell=True. Arbitrary PowerShell, CMD, or file-system deletions are strictly blocked. - Fast Deterministic Reflex Actions: Common commands like
"open calculator","open notepad","search web for ..."or"sleep"bypass the 7B LLM completely and execute immediately via theReflexRouter.
Developed and optimized for:
- OS: Windows 10 / 11 (64-bit)
- CPU: Intel Core i7-13650HX or equivalent
- GPU: NVIDIA GeForce RTX 4050 Laptop GPU (6 GB VRAM)
- RAM: 16 GB DDR5
To inspect GPU memory consumption while running Qwen 2.5 7B:
nvidia-smiDefault parameters in config.toml keep total VRAM usage under ~5.2 GB, leaving ample headroom for Windows DWM and hardware-accelerated desktop apps.
Zyven/
│
├── main.py # Application entry point and signal wiring
├── requirements.txt # Core and optional dependencies
├── config.toml # Central configuration file
├── user_state.json # Persisted window position cache
├── README.md # Complete architectural and setup guide
├── .gitignore # Git exclusions for models, logs, and environments
│
├── zyven/
│ ├── __init__.py
│ │
│ ├── core/
│ │ ├── config.py # TOML configuration loader and coordinate persistence
│ │ ├── controller.py # Central ZyvenController coordinating states and workers
│ │ ├── events.py # Event dataclasses and signal definitions
│ │ ├── logging_config.py # Human-readable logging with credential redaction
│ │ ├── router.py # Deterministic ReflexRouter for instant desktop actions
│ │ └── schemas.py # Strongly typed enums, dataclasses, and JSON validator
│ │
│ ├── ui/
│ │ ├── avatar.py # AvatarWidget with 6 states and graceful procedural fallback
│ │ ├── overlay.py # Frameless, transparent, draggable desktop companion window
│ │ └── styles.py # Dark modern cybernetic styling and state color palettes
│ │
│ ├── llm/
│ │ ├── ollama_client.py # Local Ollama HTTP client with timeout and error handling
│ │ ├── prompts.py # Central system prompt enforcing JSON response contract
│ │ └── worker.py # Asynchronous OllamaWorker QThread with bounded history
│ │
│ ├── tools/
│ │ ├── apps.py # Safe application launcher with strict allowlist
│ │ ├── browser.py # Safe browser navigation and URL-encoded web searches
│ │ ├── dispatcher.py # Policy enforcement (SAFE/DENY) and tool execution
│ │ ├── registry.py # Tool metadata and handler registry
│ │ └── system.py # Companion state actions (sleep, wake, exit)
│ │
│ ├── voice/
│ │ ├── stt.py # Whisper.cpp integration interface
│ │ ├── tts.py # Piper TTS synthesis and background Windows audio playback
│ │ ├── vad.py # Silero VAD voice activity detection interface
│ │ ├── wakeword.py # openWakeWord detector interface
│ │ └── worker.py # VoiceSubsystemWorker background thread
│ │
│ ├── memory/
│ │ └── store.py # SQLiteMemoryStore and ChromaDB interface
│ │
│ └── assets/ # Synthesized animated GIFs for all 6 avatar states
│ ├── avatar_idle.gif
│ ├── avatar_listening.gif
│ ├── avatar_thinking.gif
│ ├── avatar_acting.gif
│ ├── avatar_coding.gif
│ └── avatar_error.gif
│
├── scripts/
│ └── generate_assets.py # Script generating state-specific cybernetic GIFs
│
└── tests/
├── test_acceptance.py # 10 end-to-end acceptance tests mapping to specification
├── test_config.py # Configuration and window position persistence tests
├── test_controller.py # Coordinator, reflex handling, and sleep/wake tests
├── test_dispatcher.py # Tool execution, unknown tool rejection, and allowlist tests
├── test_memory.py # SQLite memory store, turn history, and cleanup tests
├── test_ollama.py # Mocked Ollama client, network error, and timeout tests
├── test_router.py # Reflex intent matching and LLM bypass tests
├── test_schemas.py # Response schema validation and error recovery tests
└── test_ui.py # UI creation and avatar state transition tests
git clone <repo-url> Zyven
cd Zyven
python -m venv .venv
.\.venv\Scripts\Activate.ps1
python -m pip install --upgrade pip
pip install -r requirements.txt- Download and install Ollama for Windows from https://ollama.com/download.
- Start the Ollama desktop service.
- Pull the required Qwen 2.5 7B model:
ollama pull qwen2.5:7b-instruct-q4_K_M(Optional lightweight fallback model for tight VRAM environments):
ollama pull qwen2.5:3b-instructpython main.pyConfigure ZYVEN to match your local setup:
[ollama]
url = "http://localhost:11434/api/chat"
model = "qwen2.5:7b-instruct-q4_K_M"
fallback_model = "qwen2.5:3b-instruct"
num_ctx = 2048
num_predict = 128
temperature = 0.3
keep_alive = "2m"
timeout_seconds = 30
[ui]
always_on_top = true
start_x = 1400
start_y = 750
window_width = 320
window_height = 420
avatar_size = 180
[apps]
calculator = "calc.exe"
notepad = "notepad.exe"
explorer = "explorer.exe"
mspaint = "mspaint.exe"
taskmgr = "taskmgr.exe"
[voice]
enabled = false
tts_enabled = false
piper_binary = "piper.exe"
piper_model = "models/en_US-lessac-medium.onnx"
whisper_binary = "main.exe"
whisper_model = "models/ggml-base.en.bin"
[memory]
enabled = false
storage_path = "zyven_memory.db"All tests execute without requiring an active Ollama instance or GPU compute by utilizing structured mocks and deterministic fixtures:
pytest -vExecute individual test suites:
# Acceptance test suite (covering all 10 specifications from Section 31)
pytest -v tests/test_acceptance.py
# Tool dispatcher and security allowlist tests
pytest -v tests/test_dispatcher.py
# Ollama worker context bounding and failure handling tests
pytest -v tests/test_ollama.pyZYVEN's avatar states are managed centrally via the AvatarState enum in zyven/core/schemas.py:
| State | Visual Representation | Color Accent | Trigger |
|---|---|---|---|
| IDLE | Calm breathing ring, friendly cyan eyes | Cyber Cyan (#00E5FF) |
Default resting state |
| LISTENING | Audio wave ripples, focused emerald eyes | Emerald (#00FF9D) |
Wake word detected or recording voice |
| THINKING | Orbiting violet particles, contemplative gaze | Neon Violet (#C77DFF) |
Ollama LLM inference active |
| ACTING | Pulsing solar corona, action chevrons | Solar Amber (#FFB703) |
Executing desktop tool or reflex action |
| CODING | Terminal brackets { }, matrix green eyes |
Matrix Lime (#70E000) |
Technical queries and assistance |
| ERROR | Crimson warning triangle, glitch indicator | Crimson Warning (#FF0055) |
Subsystem offline, timeout, or tool error |
Asset Fallback Guarantee: If any GIF in
zyven/assets/is missing, damaged, or unreadable,AvatarWidgetautomatically engages a procedural vector renderer usingQPainter. ZYVEN will never crash due to a missing asset file.
- Local-First Default: All reasoning is conducted on your local GPU via Ollama. No transcripts or telemetry are sent to cloud servers.
- Application Allowlist: ZYVEN can only launch executables explicitly configured in the
[apps]section ofconfig.toml. Commands likecmd.exe /c ...or arbitrary PowerShell scripts are rejected with a security violation. - No Shell Execution: All subprocess calls use explicit argument lists without
shell=True. - URL Sanitization: Web searches strictly encode queries using
urllib.parse.quote_plusbefore passing to the default system browser.
ZYVEN runs completely out-of-the-box in text mode. When ready to enable voice:
- Download the Windows Piper binary from Piper Releases.
- Download an ONNX voice model (e.g.,
en_US-lessac-medium.onnxand.json). - Set
tts_enabled = trueand updatepiper_binaryandpiper_modelinconfig.toml.
- Download
whisper.cppWindows build and a GGML model (e.g.,ggml-base.en.bin). - Install audio capture dependencies:
pip install sounddevice numpy
- Set
enabled = truein the[voice]section ofconfig.toml.
- Ollama Connection Refused:
- Verify that the Ollama service is running in your Windows system tray.
- Test via command line:
curl http://localhost:11434/api/tags.
- Model Not Found:
- Run
ollama pull qwen2.5:7b-instruct-q4_K_M.
- Run
- Window Off-Screen:
- If monitors were changed, ZYVEN automatically snaps to the bottom-right corner of the primary screen. To reset manually, delete
user_state.json.
- If monitors were changed, ZYVEN automatically snaps to the bottom-right corner of the primary screen. To reset manually, delete
- VRAM Spikes:
- In
config.toml, reducenum_ctxto1024or switchmodeltoqwen2.5:3b-instruct.
- In