Skip to content

Repository files navigation

ZYVEN — Local-First Animated Desktop AI Companion

ZYVEN is a local-first, animated desktop AI companion built for Windows using Python and PyQt6. Designed as a singular, cohesive desktop presence rather than a detached chatbot window or cloud frontend, ZYVEN combines local LLM reasoning (via Ollama and Qwen 2.5), deterministic high-speed reflex actions, and an animated avatar into a single, responsive application.


1. Product Vision & Architecture

ZYVEN is the assistant. All subsystems belong directly to ZYVEN and communicate through a centralized coordinator (ZyvenController) using PyQt6 signals and slots.

                              ZYVEN Companion (PyQt6 UI)
                                         │
                                [ User Text / Voice ]
                                         │
                                         ▼
                                  ZyvenController
                                         │
                                  ReflexRouter
                                  ┌──────┴──────────────────────────┐
                [ High Confidence Match ]             [ General Chat / Reasoning ]
                                  │                                 │
                                  ▼                                 ▼
                           ToolDispatcher                     OllamaWorker (QThread)
                     (Application Allowlist /               (Qwen 2.5 7B Q4_K_M)
                      Browser / System Control)                     │
                                  │                                 ▼
                                  │                       Validated JSON Proposal
                                  │                                 │
                                  ▼                                 ▼
                           State Transition  ◄──────────────────────┘
                          (IDLE / THINKING /
                           ACTING / ERROR)
                                  │
                                  ▼
                         Avatar / Dialogue Bubble
                                  │
                       [ Optional: Piper TTS ]

Core Design Principles

  • Single Process Event Loop: The PyQt6 Qt event loop owns the application. Expensive operations (LLM inference, speech recognition, audio playback, tool actions) run exclusively on background QThread workers. The UI never freezes.
  • VRAM Discipline: Built specifically for 6 GB VRAM GPUs (NVIDIA GeForce RTX 4050 Laptop). Context size (num_ctx = 2048), generation limits (num_predict = 128), and keep-alive intervals (keep_alive = "2m") are strictly bounded.
  • Strict Security Boundary: The LLM is never treated as trusted code. LLM output is strictly parsed as JSON proposals. Tool execution is handled solely by a ToolDispatcher with an explicit application allowlist. subprocess calls never use shell=True. Arbitrary PowerShell, CMD, or file-system deletions are strictly blocked.
  • Fast Deterministic Reflex Actions: Common commands like "open calculator", "open notepad", "search web for ..." or "sleep" bypass the 7B LLM completely and execute immediately via the ReflexRouter.

2. Target Hardware Specifications

Developed and optimized for:

  • OS: Windows 10 / 11 (64-bit)
  • CPU: Intel Core i7-13650HX or equivalent
  • GPU: NVIDIA GeForce RTX 4050 Laptop GPU (6 GB VRAM)
  • RAM: 16 GB DDR5

VRAM Discipline & Monitoring

To inspect GPU memory consumption while running Qwen 2.5 7B:

nvidia-smi

Default parameters in config.toml keep total VRAM usage under ~5.2 GB, leaving ample headroom for Windows DWM and hardware-accelerated desktop apps.


3. Project Structure

Zyven/
│
├── main.py                     # Application entry point and signal wiring
├── requirements.txt            # Core and optional dependencies
├── config.toml                 # Central configuration file
├── user_state.json             # Persisted window position cache
├── README.md                   # Complete architectural and setup guide
├── .gitignore                  # Git exclusions for models, logs, and environments
│
├── zyven/
│   ├── __init__.py
│   │
│   ├── core/
│   │   ├── config.py           # TOML configuration loader and coordinate persistence
│   │   ├── controller.py       # Central ZyvenController coordinating states and workers
│   │   ├── events.py           # Event dataclasses and signal definitions
│   │   ├── logging_config.py   # Human-readable logging with credential redaction
│   │   ├── router.py           # Deterministic ReflexRouter for instant desktop actions
│   │   └── schemas.py          # Strongly typed enums, dataclasses, and JSON validator
│   │
│   ├── ui/
│   │   ├── avatar.py           # AvatarWidget with 6 states and graceful procedural fallback
│   │   ├── overlay.py          # Frameless, transparent, draggable desktop companion window
│   │   └── styles.py           # Dark modern cybernetic styling and state color palettes
│   │
│   ├── llm/
│   │   ├── ollama_client.py    # Local Ollama HTTP client with timeout and error handling
│   │   ├── prompts.py          # Central system prompt enforcing JSON response contract
│   │   └── worker.py           # Asynchronous OllamaWorker QThread with bounded history
│   │
│   ├── tools/
│   │   ├── apps.py             # Safe application launcher with strict allowlist
│   │   ├── browser.py          # Safe browser navigation and URL-encoded web searches
│   │   ├── dispatcher.py       # Policy enforcement (SAFE/DENY) and tool execution
│   │   ├── registry.py         # Tool metadata and handler registry
│   │   └── system.py           # Companion state actions (sleep, wake, exit)
│   │
│   ├── voice/
│   │   ├── stt.py              # Whisper.cpp integration interface
│   │   ├── tts.py              # Piper TTS synthesis and background Windows audio playback
│   │   ├── vad.py              # Silero VAD voice activity detection interface
│   │   ├── wakeword.py         # openWakeWord detector interface
│   │   └── worker.py           # VoiceSubsystemWorker background thread
│   │
│   ├── memory/
│   │   └── store.py            # SQLiteMemoryStore and ChromaDB interface
│   │
│   └── assets/                 # Synthesized animated GIFs for all 6 avatar states
│       ├── avatar_idle.gif
│       ├── avatar_listening.gif
│       ├── avatar_thinking.gif
│       ├── avatar_acting.gif
│       ├── avatar_coding.gif
│       └── avatar_error.gif
│
├── scripts/
│   └── generate_assets.py      # Script generating state-specific cybernetic GIFs
│
└── tests/
    ├── test_acceptance.py      # 10 end-to-end acceptance tests mapping to specification
    ├── test_config.py          # Configuration and window position persistence tests
    ├── test_controller.py      # Coordinator, reflex handling, and sleep/wake tests
    ├── test_dispatcher.py      # Tool execution, unknown tool rejection, and allowlist tests
    ├── test_memory.py          # SQLite memory store, turn history, and cleanup tests
    ├── test_ollama.py          # Mocked Ollama client, network error, and timeout tests
    ├── test_router.py          # Reflex intent matching and LLM bypass tests
    ├── test_schemas.py         # Response schema validation and error recovery tests
    └── test_ui.py              # UI creation and avatar state transition tests

4. Installation & Setup (Windows PowerShell)

Step 1: Clone Repository & Create Virtual Environment

git clone <repo-url> Zyven
cd Zyven
python -m venv .venv
.\.venv\Scripts\Activate.ps1
python -m pip install --upgrade pip
pip install -r requirements.txt

Step 2: Install and Start Ollama

  1. Download and install Ollama for Windows from https://ollama.com/download.
  2. Start the Ollama desktop service.
  3. Pull the required Qwen 2.5 7B model:
ollama pull qwen2.5:7b-instruct-q4_K_M

(Optional lightweight fallback model for tight VRAM environments):

ollama pull qwen2.5:3b-instruct

Step 3: Run ZYVEN

python main.py

5. Configuration (config.toml)

Configure ZYVEN to match your local setup:

[ollama]
url = "http://localhost:11434/api/chat"
model = "qwen2.5:7b-instruct-q4_K_M"
fallback_model = "qwen2.5:3b-instruct"
num_ctx = 2048
num_predict = 128
temperature = 0.3
keep_alive = "2m"
timeout_seconds = 30

[ui]
always_on_top = true
start_x = 1400
start_y = 750
window_width = 320
window_height = 420
avatar_size = 180

[apps]
calculator = "calc.exe"
notepad = "notepad.exe"
explorer = "explorer.exe"
mspaint = "mspaint.exe"
taskmgr = "taskmgr.exe"

[voice]
enabled = false
tts_enabled = false
piper_binary = "piper.exe"
piper_model = "models/en_US-lessac-medium.onnx"
whisper_binary = "main.exe"
whisper_model = "models/ggml-base.en.bin"

[memory]
enabled = false
storage_path = "zyven_memory.db"

6. Running Automated Tests

All tests execute without requiring an active Ollama instance or GPU compute by utilizing structured mocks and deterministic fixtures:

pytest -v

Execute individual test suites:

# Acceptance test suite (covering all 10 specifications from Section 31)
pytest -v tests/test_acceptance.py

# Tool dispatcher and security allowlist tests
pytest -v tests/test_dispatcher.py

# Ollama worker context bounding and failure handling tests
pytest -v tests/test_ollama.py

7. Avatar State Management

ZYVEN's avatar states are managed centrally via the AvatarState enum in zyven/core/schemas.py:

State Visual Representation Color Accent Trigger
IDLE Calm breathing ring, friendly cyan eyes Cyber Cyan (#00E5FF) Default resting state
LISTENING Audio wave ripples, focused emerald eyes Emerald (#00FF9D) Wake word detected or recording voice
THINKING Orbiting violet particles, contemplative gaze Neon Violet (#C77DFF) Ollama LLM inference active
ACTING Pulsing solar corona, action chevrons Solar Amber (#FFB703) Executing desktop tool or reflex action
CODING Terminal brackets { }, matrix green eyes Matrix Lime (#70E000) Technical queries and assistance
ERROR Crimson warning triangle, glitch indicator Crimson Warning (#FF0055) Subsystem offline, timeout, or tool error

Asset Fallback Guarantee: If any GIF in zyven/assets/ is missing, damaged, or unreadable, AvatarWidget automatically engages a procedural vector renderer using QPainter. ZYVEN will never crash due to a missing asset file.


8. Security & Privacy Model

  1. Local-First Default: All reasoning is conducted on your local GPU via Ollama. No transcripts or telemetry are sent to cloud servers.
  2. Application Allowlist: ZYVEN can only launch executables explicitly configured in the [apps] section of config.toml. Commands like cmd.exe /c ... or arbitrary PowerShell scripts are rejected with a security violation.
  3. No Shell Execution: All subprocess calls use explicit argument lists without shell=True.
  4. URL Sanitization: Web searches strictly encode queries using urllib.parse.quote_plus before passing to the default system browser.

9. Voice Subsystem Setup (Optional)

ZYVEN runs completely out-of-the-box in text mode. When ready to enable voice:

Piper TTS (Voice Output)

  1. Download the Windows Piper binary from Piper Releases.
  2. Download an ONNX voice model (e.g., en_US-lessac-medium.onnx and .json).
  3. Set tts_enabled = true and update piper_binary and piper_model in config.toml.

Whisper.cpp (Voice Input)

  1. Download whisper.cpp Windows build and a GGML model (e.g., ggml-base.en.bin).
  2. Install audio capture dependencies:
    pip install sounddevice numpy
  3. Set enabled = true in the [voice] section of config.toml.

10. Troubleshooting

  • Ollama Connection Refused:
    • Verify that the Ollama service is running in your Windows system tray.
    • Test via command line: curl http://localhost:11434/api/tags.
  • Model Not Found:
    • Run ollama pull qwen2.5:7b-instruct-q4_K_M.
  • Window Off-Screen:
    • If monitors were changed, ZYVEN automatically snaps to the bottom-right corner of the primary screen. To reset manually, delete user_state.json.
  • VRAM Spikes:
    • In config.toml, reduce num_ctx to 1024 or switch model to qwen2.5:3b-instruct.

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages