A self-built home SOC: Wazuh + NetAlertX + Grafana, with an AI security analyst in the loop.
Every file-integrity alert generated by Wazuh is automatically forwarded to Google's Gemini API, analyzed by a calibrated Tier-2 SOC-analyst prompt, and surfaced — risk score, evidence, and recommended actions — on a live Grafana dashboard alongside network inventory and raw event data.
- Overview
- Architecture
- Features
- Tech stack
- Repository contents
- Build journey
- Engineering notes & real-world debugging
- Design decisions
- Roadmap
- Acknowledgements
This project is a fully self-hosted, two-VM home lab that reproduces the core of a real Security Operations Center: an endpoint being monitored, a central platform collecting and correlating events, and an AI layer that triages what those events actually mean.
What it proves end-to-end:
A file gets created, modified, or deleted on the monitored Windows endpoint → the Wazuh agent detects it in real time → the event is analyzed and stored by the Wazuh Manager and Indexer → a custom Python service picks it up, sends it to Gemini with a strict SOC-analyst prompt, and stores the structured response → Grafana pulls from all three systems (Wazuh, NetAlertX, and the AI service) into one live dashboard.
No manual steps in the middle. One file change on the endpoint updates six dashboard panels within seconds.
flowchart LR
subgraph WIN["🖥️ Windows VM — 192.168.50.40"]
A[Wazuh Agent<br/>File Integrity Monitoring]
G[Grafana OSS]
end
subgraph UB["🐧 Ubuntu VM — 192.168.50.50"]
M[Wazuh Manager]
I[(Wazuh Indexer<br/>OpenSearch)]
F[Filebeat]
N[NetAlertX<br/>Docker]
P[Python Forwarder<br/>systemd service]
end
subgraph CLOUD["☁️ Cloud"]
AI[Gemini API<br/>SOC Analyst Prompt]
end
A -- "TCP 1514 (encrypted)" --> M
M --> F
F -- "TLS" --> I
P -- "tails alerts.json" --> M
P -- "HTTPS" --> AI
AI -- "structured JSON" --> P
G -- "OpenSearch query" --> I
G -- "REST + Bearer token" --> N
G -- "REST + API key" --> P
Two virtual machines on an isolated VMware NAT network (192.168.50.0/24), fully firewalled so only the two VMs can talk to each other:
| Host | Role | Address |
|---|---|---|
| Windows 10 VM | Monitored endpoint, runs Wazuh agent + Grafana | 192.168.50.40 |
| Ubuntu 22.04 VM | Wazuh Manager/Indexer/Dashboard, NetAlertX, AI forwarder | 192.168.50.50 |
- 🔍 Real-time File Integrity Monitoring — realtime
syscheckwatching a target directory, detecting add/modify/delete within seconds - 🤖 AI-powered alert triage — every FIM event is scored 0–100, classified (Low → Critical), and given evidence-backed, plain-English recommended actions by an LLM constrained with a strict, prompt-injection-aware system prompt
- 📡 Live network inventory — NetAlertX continuously discovers and fingerprints devices on the lab subnet via ARP/mDNS/NetBIOS scanning
- 📊 Unified dashboard — one Grafana view pulling from three independently-secured backends (OpenSearch, a REST API, and a custom microservice)
- 🔐 Defense-in-depth hardening — UFW rules scoped to a single source IP per service, least-privilege OpenSearch role for Grafana (read-only, single index pattern), locally-generated API keys for every internal service boundary, hardened systemd unit (
ProtectSystem=strict,NoNewPrivileges, restricted read/write paths) - ♻️ Resilient by design — the AI forwarder survives provider-side outages (Gemini
503s) via automatic retry without losing or duplicating events, and Wazuh Manager waits on a proper systemd dependency instead of racing the Indexer on boot
| Layer | Technology |
|---|---|
| Endpoint monitoring | Wazuh Agent 4.14.7 |
| SIEM core | Wazuh Manager, Wazuh Indexer (OpenSearch 2.x), Filebeat |
| Network discovery | NetAlertX (Docker) |
| AI analysis | Google Gemini API (gemini-flash-latest), Python 3.10 (Flask, requests, waitress) |
| Visualization | Grafana OSS 13.2 — OpenSearch & Infinity data source plugins |
| Infrastructure | VMware Workstation Pro, Ubuntu 22.04 LTS, Windows 10 |
| Hardening | UFW, systemd sandboxing, least-privilege service accounts, TLS everywhere |
.
├── README.md
├── forwarder/
│ ├── wazuh_gemini_forwarder.py # the AI forwarder service
│ ├── wazuh-gemini-forwarder.service # hardened systemd unit
│ └── system_prompt.txt # SOC-analyst playbook sent to Gemini
├── config/
│ └── .env.example # environment template (no real secrets)
├── docs/
│ ├── screenshots/ # dashboard & terminal captures
│ └── screenshot-checklist.md
└── LICENSE
Two VMs built from scratch in VMware Workstation Pro, on a custom isolated NAT network (192.168.50.0/24) rather than bridging to the real home LAN — keeps the lab fully self-contained.
Once both VMs had static IPs, the firewall was locked down immediately — every subsequent service port is scoped to allow traffic from the monitored endpoint's IP only.
The Wazuh all-in-one installer (Manager + Indexer + Dashboard + Filebeat) on the Ubuntu VM:
The installer's auto-generated certificates only trust 127.0.0.1 — rebinding the Indexer to a LAN-reachable IP broke trust across three services, each traced and fixed independently.
(See Engineering notes below for the full diagnosis story across the Indexer, Filebeat, and Dashboard.)
The Windows agent enrolled and connected, then a dedicated realtime-monitored test directory was configured. A full add → modify → delete cycle was triggered and traced end-to-end through the alert pipeline.
The
modifiedevent above shows the full evidence chain Wazuh captures automatically: old/new SHA-256, MITRE ATT&CK technique (T1565.001), and PCI-DSS/GDPR compliance tags — this raw structure is exactly what gets handed to the AI analyst in step 6.
Deployed via Docker Compose with NET_RAW/NET_ADMIN capabilities for active ARP/mDNS scanning.
A Python service tails Wazuh's alert log, filters for file-integrity events, and sends each one to Gemini with a strict, prompt-injection-aware SOC-analyst system prompt, running as a hardened systemd service.
Full evidence-backed output for a real event: risk score, severity, confidence, plain-English summary, and prioritized recommended actions — generated entirely from the Wazuh alert fields, with zero hallucinated detail.
Three independently-secured data sources feeding one dashboard: Wazuh via OpenSearch (TLS + least-privilege read account), NetAlertX via REST (Bearer token), and the AI forwarder via REST (custom API key header).
The full six-panel dashboard, live, all panels updating from a single triggered file change:
TLS certificate cascade after rebinding the Indexer to a real IP
Wazuh's installer generates certificates trusting only 127.0.0.1. Rebinding the Indexer to a LAN-reachable address broke trust in three places that each needed independent fixes:
- Indexer — regenerated its own cert with
wazuh-certs-tool.sh, reusing the existing CA - Filebeat — was still pointed at
127.0.0.1:9200and trusting the old CA; both had to be updated - Dashboard — same dual issue (wrong host + stale CA)
Traced with openssl x509 -noout -ext subjectAltName, openssl s_client -verify_ip, and comparing sha256sum of each service's local CA copy against the freshly generated one.
Wazuh Manager boot race condition
After VM reboots, wazuh-manager.service would intermittently fail with Result: timeout — the manager was starting before the Indexer was ready to accept its connector registrations and eventually gave up. Fixed with a systemd drop-in (systemctl edit) adding After=/Wants=wazuh-indexer.service and a longer TimeoutStartSec, rather than relying on manual retries after every reboot.
AI provider pivot: Airia → Gemini
The original design called for Airia's hosted agent builder, which turned out to require sales-team onboarding rather than self-serve signup. Pivoted to Google's Gemini API (free tier, instant self-serve key, response_mime_type: application/json for guaranteed structured output) and rebuilt the forwarder around it — same architecture, same SOC-analyst playbook, different provider. The forwarder's retry logic also had to absorb genuine Gemini-side 503 overload periods without dropping events.
Least-privilege permission tuning
The Grafana OpenSearch read account was deliberately scoped to read on wazuh-alerts-4.x-* only. This correctly blocked Grafana's "Get Version" cluster check with a 403 until cluster_monitor was explicitly added — a good demonstration that the security model was actually restrictive, not just decorative, and that not every permission fix should default to broadening access blindly.
- Isolated NAT network instead of bridging to the host LAN — keeps the lab fully self-contained and repeatable regardless of the host's real network
- Cloud AI API over local LLM inference — evaluated running Ollama locally, but with Wazuh's Indexer, Docker, and NetAlertX already sharing an 8 GB VM, a local model would have degraded both its own output reliability and the stability of the core monitoring stack. A hosted model with
response_mime_type: application/jsongives consistent, strictly-parseable output, which the whole downstream dashboard depends on. - Python service over a no-code workflow tool — the forwarder is a genuine, auditable systemd service (log rotation, retry semantics, hardened sandboxing) rather than a black-box automation graph, intentionally chosen to demonstrate real scripting/ops capability.
- Value mappings for device status (
1/0→ Online/Offline) and severity color coding - Alerting rules (Grafana alert → email/webhook on Critical-severity AI analysis)
- Second monitored endpoint to demonstrate multi-agent correlation
- Optional local-LLM variant as a standalone follow-up project
Built by following and adapting a Wazuh/NetAlertX/Grafana/Airia guide from The Social Dork, with the AI-analysis layer independently rebuilt around the Gemini API.
First cybersecurity home-lab project — built, broken, and fixed one service at a time.




















