Experiment & Evaluation Infrastructure for Vision-Language-Action (VLA) Foundation Models.
DriveScope is an open-source experimentation, evaluation, and failure-analysis platform for Vision-Language-Action (VLA) models and physical-AI driving agents. It provides a reproducible evaluation layer sitting around scenario sequences, camera observations, model inference heads, reasoning traces, metric evaluations, physical sensor perturbations, and failure intelligence.
- Executive Summary
- Core Architecture
- End-to-End Evaluation Dataflow
- Reasoning-Action Consistency Engine
- Hardware-in-the-Loop & Physical Sensor Rig (ECE Core)
- Run Lifecycle State Machine
- Pluggable Model Adapters & Quantized CPU Path
- Dual Execution Modes
- Quickstart Guide
- Repository Structure
- Future Scope & Planned Extensions
- Testing & Quality Verification
- Project Motivation & Interview Defense Guide
- License
Evaluating Vision-Language-Action (VLA) agents requires more than eyeballing demo videos or computing offline loss metrics. A robust VLA evaluation layer must answer:
- Did the model do the right thing? (Action MAE/RMSE, steering drift, brake force calibration).
- Did the model reason correctly, and did its action match that reasoning? (Reasoning-Action Consistency).
- How robust is the decision when real physical sensors degrade? (CMOS noise, rolling shutter, glare, frame drops).
-
How fast did the model react? (Hardware capture timestamp
$\rightarrow$ inference latency$\rightarrow$ actuator lag). - Can this exact result be reproduced on any machine? (Deterministic seeds, immutable manifests, config hashes).
DriveScope decouples the evaluation and experiment orchestrator from the scenario source (datasets, live sensor rigs, CARLA simulators) and the model implementation (mock generators, rule baselines, quantized local VLMs, remote hosted APIs).
flowchart TD
subgraph Frontend["Presentation Layer (Next.js 14 + TypeScript)"]
UI_Dash["Dashboard & Overview"]
UI_Scen["Scenario Browser & Scrubber"]
UI_Exp["Experiment Builder & Config"]
UI_Rep["Replay Workbench & Telemetry"]
UI_Comp["Run Comparison & Robustness Diff"]
UI_Fail["Failure Explorer & Clustering"]
end
subgraph ControlPlane["FastAPI Control Plane (REST + WebSockets)"]
API_Scen["Scenario Service"]
API_Run["Run & Lifecycle Service"]
API_Eval["Analysis & Metric Service"]
API_Auth["Auth & Signed Links (HMAC)"]
API_WS["WebSocket Event Broadcast"]
end
subgraph Storage["Data & State Layer"]
DB[("Relational Database (SQLite / PostgreSQL)")]
BlobStore["Artifact Storage (Local Disk / MinIO S3)"]
MsgBroker[("Message Broker (In-process Async / Redis)")]
end
subgraph Execution["Worker & Compute Engines"]
Worker_Inf["Inference Worker"]
Worker_Eval["Evaluation Worker"]
end
subgraph Adapters["VLA Model Adapter Layer"]
M_Mock["MockVLA (Deterministic)"]
M_Rule["RuleBaseline (Heuristic)"]
M_VLM["LocalQuantizedVLM (int4 GGUF/ONNX)"]
M_Remote["RemoteVLA (Hosted API)"]
end
subgraph Sources["Scenario Source Abstraction"]
S_Data["DatasetScenarioSource (Prerecorded)"]
S_HW["HardwareScenarioSource (ESP32-CAM / RPi)"]
S_Sim["SimulatorScenarioSource (CARLA Bridge)"]
end
Frontend <--> ControlPlane
ControlPlane <--> DB
ControlPlane <--> BlobStore
ControlPlane <--> MsgBroker
MsgBroker <--> Execution
Worker_Inf --> Adapters
Worker_Inf <--> Sources
Worker_Eval --> API_Eval
sequenceDiagram
autonumber
participant UI as Next.js Workbench
participant API as FastAPI Control Plane
participant Worker as Execution Engine
participant Source as Scenario Source (Data / HW)
participant Model as VLA Adapter
participant Eval as Metric & Evaluation Engine
participant DB as Relational Storage
UI->>API: POST /api/v1/experiments (Model, Scenarios, Perturbations)
API->>DB: Persist Experiment & Queue Run
API->>Worker: Dispatch execution (In-process or Celery)
Worker->>Source: Load Scenario Metadata & Frame Iterator
loop Frame-by-Frame Evaluation Loop (with Frame Stride)
Source->>Worker: Emit Frame (Image URI, Hardware TS, Ego State, Ground Truth)
Worker->>Worker: Apply Calibrated Sensor Perturbation (CMOS noise, Rolling shutter)
Worker->>Model: infer(Observation)
Model-->>Worker: VLAOutput (Reasoning, Predicted Action, Confidence, Latency)
Worker->>DB: Persist InferenceTrace record
Worker-->>API: Stream WebSocket Progress (frame_idx, progress, latency)
API-->>UI: Live WebSocket progress broadcast
end
Worker->>Eval: evaluate_run(Traces, Profile)
Eval->>Eval: Compute Action Error (MAE, RMSE, Lateral Drift)
Eval->>Eval: Compute Temporal Metrics (Reaction Delay, p95 Latency)
Eval->>Eval: Compute Safety Proxies (Min TTC, Missed Hazards)
Eval->>Eval: Compute Reasoning-Action Consistency Score
Eval-->>Worker: Return Metrics list & FailureRecords
Worker->>DB: Persist Metrics, Failures & Generate Manifest (SHA-256)
Worker->>API: Update Status to COMPLETED
API-->>UI: Run Completed Event & Manifest Ready
A critical failure mode in Vision-Language-Action foundation models is semantic dissociation — where the natural language chain-of-thought correctly identifies a hazard ("pedestrian crossing from right curb"), but the action output head fails to actuate the brake or continues commanding throttle.
flowchart LR
subgraph VLA_Output["VLA Model Output Stream"]
Lang["Natural Language Reasoning: Pedestrian crossing ahead, applying emergency stop"]
Act["Actuator Head Prediction: steering 0.0, brake 0.05, throttle 0.30"]
Conf["Reported Confidence: 0.95 (High)"]
end
subgraph ConsistencyEngine["Reasoning-Action Consistency Evaluator"]
HAA["1. Hazard-Action Agreement (w1 = 0.35) - Detects if stated hazard yields brake at least 0.30"]
TA["2. Temporal Alignment (w2 = 0.25) - Measures latency between reasoning and action onset"]
DA["3. Directional Agreement (w3 = 0.25) - Validates textual steering intent vs command sign"]
CC["4. Confidence Calibration (w4 = 0.15) - Penalizes high confidence on large action error"]
end
subgraph OutputDecision["Consistency Score & Classification"]
Score["Composite Consistency Score (0.00 to 1.00)"]
Failure["Failure Detector - Flag REASONING_ACTION_CONTRADICTION"]
end
Lang --> HAA
Act --> HAA
Lang --> TA
Act --> TA
Lang --> DA
Act --> DA
Conf --> CC
Act --> CC
HAA --> Score
TA --> Score
DA --> Score
CC --> Score
HAA -.-> Failure
-
$S_{\text{hazard}}$ (Hazard-Action Agreement): Verifies if verbal hazard identification induces braking ($\text{brake} \ge 0.30$ ). -
$S_{\text{dir}}$ (Directional Agreement): Verifies if lateral intent ("turn left", "steer right") matches steering sign and angle. -
$S_{\text{temp}}$ (Temporal Alignment): Evaluates the frame latency between verbal realization and control onset. -
$S_{\text{conf}}$ (Confidence Calibration): Penalizes overconfident predictions that exhibit severe error or miss hazards.
DriveScope bridges pure software experimentation with real-world electronics and sensor physics. It integrates directly with physical camera sensor rigs (ESP32-CAM, Raspberry Pi Camera module) mounted on 2-axis Pan-Tilt servos.
flowchart TD
subgraph PhysicalRig["Physical Sensor Rig (ESP32-CAM / Raspberry Pi)"]
CMOS["Camera Sensor (OV2640 / IMX219)"]
IMU["6-DOF IMU (MPU6050 Accelerometer/Gyro)"]
Timer["Microsecond Hardware Timer (esp_timer_get_time)"]
Servos["2-Axis Pan-Tilt Servo Motors"]
FW["Firmware HTTP / MJPEG Server (drivescope_rig.ino)"]
CMOS --> FW
IMU --> FW
Timer --> FW
Servos <--> FW
end
subgraph Network["Local Network / Wi-Fi"]
Stream["MJPEG Stream + X-Hardware-Timestamp-Us"]
Telemetry["IMU Telemetry & Gyro Data"]
Control["Pan-Tilt REST / WebSocket Control"]
end
subgraph Driver["DriveScope Hardware Source"]
HSS["HardwareScenarioSource Protocol (packages/vla-sdk)"]
Sync["Hardware Clock Synchronization & Lag Tracking (Capture -> Infer -> Act)"]
Noise["Physical CMOS Noise Calibrator (Poisson Shot + Gaussian Dark Current)"]
end
FW --> Stream
FW --> Telemetry
Control --> FW
Stream --> HSS
Telemetry --> HSS
HSS --> Sync
HSS --> Noise
Rather than only applying synthetic filters, DriveScope incorporates measured physical noise profiles:
- CMOS Photon Shot Noise: Modeled via Poisson distributions parameterized by illuminance and sensor gain.
- Dark Current / Read Noise: Gaussian thermal noise from camera sensor readout circuits.
- Rolling Shutter Distortion: Progressive scanline horizontal skew caused by sensor line-readout delay during platform vibration.
stateDiagram-v2
[*] --> CREATED: Experiment Instantiated
CREATED --> QUEUED: Run Queued in Dispatcher
QUEUED --> RUNNING: Worker Picks Up Job
RUNNING --> EVALUATING: Frame Loop Completed
EVALUATING --> COMPLETED: Metrics Computed & Manifest Generated
RUNNING --> FAILED: Missing Frame / Adapter Timeout
EVALUATING --> FAILED: Evaluation Error
QUEUED --> CANCELLED: User Interrupted
RUNNING --> CANCELLED: User Interrupted
COMPLETED --> [*]
FAILED --> [*]
CANCELLED --> [*]
DriveScope enforces strict boundary contracts for all models via the VLAAdapter protocol:
| Adapter | Architecture | Execution Engine | Primary Use Case |
|---|---|---|---|
MockVLA |
Synthetic Deterministic Generator | Pure Python / CPU | CI/CD testing, reproducibility baselines, anomaly injection |
RuleBaseline |
Classical Control Heuristics | Pure Python / CPU | Interpretable rule baseline for speed regulation and braking |
LocalQuantizedVLM |
SmolVLM2 (256M / 500M / 2.2B) | int4/int8 GGUF via llama.cpp MTMDChatHandler + JSON-schema grammar |
True local CPU multimodal inference without GPUs (heuristic fallback offline) |
RemoteVLAAdapter |
OpenAI / Anthropic / Custom endpoints | HTTP / REST API | Large hosted multimodal frontier models |
SimulatorBridge (Planned) |
CARLA / Unreal Engine Bridge | Python API / RPC | Closed-loop dynamic vehicle physics testing |
DriveScope is built to run on weak hardware (like laptops and Raspberry Pis) without sacrificing enterprise-grade distributed scaling:
flowchart LR
subgraph MinimalMode["MODE=minimal (Laptop / Edge / Codespaces)"]
M_App["FastAPI App (Single Process)"]
M_DB[("SQLite (Thread-safe StaticPool)")]
M_Queue["In-Process Async Worker"]
M_Store["Local Disk ./data"]
M_App <--> M_DB
M_App <--> M_Queue
M_App <--> M_Store
end
subgraph DockerMode["MODE=docker (Production / Multi-Node)"]
D_App["FastAPI Control Plane"]
D_PG[("PostgreSQL 15")]
D_Redis[("Redis 7 Broker")]
D_Celery["Distributed Celery Workers"]
D_MinIO["MinIO S3 Object Storage"]
D_App <--> D_PG
D_App <--> D_Redis
D_Redis <--> D_Celery
D_App <--> D_MinIO
end
# 1. Clone the repository
git clone https://github.com/paul-abhirup/DriveScope.git
cd DriveScope
# 2. Install dependencies & packages in editable mode
pip install -r requirements.txt
pip install -e packages/scenario-schema
pip install -e packages/vla-sdk
# 3. Generate sample benchmark dataset
python data/generate_demo_data.py
# 4. (Optional) Download quantized VLM weights + install llama-cpp backend for real local inference
pip install llama-cpp-python
./scripts/download_models.sh --model smolvlm-2.2b --quant q4_k_m # → ./weights/
# 5. Start FastAPI Control Plane in minimal mode
MODE=minimal uvicorn services.api.main:app --reload --port 8000In a second terminal, start the Next.js frontend:
cd apps/web
npm install
npm run devVisit http://localhost:3000 to open the DriveScope Workbench.
docker compose -f infra/docker-compose.yml up --buildThis starts Next.js (port 3000), FastAPI (port 8000), PostgreSQL (port 5432), Redis (port 6379), MinIO (port 9000/9001), and Celery worker.
flowchart TD
subgraph CorePlatform["DriveScope Core (Current Platform)"]
Core_API["FastAPI Control Plane"]
Core_Eval["Multi-Factor Evaluation Engine"]
Core_Rep["Replay & Comparison Workbench"]
end
subgraph HardwareTrack["Hardware & Edge Track (ECE Extensions)"]
HW_Active["Active Pan-Tilt Subject Tracking via Servos"]
HW_Sync["Sub-Millisecond Microcontroller Sync Profiling"]
HW_Multi["Multi-Camera Stereo & Ultrasonic Sensor Array"]
end
subgraph SimulationTrack["Simulation & Closed-Loop Track"]
Sim_CARLA["CARLA Simulator Dynamic Vehicle Bridge"]
Sim_Loop["Closed-Loop Real-Time Trajectory Control"]
Sim_Fuzz["Failure-Driven Scenario Fuzzing & Mutation"]
end
subgraph AITrack["AI Intelligence & Analysis Track"]
AI_Local["Native GGUF/ONNX Quantized VLM Tokenization"]
AI_Cluster["Vector Embeddings & Semantic Failure Clustering"]
AI_Video["Automated Annotated Video MP4 Export"]
end
CorePlatform --- HardwareTrack
CorePlatform --- SimulationTrack
CorePlatform --- AITrack
- Bidirectional bridge with CARLA simulator where VLA actuator commands (
steering,brake,throttle) drive vehicle physics in real-time. - Supports closed-loop evaluation where decisions alter subsequent observations.
- Mutation engine that generates parameter variations around identified failure cases (e.g. altering rain intensity, shifting pedestrian crossing speed, or adding glare at the exact failure timestep).
- Sentence transformer embeddings for natural language reasoning coupled with action error tensors.
- UMAP /
$t$ -SNE projection maps in the Next.js Failure Explorer to identify semantic clusters of failure modes automatically.
- Backend video compositor overlaying dynamic telemetry HUDs (steering wheel gauges, brake pressure indicators, reasoning subtitle overlays, and TTC warning boxes) into shareable MP4 video packages.
- Expansion of the ESP32-CAM setup to a stereo camera rig + time-of-flight (ToF) distance sensor + 6-DOF IMU for multi-sensor fusion evaluation.
DriveScope maintains a comprehensive test suite across unit calculations, API endpoints, and reproducibility:
# Run all tests
pytest tests/ -v- ✅ Deterministic Seed Reproducibility: Matching seeds guarantee byte-identical model outputs and metric scores.
- ✅ Reasoning-Action Contradiction Detection: Explicitly validates that verbal hazards without braking trigger failure records.
- ✅ Calibrated Physical Perturbations: Verifies Poisson shot noise and rolling shutter scanline transformations.
- ✅ End-to-End Experiment Lifecycle: Validates the entire flow from scenario ingestion to experiment creation, worker execution, trace persistence, metric aggregation, and manifest export.
MIT License. See LICENSE for details.
