Skip to content

Latest commit

 

History

38 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Aegis-Eval: The Evolution of Adversarial RAG Evaluation 🛡️

Aegis-Eval is a rigorous, open-source evaluation framework designed to test Retrieval-Augmented Generation (RAG) pipelines against deterministic, type-matched adversarial attacks.

After traversing through exact-match brittleness, NLI hallucinations, conditional extraction failures, and representation limits, Aegis has evolved into a production-ready security framework enforcing Authorization-State Monotonicity.

🚀 The Aegis Journey & Architectures

V1: The Foundation

The beginning of adversarial generation and basic matching.

Adversary Model ──> Target RAG ──> Exact Match Evaluator
(Generates query)                      (String inclusion)
  • The Problem: Exact string matching is far too brittle for abstract concepts, leading to false negatives.

V2.x: Hardened RAG & Deterministic Extraction

Externalizing grounding policy and structurally extracting conditions.

                         Query + Evidence
                                │
                                ▼
                        ┌───────────────┐
                        │ Evidence Gate │ (MS-MARCO + NLI)
                        └───────┬───────┘
                                │
                    ┌───────────┼───────────┐
                    ▼           ▼           ▼
               INSUFFICIENT  CONFLICT   SUFFICIENT
                    │           │           │
                 ABSTAIN    constrained   grounded
                            generation    generation
                                │           │
                                └─────┬─────┘
                                      ▼
                      ┌──────────────────────────────┐
                      │ Syntactic Extractor (E0)     │ (Proposition Binding)
                      └──────────────┬───────────────┘
                                     │
                                PASS/REJECT
  • The Breakthrough: Ripped grounding out of the LLM and placed it into discrete NLI gates. Replaced probabilistic NLI with Proposition-Bound structural extraction (E+E0) to crush ambiguity.

V3.x: End-to-End Recovery & Representation Attacks

Eliminating the NLI trigger, but discovering the representation boundary flaws.

Generator ──> Claim Extraction ──> Asymmetric Repair Loop ──> Authorized State
  • The Discovery: V3.4 Independent Black-Box Red Teaming proved that perfect representation is impossible. Abstraction loss combined with pipeline repair bypasses leads directly to authorization amplification.

V4: Authorization-State Monotonicity (Production Readiness)

The shift from semantic extraction to capability-based security. Implementing immutable authorization capabilities bound by strict network and orchestration boundaries.

External Request
      │
      ▼
┌──────────────┐
│  Middleware  │ (Auth & Rate Limiting, Audit Logging)
└──────┬───────┘
       │
       ▼
┌──────────────┐
│ SSRF Defense │ (Connection-level DNS pinning, strict IP validation)
└──────┬───────┘
       │
       ▼
┌──────────────┐    ┌──────────────────────────────────┐
│ Orchestration│───>│          V4 Authorizer           │ (The Singleton)
└──────┬───────┘    │  Enforces monotonic transitions  │
       │            └────────────────┬─────────────────┘
       │                             │ Mints Capability
       ▼                             ▼
┌──────────────┐    ┌──────────────────────────────────┐
│   Response   │<───│         AuthorizedAnswer         │ (Runtime Validated)
└──────────────┘    └──────────────────────────────────┘
  • The Breakthrough:
    • Gate P0-P2: Replaced implicit state with a mathematically monotonic Authorization State Machine. A repair operation can never mint PASS_SUBSTANTIVE without a cryptographically bound AuthorizationGrant.
    • Gate P3: Hardened the API layer against SSRF, DNS Rebinding, and fail-closed the orchestration to prevent API-level capability forgery.

📂 Repository Git Tree

Aegis/
├── docs/
│   ├── JOURNAL_V1_The_Foundation.md
│   ├── JOURNAL_V2.1_Benchmark_Infrastructure.md
│   ├── JOURNAL_V2.2_Deterministic_Leap.md
│   ├── JOURNAL_V2.3_Llama3_Pilot.md
│   ├── JOURNAL_V2.4_Hardened_RAG.md
│   ├── JOURNAL_V2.4.1_Calibration.md
│   ├── JOURNAL_V2.5_Scientific_Validation.md
│   ├── JOURNAL_V2.6_Causal_Diagnosis.md
│   ├── JOURNAL_V2.7_Conflict_Classifier.md
│   ├── JOURNAL_V2.8_Structured_Conflict.md
│   ├── JOURNAL_V2.9_Adversarial_Safety.md
│   ├── JOURNAL_V3.0_End_to_End_Recovery.md
│   ├── JOURNAL_V3.1_Deterministic_Reconstruction.md
│   ├── JOURNAL_V3.2_Factorial_Recovery.md
│   ├── JOURNAL_V3.3_Representation_Attacks.md
│   └── JOURNAL_V3.4_BlackBox_RedTeam.md
├── experiments/
│   └── v2.3/llama3-8b/
├── reports/
│   ├── benchmark-v3.4/                       # Immutable V3.4 forensic data
│   ├── benchmark-v4.1-p2/                    # Gate P2 Independent Generalization
│   └── benchmark-p3-closure/                 # Gate P3 API & Network Security Reports
├── scripts/
│   └── aegis_cli.py                          # Unified CLI
├── src/
│   └── aegis_eval/
│       ├── data/                             # Manifest structures and DB schemas
│       ├── evaluator/                        # NLI Cross-encoders, Aggregators, Metrics
│       ├── hardened_rag/                     # V2/V3 Evidence Gates & Verification mechanisms
│       ├── targets/                          # Multi-Model target integration contracts
│       └── v4/                               # V4 Production Security Architecture
│           ├── api/                          # SSRF Protections and Middleware
│           ├── state.py                      # Monotonic Authorization States
│           └── authorizer.py                 # Core Authorizer Singleton
└── tests/
    └── security/                             # Comprehensive V4/P1/P2/P3 Security Suites

📚 Essential Reading (The Aegis Lore)

📈 Current Status: V4 Production Readiness Roadmap

Aegis is currently marching through the rigorous Production Readiness gates.

  • [x] Gate P0: Security Architecture (Monotonicity and Capability Enforcement)
  • [x] Gate P1: Security Testing & Regression (Frozen V3.4 failure suites)
  • [x] Gate P2: Independent Generalization (Single-Model-Family Evaluation)
  • [x] Gate P3: API & Network Security (SSRF, DNS Rebinding, Rate Limits, Auditability)
  • [x] Gate P4: Supply-Chain Reproducibility (Dependency Pinning & Security Policies)
  • [ ] Gate P5: Operational Security (Up Next)

🛠️ Getting Started (V4 CLI)

Aegis-Eval operates via a unified CLI (scripts/aegis_cli.py).

1. Environment Setup

python -m venv .venv
.\.venv\Scripts\Activate.ps1
pip install -e ".[dev]"

2. Run an Offline Multi-Model Evaluation

Generate the raw responses from a target:

python scripts/aegis_cli.py generate --target http://127.0.0.1:8000/query --output runs/my-model-generation.json

Evaluate the responses offline:

$env:DATABASE_URL="sqlite:///aegis_eval.db"
python scripts/aegis_cli.py evaluate --responses runs/my-model-generation.json --queries reports/benchmark-v2.2.0/adversarial-v2.2.0.json

About

Security research framework for authorization-monotone RAG systems, adversarial evaluation, semantic-boundary testing, and fail-closed AI authorization.

Topics

Resources

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages