Attributed Agentic RAG · Low-Resource NLP · Multimodal Grounding
Production AI reliability requires deterministic out-of-band verification, fine-grained atomic claim decomposition, and real-time multimodal perception grounding.
My work bridges foundation model capabilities and production engineering constraints across two core areas:
-
Attributed Agentic RAG & Fact Verification: Eliminating multi-hop hallucination loops via fine-grained atomic claim decomposition, calibrated DeBERTa-v3 cross-encoders (
$\tau \ge 0.82$ ), and formal 3-state selective abstention. - High-Throughput PyTorch Inference Acceleration: Accelerating autoregressive token generation via rejection-sampling speculative decoding (1.92x speedup) and single-GPU 70B parameter layer streaming.
- Autonomous Agent Security & Deterministic Reference Monitors: Gating multi-agent tool execution loops with out-of-band AST taint tracking and join semi-lattice information flow control (0.0% ASR on AgentDojo).
+-------------------------------------------------------------------------------------------------------+
| APPLIED AI SYSTEMS ARCHITECTURES |
+---------------------------------------------------+---------------------------------------------------+
| 1. WARRANT | 2. STARK VISION |
| Attributed Agentic RAG Engine | Real-Time Multimodal Grounding Engine |
| Atomic Claim Decomp & DeBERTa-v3 NLI Verification | Continuous Whisper STT + Swin-T + SAM 2 Memory |
| [Zero Hallucination / 37 Automated Tests] | [30 FPS Mask Propagation / Consumer Edge GPU] |
+---------------------------------------------------+---------------------------------------------------+
- The Problem: Standard RAG pipelines suffer from silent multi-hop hallucinations when questions require complex cross-document reasoning or when retrieved context is incomplete or adversarial.
-
Systems Architecture:
- Atomic Claim Decomposition: Parses complex model outputs into discrete, verifiable factual propositions.
- Deterministic Regex Guard: Fast-path pre-validation of numerical, temporal, and entity constraints before neural evaluation.
-
Calibrated DeBERTa-v3 Cross-Encoder: NLI cross-encoder scoring claim entailment against retrieved context at an empirically calibrated threshold (
$\tau \ge 0.82$ ). -
Cyclic LangGraph State Machine: Enforces a 3-state selective prediction contract (
FULL_PASS,PARTIAL_PASSwith claim pruning, orABSTAIN), formally refusing ungrounded extrapolation. - Hybrid Dense/Sparse Retrieval: Qdrant vector indexing combining dense semantic embeddings with sparse BM25 Reciprocal Rank Fusion (RRF) and sub-token citation DAGs.
-
Empirical Benchmarks:
- Zero Hallucination Extrapolations: Enforces formal abstention across 200 HotpotQA evaluation questions rather than emitting unsupported claims.
- Zero-VRAM CPU Verifier: DeBERTa cross-encoder evaluates on pure CPU in 689 ms, preserving GPU memory entirely for high-throughput generation.
- Automated Reliability: Backed by an 37-test automated regression suite covering edge-case token splits, contradiction pruning, and cyclical graph states.
- Tech Stack: Python 3.11, PyTorch, Hugging Face Transformers, DeBERTa-v3, Qdrant, LangGraph, FlashRank, FastAPI, Next.js 14, Docker.
- The Problem: Existing referring video object segmentation models struggle with latency and temporal drift when operators use continuous, real-time vocal commands rather than static text queries.
-
Systems Architecture:
- Decoupled Dual-Loop Pipeline: Streams continuous natural vocal instructions through OpenAI Whisper, extracting acoustic tokens with minimal audio chunk latency.
- Open-Vocabulary Spatial Grounder: Integrates Grounding DINO (Swin-T backbone + text-visual cross-attention) to extract open-vocabulary referring expressions and predict geometric bounding box prompts.
-
Automated Ambiguity Detection: Formulates a confidence-delta metric (
$\Delta\mathrm{score} < 0.15$ ) to identify semantic multi-target conflicts before spatial initialization, preventing false-positive tracking drift. - SAM 2 Temporal Memory Propagation: Injects predicted bounding boxes as spatial prompts into Meta's Segment Anything Model 2 (SAM 2) memory-attention mechanism, maintaining pixel-accurate object masks across camera occlusions at a sustained 30 FPS on an NVIDIA RTX 4060 (8 GB VRAM).
-
Engineering Highlights:
- Supported by a custom 13-test automated validation suite verifying streaming chunk boundaries, ambiguity thresholds, and temporal memory propagation.
- Optimized for edge inference on consumer GPU hardware with low-latency OpenCV video streaming.
- Tech Stack: Python, PyTorch, CUDA, Grounding DINO, Meta SAM 2, OpenAI Whisper, OpenCV.
Open-Source ML Systems Contributor | Awras AI
Low-Resource Language Model Pre-training & SFT Infrastructure (2025 -- Present)
- Engineered distributed data cleaning, deduplication, and quality filters for 100K+ token Algerian Darija and dialectal Arabic datasets, eliminating token fragmentation and dialect representation bias.
- Built end-to-end data processing pipelines and supervised fine-tuning (SFT) workflows on Transformer architectures, standardizing semantic consistency and factual accuracy evaluation benchmarks.
AI Research Fellow | School of AI Algiers
LLM-Guided Reinforcement Learning for MuJoCo Humanoid-v4 (2023 -- Present)
- Developed an LLM-assisted RL framework coupling a foundation model with Soft Actor-Critic (SAC) to automate iterative reward-function synthesis and refinement for obstacle navigation in MuJoCo Humanoid-v4.
- Achieved a 32% increase in episodic return (3,290 to 4,350), boosted obstacle avoidance success from 38% to 69%, and reduced collision rates from 51% to 24% across 3 random seeds compared to hand-tuned baselines.
- Research Paper Poster: Download Poster (PDF)
Industrial Telecommunications Applied AI (Jul 2025 -- Aug 2025 · Algiers, Algeria)
- Engineered machine learning and computer vision pipelines for industrial telecom infrastructure datasets, executing automated feature extraction, data preprocessing, and model validation.
- Conducted comparative performance benchmarks across statistical ML and deep neural network baselines to evaluate operational classification accuracy and inference efficiency.
- Deep Learning Specialization — DeepLearning.AI & Andrew Ng
- Focus: Deep Neural Networks, Convolutional Neural Networks (CNNs), Sequence Models & Attention Mechanisms, Hyperparameter Tuning & Optimization.
- Machine Learning Specialization — Stanford Online & DeepLearning.AI
- Focus: Supervised Learning, Advanced Learning Algorithms, Unsupervised Learning, Recommender Systems, Reinforcement Learning.
| Engineering Domain | Production Technologies & Tooling |
|---|---|
| Deep Learning & Vision | PyTorch, Hugging Face Transformers, SAM 2 (Segment Anything), Swin-T, Whisper STT, OpenCV, Torchvision, Scikit-Learn |
| Agentic Systems & NLP | Attributed RAG, DeBERTa-v3 NLI, LangGraph (Cyclic State Machines), Qdrant (Hybrid Dense/BM25 RRF), FlashRank Re-ranking, SpaCy |
| AI Security & Guardrails | AgentDojo Benchmark, Dynamic AST Taint Tracking (ast, sqlglot), Join Semi-Lattices, HMAC-SHA256 HITL Gating, PyRIT |
| Backend & Distributed Systems | Python 3.11+, C++, FastAPI, Pydantic v2, PostgreSQL, Redis, Docker, Linux (Ubuntu/POSIX), Git/GitHub CI/CD |
| Reliability & Testing | Pytest (37+ Automated Regression Test Suites), HotpotQA Multi-Hop Evaluation, Semantic Abstention Contracts |
| Frontend & Telemetry | Next.js 14, TypeScript, Tailwind CSS, Dynamic Lineage DAGs, WebSockets, Vercel |
- University of Science and Technology Houari Boumediene (USTHB) | Bab Ezzouar, Algiers
- State Engineering Degree (Diplôme d'Ingénieur d'État) in Computer Science — Artificial Intelligence Specialization
- Coursework: Deep Learning, Machine Learning, Computer Vision, Natural Language Processing, Distributed Systems, High-Performance Computing, Advanced Algorithms.
- Micro Club USTHB (2024 -- Present): Member & AI Workshop Contributor — Leading technical sessions on open-source machine learning pipelines, algorithmic problem solving, and student hackathons.
- Google Developer Groups (GDG) Algiers (2024 -- Present): Active member & technical contributor in AI meetups and DevFests.