基于开源题库(六大关卡 95 题)+ 纯规则判分引擎(不使用 LLM-as-Judge)+ GitHub Actions CI 复跑 + BYOK 自带密钥试测页,接入任意 OpenAI 兼容 API → 自动完成「安全常识 → 越狱拒答与钓鱼识别 → 恶意代码判断 → 代码漏洞识别 → 密码学与编码 → 日志研判与应急响应」六关测评,产出安全评分与公开排行榜。
-
Updated
Sep 19, 2026 - JavaScript
基于开源题库(六大关卡 95 题)+ 纯规则判分引擎(不使用 LLM-as-Judge)+ GitHub Actions CI 复跑 + BYOK 自带密钥试测页,接入任意 OpenAI 兼容 API → 自动完成「安全常识 → 越狱拒答与钓鱼识别 → 恶意代码判断 → 代码漏洞识别 → 密码学与编码 → 日志研判与应急响应」六关测评,产出安全评分与公开排行榜。
Auditable benchmark records evaluating Grounded DI across independent benchmark suites and domains.
Local benchmark audit artifacts for legal-work-product tasks derived from the Harvey LAB benchmark framework. Grounded DI scored a provisional 501/501 on 8/10/2026.
To associate your repository with the aibenchmark topic, visit your repo's landing page and select "manage topics."