Auditable DeepSWE v1.1 evaluation of Agent Claim Network: coding-harness pass rate and cross-agent claim reuse|ACN 在 DeepSWE v1.1 上的可复现评测:coding harness 通过率与跨 agent 的 claim 复用收益。
-
Updated
Sep 17, 2026 - Python
Auditable DeepSWE v1.1 evaluation of Agent Claim Network: coding-harness pass rate and cross-agent claim reuse|ACN 在 DeepSWE v1.1 上的可复现评测:coding harness 通过率与跨 agent 的 claim 复用收益。
Public evidence for AIC coding-agent evaluations: frozen patches, canonical verifier scores, integrity hashes, cost, timing, and task context.
DeepSWE v1.1: Perfect score (113/113) — all tasks solved with reward=1.0
Reproducible analysis of reasoning-effort saturation on DeepSWE v1.1
DeepSWE Launcher ⚡ lightweight DeepSWE benchmark launcher: batching, resume, multi-model profiles, live dashboard (terminal + WebUI), Excel reports
Forensic audit scanner for SWE-bench & DeepSWE v1.1 containers detecting .git reflog leakage and test harness spoofing.
Daily mirror of the DeepSWE live leaderboard with TrueIQ benchmarking analysis.
To associate your repository with the deepswe topic, visit your repo's landing page and select "manage topics."