GOAI 2026 Track 3 Open Exploration submission (Team Guild, No. 12). An exploration environment that measures whether rewording a harmful act changes an LLM judge's moral verdict. Three controls show the euphemism effect belongs to the older model generation, not to lineage. 16 live models, 101 tests, byte-identical offline reproduction.
reproducible-research ai-safety framing-bias modelscope llm-evaluation llm-as-a-judge evalscope goai-2026 evaluation-robustness exploration-environment
-
Updated
Sep 3, 2026 - Python