Jev vs GPT-4.1 as synthetic survey respondents on Twin-2K-500. How you ask mattered more than which model you used.
benchmark calibration behavioral-economics digital-twin digital-twins probability-elicitation llm llm-evaluation survey-simulation jev synthetic-survey twin-2k-500
-
Updated
Sep 23, 2026 - Python