Shared-endpoint LLM judges can drift between repeated runs
A preregistered arXiv study of black-box LLM observers on shared endpoints reports weak agreement in same-window repeat rankings and lower agreement when byte-identical requests were replayed the next day. Why it matters: If an LLM score decides whether content or code passes a gate, a model label alone does not show that the gate will make the same decision tomorrow.
Try this: Replay a fixed prompt sample through one LLM-judge model today and the next day. Save the agreement table, raw outputs, model identifier, and the threshold at which human review takes over.