← Back to issue5 / 13 · Week of Aug 31, 2026

Shared-endpoint LLM judges can drift between repeated runs

A preregistered arXiv study of black-box LLM observers on shared endpoints reports weak agreement in same-window repeat rankings and lower agreement when byte-identical requests were replayed the next day. Why it matters: If an LLM score decides whether content or code passes a gate, a model label alone does not show that the gate will make the same decision tomorrow.

Try this: Replay a fixed prompt sample through one LLM-judge model today and the next day. Save the agreement table, raw outputs, model identifier, and the threshold at which human review takes over.

Source
arXiv — Clean Engineering, Unstable Measurement
View source →

Get the field brief every week.

Important AI developments, useful explanations, and practical resources in one weekly read. Context to understand what matters, with links to the original sources and deeper reading.

Subscribe free →
Free weekly·No spam·Unsubscribe anytime