OSReward tests computer-use judges against failures
OSReward benchmarks vision-language judges on computer-use trajectories with human-verified verdicts. Its authors report that evaluated judges systematically mark some failed runs as successful and release benchmark data plus reward models. Why it matters: A model's completion message is a weak success signal when the agent touched a browser or desktop. Cheap judging can shrink review time while quietly accepting a broken state.
Try this: For one computer-use task, require an external proof object—a saved file, API response, or UI assertion—and compare it with the judge's verdict on failed runs.