A researcher grading their own work is a well-known failure mode, not because people are dishonest, but because it is genuinely hard to see the flaws in something you built and want to work. So self-report alone is not how we settle anything material.

Most research or engineering results that could affect what the system does — a new model family, a new evaluation contract, a runtime change — go through review by someone who did not build the thing being reviewed, before final approval is given. That review checks the underlying evidence directly: hashes, data splits, predictions and metrics, not just the summary. A limited set of exceptions, scoped and disclosed case by case, are closed directly through a formal governance decision instead — formal approval remains required either way.

This process regularly returns findings, and sometimes a full rejection. That is not a failure of the process — it is the process working. A result that survives independent review is trusted more precisely because it had a real chance of being turned down.

We treat this discipline as part of the product, not internal bureaucracy. It is one of the reasons we can publish research conclusions on this site with confidence, including the ones that didn't go our way.