Dataset construction and training methodology, which had evolved slightly differently across model families, were unified into one canonical, leakage-audited approach. This reduced the risk that two models were quietly being compared under two different sets of rules.