The Validation Split That Flattered Every Number

The classifier reported a macro F1 of 0.975. It was, in the narrow sense, telling the truth: that number came out of a validation set the model never trained on, computed correctly, reproducibly. What it measured was “a new frame of a truck I have already seen.” Roughly 94% of the validation images had another frame of the same truck pass sitting in the training set. Different filename, different moment, same vehicle, same load, same lighting, half a second apart. ...

August 9, 2026 · 7 min · Leandro Garcia