Suppose we define the set of nearly optimal models
This is a Rashomon set: a collection of models that are essentially equivalent according to the evaluation criterion.
Many good models
Good predictive performance need not identify a unique model.
Different mechanisms
Two members of the Rashomon set may rely on completely different features.
Why this matters
The benchmark tells us that the model works.
It does not tell us how it works.