There is no ranking here, and that is deliberate

Comparing two methods on running time needs both to run on the same machine. Correcting across machines by a fitted factor was tried and falsified: the factor fitted on one method did not transfer to another on the same hardware, so a ranking built on it would have ordered methods by which machine happened to suit them. A submission gets a verdict on its own paper, against that paper's own published values, and is not ranked against anybody else's.