This information won't be illustrative because there's nothing to compare it to. It's impossible, by definition, to investigate the validity of a closed system from within. Ideally, after obtaining the final results for three different rating systems, you should compare them to some abstract standard, and the calculation system that produces the most reliable results could claim to be the best. But, there is no standard because it would itself be the optimal calculation system, and others wouldn't be needed...
Therefore, there's no point in waiting for the series of tournaments to finish. It's better not to have illusions than to lose them later... :)
After all, in principle, we have almost determined which components the formula should take into account... to put it all together, we need some coefficients (arbitrary), but we can conduct a separate analysis....
However, Fireball promised a ready-made formula....
And I support him in the idea that we can select appropriate coefficients for each comparison criterion!!!