Auto-translated
The coefficients must be in place. They should be based on three parameters, in descending order of priority:
1) The quality of the template in terms of "skill over randomness" (expert assessment by top players; for example, the 20th-anniversary tournament where Mirror matches and 6LM10A were more popular than UG, and even more popular than the very popular JC).
2) The length of the game (this should be a factor, but it shouldn't have an excessive impact; we need to find reasonable numbers using our brains).
3) The size of the template, if the first point hasn't already established a clear preference (but without stupid – I can't think of a better word – imbalances like today's UG or, conversely, M-decks that last for 5 weeks and end with scores around 7 points).
And the weight of each subsequent parameter's influence on the rating should be significantly lower overall. That's the idea.
Everything needs to be tested to see how it works in specific (and, most importantly, borderline) situations. Ideally, we should recalculate everything retroactively after a preliminary warning.
The last few months have been irrevocably wasted (from a rational point of view) due to the abuse of UG.
------
Another idea for the first point: separate the ranked pool of templates and calculate a separate rating for it, similar to Hearthstone's "Standard" mode.
And include "games without rules" / random templates in the overall "Wild" rating (in the case of Heroes, this would include all games, which is essentially what we have now, with the caveat that the current system has issues with head-to-head matchups).
But this requires a more careful implementation and a dedicated team of reasonable experts to add templates to the pool, so that the system doesn't become stagnant.