Why Ratings Feel Unfair: What Elo Can and Can't Know
Quick Answer: Because a single number cannot see what actually decided the match — your partner's form, the sun behind the glass at 7pm, the hamstring you did not mention. A rating is not a verdict on you. It is a forecast with error bars, and most of what feels unfair is uncertainty the system never showed you.
There is a particular walk from the court to the car park. You played well. The number went the other way. And somewhere between the glass door and the boot of the car you decide the whole thing is rigged.
It is not rigged. It is also not innocent. Both of those are true, and the interesting article is the one that holds them at the same time.
So: here is exactly what a rating system knows about you, exactly what it cannot know, and where the line falls between "this is unfair" and "this is uncertain." They are not the same complaint, and only one of them has a fix.
What a scoreline actually contains
A rating engine watching your Tuesday receives four facts: who was on which side, who won, roughly by how much, and when. That is the entire evidence file.
From that it has to answer a much harder question than the one it was asked. Not "who won?" — it already knows — but "how good is each of these four people?" Four unknowns, one equation, played on a court where you often cannot tell whose shot won the point.
Now list what is missing from those four facts.
Your partner's night. Padel's most common grievance, and the most legitimate. Most systems score a pair as one object: Playtomic's algorithm "calculates the average level per team, compares the expected outcome with the actual result, and adjusts each player's level accordingly" (The Playtomic Levels: ups & downs). You played the best forty minutes of your year next to someone who could not find a volley. The file records one thing: the pair underperformed.
Chemistry. You and Mara have played eleven matches together and never once discussed who takes the middle, because she goes when you go. Put either of you with a better player you have never met and the pair is worse. Real, repeatable, and completely invisible to a number attached to one person.
The conditions. Outdoor court, 7pm, the sun sitting exactly behind the back glass at one end. Cold night, dead balls. New balls that fly. The court with the wall you cannot read.
Your body. A tight calf you did not mention, because nobody wants to be the person who mentions it.
Everything you meant to do. You tried the right shot and executed it badly; they tried the wrong shot and shanked it into a winner. Both go in as "point."
None of that is recoverable from a scoreline. The system is not refusing to look. There is nothing there to look at.
Unfair, or just uncertain?
This is the distinction that dissolves most of the anger, so it is worth being precise.
A rating can be wrong in two completely different ways.
Bias means the system is systematically wrong about a kind of player — it consistently underrates left-handers, or defenders, or people who play in a weaker city. Bias is a defect. It does not average out. It deserves the complaint.
Variance means the system is right on average and wrong tonight. It gave you a number that will be about correct over thirty matches, and any single match can throw it around. Variance is not a defect. It is what estimation is.
Almost every "my rating is unfair" story, examined honestly, is variance. The tell is that the grievances point in every direction at once: the same system is accused of overrating the sandbagger and underrating you, in the same club, in the same week. A biased system produces complaints that all lean the same way. A noisy one produces complaints that cancel out.
Which is not a reason to feel better, exactly — being on the wrong end of variance is genuinely annoying. It is a reason to stop reading single results as judgments, because a single result was never evidence about you. It was one draw from a distribution.
One honest amplifier is worth naming: variance is far larger for new accounts than anybody expects. Playtomic damps every change by a reliability percentage and calls it "the most important factor affecting your level changes," so early swings are the design working rather than failing. If your number lurched 0.3 in a month, that is reliability, not judgment — and the level change simulator shows it on your own numbers: hold the match constant, move the reliability dial, and watch the magnitude change while the direction does not.
A forecast, not a verdict
Here is the reframe that does the most work, and it costs nothing to adopt.
Nobody is insulted by a weather forecast. "70% chance of rain" is not a moral claim about Tuesday. If it stays dry you do not accuse the forecaster of bias; you understand that 30% happens roughly three times in ten, and that a forecaster who was never wrong would be lying about their confidence.
A rating is the same object. Your 3.2 is not a statement that you are worse than a 3.4. It is a claim that in a large number of matches, roughly, the 3.4 wins more often. It is a probability wearing a decimal point.
Chess worked this out decades ago and built it into the arithmetic. Glicko's own documentation says it plainly: "it is usually more informative to summarize a player's strength in the form of an interval (rather than merely report a rating)" — the 95% interval running two rating deviations either side, so a player on 1850 with a deviation of 50 is really somewhere between 1750 and 1950 (Mark E. Glickman, The Glicko system).
Read that with padel eyes. The honest version of your level is not 3.2. It is around 3.2, and the width of the "around" depends entirely on how much the system has seen. Every platform computes that width. Almost none show it to you, which is where a defensible number starts to feel like an indictment.
What a system owes you when it does not know
Uncertainty is not the failure. Hiding it is.
Four things separate a rating you can live with from one you resent, and none of them require better math — only more honesty about the math already there.
Show the error bars. A provisional badge, a confidence ring, a wider band while the estimate settles. Pickleball's DUPR does the simple version well, publishing a Reliability Score as a percentage and stating that "players who have a score of at least 60% have a reliable rating" (DUPR Reliability Score). You then know at a glance whether to argue with your own number.
Use every scrap of the result. A 7-5 and a 6-0 are different evidence, and a system that treats them identically discards most of what it watched. SquashLevels reads "a combination of points scores and games scores to assess the result" (What are Levels?). Margin awareness is also the direct answer to the partner-drag grievance: if a narrow loss beside a struggling partner can still move you up, most of the unfairness evaporates on the spot.
Explain the move in one sentence. "You were expected to take 40% of the games and you took 55%, so you gained." A number that moves without explanation is not a rating, it is a grievance with a decimal point. Cheapest fix on the list, most often skipped.
Do not pretend the estimate is a person. Your rating is a system's opinion of your recent results — not your ceiling, your potential, or your worth on a court. DUPR is refreshingly blunt about the scope: "the algorithm isn't trying to measure your very best day on the court; it's predicting your typical performance" (Cracking the DUPR code).
To feel the machinery rather than read about it, put four ratings and a real scoreline into the padel Elo calculator and change one thing at a time. It stops looking arbitrary in about ninety seconds. The deeper version is how Elo ratings work, and why your level dropped walks the four honest causes of the night that brought you here.
The unfairness that is actually real
Everything above has been a defense of the arithmetic. Here is the part where the complaint lands.
The math is not what hurts. The stakes are.
Your level is public. Strangers read it before they accept you. It decides which matches you can join and, in some competitions, which category you belong to. So an estimate with honest error bars gets treated as a credential, and people start behaving the way people always behave around credentials — protecting them.
That behavior is documented, not theoretical. The one substantial neutral critique of Playtomic's system on the English web observes that inaccurate matchmaking means "many players end up picking and choosing their open matches more and more carefully" (Asa Jay Ackley, Is Playtomic's rating system flawed?, September 2025). Follow that thread and everything else appears: the under-declared signup, the weak partner chosen for the upside, the friendly Tuesday quietly ducked in case it costs something.
None of that is about the formula. Every one of those behaviors is a response to consequences, and it would appear around any sufficiently public number, however well computed. A perfect rating with high stakes still produces anxious players; an imperfect rating with no stakes produces almost none.
Which is the honest answer to "why does this feel unfair." Not because a system read your Tuesday wrong — because a private estimate was handed a public job, and you are now watching a decimal that was never designed to carry your reputation. Taken to its conclusion, that is the case against one global number.
Where a crew number sits
Rivals took the other branch of that fork, and it is a boundary rather than a boast. Its rating is crew-local: it exists inside your group of friends and nowhere else, so there is no stranger reading it and nothing to protect. It is margin-aware, so a 7-5 and a 6-0 are different evidence and a narrow loss to the strongest pair can still move you up. It shows a provisional standing instead of a confident-looking number after two matches. And every change explains itself in one sentence you can tap, which is the only real cure for suspicion: nobody distrusts a number they can audit.
It will not get you into an open match at a club or enter you in a category — that is what a platform level or a federation rating is for, and those systems carry stakes precisely because that is their job. The two solve different problems. Free covers one crew of four with nothing inside it metered.
Hold your rating the way you hold a forecast. Useful, revisable, occasionally wrong on a Tuesday, and not about you. Then go and play the ten matches that narrow it.
FAQ
Is the Playtomic rating system flawed?
It has a documented weakness rather than a flaw: it scores each pair as an average, so your partner's night moves your level, and its published inputs do not include your margin of victory. Both are honest consequences of what a scoreline contains. The larger problem is not accuracy but stakes — a public number changes how people choose matches.
What does my padel rating actually mean?
It is a system's estimate of how you have recently performed against the specific people it watched you play, expressed on that system's own scale. It is a probability, not a ranking of worth: a 3.4 beating a 3.2 more often than not. Treat it as a range of about half a unit rather than a decimal.
Why does my padel level go up and down so much?
Because the system's confidence in you is low, so every result moves it a long way. Volatility is a statement about the estimate, not about your padel — and it shrinks as your match history grows. A swing bigger than a tenth of a level usually means the estimate is still settling.
Can a rating system be biased against certain players?
It can, and that is a different complaint from noise. Bias means being systematically wrong about a type of player and never correcting; it deserves a hearing. Most rating anger, though, is variance — the same system accused of both overrating and underrating people in the same club in the same week, which is the signature of noise rather than prejudice.
Is it worth rating friendly matches at all?
Yes, if the number stays inside the group. The value of a rating is settling arguments with evidence; the cost is anxiety, and almost all of the anxiety comes from the number being public. A rating your three friends can see, and audit, keeps the value and drops the cost.