What a 9.2/10 actually means, why vote counts change everything, and how a weighted average produces a more useful order.
A film rated 9.4 by two hundred people and a film rated 8.6 by nine hundred thousand are not comparable, but almost every interface displays them identically. The small sample is dominated by people who sought the film out — fans of the director, of the genre, of the lead. They were always going to like it. The large sample includes everyone, including the people who wandered in and were bored.
So the 9.4 does not mean “better”. It means “fewer, more favourable, more self-selected voters”.
The standard fix is to pull every rating toward the global mean, in proportion to how few votes it has. A film with a handful of votes sits near the average no matter how enthusiastic those votes are; as the count grows, its own rating takes over. It is the same idea as not believing a poll of twelve people.
Concretely, this site uses a prior of around eight hundred votes: below that, a rating is treated with suspicion. A 9.0 from forty voters lands near 6.7 after weighting — not because the film is bad, but because we do not yet know. That is what the top-rated ranking is built on, and it is why it looks different from a raw sort.
None of this touches the actual question. A rating estimates how a large population received a film on average. You are one person, on one specific evening, with a specific amount of time, in a specific mood, possibly with other people in the room. The correlation between “generally well received” and “right for you tonight” is real but weak — which is exactly the gap that makes catalogue browsing so frustrating.
This is why the app’s default order is not “highest rated” but a blend: a weighted rating, a bonus for titles whose audience is neither tiny nor enormous, a penalty for raw popularity, and a nudge for recognised directors and casts. The goal is not to predict quality — it is to build a pool worth running duels on.
The practical approach: set a floor, never a target. Filter out what is likely to be a waste of an evening, then let your own duels decide among what remains. A rating is good at excluding; it is poor at choosing.
See the difference in practice: the weighted top 30 against the gems nobody voted for.