How a dub is scored
In a private room the funniest version wins. In a ranked match nobody votes — the server measures the performance against the one it is standing in for.
A vote is the right way to settle a party. It is the wrong way to settle a rating: in a match of two the ballot is 1–1 whatever anybody did, and in a bigger one it measures who is funniest to the people in that particular room. So ranked matches are decided by measurement, and this is what is being measured.
The six components
| Component | Points | What it asks |
|---|---|---|
| Intonation | 28 | Do you rise and fall where the original rises and falls? |
| Timing | 24 | Do your lines land where the lines land? |
| Dynamics | 16 | Do you get loud and quiet in the same places? |
| Rhythm | 12 | Is the pace of the delivery the same shape? |
| Expression | 20 | Was the reading actually menacing, or merely shaped like it? |
| Likeness | 10 | A bonus for happening to sound like them. |
Five of those are arithmetic on the audio. Expression is the one judgement a measurement cannot make — whether a line was read as sarcastic rather than just contoured like sarcasm — and comes from a language model. Likeness is a bonus on top rather than part of the hundred, because vocal texture is the one thing on the list nobody can practise their way into, and charging for it would be scoring anatomy instead of performance.
It is not listening for your voice
Every measurement is normalised so that register and microphone gain fall out. Pitch is compared as a contour with your own median removed; energy as a shape with your level removed. A woman performing a gravelly male character is judged on whether she rises where he rises, which is the only question the game is asking. Nothing about the scoring wants you to sound like the actor.
Late is fine. Scattered is not.
Timing is split into two measurements that an average of offsets cannot tell apart, and they are treated completely differently:
- —Bias is how late you are on every line. One constant shift, barely perceptible, and the kind of thing an editor fixes by dragging the whole track. Under 120ms it costs nothing at all, and even a fifth of a second costs very little.
- —Jitter is the scatter around that — late here, early there, no two lines agreeing. Nothing fixes it and it is what actually sounds wrong, so its tolerance is about a third as wide: free under 60ms, and a real deduction by 180ms.
The results screen tells you which of the two you were, in words. “Consistently 200ms late” is a note about your headphones; “scattered by about 200ms” is a note about your reading, and they are the same 200ms.
Two rules that shape every curve
Nothing gates. Each component has a free zone, a smooth falloff and a floor, and the components add rather than multiply. Somebody slightly off who nails everything else beats somebody mechanically exact and lifeless — which is the behaviour a human judge has and a threshold does not.
And nobody is punished for giving more. Every measurement splits into shape — do you move the way the original moves — and size, how big that movement is. Size is only penalised when it is smaller than the original. A huge, committed, over-the-top take that follows the contours scores full marks. A timid one tracing the same curve at half scale does not. Being flat is the mistake; being big never is.
There is a floor under every line you actually deliver, and it is not zero. A line performed badly still scores something, because otherwise the safest way to protect a rating would be to stay silent.
What used to be scored and is not
Word accuracy held eight points and was removed. It needed speech recognition on both sides, it penalised accents and non-native speakers for pronunciation rather than performance, and intonation and dynamics already describe how a line was delivered. Its points went to those two, which is why they are the largest on the table.
So how do you actually score well
- Listen to the original once for shape, not for words. You are copying a curve.
- Fix your latency before you fix your acting. If you are consistently late, move your monitoring or your headphones — the score forgives a constant offset far more readily than an inconsistent one.
- Start lines on the beat rather than catching up mid-line. Jitter is what costs, and most of it comes from starting late and rushing to fit.
- Commit. Bigger than the original is free; smaller than it is not.
- Perform the quiet lines quietly. Dynamics is a shape, and flattening everything to one level is the single most common way to lose those sixteen points.
None of this makes a performance funny, and in a private room funny is the whole game. Ranked is the other thing: the same clip, everybody voicing every part, and a number at the end that says who got closest.
Play a ranked match
Queue on your own, get dealt the same scene as three strangers, and voice every character in it. The server scores the performances and the ladder does the rest.
Find a match