Per-use scoring
Voice weights jitter and loss heavily and barely cares about download speed. Remote work weights upload. Gaming weights latency. The weights differ because the requirements do.
A connection is not simply good or bad — it is good or bad for something. The same line can be excellent for streaming and unusable for calls, and a single headline score hides exactly that. This test scores each use separately and tells you which measurement held each one back.
Voice weights jitter and loss heavily and barely cares about download speed. Remote work weights upload. Gaming weights latency. The weights differ because the requirements do.
The lowest of the individual scores, not the average. Averaging a connection that is perfect for browsing and hopeless for calls produces “good”, which is the one answer that helps nobody.
Each score lists what went into it. A score built from two measurements should not look like one built from six, so the ones that could not be measured are named rather than quietly assumed.
Every score names its weakest contributing measurement. That is the thing to fix, and it is frequently not the one people expect.
Each measurement is mapped to a 0–100 figure on a curve that flattens at the top, because past a point more does not help: the difference between 500 Mbps and 1 Gbps changes nothing about a voice call, while the difference between 5 ms and 50 ms of jitter changes everything.
Those figures are then weighted per use case and combined. A measurement that could not be taken is excluded and the remaining weights are renormalised — it is never counted as zero, and never as full marks. Both would be inventions, in opposite directions, and the spec this tool was built to is explicit that nothing may be fabricated.
If nothing at all could be measured, there is no score. The tool says so rather than producing a number.
The thresholds are visible in the findings below each run: every one states the measurement it rests on, so you can disagree with the grading and still trust the numbers.
Because they depend on different things. Browsing is forgiving of jitter and latency and mostly wants download capacity. A voice call needs small packets to arrive at a steady pace and is barely affected by raw speed. A connection with 200 Mbps down and 40 ms of jitter is excellent for one and poor for the other.
Because a connection is as good as its worst dimension for the thing you are trying to do. Averaging is how a line that cannot hold a call gets described as 'good' — which is accurate on paper and useless to the person whose calls keep dropping.
Because they were not measured on that run — the upload stage may not have run, or a probe may have failed. They are excluded from the score rather than guessed at. A missing measurement scored as zero would report a failure that did not happen; scored as full marks it would hide a real one.
It includes one, but speed is a small part of it. The measurements that decide whether a connection is pleasant to work on — latency under load, jitter, stability — are invisible to a plain speed test, which is why connections that test fast often still feel bad.
Whatever the weakest dimension says, on the use case you care about. It is usually not bandwidth. In practice the most common fixes are enabling queue management on the router for latency under load, and moving off Wi-Fi for jitter and loss.