Glossary
Strict definitions for every number shown on the leaderboard and in reports. The three score languages (median / rank key / pass rate) mean different things and are not interchangeable — see the methodology.
Median Score
The (weighted) median of all valid public detection scores for a gateway within the 90-day window. It answers "how does this gateway typically behave?". The median is naturally resistant to single-check pollution — one extreme score cannot move it.
Rank Key
The only sorting input of the leaderboard: Bayesian-shrunk 90% confidence lower bound × critical-rate Wilson upper-bound penalty. It answers "how much evidence backs this score?". More detections and fewer critical findings bring the rank key closer to the median score.
Lower Confidence Bound
The lower end of the score interval at 90% confidence (Wilson / normal approximation). Fewer samples → wider interval. This is the mathematical reason a brand-new gateway cannot top the board overnight.
Shrinkage
A statistical technique that pulls observed scores toward a prior (the median baseline) when sample size is small. Few samples → strong shrinkage → conservative rank key. This is the core anti-gaming mechanism against small-sample manipulation.
Critical Rate
The share of detections concluded as high_risk (untrusted). Its Wilson 97.5% upper bound enters the rank key as a penalty — a single critical finding is enough to visibly depress the rank key.
Insufficient Evidence
The honest label applied when public valid samples fall below the display threshold, rendered as a neutral gray "—". Insufficient evidence ≠ problematic — it is never rendered red or green. This is one of the site’s red lines.
Three Pillars
The three scoring dimensions of the detection engine: A — safety & abuse controls (40%), B — correctness & quality (35%), C — performance & throughput (20%), plus stability (5%). Every check belongs to a pillar with a weight.
Pass Rate
The share of detections that concluded "trusted / basically trusted", taken directly from the engine’s archived data. It reflects consistency and is a third language — distinct from both the median score and the rank key.
Official Sample
A detection funded by the platform and run with a platform-hosted key (isOfficial=true), weighted 3× in aggregation. Monthly sampling audits are the public source of official samples; both seeds and results are published.
New-Account Downweighting
Samples submitted by accounts younger than 7 days are weighted 0.3×. When more than 3 new accounts submit samples for the same domain, that batch is zero-weighted and escalated to manual review (anti-gaming v2).
Hosted Key
A gateway API key that, after merchant authorization, is stored encrypted in the detection engine’s secretbox. The platform database keeps no copy; tasks are submitted by endpoint reference and the key can be rotated or revoked at any time.
Availability Rate
The success ratio of liveness / lightweight probes from an independent monitoring source (7-day weighted). It is independent of detection scores — availability says "alive", scores say "quality".
The Chinese glossary adds per-term detail pages with links into the full methodology.