Methodology
Every number on this site comes from the same public pipeline: four detection phases, 21 checks, three weighted pillars — then a Bayesian ranking that rewards evidence, not marketing.
Detection: four phases, 21 checks
- Phase 1 — Real-model verification: does the declared model actually answer, with the declared identity?
- Phase 2 — Metering cross-check: reported token counts vs. character-level ground truth; declared pricing vs. charged amounts.
- Phase 3 — Capability baseline: core competence probes against the model’s expected behavior.
- Phase 4 — Stability & throughput: consistency under repetition and load.
Checks are grouped into pillars A (safety & abuse controls, 40%), B (correctness & quality, 35%), C (performance & throughput, 20%) and stability (5%). The live checklist is served by the detection engine — see the Chinese methodology for the current per-check list with evidence fields.
Ranking: evidence over optimism
The rank key is rank v2 and reproducible from public parameters: shrunk median − 90% confidence lower bound, multiplied by a critical-rate Wilson upper-bound penalty. Small samples shrink toward the prior; a single critical finding (model substitution, metering fraud) measurably depresses the key.
Full math and the anti-gaming rules (new-account downweighting, official-sample 3× weight with published seeds) are in the Chinese ranking spec.
Red lines we do not cross
- Insufficient evidence ≠ problematic. Below the display threshold a gateway is shown as a neutral gray dash — never red, never green.
- Payment never affects scores, verdicts, ranking or display weight. Subscriptions buy detection capacity and engineering benefits only; architecture-assertion tests enforce this in CI.
- No fabricated data. If the board is empty, it stays empty until real public detections exist.
The complete methodology series (scoring, ranking, thresholds, threat model, limits) is published in Chinese: /methodology.