Methodology

Every number on this site comes from the same public pipeline: four detection phases, 21 checks, three weighted pillars — then a Bayesian ranking that rewards evidence, not marketing.

Detection: four phases, 21 checks

  • Phase 1 — Real-model verification: does the declared model actually answer, with the declared identity?
  • Phase 2 — Metering cross-check: reported token counts vs. character-level ground truth; declared pricing vs. charged amounts.
  • Phase 3 — Capability baseline: core competence probes against the model’s expected behavior.
  • Phase 4 — Stability & throughput: consistency under repetition and load.

Checks are grouped into pillars A (safety & abuse controls, 40%), B (correctness & quality, 35%), C (performance & throughput, 20%) and stability (5%). The live checklist is served by the detection engine — see the Chinese methodology for the current per-check list with evidence fields.

Ranking: evidence over optimism

The rank key is rank v2 and reproducible from public parameters: shrunk median − 90% confidence lower bound, multiplied by a critical-rate Wilson upper-bound penalty. Small samples shrink toward the prior; a single critical finding (model substitution, metering fraud) measurably depresses the key.

Full math and the anti-gaming rules (new-account downweighting, official-sample 3× weight with published seeds) are in the Chinese ranking spec.

Red lines we do not cross

  • Insufficient evidence ≠ problematic. Below the display threshold a gateway is shown as a neutral gray dash — never red, never green.
  • Payment never affects scores, verdicts, ranking or display weight. Subscriptions buy detection capacity and engineering benefits only; architecture-assertion tests enforce this in CI.
  • No fabricated data. If the board is empty, it stays empty until real public detections exist.

The complete methodology series (scoring, ranking, thresholds, threat model, limits) is published in Chinese: /methodology.