effiqefficiency × iqMethodology

← Back to explorer

Methodology

effiq ranks model reasoning variants for token-efficiency: high capability per real or estimated task dollar. The default gate keeps Artificial Analysis Intelligence Index ≥ 40. Users can move that floor and reweight metrics.

Sources

WhatLLM and LLM Stats scrapers are configured but skipped until robots/terms checks pass. Their derived scores never override AA or official provider prices.

Efficiency Score

Among variants that pass the intelligence floor:

  1. Percentile-rank intelligence, coding, agentic, and throughput.
  2. Log-normalize and invert task cost and latency.
  3. Combine with user weights (default 35/15/10/30/5/5).
  4. Apply evidence-coverage and approximation penalties.

Raw capability per dollar (domain score ÷ task cost) is shown separately so near-zero prices cannot silently dominate.

Usage profiles

Eight profiles change domain evidence mix, default weights, and workload token assumptions: General, Coding, Agents, Math & Science, Finance, Research, Writing & Literature, Multimodal.

Approximations

  1. Exact measured variant
  2. Same-family interpolation between efforts
  3. Nearest-effort extrapolation
  4. Family aggregate + effort curve
  5. Insufficient data (no fabricated value)

Conservative ranking uses the less favorable bound of an estimate (higher cost / lower capability) when confidence is limited.

Sync

npm run sync refreshes the matrix. On deployment, a GitHub Actions scheduled workflow runs daily at 04:00 UTC and refreshes the matrix automatically.