Skip to content
effiqefficiency × iq

Project

About effiq

effiq (efficiency × IQ) ranks LLM reasoning variants by capability per real or estimated task dollar. It is a free, open-source, static site.

Who maintains effiq?

effiq is built and maintained by smsheese (GitHub). The project started as a practical answer to one question: which reasoning variant delivers enough capability for the least task dollar. The source code is public, and the daily matrix refresh runs from the same repository. You can reach the maintainer by opening an issue or a discussion on the GitHub repository — that is the fastest route to a response.

Why capability per task dollar?

Benchmark leaderboards rank models by raw capability. They do not answer what a completed task costs. Two variants can sit close together on intelligence while sitting far apart on price, especially once reasoning effort is taken into account. effiq merges public benchmark indexes with published provider prices, computes a task cost per reasoning variant, and ranks the result. Measured metrics stay labeled apart from estimates, and approximations keep their labels in the explorer.

How do I report a problem?

Wrong price, missing model, or a score that looks off — open an issue on github.com/smsheese/effiq, or use the Feedback button in the header. The scoring guide below documents how scores are computed, and the health endpoint reports matrix freshness.

License and attribution

effiq is MIT-licensed. Benchmark indexes and measured task costs come from Artificial Analysis. Provider catalog and live routing metrics come from OpenRouter. Cursor variant pricing comes from Cursor public docs. OpenCode Go pricing comes from OpenCode Go docs. Full attribution lives in the repository README.

Scoring guide

How effiq ranks models

effiq ranks LLM reasoning variants by capability per task dollar. A reasoning variant is one model at one effort level, such as medium or high. The product question is not "which model is smartest" — it is which variant delivers enough capability for the least real or estimated task dollar.

The data matrix behind every number on this page was refreshed (UTC).

Defaults

What are the default ranking settings?

The default ranking settings are an intelligence floor of 40, Effiq Score weights of 35% intelligence, 15% coding, 10% agentic, 30% task cost, 5% latency, and 5% throughput, conservative bounds on, and approximations on. The intelligence floor removes every reasoning variant whose Artificial Analysis Intelligence Index falls below 40, so weak models never enter the table. Conservative bounds mean weak estimates use the worse direction: costs read higher and capability reads lower. With approximations on, rows that rely on interpolated or family estimates stay in the table when their confidence is high enough. These four values are only defaults. Every one of them is adjustable in the explorer, and the table recomputes the moment you move a slider. The table below states each default and what it controls.

SettingDefaultMeaning
Intelligence floor40Variants below this Artificial Analysis Intelligence Index leave the table.
Effiq Score weights35 / 15 / 10 / 30 / 5 / 5Intelligence, coding, agentic, task cost, latency, throughput.
Conservative boundsOnWeak estimates use a worse bound: higher cost, lower capability.
ApproximationsOnIf confidence is high enough, rows with interpolated or family estimates stay in the table.

Evidence

Which sources does effiq use?

effiq uses four primary sources: Artificial Analysis, OpenRouter, Cursor published pricing, and OpenCode Go published pricing, plus CursorBench snapshots for measured coding capability. Artificial Analysis supplies the intelligence, coding, and agentic indexes and the only measured cost-per-task numbers in the matrix. OpenRouter supplies live provider prices, context lengths, and endpoint latency and throughput percentiles for every variant it routes. Cursor pricing supplies subscription-pool variants with their effort and thinking axes. OpenCode Go supplies published per-token prices for 28 open coding models. The nightly merge job folds all of these into one canonical matrix. Derived scores from WhatLLM and LLM Stats stay off until their robots and terms checks pass, and they can never override the primary sources.

SourceWhat it supplies
Artificial AnalysisIntelligence, coding, and agentic indexes. Measured cost per task. Throughput and time to first token.
OpenRouterLive provider prices, context length, and endpoint latency and throughput percentiles.
Cursor PricingSubscription-pool variants, effort and thinking axes, and published token prices.
OpenCode GoPublished per-token prices for 28 open coding models on the $10/month OpenCode Zen plan, with Zen endpoint shapes.
CursorBenchOfficial Cursor agent benchmarks (CursorBench 3.2). Measured coding capability, cost per task, tokens, and steps across 60 configurations.

WhatLLM and LLM Stats adapters stay off until robots and terms checks pass. Their derived scores never override Artificial Analysis or official provider prices.

Ranking

What is the Effiq Score?

The Effiq Score is a 0-100 number that ranks reasoning variants by capability per task dollar under your current settings. Only variants that pass the intelligence floor enter the score. The computation happens in four steps. First, effiq percentile-ranks intelligence, coding, agentic capability, and throughput among the gated set. Second, it log-normalizes task cost and latency and inverts them, so a lower cost and a faster response raise the score instead of lowering it. Third, it combines those normalized values with the metric weights you have set — the default mix is 35% intelligence, 15% coding, 10% agentic, 30% task cost, 5% latency, and 5% throughput. Fourth, it applies evidence-coverage and approximation penalties when the evidence behind a row is thin. Because the inputs change with your profile and sliders, the score is not a fixed leaderboard position.

  1. 1Percentile-rank intelligence, coding, agentic, and throughput among the gated set.
  2. 2Log-normalize task cost and latency, then invert them so lower cost and lower latency raise the score.
  3. 3Combine the normalized values with the user weights.
  4. 4Apply evidence-coverage and approximation penalties if the evidence is thin.

Default weight mix

MetricDefault weightWhat a higher weight favors
Intelligence35%General reasoning on the Artificial Analysis Intelligence Index.
Coding15%Coding benchmarks, including CursorBench when present.
Agentic10%Tool use and multi-step agent benchmarks.
Task cost30%Lower dollar cost per task.
Latency5%Faster first-token and round-trip response.
Throughput5%More output tokens per second.

Capability per dollar

The explorer also shows raw capability per dollar. That value is the domain score divided by task cost. It is not the Effiq Score. Near-zero prices cannot hide other ranking factors in silence.

Workloads

Which usage profiles can I choose?

The explorer includes eight usage profiles: General, Coding, Agents, Math and Science, Finance, Research, Writing and Literature, and Multimodal. A usage profile is a named preset that changes three things at once: the domain evidence mix, the default metric weights, and the token workload used to estimate task cost. A coding workload, for example, prices a coding agent loop, while a research workload prices a long-context retrieval and synthesis task. This matters because a variant can win on one profile and lose on another: a cheap generalist can beat a costly specialist at document analysis and then lose to it on hard reasoning. The profile table below lists each profile, its focus, and its workload label. Pick a profile first, then fine-tune the sliders if the preset is close but not exact.

ProfileFocusWorkload label
GeneralBroad capability with balanced cost.Standard mixed task
CodingSoftware engineering strength and agent loops.Coding agent loop
AgentsTool use, multi-step reliability, and cache.Multi-step agent
Math and ScienceHard reasoning and science questions.Hard reasoning
FinanceNumeric reasoning and document extraction.Document analysis
ResearchLong context, retrieval, and synthesis.Long-context research
Writing and LiteratureStyle, instruction following, and long output.Long-form writing
MultimodalImage, audio, or video plus text.Vision + text

Gaps

How does effiq estimate missing reasoning variants?

Missing reasoning variants follow a five-step estimation ladder, and every estimate keeps its label in the explorer. Many providers publish prices or benchmarks for only some effort levels of the same model. Rather than dropping those variants, effiq fills the gap in a fixed order: use direct measured evidence when it exists for that exact effort; otherwise interpolate between two measured efforts in the same family; otherwise extrapolate from the nearest measured effort; otherwise estimate from the family aggregate plus an effort curve. If even that fails, the field stays empty — effiq never fabricates a value. Each fallback step lowers confidence, and the confidence value feeds the min-confidence filter in the explorer. With conservative ranking on, the estimate uses the less favorable bound: a higher cost or a lower capability number.

  1. 1

    Exact measured variant

    Direct measured evidence for that effort.

  2. 2

    Same-family interpolation

    Estimate between two measured efforts in the same family.

  3. 3

    Nearest-effort extrapolation

    Estimate from the closest measured effort.

  4. 4

    Family aggregate plus effort curve

    Estimate from the family pattern.

  5. 5

    Insufficient data

    No fabricated value.

Cursor

What changes in the Cursor section?

In the Cursor section, CursorBench evidence carries 2.5× weight in domain capability, and measured CursorBench task cost is preferred wherever it exists. The Cursor section is a provider-scoped view of the same ranking engine: it limits the table to variants available on Cursor's subscription pool, which publish their own effort and thinking axes and their own token prices. Because the visitor has chosen Cursor as the channel, the official Cursor agent benchmark (CursorBench 3.2, measured across 60 configurations) becomes the dominant coding signal instead of a shared one, so rankings inside the section can differ from the all-providers view. A banner in the explorer states when this section is active, so the changed weighting is never silent.

  • CursorBench capability scores use 2.5× weight in domain capability.
  • Measured CursorBench task cost is preferred for Cursor models that have it.
  • The banner in the explorer states when this Cursor section is active.

Freshness

How fresh is the data?

A GitHub Actions workflow refreshes the matrix daily at 04:00 UTC. The nightly job re-reads every source, re-merges them into the canonical matrix, and republishes the static site, so a price change at any provider reaches the rankings within a day. The most recent refresh happened (UTC). Freshness is machine-readable as well: the health endpoint reports matrix age and sync status as JSON, the matrix files carry a generatedAt timestamp, and the JSON-LD on both pages states dateModified. If a refresh fails, the site keeps serving the previous matrix rather than an empty table.

Open the model explorer