Your comparison
Your workload. Your price.
Start with the highest-scoring model that fits your budget and estimated capacity. Then compare the bill, the limits and the evidence.
The benchmark
Capability has a price.
Compare both.
See where each model sits on coding performance and cost per task. Switch pricing views to explore the difference a plan could make.
Higher means a better benchmark score. Further left means lower cost per task. The cost axis is logarithmic: each step is 10×. Hollow points without an eligible subscription stay at their API cost. Use arrow keys to explore and Enter for details.
Plans: all on
The plan index
The plan behind the price.
Each model’s lowest estimated cost per task, assuming full use of the selected plan’s quota. Use your comparison above for the bill at your actual workload.
The allowance
How far does the quota go?
How many days of steady-state quota would cover 113 benchmark tasks? These figures describe allowance, not completion time; resets, starting balance and run duration are not modeled.
Behind the numbers
No mystery math.
We compare benchmark costs with plan prices and quotas. Your cost per task depends on how much you use.
Cost per task = monthly price ÷ tasks completed
How we calculate the estimates
Full-use estimates assume you consume the entire modeled quota. Quotas measured on one model may differ on another; reasoning settings, caching, tools and provider rates also affect costs.
Missing evidence and unpriced plans
A note on uncertainty
Where the estimates disagree.
Two conversions of the same quota disagree. These estimates share source data and are not independent validation. The gap is unresolved; it should lower confidence in the conversion.
See individual comparisons
The source material
Follow the numbers.
Explore the benchmarks, quota measurements and pricing references behind this comparison. Full notes in SOURCES.md.