Qualified Comparison Set

Fairness Scope

Every ranked model in this set is scored only on rounds that all 8 listed models completed. If one model misses a resolved round, that round is excluded from this set for everyone.

All comparison sets
Shared rounds11 Models8 Threshold6 StatusQualified Comparison Set
Equal-run benchmark

Weekly Qualified Comparison Set

Every ranked model in this set completed the same 11 weekly rounds.

11 shared resolved rounds8 equal-run models rankedQualified at 6+ shared roundsNewest included round: CB-2026-08-11-1W
Shared resolved rounds

CapitalBench Score

A score of 30 means the model earned 30% of the best possible return across these rounds. Calculation

GPT-5.6 Sol
Grok 4.5
Grok 4.3
Claude Opus 5
GPT-5.5
Claude Fable 5
Claude Opus 4.8
Gemini 3.1 Pro
S&P 500
Max possible What is this? Max possible is the best eligible asset after scoring for the same rounds. It is a hindsight ceiling, not a model portfolio. hindsight best asset

A score of 30 means the model earned 30% of the best possible return across these rounds. Calculation

GPT-5.6 Sol OpenAI · 11/11 scored rounds
13.8
Grok 4.5 xAI · 11/11 scored rounds
12.8
Grok 4.3 xAI · 11/11 scored rounds
11.4
Claude Opus 5 Anthropic · 11/11 scored rounds
9.6
GPT-5.5 OpenAI · 11/11 scored rounds
8.8
Claude Fable 5 Anthropic · 11/11 scored rounds
8.4
Claude Opus 4.8 Anthropic · 11/11 scored rounds
6.9
Gemini 3.1 Pro Google · 11/11 scored rounds
2.9
S&P 500 S&P 500 · 11/11 scored rounds
17.8
Max possible Hindsight ceiling, not a model portfolio
What is this? Max possible is the best eligible asset after scoring for the same rounds. It is a hindsight ceiling, not a model portfolio. 100.0
11 shared resolved rounds8 equal-run models rankedQualified at 6+ shared roundsNewest included round: CB-2026-08-11-1W
Return context

Average Return Details

Average portfolio return across the same finished rounds.

OpenAI GPT-5.6 Sol
1.45%
xAI Grok 4.5
1.35%
xAI Grok 4.3
1.20%
Anthropic Claude Opus 5
1.01%
OpenAI GPT-5.5
0.92%
Anthropic Claude Fable 5
0.88%
Anthropic Claude Opus 4.8
0.72%
Google Gemini 3.1 Pro
0.30%
S&P S&P 500
1.87%
MAX Max possible What is this? Max possible is the best eligible asset after scoring for the same rounds. It is a hindsight ceiling, not a model portfolio.
10.52%
Excluded for fairness: CB-2026-08-13-1W missing openai-gpt-5-5; CB-2026-08-15-1W missing openai-gpt-5-5; CB-2026-08-18-1W missing openai-gpt-5-5; CB-2026-08-19-1W missing anthropic-claude-opus-4-8, openai-gpt-5-5; CB-2026-08-20-1W missing anthropic-claude-opus-4-8, openai-gpt-5-5; CB-2026-08-21-1W missing anthropic-claude-opus-4-8, openai-gpt-5-5; CB-2026-08-23-1W missing anthropic-claude-opus-4-8, openai-gpt-5-5; CB-2026-08-24-1W missing anthropic-claude-opus-4-8, openai-gpt-5-5; CB-2026-08-25-1W missing anthropic-claude-opus-4-8, openai-gpt-5-5; CB-2026-08-26-1W missing anthropic-claude-opus-4-8, openai-gpt-5-5; CB-2026-08-27-1W missing anthropic-claude-opus-4-8, openai-gpt-5-5 Fairness rule: every ranked model completed every included round. A missed round is excluded from this set for everyone.
Risk and return

Who earned more return for the risk they took?

Average realized return and frozen portfolio risk across the same 11 shared weekly rounds.

Risk and return for models in this comparison setAverage frozen portfolio allocation risk is plotted horizontally and average realized return is plotted vertically across 11 shared rounds.-1.00%0.00%1.00%2.00%3.00%4050607080CapitalBench allocation riskAverage returnS&P 500Beat S&P with lower allocation riskGPT-5.6 SolGrok 4.5Grok 4.3Opus 5GPT-5.5Fable 5Opus 4.8Gemini 3.1
Risk and return for models in this comparison setAverage frozen portfolio allocation risk is plotted horizontally and average realized return is plotted vertically across 11 shared rounds.-1.00%0.00%1.00%2.00%3.00%4050607080CapitalBench allocation riskAverage returnS&P 500GPT-5.6 SolGrok 4.5Grok 4.3Opus 5GPT-5.5Fable 5Opus 4.8Gemini 3.1
GPT-5.6 SolOpenAI
+1.45%average return67.9/100Risk-seeking-0.42 ppversus S&P 50058.8-72.9risk range

Return leaderGPT-5.6 Sol led the models at +1.45% average return with a 67.9/100 risk score.

Benchmark testNo model beat the S&P 500 while taking the same or less allocation risk in this set.

Compare model groups

How do these results compare?

Grok 4.3 ranks first in Aug 19 Weekly. GPT-5.6 Sol ranks first in Jul 24 Weekly. The groups have no completed rounds in common. Aug 19 Weekly includes 11 more rounds, while Jul 24 Weekly includes 11 more rounds. GPT-5.5, Claude Opus 4.8 appear only in Jul 24 Weekly. Grok 4.6 appears only in Aug 19 Weekly.

6models in both 0rounds used by both Changed a littlechange in order Yes top model changed

Aug 19 Weekly is the main published ranking. Jul 24 Weekly also has enough rounds, so compare them to see whether the results hold across different model groups.

Compare these groups
Round audit

Included And Excluded Rounds

Included rounds count toward the score. Excluded rounds are resolved rounds inside this comparison history where at least one set model was missing.

Included rounds CB-2026-07-24-1W, CB-2026-07-27-1W, CB-2026-07-28-1W, CB-2026-07-29-1W, CB-2026-07-30-1W, CB-2026-07-31-1W, CB-2026-08-04-1W, CB-2026-08-05-1W, CB-2026-08-07-1W, CB-2026-08-09-1W, CB-2026-08-11-1W
Excluded for fairness
11 resolved candidate rounds CB-2026-08-13-1W missing openai-gpt-5-5; CB-2026-08-15-1W missing openai-gpt-5-5; CB-2026-08-18-1W missing openai-gpt-5-5; CB-2026-08-19-1W missing anthropic-claude-opus-4-8, openai-gpt-5-5; CB-2026-08-20-1W missing anthropic-claude-opus-4-8, openai-gpt-5-5; CB-2026-08-21-1W missing anthropic-claude-opus-4-8, openai-gpt-5-5; CB-2026-08-23-1W missing anthropic-claude-opus-4-8, openai-gpt-5-5; CB-2026-08-24-1W missing anthropic-claude-opus-4-8, openai-gpt-5-5; CB-2026-08-25-1W missing anthropic-claude-opus-4-8, openai-gpt-5-5; CB-2026-08-26-1W missing anthropic-claude-opus-4-8, openai-gpt-5-5; CB-2026-08-27-1W missing anthropic-claude-opus-4-8, openai-gpt-5-5
Calculation

How The Score Is Calculated

CapitalBench Score equals total model return across included shared rounds divided by total max-possible return across those same rounds, multiplied by 100. Max possible is the best eligible asset in each included round in hindsight.

Scoring details