Benchmark Forming

Fairness Scope

Every ranked model in this set is scored only on rounds that all 8 listed models completed. If one model misses a resolved round, that round is excluded from this set for everyone.

All comparison sets
Shared rounds3 Models8 Threshold6 StatusBenchmark Forming
Equal-run benchmark

Weekly Benchmark Forming

Every ranked model in this set completed the same 3 weekly rounds.

3 shared resolved rounds8 equal-run models ranked3 more shared rounds to qualifyNewest included round: CB-2026-08-18-1W
Shared resolved rounds

CapitalBench Score

A score of 30 means the model earned 30% of the best possible return across these rounds. Calculation

Grok 4.3
Claude Fable 5
GPT-5.6 Sol
Claude Opus 5
Gemini 3.1 Pro
Grok 4.5
Claude Opus 4.8
Grok 4.6
S&P 500
Max possible What is this? Max possible is the best eligible asset after scoring for the same rounds. It is a hindsight ceiling, not a model portfolio. hindsight best asset

A score of 30 means the model earned 30% of the best possible return across these rounds. Calculation

Grok 4.3 xAI · 3/3 scored rounds
8.6
Claude Fable 5 Anthropic · 3/3 scored rounds
5.7
GPT-5.6 Sol OpenAI · 3/3 scored rounds
4.8
Claude Opus 5 Anthropic · 3/3 scored rounds
4.8
Gemini 3.1 Pro Google · 3/3 scored rounds
4.3
Grok 4.5 xAI · 3/3 scored rounds
4.3
Claude Opus 4.8 Anthropic · 3/3 scored rounds
-0.0
Grok 4.6 xAI · 3/3 scored rounds
-2.0
S&P 500 S&P 500 · 3/3 scored rounds
-4.1
Max possible Hindsight ceiling, not a model portfolio
What is this? Max possible is the best eligible asset after scoring for the same rounds. It is a hindsight ceiling, not a model portfolio. 100.0
3 shared resolved rounds8 equal-run models ranked3 more shared rounds to qualifyNewest included round: CB-2026-08-18-1W
Return context

Average Return Details

Average portfolio return across the same finished rounds.

xAI Grok 4.3
2.35%
Anthropic Claude Fable 5
1.55%
OpenAI GPT-5.6 Sol
1.31%
Anthropic Claude Opus 5
1.30%
Google Gemini 3.1 Pro
1.18%
xAI Grok 4.5
1.17%
Anthropic Claude Opus 4.8
-0.01%
xAI Grok 4.6
-0.53%
S&P S&P 500
-1.12%
MAX Max possible What is this? Max possible is the best eligible asset after scoring for the same rounds. It is a hindsight ceiling, not a model portfolio.
27.24%
Excluded for fairness: CB-2026-08-19-1W missing anthropic-claude-opus-4-8; CB-2026-08-20-1W missing anthropic-claude-opus-4-8; CB-2026-08-21-1W missing anthropic-claude-opus-4-8; CB-2026-08-23-1W missing anthropic-claude-opus-4-8; CB-2026-08-24-1W missing anthropic-claude-opus-4-8; CB-2026-08-25-1W missing anthropic-claude-opus-4-8; CB-2026-08-26-1W missing anthropic-claude-opus-4-8; CB-2026-08-27-1W missing anthropic-claude-opus-4-8 Fairness rule: every ranked model completed every included round. A missed round is excluded from this set for everyone. Forming: this set becomes the Current Weekly Benchmark at 6 shared resolved rounds.
Risk and return

Who earned more return for the risk they took?

Average realized return and frozen portfolio risk across the same 3 shared weekly rounds.

Risk and return for models in this comparison setAverage frozen portfolio allocation risk is plotted horizontally and average realized return is plotted vertically across 3 shared rounds.-2.00%0.00%2.00%4.00%5060708090CapitalBench allocation riskAverage returnS&P 500Beat S&P with lower allocation riskGrok 4.3Fable 5GPT-5.6 SolOpus 5Gemini 3.1Grok 4.5Opus 4.8Grok 4.6
Risk and return for models in this comparison setAverage frozen portfolio allocation risk is plotted horizontally and average realized return is plotted vertically across 3 shared rounds.-2.00%0.00%2.00%4.00%5060708090CapitalBench allocation riskAverage returnS&P 500Grok 4.3Fable 5GPT-5.6 SolOpus 5Gemini 3.1Grok 4.5Opus 4.8Grok 4.6
Grok 4.3xAI
+2.35%average return72.2/100Risk-seeking+3.47 ppversus S&P 50052.9-87.5risk range

Return leaderGrok 4.3 led the models at +2.35% average return with a 72.2/100 risk score.

Benchmark testClaude Opus 4.8 beat the S&P 500 while taking no more allocation risk.

Compare model groups

How do these results compare?

Grok 4.3 ranks first in both groups. The groups share 3 completed rounds. Aug 19 Weekly includes 8 more rounds. Claude Opus 4.8 appears only in Aug 13 Weekly.

7models in both 3rounds used by both Hardly changedchange in order No top model changed

Use Aug 19 Weekly as the more reliable ranking because it has 11 completed rounds. Aug 13 Weekly has 3 and needs 3 more before it has enough evidence to become the main ranking.

Compare these groups
Round audit

Included And Excluded Rounds

Included rounds count toward the score. Excluded rounds are resolved rounds inside this comparison history where at least one set model was missing.

3 more to qualify
Included rounds CB-2026-08-13-1W, CB-2026-08-15-1W, CB-2026-08-18-1W
Excluded for fairness
8 resolved candidate rounds CB-2026-08-19-1W missing anthropic-claude-opus-4-8; CB-2026-08-20-1W missing anthropic-claude-opus-4-8; CB-2026-08-21-1W missing anthropic-claude-opus-4-8; CB-2026-08-23-1W missing anthropic-claude-opus-4-8; CB-2026-08-24-1W missing anthropic-claude-opus-4-8; CB-2026-08-25-1W missing anthropic-claude-opus-4-8; CB-2026-08-26-1W missing anthropic-claude-opus-4-8; CB-2026-08-27-1W missing anthropic-claude-opus-4-8
Calculation

How The Score Is Calculated

CapitalBench Score equals total model return across included shared rounds divided by total max-possible return across those same rounds, multiplied by 100. Max possible is the best eligible asset in each included round in hindsight.

Scoring details