Grok 4.3 leads down environments at +2.35% across 4 tests; Grok 4.3 leads up environments at +1.68% across 6 tests.
Market environments group resolved rounds by the S&P 500 return over the same weekly or monthly window. Models are compared only on shared rounds; high confidence requires at least six observations and stable leadership.
Medium confidenceMath: deterministicData through Sep 4, 2026
If the weekly model allocations were averaged into one consensus portfolio, it returned +0.97% versus +0.11% for the S&P 500 and +9.45% for the hindsight best asset.
Consensus means the average of model allocations in the same round. CapitalBench Score compares that return with the hindsight-best eligible asset for that exact scoring window.
High confidenceMath: deterministicData through Sep 4, 2026
If the monthly model allocations were averaged into one consensus portfolio, it returned +1.87% versus +0.05% for the S&P 500 and +27.90% for the hindsight best asset.
Consensus means the average of model allocations in the same round. CapitalBench Score compares that return with the hindsight-best eligible asset for that exact scoring window.
High confidenceMath: deterministicData through Sep 4, 2026
Every ranked model in this set completed the same 11 weekly rounds.
11 shared resolved rounds7 equal-run models rankedQualified at 6+ shared roundsNewest included round: CB-2026-08-27-1W
Shared resolved rounds
CapitalBench Score
A score of 30 means the model earned 30% of the best possible return across these rounds.Calculation
-1.20.050.0100.0
7.4
7.2
7.1
6.4
6.1
3.9
2.6
-1.2
100.0
Grok 4.3
GPT-5.6 Sol
Claude Fable 5
Grok 4.5
Claude Opus 5
Gemini 3.1 Pro
Grok 4.6
S&PS&P 500
MAXMax possibleWhat is this?Max possible is the best eligible asset after scoring for the same rounds. It is a hindsight ceiling, not a model portfolio.hindsight best asset
A score of 30 means the model earned 30% of the best possible return across these rounds.Calculation
Grok 4.3xAI · 11/11 scored rounds
7.4
GPT-5.6 SolOpenAI · 11/11 scored rounds
7.2
Claude Fable 5Anthropic · 11/11 scored rounds
7.1
Grok 4.5xAI · 11/11 scored rounds
6.4
Claude Opus 5Anthropic · 11/11 scored rounds
6.1
Gemini 3.1 ProGoogle · 11/11 scored rounds
3.9
Grok 4.6xAI · 11/11 scored rounds
2.6
S&PS&P 500S&P 500 · 11/11 scored rounds
-1.2
MAXMax possibleHindsight ceiling, not a model portfolio
What is this?Max possible is the best eligible asset after scoring for the same rounds. It is a hindsight ceiling, not a model portfolio. 100.0
11 shared resolved rounds7 equal-run models rankedQualified at 6+ shared roundsNewest included round: CB-2026-08-27-1W
Return context
Average Return Details
Average portfolio return across the same finished rounds.
Grok 4.3
1.10%
GPT-5.6 Sol
1.07%
Claude Fable 5
1.06%
Grok 4.5
0.96%
Claude Opus 5
0.90%
Gemini 3.1 Pro
0.58%
Grok 4.6
0.39%
S&PS&P 500
-0.18%
MAXMax possibleWhat is this?Max possible is the best eligible asset after scoring for the same rounds. It is a hindsight ceiling, not a model portfolio.
14.89%
Fairness rule: every ranked model completed every included round. A missed round is excluded from this set for everyone.
Equal-run benchmark
Current Monthly Benchmark
Every ranked model in this set completed the same 8 monthly rounds.
8 shared resolved rounds8 equal-run models rankedQualified at 3+ shared roundsNewest included round: CB-2026-08-05-1M
Shared resolved rounds
CapitalBench Score
A score of 30 means the model earned 30% of the best possible return across these rounds.Calculation
0.025.050.075.0100.0
13.0
11.9
10.7
10.5
10.0
9.2
6.8
5.6
9.9
100.0
Grok 4.5
Claude Fable 5
Gemini 3.1 Pro
GPT-5.5
GPT-5.6 Sol
Claude Opus 5
Claude Opus 4.8
Grok 4.3
S&PS&P 500
MAXMax possibleWhat is this?Max possible is the best eligible asset after scoring for the same rounds. It is a hindsight ceiling, not a model portfolio.hindsight best asset
A score of 30 means the model earned 30% of the best possible return across these rounds.Calculation
Grok 4.5xAI · 8/8 scored rounds
13.0
Claude Fable 5Anthropic · 8/8 scored rounds
11.9
Gemini 3.1 ProGoogle · 8/8 scored rounds
10.7
GPT-5.5OpenAI · 8/8 scored rounds
10.5
GPT-5.6 SolOpenAI · 8/8 scored rounds
10.0
Claude Opus 5Anthropic · 8/8 scored rounds
9.2
Claude Opus 4.8Anthropic · 8/8 scored rounds
6.8
Grok 4.3xAI · 8/8 scored rounds
5.6
S&PS&P 500S&P 500 · 8/8 scored rounds
9.9
MAXMax possibleHindsight ceiling, not a model portfolio
What is this?Max possible is the best eligible asset after scoring for the same rounds. It is a hindsight ceiling, not a model portfolio. 100.0
8 shared resolved rounds8 equal-run models rankedQualified at 3+ shared roundsNewest included round: CB-2026-08-05-1M
Return context
Average Return Details
Average portfolio return across the same finished rounds.
Grok 4.5
3.82%
Claude Fable 5
3.49%
Gemini 3.1 Pro
3.16%
GPT-5.5
3.08%
GPT-5.6 Sol
2.95%
Claude Opus 5
2.72%
Claude Opus 4.8
2.02%
Grok 4.3
1.65%
S&PS&P 500
2.91%
MAXMax possibleWhat is this?Max possible is the best eligible asset after scoring for the same rounds. It is a hindsight ceiling, not a model portfolio.
29.43%
Fairness rule: every ranked model completed every included round. A missed round is excluded from this set for everyone.
Fair model comparisons
Benchmark Comparison Sets
Sets are living groups. Older sets keep adding shared rounds, while newer model rosters become current automatically
after enough shared results.