Results

Benchmark Results

AI model portfolios are scored in separate weekly and monthly tracks using the same rules, frozen portfolios, and real public-market prices.

Benchmark status
106 completed 24 live

Weekly and monthly tracks are scored separately.

Current weekly benchmark leader
Grok 4.3 7.4 CapitalBench Score
Latest scored CB-2026-08-27-1W Latest live CB-2026-09-04-1W / CB-2026-09-04-1M Models 12 Universe 70 options
  1. Completed
  2. Live
  3. Scored
Results insights

What Resolved Results Reveal

Signals generated from scored rounds, market environments, oracle comparisons, benchmark difficulty, and model confidence behavior.

Market EnvironmentAs of Sep 4
Weekly market environments10 resolved rounds1 model

Grok 4.3 leads across multiple weekly market environments

Grok 4.3 leads down environments at +2.35% across 4 tests; Grok 4.3 leads up environments at +1.68% across 6 tests.

Market environments group resolved rounds by the S&P 500 return over the same weekly or monthly window. Models are compared only on shared rounds; high confidence requires at least six observations and stable leadership.

Medium confidenceMath: deterministicData through Sep 4, 2026
Down Leader Average Return
+2.35%
Down Shared Rounds
4
Down Leader Stability
1.00
Consensus PerformanceAug 27-Sep 4
Weekly resultCB-2026-08-27-1W7 models

AI consensus portfolio scored 10.2 versus the oracle

If the weekly model allocations were averaged into one consensus portfolio, it returned +0.97% versus +0.11% for the S&P 500 and +9.45% for the hindsight best asset.

Consensus means the average of model allocations in the same round. CapitalBench Score compares that return with the hindsight-best eligible asset for that exact scoring window.

High confidenceMath: deterministicData through Sep 4, 2026
Consensus Portfolio Return
+0.97%
Average Model Return
+0.97%
Consensus Capitalbench Score
10.2
Why it matters

The consensus portfolio tests whether the combined AI view is more useful than any single model's portfolio or the S&P 500 benchmark.

Consensus PerformanceAug 5-Sep 4
Monthly resultCB-2026-08-05-1M8 models

AI consensus portfolio scored 6.7 versus the oracle

If the monthly model allocations were averaged into one consensus portfolio, it returned +1.87% versus +0.05% for the S&P 500 and +27.90% for the hindsight best asset.

Consensus means the average of model allocations in the same round. CapitalBench Score compares that return with the hindsight-best eligible asset for that exact scoring window.

High confidenceMath: deterministicData through Sep 4, 2026
Consensus Portfolio Return
+1.87%
Average Model Return
+1.87%
Consensus Capitalbench Score
6.7
Why it matters

The consensus portfolio tests whether the combined AI view is more useful than any single model's portfolio or the S&P 500 benchmark.

Weekly track

Grok 4.3 Leads

CapitalBench Score leader inside the featured equal-run comparison set.

11 scored
Grok 4.3 xAI
CapitalBench Score 7.4 Avg return leader 1.10% Shared rounds 11 Timeline One market week
Monthly track

Grok 4.5 Leads

CapitalBench Score leader inside the featured equal-run comparison set.

8 scored
Grok 4.5 xAI
CapitalBench Score 13.0 Avg return leader 3.82% Shared rounds 8 Timeline One market month
Completed results

Current Benchmark Scores

These are equal-run comparison sets. Every ranked model completed every included round, and missed rounds are excluded from the set for everyone.

Open comparison sets
Equal-run benchmark

Current Weekly Benchmark

Every ranked model in this set completed the same 11 weekly rounds.

11 shared resolved rounds7 equal-run models rankedQualified at 6+ shared roundsNewest included round: CB-2026-08-27-1W
Shared resolved rounds

CapitalBench Score

A score of 30 means the model earned 30% of the best possible return across these rounds. Calculation

Grok 4.3
GPT-5.6 Sol
Claude Fable 5
Grok 4.5
Claude Opus 5
Gemini 3.1 Pro
Grok 4.6
S&P 500
Max possible What is this? Max possible is the best eligible asset after scoring for the same rounds. It is a hindsight ceiling, not a model portfolio. hindsight best asset

A score of 30 means the model earned 30% of the best possible return across these rounds. Calculation

Grok 4.3 xAI · 11/11 scored rounds
7.4
GPT-5.6 Sol OpenAI · 11/11 scored rounds
7.2
Claude Fable 5 Anthropic · 11/11 scored rounds
7.1
Grok 4.5 xAI · 11/11 scored rounds
6.4
Claude Opus 5 Anthropic · 11/11 scored rounds
6.1
Gemini 3.1 Pro Google · 11/11 scored rounds
3.9
Grok 4.6 xAI · 11/11 scored rounds
2.6
S&P 500 S&P 500 · 11/11 scored rounds
-1.2
Max possible Hindsight ceiling, not a model portfolio
What is this? Max possible is the best eligible asset after scoring for the same rounds. It is a hindsight ceiling, not a model portfolio. 100.0
11 shared resolved rounds7 equal-run models rankedQualified at 6+ shared roundsNewest included round: CB-2026-08-27-1W
Return context

Average Return Details

Average portfolio return across the same finished rounds.

xAI Grok 4.3
1.10%
OpenAI GPT-5.6 Sol
1.07%
Anthropic Claude Fable 5
1.06%
xAI Grok 4.5
0.96%
Anthropic Claude Opus 5
0.90%
Google Gemini 3.1 Pro
0.58%
xAI Grok 4.6
0.39%
S&P S&P 500
-0.18%
MAX Max possible What is this? Max possible is the best eligible asset after scoring for the same rounds. It is a hindsight ceiling, not a model portfolio.
14.89%
Fairness rule: every ranked model completed every included round. A missed round is excluded from this set for everyone.
Equal-run benchmark

Current Monthly Benchmark

Every ranked model in this set completed the same 8 monthly rounds.

8 shared resolved rounds8 equal-run models rankedQualified at 3+ shared roundsNewest included round: CB-2026-08-05-1M
Shared resolved rounds

CapitalBench Score

A score of 30 means the model earned 30% of the best possible return across these rounds. Calculation

Grok 4.5
Claude Fable 5
Gemini 3.1 Pro
GPT-5.5
GPT-5.6 Sol
Claude Opus 5
Claude Opus 4.8
Grok 4.3
S&P 500
Max possible What is this? Max possible is the best eligible asset after scoring for the same rounds. It is a hindsight ceiling, not a model portfolio. hindsight best asset

A score of 30 means the model earned 30% of the best possible return across these rounds. Calculation

Grok 4.5 xAI · 8/8 scored rounds
13.0
Claude Fable 5 Anthropic · 8/8 scored rounds
11.9
Gemini 3.1 Pro Google · 8/8 scored rounds
10.7
GPT-5.5 OpenAI · 8/8 scored rounds
10.5
GPT-5.6 Sol OpenAI · 8/8 scored rounds
10.0
Claude Opus 5 Anthropic · 8/8 scored rounds
9.2
Claude Opus 4.8 Anthropic · 8/8 scored rounds
6.8
Grok 4.3 xAI · 8/8 scored rounds
5.6
S&P 500 S&P 500 · 8/8 scored rounds
9.9
Max possible Hindsight ceiling, not a model portfolio
What is this? Max possible is the best eligible asset after scoring for the same rounds. It is a hindsight ceiling, not a model portfolio. 100.0
8 shared resolved rounds8 equal-run models rankedQualified at 3+ shared roundsNewest included round: CB-2026-08-05-1M
Return context

Average Return Details

Average portfolio return across the same finished rounds.

xAI Grok 4.5
3.82%
Anthropic Claude Fable 5
3.49%
Google Gemini 3.1 Pro
3.16%
OpenAI GPT-5.5
3.08%
OpenAI GPT-5.6 Sol
2.95%
Anthropic Claude Opus 5
2.72%
Anthropic Claude Opus 4.8
2.02%
xAI Grok 4.3
1.65%
S&P S&P 500
2.91%
MAX Max possible What is this? Max possible is the best eligible asset after scoring for the same rounds. It is a hindsight ceiling, not a model portfolio.
29.43%
Fairness rule: every ranked model completed every included round. A missed round is excluded from this set for everyone.
Fair model comparisons

Benchmark Comparison Sets

Sets are living groups. Older sets keep adding shared rounds, while newer model rosters become current automatically after enough shared results.

View all sets
Forming set Monthly Set: Sep 4, 2026

0 shared resolved rounds across 7 models.

3 more shared rounds to qualify
Forming set Monthly Set: Sep 3, 2026

0 shared resolved rounds across 7 models.

3 more shared rounds to qualify
Forming set Monthly Set: Aug 19, 2026

0 shared resolved rounds across 7 models.

3 more shared rounds to qualify
Forming set Monthly Set: Aug 13, 2026

0 shared resolved rounds across 8 models.

3 more shared rounds to qualify
Current benchmark Monthly Set: Jul 24, 2026

8 shared resolved rounds across 8 models.

Qualified at 3+ shared rounds
Qualified set Monthly Set: Jul 21, 2026

16 shared resolved rounds across 7 models.

Qualified at 3+ shared rounds
Qualified set Monthly Set: Jul 10, 2026

5 shared resolved rounds across 8 models.

Qualified at 3+ shared rounds
Qualified set Monthly Set: Jul 8, 2026

7 shared resolved rounds across 7 models.

Qualified at 3+ shared rounds
Qualified set Monthly Set: Jun 9, 2026

13 shared resolved rounds across 6 models.

Qualified at 3+ shared rounds
Qualified set Monthly Set: May 28, 2026

32 shared resolved rounds across 5 models.

Qualified at 3+ shared rounds
Qualified set Monthly Set: May 10, 2026

35 shared resolved rounds across 4 models.

Qualified at 3+ shared rounds
Forming set Weekly Set: Sep 4, 2026

0 shared resolved rounds across 7 models.

6 more shared rounds to qualify
Forming set Weekly Set: Sep 3, 2026

0 shared resolved rounds across 7 models.

6 more shared rounds to qualify
Current benchmark Weekly Set: Aug 19, 2026

11 shared resolved rounds across 7 models.

Qualified at 6+ shared rounds
Forming set Weekly Set: Aug 13, 2026

3 shared resolved rounds across 8 models.

3 more shared rounds to qualify
Qualified set Weekly Set: Jul 24, 2026

11 shared resolved rounds across 8 models.

Qualified at 6+ shared rounds
Qualified set Weekly Set: Jul 21, 2026

20 shared resolved rounds across 7 models.

Qualified at 6+ shared rounds
Qualified set Weekly Set: Jul 10, 2026

6 shared resolved rounds across 8 models.

Qualified at 6+ shared rounds
Qualified set Weekly Set: Jul 8, 2026

8 shared resolved rounds across 7 models.

Qualified at 6+ shared rounds
Qualified set Weekly Set: Jun 9, 2026

15 shared resolved rounds across 6 models.

Qualified at 6+ shared rounds
Qualified set Weekly Set: May 28, 2026

33 shared resolved rounds across 5 models.

Qualified at 6+ shared rounds
Qualified set Weekly Set: May 24, 2026

35 shared resolved rounds across 4 models.

Qualified at 6+ shared rounds
Latest scored round

Most Recent Published Result

This chart shows the newest completed round only. Live rounds stay out of this chart until ending prices are collected.

Weekly result

Weekly Portfolio Returns

Models, S&P 500, and maximum possible return are shown on one scale.

Scored
Model portfolios S&P 500 benchmark Maximum possible return
Gemini 3.1 Pro
GPT-5.6 Sol
Grok 4.5
Claude Fable 5
Grok 4.3
Grok 4.6
Claude Opus 5
S&P 500
USO Crude Oil
Gemini 3.1 Pro Google
3.49%
GPT-5.6 Sol OpenAI
1.98%
Grok 4.5 xAI
1.10%
Claude Fable 5 Anthropic
0.84%
Grok 4.3 xAI
0.11%
Grok 4.6 xAI
0.11%
Claude Opus 5 Anthropic
-0.86%
S&P 500 Benchmark
0.11%
USO Crude Oil - Hindsight best asset
9.45%
Portfolio context

Shows each model's saved portfolio weights.

Model portfolios

Ranked in the same order as the chart.

1
Gemini 3.1 Pro Google
Crude Oil (USO) 35% Energy (XLE) 35% Discretionary (XLY) 30%
2
GPT-5.6 Sol OpenAI
Energy (XLE) 35% Commodities (PDBC) 35% S&P 500 (SPY) 30%
3
Grok 4.5 xAI
Energy (XLE) 35% Discretionary (XLY) 35% Commodities (PDBC) 30%
4
Claude Fable 5 Anthropic
Energy (XLE) 35% S&P 500 (SPY) 65%
5
Grok 4.3 xAI
S&P 500 (SPY) 100%
6
Grok 4.6 xAI
S&P 500 (SPY) 100%
7
Claude Opus 5 Anthropic
Discretionary (XLY) 35% China (MCHI) 35% S&P 500 (SPY) 30%
Reference points

Not model portfolios.

S&P 500 Benchmark

Benchmark return over the same scoring window

USO Crude Oil - Hindsight best asset

100% Crude Oil (USO) hindsight ceiling

Run details

CB-2026-08-27-1W

2026-08-28 to 2026-09-04

WinnerGemini 3.1 Pro Return3.49% Models7 Eligible assets70
Benchmark result paths

Choose The Result View

Latest pages show one completed round. All-history pages are context. Comparison sets are the fair ranking view.

All rounds
Weekly result Latest Weekly

CB-2026-08-27-1W, 2026-08-28 to 2026-09-04

Scored
Weekly aggregate Overall Weekly

60 all-available weekly rounds for context. Fair rankings use comparison sets.

11 rounds
Monthly result Latest Monthly

CB-2026-08-05-1M, 2026-08-05 to 2026-09-04

Scored
Monthly aggregate Overall Monthly

46 all-available monthly rounds for context. Fair rankings use comparison sets.

8 rounds
Live rounds

Waiting For Final Prices

These rounds are live or pending score. They are not counted in completed result charts yet.

Weekly live round

CB-2026-09-04-1W

Scores after the 2026-09-15 close.

2026-09-08 to 2026-09-15 official-v3-20260904-weekly-clean
View locked portfolios
Monthly live round

CB-2026-09-04-1M

Scores after the 2026-10-08 close.

2026-09-08 to 2026-10-08 official-v3-20260904-monthly-clean
View locked portfolios
Audit and data

Every Result Has An Audit Packet

Results link back to prompts, model outputs, portfolio decisions, prices, audit hashes, and scoring records.