CapitalBench

The benchmark for AI capital allocation

Each AI model gets the same market brief, builds one portfolio in one response, and uses no browsing or tools. Real market prices determine the result.

See how AI models perform against each other, how they invest and take risk, and how they perform in the real market.

Read the CapitalBench Manifesto
Benchmark results

Which models are performing best?

Monthly and weekly tracks stay separate.

Current Monthly Benchmark

CapitalBench Score

A score of 30 means the model earned 30% of the best possible return across these rounds. Calculation

Grok 4.5
Claude Fable 5
Gemini 3.1 Pro
GPT-5.5
GPT-5.6 Sol
Claude Opus 5
Claude Opus 4.8
Grok 4.3
S&P 500
Max possible What is this? Max possible is the best eligible asset after scoring for the same rounds. It is a hindsight ceiling, not a model portfolio. hindsight best asset

A score of 30 means the model earned 30% of the best possible return across these rounds. Calculation

Grok 4.5 xAI · 8/8 scored rounds
13.0
Claude Fable 5 Anthropic · 8/8 scored rounds
11.9
Gemini 3.1 Pro Google · 8/8 scored rounds
10.7
GPT-5.5 OpenAI · 8/8 scored rounds
10.5
GPT-5.6 Sol OpenAI · 8/8 scored rounds
10.0
Claude Opus 5 Anthropic · 8/8 scored rounds
9.2
Claude Opus 4.8 Anthropic · 8/8 scored rounds
6.8
Grok 4.3 xAI · 8/8 scored rounds
5.6
S&P 500 S&P 500 · 8/8 scored rounds
9.9
Max possible Hindsight ceiling, not a model portfolio
What is this? Max possible is the best eligible asset after scoring for the same rounds. It is a hindsight ceiling, not a model portfolio. 100.0
8 shared resolved rounds8 equal-run models rankedQualified at 3+ shared roundsNewest included round: CB-2026-08-05-1M
Return context

Average Return Details

Average portfolio return across the same finished rounds.

xAI Grok 4.5
3.82%
Anthropic Claude Fable 5
3.49%
Google Gemini 3.1 Pro
3.16%
OpenAI GPT-5.5
3.08%
OpenAI GPT-5.6 Sol
2.95%
Anthropic Claude Opus 5
2.72%
Anthropic Claude Opus 4.8
2.02%
xAI Grok 4.3
1.65%
S&P S&P 500
2.91%
MAX Max possible What is this? Max possible is the best eligible asset after scoring for the same rounds. It is a hindsight ceiling, not a model portfolio.
29.43%
Fairness rule: every ranked model completed every included round. A missed round is excluded from this set for everyone.

Monthly and weekly are separate comparison tracks. Scores are never mixed across horizons. Read scoring rules

Model Risk Benchmark

Which AI models take the most risk?

Compare how aggressively each model allocates capital across its official portfolios.

Updated through September 8, 2026Higher means more risk-seeking, not better.
GPT-5.5OpenAI
79.7/100Risk-seeking
98portfolios34.1%largest position3.8%defensive assets
GPT-5.6 SolOpenAI
73.6/100Risk-seeking
69portfolios36.1%largest position8.1%defensive assets
Grok 4.3xAI
73.2/100Risk-seeking
130portfolios46.6%largest position6.2%defensive assets
Grok 4.5xAI
72.9/100Risk-seeking
75portfolios34.8%largest position6.3%defensive assets
Claude Opus 5Anthropic
71.4/100Risk-seeking
54portfolios35.3%largest position6.5%defensive assets
Claude Opus 4.7Anthropic
71.3/100Risk-seeking
70portfolios31.8%largest position16.6%defensive assets
Gemini 3.1 ProGoogle
71.0/100Risk-seeking
130portfolios40.5%largest position14.2%defensive assets
Claude Fable 5Anthropic
70.8/100Risk-seeking
84portfolios31.3%largest position12.1%defensive assets
Grok 4.6xAI
70.5/100Risk-seeking
32portfolios63.1%largest position4.8%defensive assets
Claude Opus 4.8Anthropic
69.2/100Risk-seeking
99portfolios35.3%largest position11.7%defensive assets
Portfolio Difference

Which AI models invest most differently from the group?

Compare each model's portfolio with the choices made by the other models in the same rounds.

Updated through September 8, 2026Overall: 50% monthly / 50% weekly
Grok 4.6xAI
60.7/100
30same-round comparisonsEstablished sample
Grok 4.5xAI
48.9/100
30same-round comparisonsEstablished sample

Different does not mean better or worse. The score compares portfolio outputs; it does not prove copying, influence, or intent.

Explore Portfolio Difference
Performance by market

Who leads when markets rise or fall?

Monthly leaders vs. the S&P 500.

Monthly snapshot Updated Sep 4 4 of 5 market types comparable
Market fell S&P < -1.0%
Leader Grok 4.3
Average return -1.22%
Versus S&P 500 -1.90%
+0.68 pts vs S&P 6 results · Some evidence
Market rose S&P > +1.0%
Leader Grok 4.5
Average return +4.11%
Versus S&P 500 +3.90%
+0.22 pts vs S&P 6 results · Some evidence
AI positioning

What are AI models doing right now?

Live allocations before the next official score.

As of September 4, 2026
Current risk appetite As of September 4, 2026 75.5/100 Risk-seeking / Broad risk seeking
Consensus allocation As of September 4, 2026 20.0% Cybersecurity (CIBR) average live weight
Risk shift As of September 4, 2026 -3.0 Change vs Sep 3 portfolios
Model agreement As of September 4, 2026 Tight 4.0 point dispersion
Current risk appetite 75.5/100 As of September 4, 2026 / Risk-seeking
As of September 4, 2026 Broad risk seeking Combined view of monthly and weekly model portfolios for this date.
Unscored portfolios 68.9/100 Separate read across every open portfolio before official scoring.
Largest current allocations
Cybersecurity (CIBR) 20.0% S&P 500 (SPY) 13.2% Software (IGV) 12.5% Agriculture Commodities (DBA) 11.8% Aerospace and Defense (ITA) 7.5% Materials Sector (XLB) 7.1%
Regime mix
Broad and cyclical equity 39.3% Growth and technology 32.5% Real assets and inflation 23.6% Defensive equity 4.6%
Trust and proof

What makes each model test comparable?

Same brief. Same choices. One response. No tools. Real market prices.

  1. Step 1 Same report

    Every model reads the same market report.

  2. Step 2 Same choices

    Every model chooses from the same 70 assets.

  3. Step 3 One response, then locked

    No browsing, tools, or follow-up prompts. The model's submitted portfolio is frozen before results are known.

  4. Step 4 Fixed wait window

    The frozen portfolio sits untouched for 7 days or 1 month.

  5. Step 5 Prices score it

    Real ending prices decide which model did best.

Benchmark universe

What can models choose from?

The active roster, asset menu, horizons, and open rounds.

Models 12
Asset choices 70
Round lengths 2
Open rounds 24
Protocol Single-turn Non-agentic calls
Latest official results

What happened in the latest scored rounds?

Finished monthly and weekly rounds scored against real market returns.

Monthly result1 of 46
Monthly official result

Monthly result scored Sep 4

Same-window returns, ranked after final prices.

Scored
Model portfolios S&P 500 benchmark Maximum possible return
Claude Fable 5
Grok 4.5
GPT-5.5
GPT-5.6 Sol
Claude Opus 5
Gemini 3.1 Pro
Grok 4.3
Claude Opus 4.8
S&P 500
ETHA Ethereum ETF
Claude Fable 5 Anthropic
4.16%
Grok 4.5 xAI
3.70%
GPT-5.5 OpenAI
3.62%
GPT-5.6 Sol OpenAI
3.29%
Claude Opus 5 Anthropic
0.88%
Gemini 3.1 Pro Google
0.21%
Grok 4.3 xAI
0.05%
Claude Opus 4.8 Anthropic
-0.94%
S&P 500 Benchmark
0.05%
ETHA Ethereum ETF - Hindsight best asset
27.90%
Portfolio context

Shows each model's saved portfolio weights.

Model portfolios

Ranked in the same order as the chart.

1
Claude Fable 5 Anthropic
Semiconductors (SMH) 25% Software (IGV) 25% Taiwan (EWT) 20% Biotechnology (XBI) 20% S&P 500 (SPY) 10%
2
Grok 4.5 xAI
Semiconductors (SMH) 30% Biotechnology (XBI) 25% Taiwan (EWT) 20% S&P 500 (SPY) 25%
3
GPT-5.5 OpenAI
Biotechnology (XBI) 30% Semiconductors (SMH) 25% Taiwan (EWT) 20% Cybersecurity (CIBR) 15% Financials (XLF) 10%
4
GPT-5.6 Sol OpenAI
Semiconductors (SMH) 50% Biotechnology (XBI) 50%
5
Claude Opus 5 Anthropic
S&P 500 (SPY) 40% Healthcare (XLV) 20% Financials (XLF) 20% Equal-Weight S&P 500 (RSP) 20%
6
Gemini 3.1 Pro Google
Technology (XLK) 30% Industrials (XLI) 30% Gold (IAU) 20% Healthcare (XLV) 20%
7
Grok 4.3 xAI
S&P 500 (SPY) 100%
8
Claude Opus 4.8 Anthropic
S&P 500 (SPY) 50% Technology (XLK) 30% Industrials (XLI) 20%
Reference points

Not model portfolios.

S&P 500 Benchmark

Benchmark return over the same scoring window

ETHA Ethereum ETF - Hindsight best asset

100% Ethereum ETF (ETHA) hindsight ceiling

Official scored round

Monthly result scored Sep 4

Audit ID: CB-2026-08-05-1M

ScoredSep 4WindowAug 5 to Sep 4Models8Asset choices70LeaderClaude Fable 5HorizonMonthly
Live dashboard

What is still in progress?

Open portfolios, interim returns, and upcoming score dates.

Open rounds 24 19 monthly / 5 weekly
Frozen portfolios 174 39 assets currently held
Latest close Sep 4 Live returns update before final scoring
Next score Sep 8 Official results publish after ending prices
Audit packet

How can you verify the benchmark?

Round packets expose the report, prompt, portfolios, prices, hashes, and result status.