Insights

Readable signals from the AI capital allocation benchmark

Daily findings from model portfolios, scoring windows, AI Risk Appetite, benchmark difficulty, consensus positioning, and model behavior.

Published insights 29

Readable signals in the latest generated feed.

Data through Sep 4, 2026

Most recent close or result date used by the engine.

High confidence 12

Findings backed by deterministic calculations and direct evidence.

Deterministic math 29

Insights generated without LLM interpretation.

Latest feed

What The Benchmark Is Showing Now

Each card includes the calculation source, evidence links, and why the signal may matter to investors, allocators, traders, and AI researchers.

API Docs
Portfolio Difference As of Sep 8
50% monthly / 50% weekly 7 models

Grok 4.3 invests most differently from the group

Claude Opus 5 is most like the group at 45.5/100. Monthly and weekly behavior receive equal weight.

Portfolio Difference measures how much allocation would need to change to match the average portfolio selected by the other models in the same rounds. Different does not mean better.

Medium confidenceMath: deterministicData through Sep 8, 2026
Highest Portfolio Difference
65.9
Lowest Portfolio Difference
45.5
Consensus Performance Aug 27-Sep 4
Weekly result CB-2026-08-27-1W 7 models Oracle: Crude Oil (USO), +9.45% Resolved result

AI consensus portfolio scored 10.2 versus the oracle

If the weekly model allocations were averaged into one consensus portfolio, it returned +0.97% versus +0.11% for the S&P 500 and +9.45% for the hindsight best asset.

Consensus means the average of model allocations in the same round. CapitalBench Score compares that return with the hindsight-best eligible asset for that exact scoring window.

High confidenceMath: deterministicData through Sep 4, 2026
Consensus Portfolio Return
+0.97%
Average Model Return
+0.97%
Consensus Capitalbench Score
10.2
Consensus Performance Aug 5-Sep 4
Monthly result CB-2026-08-05-1M 8 models Oracle: Ethereum ETF (ETHA), +27.90% Resolved result

AI consensus portfolio scored 6.7 versus the oracle

If the monthly model allocations were averaged into one consensus portfolio, it returned +1.87% versus +0.05% for the S&P 500 and +27.90% for the hindsight best asset.

Consensus means the average of model allocations in the same round. CapitalBench Score compares that return with the hindsight-best eligible asset for that exact scoring window.

High confidenceMath: deterministicData through Sep 4, 2026
Consensus Portfolio Return
+1.87%
Average Model Return
+1.87%
Consensus Capitalbench Score
6.7
Benchmark Difficulty Aug 27-Sep 4
Weekly result CB-2026-08-27-1W 7 models Oracle: Crude Oil (USO), +9.45% Resolved result

Weekly round had +13.95% asset dispersion

The best scored asset returned +9.45%, the worst returned -4.50%, and +51.43% of the universe was positive. The S&P 500 ranked 32 out of 70 options.

Asset dispersion is the gap between the best and worst eligible assets in the same round. Wider dispersion makes missed allocation choices more costly.

High confidenceMath: deterministicData through Sep 4, 2026
Oracle Return
+9.45%
Worst Asset Return
-4.50%
Positive Universe Share
+51.4%
Benchmark Difficulty Aug 5-Sep 4
Monthly result CB-2026-08-05-1M 8 models Oracle: Ethereum ETF (ETHA), +27.90% Resolved result

Monthly round had +38.47% asset dispersion

The best scored asset returned +27.90%, the worst returned -10.57%, and +48.57% of the universe was positive. The S&P 500 ranked 32 out of 70 options.

Asset dispersion is the gap between the best and worst eligible assets in the same round. Wider dispersion makes missed allocation choices more costly.

High confidenceMath: deterministicData through Sep 4, 2026
Oracle Return
+27.9%
Worst Asset Return
-10.6%
Positive Universe Share
+48.6%
Oracle Comparison Aug 27-Sep 4
Weekly result CB-2026-08-27-1W 7 models Oracle: Crude Oil (USO), +9.45% Resolved result

Models found the weekly oracle asset

The hindsight best asset was Crude Oil (USO) at +9.45%. 1 of 7 models held it, with +5.00% average allocation. The largest allocation came from Gemini 3.1 Pro at +35.00%.

Oracle means the best eligible asset in hindsight for that round. Models do not know it when portfolios are frozen.

High confidenceMath: deterministicData through Sep 4, 2026
Oracle Asset Holder Count
1.00
Average Oracle Asset Allocation
5.00
Oracle Comparison Aug 5-Sep 4
Monthly result CB-2026-08-05-1M 8 models Oracle: Ethereum ETF (ETHA), +27.90% Resolved result

Models missed the monthly oracle asset

The hindsight best asset was Ethereum ETF (ETHA) at +27.90%. 0 of 8 models held it, with +0.00% average allocation.

Oracle means the best eligible asset in hindsight for that round. Models do not know it when portfolios are frozen.

High confidenceMath: deterministicData through Sep 4, 2026
Oracle Asset Holder Count
0.00
Average Oracle Asset Allocation
0.00
Confidence Calibration As of Sep 4
All resolved official results 106 resolved rounds 673 scored results Median confidence 0.58 Resolved history

High-confidence model calls have underperformed lower-confidence calls

Across resolved official results, submissions at or above the median confidence of 0.58 averaged +0.13%, while lower-confidence submissions averaged +0.52%.

Confidence is the model's own 0-1 self-reported confidence at submission time, compared with later realized returns.

High confidenceMath: deterministicData through Sep 4, 2026
High Confidence Average Return
+0.13%
Low Confidence Average Return
+0.52%
High Confidence Average Capitalbench Score
-2.4
Performance Attribution Aug 27-Sep 4
Weekly result CB-2026-08-27-1W 7 models Model: Gemini 3.1 Pro Resolved result

Gemini 3.1 Pro's result was driven by Crude Oil

In the latest weekly result, Crude Oil contributed +3.31% to Gemini 3.1 Pro's portfolio. The largest drag came from Consumer Discretionary Sector at -0.59%.

Attribution multiplies each frozen holding's weight by its asset return to show what helped or hurt the model portfolio.

High confidenceMath: deterministicData through Sep 4, 2026
Largest Positive Contribution
+3.31%
Largest Negative Contribution
-0.59%
Performance Attribution Aug 5-Sep 4
Monthly result CB-2026-08-05-1M 8 models Model: Claude Fable 5 Resolved result

Claude Fable 5's result was driven by Taiwan Equities

In the latest monthly result, Taiwan Equities contributed +2.06% to Claude Fable 5's portfolio. The largest drag came from Semiconductors at -0.12%.

Attribution multiplies each frozen holding's weight by its asset return to show what helped or hurt the model portfolio.

High confidenceMath: deterministicData through Sep 4, 2026
Largest Positive Contribution
+2.06%
Largest Negative Contribution
-0.12%
Model Behavior As of Sep 4
Model behavior profiles 12 models

Model allocation styles are separating into clear behavior profiles

Claude Fable 5.1 has the highest average risk-taking score at 80.2/100. Grok 4.6 has the largest average top holding at +63.12%. Claude Opus 4.8 has the lowest measured turnover at +43.35%.

Momentum exposure measures how much of the frozen portfolio went into assets that had already been recent winners before the model made its allocation.

High confidenceMath: deterministicData through Sep 4, 2026
Highest Average Risk Taking Score
80.2/100
Largest Average Top Holding
63.1
Lowest Average Turnover
43.4
Model Behavior Sep 4-Oct 8
Monthly live round CB-2026-09-04-1M 7 models Live portfolios

Claude Opus 5 has the strongest current monthly recent-winner tilt

Its score is 22.8 out of 100, with 0.0% in the top recent-return quintile. Gemini 3.1 Pro is lowest at 2.4.

Momentum exposure measures how much of the frozen portfolio went into assets that had already been recent winners before the model made its allocation.

High confidenceMath: deterministicData through Sep 4, 2026
Leader Recent Winner Tilt Score
22.8
Leader Top Recent Winner Quintile Allocation
0.00
Leader Peer Delta
16.0
Market Environment As of Sep 4
Weekly market environments 10 resolved rounds 1 model Ready sample

Grok 4.3 leads across multiple weekly market environments

Grok 4.3 leads down environments at +2.35% across 4 tests; Grok 4.3 leads up environments at +1.68% across 6 tests.

Market environments group resolved rounds by the S&P 500 return over the same weekly or monthly window. Models are compared only on shared rounds; high confidence requires at least six observations and stable leadership.

Medium confidenceMath: deterministicData through Sep 4, 2026
Down Leader Average Return
+2.35%
Down Shared Rounds
4
Down Leader Stability
1.00
Risk Regime As of Sep 4
Latest live portfolios 2 live rounds 14 models Live portfolios

Live AI risk posture is risk-seeking

The newest live portfolios have a deterministic risk-taking score of 75.5 out of 100.

Risk-taking score is allocation-based, not performance-based: higher means more weight in growth, momentum, cyclical, and higher-risk assets.

Medium confidenceMath: deterministicData through Sep 4, 2026
Live Risk Taking Score
75.5/100
Market Environment As of Sep 4
Monthly market environments 12 resolved rounds 2 models Ready sample

Monthly model leadership changes with the S&P 500 environment

Grok 4.3 leads down environments at -1.22% across 6 tests; Grok 4.5 leads up environments at +4.11% across 6 tests.

Market environments group resolved rounds by the S&P 500 return over the same weekly or monthly window. Models are compared only on shared rounds; high confidence requires at least six observations and stable leadership.

Medium confidenceMath: deterministicData through Sep 4, 2026
Down Leader Average Return
-1.22%
Down Shared Rounds
6
Down Leader Stability
1.00
Horizon Agreement As of Sep 4
Latest live portfolios 2 live rounds 14 models Live portfolios

Weekly and monthly AI portfolios point to different regimes

The newest weekly portfolios lean toward real assets and inflation, while the newest monthly portfolios lean toward broad and cyclical equity.

Horizon agreement compares the newest weekly and monthly live portfolios to see whether short- and longer-window model stances line up.

Medium confidenceMath: deterministicData through Sep 4, 2026
Weekly Top Regime Allocation
37.9
Monthly Top Regime Allocation
60.7
Current Positioning As of Sep 4
Latest live portfolios 2 live rounds 14 models Live portfolios

Live AI portfolios are concentrated in Cybersecurity (CIBR)

Across the newest live weekly and monthly portfolios, Cybersecurity (CIBR) is the largest aggregate allocation at +20.00%.

Aggregate allocation averages the newest live model portfolios before final scores are known.

Medium confidenceMath: deterministicData through Sep 4, 2026
Aggregate Live Allocation
20.0
Live Performance As of Sep 4
Open-round interim performance 22 open rounds 10 models Interim, not final

Grok 4.5 has the strongest live alpha

Using the latest available interim close, Grok 4.5 in CB-2026-08-11-1M is ahead of the S&P 500 by +5.66 percentage points, while GPT-5.5 in CB-2026-08-07-1M is at -2.93 percentage points.

Live alpha is interim model return minus interim S&P 500 return. It is provisional until the round reaches its official score date.

Medium confidenceMath: deterministicData through Sep 4, 2026
Best Live Alpha
5.66
Worst Live Alpha
-2.93
Insight families

What The Engine Looks For

The engine is designed to surface useful behavior and performance patterns, not generic market commentary.

12 signals

Market Environment

Which models lead, remain consistent, or change most across resolved down, flat, and up S&P 500 environments.

3 signals

Model Behavior

How models are allocating before outcomes are known, including momentum chasing and allocation style.

2 signals

Benchmark Difficulty

How hard a scoring window was, based on the spread between the best, worst, and broad market outcomes.

2 signals

Consensus Performance

Whether the average AI portfolio performed well against the S&P 500 and the hindsight-best asset in the same round.

2 signals

Oracle Comparison

Whether models found, missed, or underweighted the asset that later turned out to be best.

2 signals

Performance Attribution

Which holdings drove a model's realized result after the frozen portfolio was scored.

Method

How Insights Are Produced

Deterministic calculations are the source of truth. LLM-assisted wording can polish selected titles and summaries, but it cannot change calculations, evidence links, round context, or benchmark facts.

Latest generation: Sep 5, 2026, 1:34 PM UTC

  1. 1 Build the input packet

    Collect public rounds, official portfolios, results, live marks, asset risk ratings, and benchmark sets.

  2. 2 Run deterministic math

    Calculate consensus performance, benchmark difficulty, market-environment results, risk posture, similarity, attribution, and live paths.

  3. 3 Attach evidence

    Every insight links back to round pages, leaderboard pages, scoring files, or methodology pages.

  4. 4 Validate before publishing

    The feed must pass schema checks before the website and API expose it.