Readable signals in the latest generated feed.
Readable signals from the AI capital allocation benchmark
Daily findings from model portfolios, scoring windows, AI Risk Appetite, benchmark difficulty, consensus positioning, and model behavior.
Most recent close or result date used by the engine.
Findings backed by deterministic calculations and direct evidence.
Insights generated without LLM interpretation.
What The Benchmark Is Showing Now
Each card includes the calculation source, evidence links, and why the signal may matter to investors, allocators, traders, and AI researchers.
Grok 4.3 invests most differently from the group
Claude Opus 5 is most like the group at 45.5/100. Monthly and weekly behavior receive equal weight.
Portfolio Difference measures how much allocation would need to change to match the average portfolio selected by the other models in the same rounds. Different does not mean better.
- Highest Portfolio Difference
- 65.9
- Lowest Portfolio Difference
- 45.5
AI consensus portfolio scored 10.2 versus the oracle
If the weekly model allocations were averaged into one consensus portfolio, it returned +0.97% versus +0.11% for the S&P 500 and +9.45% for the hindsight best asset.
Consensus means the average of model allocations in the same round. CapitalBench Score compares that return with the hindsight-best eligible asset for that exact scoring window.
- Consensus Portfolio Return
- +0.97%
- Average Model Return
- +0.97%
- Consensus Capitalbench Score
- 10.2
AI consensus portfolio scored 6.7 versus the oracle
If the monthly model allocations were averaged into one consensus portfolio, it returned +1.87% versus +0.05% for the S&P 500 and +27.90% for the hindsight best asset.
Consensus means the average of model allocations in the same round. CapitalBench Score compares that return with the hindsight-best eligible asset for that exact scoring window.
- Consensus Portfolio Return
- +1.87%
- Average Model Return
- +1.87%
- Consensus Capitalbench Score
- 6.7
Weekly round had +13.95% asset dispersion
The best scored asset returned +9.45%, the worst returned -4.50%, and +51.43% of the universe was positive. The S&P 500 ranked 32 out of 70 options.
Asset dispersion is the gap between the best and worst eligible assets in the same round. Wider dispersion makes missed allocation choices more costly.
- Oracle Return
- +9.45%
- Worst Asset Return
- -4.50%
- Positive Universe Share
- +51.4%
Monthly round had +38.47% asset dispersion
The best scored asset returned +27.90%, the worst returned -10.57%, and +48.57% of the universe was positive. The S&P 500 ranked 32 out of 70 options.
Asset dispersion is the gap between the best and worst eligible assets in the same round. Wider dispersion makes missed allocation choices more costly.
- Oracle Return
- +27.9%
- Worst Asset Return
- -10.6%
- Positive Universe Share
- +48.6%
Models found the weekly oracle asset
The hindsight best asset was Crude Oil (USO) at +9.45%. 1 of 7 models held it, with +5.00% average allocation. The largest allocation came from Gemini 3.1 Pro at +35.00%.
Oracle means the best eligible asset in hindsight for that round. Models do not know it when portfolios are frozen.
- Oracle Asset Holder Count
- 1.00
- Average Oracle Asset Allocation
- 5.00
Models missed the monthly oracle asset
The hindsight best asset was Ethereum ETF (ETHA) at +27.90%. 0 of 8 models held it, with +0.00% average allocation.
Oracle means the best eligible asset in hindsight for that round. Models do not know it when portfolios are frozen.
- Oracle Asset Holder Count
- 0.00
- Average Oracle Asset Allocation
- 0.00
High-confidence model calls have underperformed lower-confidence calls
Across resolved official results, submissions at or above the median confidence of 0.58 averaged +0.13%, while lower-confidence submissions averaged +0.52%.
Confidence is the model's own 0-1 self-reported confidence at submission time, compared with later realized returns.
- High Confidence Average Return
- +0.13%
- Low Confidence Average Return
- +0.52%
- High Confidence Average Capitalbench Score
- -2.4
Gemini 3.1 Pro's result was driven by Crude Oil
In the latest weekly result, Crude Oil contributed +3.31% to Gemini 3.1 Pro's portfolio. The largest drag came from Consumer Discretionary Sector at -0.59%.
Attribution multiplies each frozen holding's weight by its asset return to show what helped or hurt the model portfolio.
- Largest Positive Contribution
- +3.31%
- Largest Negative Contribution
- -0.59%
Claude Fable 5's result was driven by Taiwan Equities
In the latest monthly result, Taiwan Equities contributed +2.06% to Claude Fable 5's portfolio. The largest drag came from Semiconductors at -0.12%.
Attribution multiplies each frozen holding's weight by its asset return to show what helped or hurt the model portfolio.
- Largest Positive Contribution
- +2.06%
- Largest Negative Contribution
- -0.12%
Model allocation styles are separating into clear behavior profiles
Claude Fable 5.1 has the highest average risk-taking score at 80.2/100. Grok 4.6 has the largest average top holding at +63.12%. Claude Opus 4.8 has the lowest measured turnover at +43.35%.
Momentum exposure measures how much of the frozen portfolio went into assets that had already been recent winners before the model made its allocation.
- Highest Average Risk Taking Score
- 80.2/100
- Largest Average Top Holding
- 63.1
- Lowest Average Turnover
- 43.4
Claude Opus 5 has the strongest current monthly recent-winner tilt
Its score is 22.8 out of 100, with 0.0% in the top recent-return quintile. Gemini 3.1 Pro is lowest at 2.4.
Momentum exposure measures how much of the frozen portfolio went into assets that had already been recent winners before the model made its allocation.
- Leader Recent Winner Tilt Score
- 22.8
- Leader Top Recent Winner Quintile Allocation
- 0.00
- Leader Peer Delta
- 16.0
GPT-6 Astra has the strongest current weekly recent-winner tilt
Its score is 36.2 out of 100, with 0.0% in the top recent-return quintile. Grok 4.5 is lowest at 3.7.
Momentum exposure measures how much of the frozen portfolio went into assets that had already been recent winners before the model made its allocation.
- Leader Recent Winner Tilt Score
- 36.2
- Leader Top Recent Winner Quintile Allocation
- 0.00
- Leader Peer Delta
- 21.9
Grok 4.3 leads across multiple weekly market environments
Grok 4.3 leads down environments at +2.35% across 4 tests; Grok 4.3 leads up environments at +1.68% across 6 tests.
Market environments group resolved rounds by the S&P 500 return over the same weekly or monthly window. Models are compared only on shared rounds; high confidence requires at least six observations and stable leadership.
- Down Leader Average Return
- +2.35%
- Down Shared Rounds
- 4
- Down Leader Stability
- 1.00
Live AI risk posture is risk-seeking
The newest live portfolios have a deterministic risk-taking score of 75.5 out of 100.
Risk-taking score is allocation-based, not performance-based: higher means more weight in growth, momentum, cyclical, and higher-risk assets.
- Live Risk Taking Score
- 75.5/100
Monthly model leadership changes with the S&P 500 environment
Grok 4.3 leads down environments at -1.22% across 6 tests; Grok 4.5 leads up environments at +4.11% across 6 tests.
Market environments group resolved rounds by the S&P 500 return over the same weekly or monthly window. Models are compared only on shared rounds; high confidence requires at least six observations and stable leadership.
- Down Leader Average Return
- -1.22%
- Down Shared Rounds
- 6
- Down Leader Stability
- 1.00
Weekly and monthly AI portfolios point to different regimes
The newest weekly portfolios lean toward real assets and inflation, while the newest monthly portfolios lean toward broad and cyclical equity.
Horizon agreement compares the newest weekly and monthly live portfolios to see whether short- and longer-window model stances line up.
- Weekly Top Regime Allocation
- 37.9
- Monthly Top Regime Allocation
- 60.7
Grok 4.5 leads when the S&P 500 is positive
The model averaged +4.11% across 6 shared-cohort rounds drawn from 24 resolved monthly up environments. The sample meets publication thresholds.
Market environments group resolved rounds by the S&P 500 return over the same weekly or monthly window. Models are compared only on shared rounds; high confidence requires at least six observations and stable leadership.
- Average Model Return
- +4.11%
- Average S&P 500 Return
- +3.90%
- Capitalbench Score
- 14.0
Grok 4.3 leads when the S&P 500 is positive
The model averaged +1.68% across 6 shared-cohort rounds drawn from 22 resolved weekly up environments. The sample meets publication thresholds.
Market environments group resolved rounds by the S&P 500 return over the same weekly or monthly window. Models are compared only on shared rounds; high confidence requires at least six observations and stable leadership.
- Average Model Return
- +1.68%
- Average S&P 500 Return
- +3.40%
- Capitalbench Score
- 14.7
Grok 4.3 leads when the S&P 500 is negative
The model averaged +2.35% across 4 shared-cohort rounds drawn from 17 resolved weekly down environments. The sample meets publication thresholds.
Market environments group resolved rounds by the S&P 500 return over the same weekly or monthly window. Models are compared only on shared rounds; high confidence requires at least six observations and stable leadership.
- Average Model Return
- +2.35%
- Average S&P 500 Return
- -1.01%
- Capitalbench Score
- 23.5
Grok 4.3 leads when the S&P 500 is negative
The model averaged -1.22% across 6 shared-cohort rounds drawn from 8 resolved monthly down environments. The sample meets publication thresholds.
Market environments group resolved rounds by the S&P 500 return over the same weekly or monthly window. Models are compared only on shared rounds; high confidence requires at least six observations and stable leadership.
- Average Model Return
- -1.22%
- Average S&P 500 Return
- -1.90%
- Capitalbench Score
- -6.7
Live AI portfolios are concentrated in Cybersecurity (CIBR)
Across the newest live weekly and monthly portfolios, Cybersecurity (CIBR) is the largest aggregate allocation at +20.00%.
Aggregate allocation averages the newest live model portfolios before final scores are known.
- Aggregate Live Allocation
- 20.0
Grok 4.5 has the strongest weekly score floor
Its lowest CapitalBench Score across 3 tested market directions is 7.2, with at least 4 model observations in each included direction.
Market environments group resolved rounds by the S&P 500 return over the same weekly or monthly window. Models are compared only on shared rounds; high confidence requires at least six observations and stable leadership.
- Score Floor
- 7.2
- Average Model Return
- +0.84%
- Directions Covered
- 3
Claude Opus 4.8 has the strongest monthly score floor
Its lowest CapitalBench Score across 3 tested market directions is -14.2, with at least 6 model observations in each included direction.
Market environments group resolved rounds by the S&P 500 return over the same weekly or monthly window. Models are compared only on shared rounds; high confidence requires at least six observations and stable leadership.
- Score Floor
- -14.2
- Average Model Return
- -0.45%
- Directions Covered
- 3
Grok 4.5 has the strongest live alpha
Using the latest available interim close, Grok 4.5 in CB-2026-08-11-1M is ahead of the S&P 500 by +5.66 percentage points, while GPT-5.5 in CB-2026-08-07-1M is at -2.93 percentage points.
Live alpha is interim model return minus interim S&P 500 return. It is provisional until the round reaches its official score date.
- Best Live Alpha
- 5.66
- Worst Live Alpha
- -2.93
Gemini 3.1 Pro changes most between monthly up and down environments
The model averaged -5.08% in down environments and +3.25% in up environments, a 8.3 percentage-point gap.
Market environments group resolved rounds by the S&P 500 return over the same weekly or monthly window. Models are compared only on shared rounds; high confidence requires at least six observations and stable leadership.
- Average Return Gap
- 8.33
- Down Average Return
- -5.08%
- Up Average Return
- +3.25%
Claude Opus 4.8 changes most between weekly up and down environments
The model averaged -0.40% in down environments and +1.01% in up environments, a 1.4 percentage-point gap.
Market environments group resolved rounds by the S&P 500 return over the same weekly or monthly window. Models are compared only on shared rounds; high confidence requires at least six observations and stable leadership.
- Average Return Gap
- 1.41
- Down Average Return
- -0.40%
- Up Average Return
- +1.01%
GPT-5.6 Sol leads when the S&P 500 is flat
The model averaged +0.84% across 6 shared-cohort rounds drawn from 21 resolved weekly flat environments. The sample meets publication thresholds.
Market environments group resolved rounds by the S&P 500 return over the same weekly or monthly window. Models are compared only on shared rounds; high confidence requires at least six observations and stable leadership.
- Average Model Return
- +0.84%
- Average S&P 500 Return
- +0.14%
- Capitalbench Score
- 8.4
Claude Opus 4.8 leads when the S&P 500 is flat
The model averaged -2.79% across 11 shared-cohort rounds drawn from 14 resolved monthly flat environments. The sample meets publication thresholds.
Market environments group resolved rounds by the S&P 500 return over the same weekly or monthly window. Models are compared only on shared rounds; high confidence requires at least six observations and stable leadership.
- Average Model Return
- -2.79%
- Average S&P 500 Return
- +0.25%
- Capitalbench Score
- -14.2
What The Engine Looks For
The engine is designed to surface useful behavior and performance patterns, not generic market commentary.
Market Environment
Which models lead, remain consistent, or change most across resolved down, flat, and up S&P 500 environments.
Model Behavior
How models are allocating before outcomes are known, including momentum chasing and allocation style.
Benchmark Difficulty
How hard a scoring window was, based on the spread between the best, worst, and broad market outcomes.
Consensus Performance
Whether the average AI portfolio performed well against the S&P 500 and the hindsight-best asset in the same round.
Oracle Comparison
Whether models found, missed, or underweighted the asset that later turned out to be best.
Performance Attribution
Which holdings drove a model's realized result after the frozen portfolio was scored.
How Insights Are Produced
Deterministic calculations are the source of truth. LLM-assisted wording can polish selected titles and summaries, but it cannot change calculations, evidence links, round context, or benchmark facts.
Latest generation: Sep 5, 2026, 1:34 PM UTC
- 1 Build the input packet
Collect public rounds, official portfolios, results, live marks, asset risk ratings, and benchmark sets.
- 2 Run deterministic math
Calculate consensus performance, benchmark difficulty, market-environment results, risk posture, similarity, attribution, and live paths.
- 3 Attach evidence
Every insight links back to round pages, leaderboard pages, scoring files, or methodology pages.
- 4 Validate before publishing
The feed must pass schema checks before the website and API expose it.