Model behavior patterns

How The AI Allocators Differ

A peer-normalized comparison of each model's typical allocation style across eligible official frozen portfolios, with current open positioning shown separately from historical behavior.

Most risk-seeking Claude Fable 5.1 Highest average risk-taking score Most concentrated Grok 4.6 Highest concentration across saved portfolios Most defensive Gemini 3.1 Pro Largest defensive allocation Lowest turnover Claude Opus 5 Smallest average portfolio turnover Most like the group Claude Opus 5 Lowest Portfolio Difference Most different Grok 4.3 Highest Portfolio Difference
Portfolio Difference

How Differently Does Each Model Invest?

The score is the share of a model's allocation that would need to change to match the average portfolio selected by the other models in the same rounds.

Grok 4.3
65.9 / 100
65.9% would need to change 30 same-round comparisons Established sample
GPT-6 Astra
64.6 / 100
64.6% would need to change 2 same-round comparisons Early sample
Grok 4.6
60.7 / 100
60.7% would need to change 30 same-round comparisons Established sample
Gemini 3.1 Pro
58.9 / 100
58.9% would need to change 30 same-round comparisons Established sample
Claude Fable 5.1
54.0 / 100
54.0% would need to change 4 same-round comparisons Early sample
Grok 4.5
48.9 / 100
48.9% would need to change 30 same-round comparisons Established sample
Claude Opus 5
45.5 / 100
45.5% would need to change 30 same-round comparisons Established sample

Combined gives monthly and weekly behavior equal weight. The measured model is excluded when the other models' average portfolio is calculated. This measures portfolio difference, not copying or influence.

Recent-winner tilt

Which Models Follow Recent Winners?

Higher scores mean the model placed more weight on assets that had already outperformed before its portfolio was frozen. The score measures behavior, not whether that decision was correct.

Grok 4.3
30.7 / 100
+19.2 vs peers 5% in top 20% 30 portfolios
Grok 4.6
28.4 / 100
+8.3 vs peers 0% in top 20% 30 portfolios
GPT-6 Astra
28.2 / 100
+17.6 vs peers 0% in top 20% 2 portfolios
Grok 4.5
10.9 / 100
-6.0 vs peers 0% in top 20% 30 portfolios

Combined gives monthly and weekly behavior equal weight. Monthly uses the prior 21 trading sessions; weekly uses the prior 5. S&P 500 and cash receive the neutral score of 50.

Comparison matrix

Distinct Behavior By Model

Every label, sentence, and pill comes from the same deterministic evidence record. Realized investment results are deliberately excluded from allocation-style classification.

Calculation method
Model Distinct behavior Evidence Common exposures
Claude Fable 5.1 Anthropic Provisional
Emerging allocation profile

Emerging pattern across 4 official portfolios. Portfolios averaged 3.0 holdings, a 35.0% largest position, and 82.5% turnover.

Only 4 peer-matched portfolios across 2 independent decision dates are available; stable labels require 8 and 6, respectively.
Pattern still forming3.0 holdings · 35% top83% turnoverNow: CIBR 18% Recent winners 13
Cybersecurity (CIBR) 17.5% avg Semiconductors (SMH) 16.3% avg S&P 500 (SPY) 15.0% avg
GPT-5.5 OpenAI Established pattern
High-risk steady allocator

Risk taking averaged 79.7/100, with a median 7.0 points above same-round peers; the difference had the same direction in 83% of 98 matched portfolios. Portfolios averaged 4.8 holdings, a 34.1% largest position, and 49.1% turnover.

Risk 80/100 · +7 pts vs peers4.8 holdings · 34% top49% turnoverHistorical · retired Recent winners 63
Semiconductors (SMH) 19.1% avg Crude Oil (USO) 8.3% avg Biotechnology (XBI) 7.3% avg
GPT-5.6 Sol OpenAI Moderate evidence
High-risk tactical allocator

Risk taking averaged 73.6/100, with a median 4.3 points above same-round peers; the difference had the same direction in 68% of 69 matched portfolios. Portfolios averaged 3.5 holdings, a 36.1% largest position, and 65.6% turnover.

Risk 74/100 · +4 pts vs peers3.5 holdings · 36% top66% turnoverHistorical · retired Recent winners 14
Semiconductors (SMH) 10.4% avg Cybersecurity (CIBR) 10.3% avg Energy Sector (XLE) 10.0% avg
Grok 4.3 xAI Moderate evidence
Distinctive allocator

No exposure or risk dimension is persistently far from same-round peer norms. Portfolios averaged 3.3 holdings, a 46.6% largest position, and 54.2% turnover.

Near peer mix3.3 holdings · 47% top54% turnoverNow: SPY 40% Recent winners 31
S&P 500 (SPY) 21.2% avg Semiconductors (SMH) 9.5% avg Energy Sector (XLE) 8.1% avg
Grok 4.5 xAI Moderate evidence
Group-aligned allocator

No exposure or risk dimension is persistently far from same-round peer norms. Portfolios averaged 3.8 holdings, a 34.8% largest position, and 57.3% turnover.

Near peer mix3.8 holdings · 35% top57% turnoverNow: SPY 11% Recent winners 11
Energy Sector (XLE) 13.1% avg Cybersecurity (CIBR) 7.1% avg Financials Sector (XLF) 6.8% avg
Claude Opus 5 Anthropic Moderate evidence
Group-aligned allocator

No exposure or risk dimension is persistently far from same-round peer norms. Portfolios averaged 3.7 holdings, a 35.3% largest position, and 50.8% turnover.

Near peer mix3.7 holdings · 35% top51% turnoverNow: SPY 20% Recent winners 18
S&P 500 (SPY) 20.6% avg Healthcare Sector (XLV) 8.7% avg Financials Sector (XLF) 7.9% avg
Claude Opus 4.7 Anthropic Established pattern
Defensive steady allocator

Defensive assets averaged 16.6%, with a median 15.0 percentage points above same-round peers; the difference had the same direction in 76% of 70 matched portfolios. Portfolios averaged 4.9 holdings, a 31.8% largest position, and 50.1% turnover.

Defensive 17% · +15pp vs peers4.9 holdings · 32% top50% turnoverHistorical · retired Recent winners 75
Semiconductors (SMH) 17.6% avg Healthcare Sector (XLV) 11.2% avg Financials Sector (XLF) 7.4% avg
Gemini 3.1 Pro Google Moderate evidence
Concentrated allocator

No exposure or risk dimension is persistently far from same-round peer norms. Portfolios averaged 3.2 holdings, a 40.5% largest position, and 62.3% turnover.

Near peer mix3.2 holdings · 40% top62% turnoverNow: XLU 11% Recent winners 7
Semiconductors (SMH) 13.5% avg S&P 500 (SPY) 11.2% avg Healthcare Sector (XLV) 9.2% avg
Claude Fable 5 Anthropic Moderate evidence
Group-aligned allocator

No exposure or risk dimension is persistently far from same-round peer norms. Portfolios averaged 4.2 holdings, a 31.3% largest position, and 57.6% turnover.

Near peer mix4.2 holdings · 31% top58% turnoverHistorical · retired Recent winners 18
Energy Sector (XLE) 11.1% avg Semiconductors (SMH) 9.9% avg S&P 500 (SPY) 7.9% avg
Grok 4.6 xAI Moderate evidence
Steady allocator

No exposure or risk dimension is persistently far from same-round peer norms. Portfolios averaged 2.2 holdings, a 63.1% largest position, and 52.3% turnover.

Near peer mix2.2 holdings · 63% top52% turnoverNow: SPY 53% Recent winners 28
S&P 500 (SPY) 48.4% avg Cybersecurity (CIBR) 11.6% avg Aerospace and Defense (ITA) 8.6% avg
GPT-6 Astra OpenAI Provisional
Emerging allocation profile

Emerging pattern across 2 official portfolios. Portfolios averaged 2.5 holdings, a 50.0% largest position, and n/a turnover.

Only 2 peer-matched portfolios across 1 independent decision date are available; stable labels require 8 and 6, respectively.
Pattern still forming2.5 holdings · 50% topTurnover buildingNow: SPY 48% Recent winners 28
S&P 500 (SPY) 47.5% avg Agriculture Commodities (DBA) 17.5% avg Copper (CPER) 17.5% avg
Claude Opus 4.8 Anthropic Moderate evidence
Risk-conscious steady allocator

Risk taking averaged 69.2/100, with a median 4.4 points below same-round peers; the difference had the same direction in 74% of 99 matched portfolios. Portfolios averaged 4.3 holdings, a 35.3% largest position, and 43.4% turnover.

Risk 69/100 · −4 pts vs peers4.3 holdings · 35% top43% turnoverHistorical · retired Recent winners 50
S&P 500 (SPY) 18.9% avg Healthcare Sector (XLV) 12.9% avg Financials Sector (XLF) 11.4% avg
Key numbers

Behavior Metrics In One Table

These are cumulative allocation-behavior measures across eligible official saved portfolios. Performance remains available on the weekly and monthly leaderboards, but does not determine these behavior labels.

Model RiskHoldingsTop holdingHigh riskDefensivePortfolio differenceTurnover
Claude Fable 5.1 80.2 / 1003.0035.0%76.3%0.0%54.0 / 10082.5%
GPT-5.5 79.7 / 1004.8234.1%85.3%3.8%55.9 / 10049.1%
GPT-5.6 Sol 73.6 / 1003.4936.1%77.0%8.1%51.5 / 10065.6%
Grok 4.3 73.2 / 1003.2546.6%60.1%6.2%65.9 / 10054.2%
Grok 4.5 72.9 / 1003.7934.8%76.5%6.3%48.9 / 10057.3%
Claude Opus 5 71.4 / 1003.7235.3%55.9%6.5%45.5 / 10050.8%
Claude Opus 4.7 71.3 / 1004.9031.8%63.6%16.6%55.0 / 10050.1%
Gemini 3.1 Pro 71.0 / 1003.2240.5%63.5%14.2%58.9 / 10062.3%
Claude Fable 5 70.8 / 1004.1931.3%60.4%12.1%51.2 / 10057.6%
Grok 4.6 70.5 / 1002.2263.1%40.9%4.8%60.7 / 10052.3%
GPT-6 Astra 70.1 / 1002.5050.0%52.5%0.0%64.6 / 100n/a
Claude Opus 4.8 69.2 / 1004.3435.3%46.5%11.7%77.8 / 10043.4%
Comparative findings

What Stands Out

Each finding is tied to model IDs and metric keys in the generated report.

Claude Fable 5.1, Grok 4.6

Claude Fable 5.1 and Grok 4.6 are different in different ways

Claude Fable 5.1 stands out by risk appetite at 80.2 / 100, while Grok 4.6 stands out by portfolio structure with a 63.1% average largest holding.

risk taking scoreaverage top allocation pct
Gemini 3.1 Pro, Claude Opus 5

Gemini 3.1 Pro and Claude Opus 5 look more risk-managed than the aggressive cohort

Gemini 3.1 Pro has the highest defensive allocation at 14.2%. Claude Opus 5 has the lowest measured turnover at 50.8%.

defensive pctaverage turnover pct
Claude Opus 5

Claude Opus 5 invests most like the group

Claude Opus 5 has the lowest Portfolio Difference at 45.5 / 100. That is the share of allocation that would need to change to match the average portfolio of the other models.

portfolio difference
Methodology

How Behavior Labels And Pills Are Determined

The report is rebuilt from eligible official frozen portfolios during every publication build. No model receives a manually assigned caption, and the model's own descriptive wording cannot assign its label.

For each model and round, CapitalBench subtracts the median behavior-metric value of the other models in that same round. A behavior signal must exceed its published materiality floor, point in the same direction in at least 65% of matched portfolios, and have at least 8 matched portfolios across 6 independent decision dates.

Qualifying signals are ordered by absolute median peer difference divided by their materiality floor, then by persistence and a stable metric key. The strongest exposure or risk signal supplies the label modifier; peer-normalized construction, turnover, or Portfolio Difference supplies the allocation-style noun.

Evidence is “established” only after 16 decision dates and 75% persistence. Opposite material weekly and monthly signals are marked horizon-dependent; a sufficiently sampled reversal under the newest methodology is marked evolving. The four pills always report signature, construction, tempo, and current open positioning (or lifecycle for a retired model). “Typical” uses all eligible history; “Now” uses only currently open portfolios.

Realized returns, ranks, ineligible or pilot runs, market-briefing prose, and free-form rationale wording are not classification inputs. Structured candidate-ledger, forecast, confidence, and key-risk fields are retained as decision-process context when coverage exists, but they do not override allocation evidence. Page-level “most” leader cards use active models only; retired profiles remain available as historical evidence.

Portfolio Difference is one-half of the summed absolute difference between a model's allocation and the average allocation selected by every other model in the same round. The measured model is left out of that average. A score of 42 means about 42% of allocation would need to change to match the group portfolio. Combined scores give monthly and weekly results equal weight and require observations from both horizons. The score measures output difference; it does not show that one model copied or influenced another.

Recent-winner tilt is a separate behavior measure and does not change the archetype. Each eligible asset receives a 0–100 percentile from its return before the decision cutoff, and the model's portfolio weights produce one allocation-weighted score. Current weekly portfolios use 5 trading sessions relative to SPY; monthly portfolios use 21. S&P 500 and cash are neutral at 50, and future returns never enter the calculation.

Method version: capitalbench_behavior_evidence_v2
Peer baseline: leave-one-model-out same-round peer median
Wording provenance: deterministic_source_of_truth
Prompt contract: capitalbench_model_patterns_prompt_v2

Read the full CapitalBench benchmark methodology
Qualification thresholds Published materiality floors

A median same-round peer difference must meet the relevant floor before persistence can qualify it.

  • Risk taking≥ 4 score points
  • Technology≥ 5 percentage points
  • Real assets≥ 5 percentage points
  • International assets≥ 4 percentage points
  • Defensive assets≥ 4 percentage points
  • Cash and duration≥ 4 percentage points
  • S&P 500 core≥ 5 percentage points
  • Largest holding≥ 5 percentage points
  • Holding count≥ 0.5 holdings
0-100 Risk-taking score

Average allocation-weighted risk appetite across all official saved portfolios. Higher means more growth, momentum, cyclical, and high-risk exposure.

count Avg holdings

Average number of non-zero assets in the model's official saved portfolios.

percentage_points Avg top holding

Average size of the largest single holding in each official saved portfolio.

percentage_points High-risk allocation

Average allocation to assets rated as higher risk by the CapitalBench asset risk model.

percentage_points Defensive allocation

Average allocation to cash, bonds, defensive sectors, and other lower-risk ballast.

percentage_points Technology allocation

Average allocation to technology, semiconductors, Nasdaq-style growth, and AI-linked technology exposure.

percentage_points Cash/duration allocation

Average allocation to cash-like assets and duration-sensitive bond exposure.

percentage_points International allocation

Average allocation to non-U.S. country, regional, or international equity exposure.

percentage_points Real assets allocation

Average allocation to commodities, crypto, energy, gold, and other inflation-linked or real-asset groups.

percentage_points S&P 500 core allocation

Average allocation to the S&P 500 benchmark option across official saved portfolios.

0-1 Peer overlap

Legacy API-only average cosine similarity between this model's allocation weights and peer model portfolios in the same rounds.

score_100 Portfolio Difference

The percentage of allocation that would need to change to match the average portfolio selected by the other models in the same rounds.

percentage_points Avg turnover

Average one-half summed absolute allocation change between consecutive same-track portfolios.

score_100 Recent-winner tilt

Equal-weighted monthly and weekly allocation-weighted percentile rank of assets' pre-decision recent returns. A score of 50 is neutral; higher values lean toward recent winners.

percentage_points Top recent-winner allocation

Equal-weighted monthly and weekly portfolio allocation to assets in the top 20% of the applicable pre-decision recent-return window.

rank Avg rank

Average finishing rank across resolved rounds. Lower is better.

points Avg CapitalBench Score

Average model score versus the hindsight-best eligible asset in each resolved round.