How The AI Allocators Differ
A peer-normalized comparison of each model's typical allocation style across eligible official frozen portfolios, with current open positioning shown separately from historical behavior.
How Differently Does Each Model Invest?
The score is the share of a model's allocation that would need to change to match the average portfolio selected by the other models in the same rounds.
Combined gives monthly and weekly behavior equal weight. The measured model is excluded when the other models' average portfolio is calculated. This measures portfolio difference, not copying or influence.
Which Models Follow Recent Winners?
Higher scores mean the model placed more weight on assets that had already outperformed before its portfolio was frozen. The score measures behavior, not whether that decision was correct.
Combined gives monthly and weekly behavior equal weight. Monthly uses the prior 21 trading sessions; weekly uses the prior 5. S&P 500 and cash receive the neutral score of 50.
Distinct Behavior By Model
Every label, sentence, and pill comes from the same deterministic evidence record. Realized investment results are deliberately excluded from allocation-style classification.
Emerging pattern across 4 official portfolios. Portfolios averaged 3.0 holdings, a 35.0% largest position, and 82.5% turnover.
Only 4 peer-matched portfolios across 2 independent decision dates are available; stable labels require 8 and 6, respectively.Risk taking averaged 79.7/100, with a median 7.0 points above same-round peers; the difference had the same direction in 83% of 98 matched portfolios. Portfolios averaged 4.8 holdings, a 34.1% largest position, and 49.1% turnover.
Risk taking averaged 73.6/100, with a median 4.3 points above same-round peers; the difference had the same direction in 68% of 69 matched portfolios. Portfolios averaged 3.5 holdings, a 36.1% largest position, and 65.6% turnover.
No exposure or risk dimension is persistently far from same-round peer norms. Portfolios averaged 3.3 holdings, a 46.6% largest position, and 54.2% turnover.
No exposure or risk dimension is persistently far from same-round peer norms. Portfolios averaged 3.8 holdings, a 34.8% largest position, and 57.3% turnover.
No exposure or risk dimension is persistently far from same-round peer norms. Portfolios averaged 3.7 holdings, a 35.3% largest position, and 50.8% turnover.
Defensive assets averaged 16.6%, with a median 15.0 percentage points above same-round peers; the difference had the same direction in 76% of 70 matched portfolios. Portfolios averaged 4.9 holdings, a 31.8% largest position, and 50.1% turnover.
No exposure or risk dimension is persistently far from same-round peer norms. Portfolios averaged 3.2 holdings, a 40.5% largest position, and 62.3% turnover.
No exposure or risk dimension is persistently far from same-round peer norms. Portfolios averaged 4.2 holdings, a 31.3% largest position, and 57.6% turnover.
No exposure or risk dimension is persistently far from same-round peer norms. Portfolios averaged 2.2 holdings, a 63.1% largest position, and 52.3% turnover.
Emerging pattern across 2 official portfolios. Portfolios averaged 2.5 holdings, a 50.0% largest position, and n/a turnover.
Only 2 peer-matched portfolios across 1 independent decision date are available; stable labels require 8 and 6, respectively.Risk taking averaged 69.2/100, with a median 4.4 points below same-round peers; the difference had the same direction in 74% of 99 matched portfolios. Portfolios averaged 4.3 holdings, a 35.3% largest position, and 43.4% turnover.
Behavior Metrics In One Table
These are cumulative allocation-behavior measures across eligible official saved portfolios. Performance remains available on the weekly and monthly leaderboards, but does not determine these behavior labels.
| Model | Risk | Holdings | Top holding | High risk | Defensive | Portfolio difference | Turnover |
|---|---|---|---|---|---|---|---|
| Claude Fable 5.1 | 80.2 / 100 | 3.00 | 35.0% | 76.3% | 0.0% | 54.0 / 100 | 82.5% |
| GPT-5.5 | 79.7 / 100 | 4.82 | 34.1% | 85.3% | 3.8% | 55.9 / 100 | 49.1% |
| GPT-5.6 Sol | 73.6 / 100 | 3.49 | 36.1% | 77.0% | 8.1% | 51.5 / 100 | 65.6% |
| Grok 4.3 | 73.2 / 100 | 3.25 | 46.6% | 60.1% | 6.2% | 65.9 / 100 | 54.2% |
| Grok 4.5 | 72.9 / 100 | 3.79 | 34.8% | 76.5% | 6.3% | 48.9 / 100 | 57.3% |
| Claude Opus 5 | 71.4 / 100 | 3.72 | 35.3% | 55.9% | 6.5% | 45.5 / 100 | 50.8% |
| Claude Opus 4.7 | 71.3 / 100 | 4.90 | 31.8% | 63.6% | 16.6% | 55.0 / 100 | 50.1% |
| Gemini 3.1 Pro | 71.0 / 100 | 3.22 | 40.5% | 63.5% | 14.2% | 58.9 / 100 | 62.3% |
| Claude Fable 5 | 70.8 / 100 | 4.19 | 31.3% | 60.4% | 12.1% | 51.2 / 100 | 57.6% |
| Grok 4.6 | 70.5 / 100 | 2.22 | 63.1% | 40.9% | 4.8% | 60.7 / 100 | 52.3% |
| GPT-6 Astra | 70.1 / 100 | 2.50 | 50.0% | 52.5% | 0.0% | 64.6 / 100 | n/a |
| Claude Opus 4.8 | 69.2 / 100 | 4.34 | 35.3% | 46.5% | 11.7% | 77.8 / 100 | 43.4% |
What Stands Out
Each finding is tied to model IDs and metric keys in the generated report.
Claude Fable 5.1 and Grok 4.6 are different in different ways
Claude Fable 5.1 stands out by risk appetite at 80.2 / 100, while Grok 4.6 stands out by portfolio structure with a 63.1% average largest holding.
Gemini 3.1 Pro and Claude Opus 5 look more risk-managed than the aggressive cohort
Gemini 3.1 Pro has the highest defensive allocation at 14.2%. Claude Opus 5 has the lowest measured turnover at 50.8%.
Claude Opus 5 invests most like the group
Claude Opus 5 has the lowest Portfolio Difference at 45.5 / 100. That is the share of allocation that would need to change to match the average portfolio of the other models.
How Behavior Labels And Pills Are Determined
The report is rebuilt from eligible official frozen portfolios during every publication build. No model receives a manually assigned caption, and the model's own descriptive wording cannot assign its label.
For each model and round, CapitalBench subtracts the median behavior-metric value of the other models in that same round. A behavior signal must exceed its published materiality floor, point in the same direction in at least 65% of matched portfolios, and have at least 8 matched portfolios across 6 independent decision dates.
Qualifying signals are ordered by absolute median peer difference divided by their materiality floor, then by persistence and a stable metric key. The strongest exposure or risk signal supplies the label modifier; peer-normalized construction, turnover, or Portfolio Difference supplies the allocation-style noun.
Evidence is “established” only after 16 decision dates and 75% persistence. Opposite material weekly and monthly signals are marked horizon-dependent; a sufficiently sampled reversal under the newest methodology is marked evolving. The four pills always report signature, construction, tempo, and current open positioning (or lifecycle for a retired model). “Typical” uses all eligible history; “Now” uses only currently open portfolios.
Realized returns, ranks, ineligible or pilot runs, market-briefing prose, and free-form rationale wording are not classification inputs. Structured candidate-ledger, forecast, confidence, and key-risk fields are retained as decision-process context when coverage exists, but they do not override allocation evidence. Page-level “most” leader cards use active models only; retired profiles remain available as historical evidence.
Portfolio Difference is one-half of the summed absolute difference between a model's allocation and the average allocation selected by every other model in the same round. The measured model is left out of that average. A score of 42 means about 42% of allocation would need to change to match the group portfolio. Combined scores give monthly and weekly results equal weight and require observations from both horizons. The score measures output difference; it does not show that one model copied or influenced another.
Recent-winner tilt is a separate behavior measure and does not change the archetype. Each eligible asset receives a 0–100 percentile from its return before the decision cutoff, and the model's portfolio weights produce one allocation-weighted score. Current weekly portfolios use 5 trading sessions relative to SPY; monthly portfolios use 21. S&P 500 and cash are neutral at 50, and future returns never enter the calculation.
Method version: capitalbench_behavior_evidence_v2
Peer baseline: leave-one-model-out same-round peer median
Wording provenance: deterministic_source_of_truth
Prompt contract: capitalbench_model_patterns_prompt_v2
A median same-round peer difference must meet the relevant floor before persistence can qualify it.
- Risk taking≥ 4 score points
- Technology≥ 5 percentage points
- Real assets≥ 5 percentage points
- International assets≥ 4 percentage points
- Defensive assets≥ 4 percentage points
- Cash and duration≥ 4 percentage points
- S&P 500 core≥ 5 percentage points
- Largest holding≥ 5 percentage points
- Holding count≥ 0.5 holdings
Average allocation-weighted risk appetite across all official saved portfolios. Higher means more growth, momentum, cyclical, and high-risk exposure.
Average number of non-zero assets in the model's official saved portfolios.
Average size of the largest single holding in each official saved portfolio.
Average allocation to assets rated as higher risk by the CapitalBench asset risk model.
Average allocation to cash, bonds, defensive sectors, and other lower-risk ballast.
Average allocation to technology, semiconductors, Nasdaq-style growth, and AI-linked technology exposure.
Average allocation to cash-like assets and duration-sensitive bond exposure.
Average allocation to non-U.S. country, regional, or international equity exposure.
Average allocation to commodities, crypto, energy, gold, and other inflation-linked or real-asset groups.
Average allocation to the S&P 500 benchmark option across official saved portfolios.
Legacy API-only average cosine similarity between this model's allocation weights and peer model portfolios in the same rounds.
The percentage of allocation that would need to change to match the average portfolio selected by the other models in the same rounds.
Average one-half summed absolute allocation change between consecutive same-track portfolios.
Equal-weighted monthly and weekly allocation-weighted percentile rank of assets' pre-decision recent returns. A score of 50 is neutral; higher values lean toward recent winners.
Equal-weighted monthly and weekly portfolio allocation to assets in the top 20% of the applicable pre-decision recent-return window.
Average finishing rank across resolved rounds. Lower is better.
Average model score versus the hindsight-best eligible asset in each resolved round.