Choose two groups

Which results do you want to compare?

A round is one timed market test. A model group uses only rounds completed by every model being ranked together.

Bottom line

What changed?

Grok 4.3 ranks first in both groups. The groups share 3 completed rounds. Aug 19 Weekly includes 8 more rounds. Claude Opus 4.8 appears only in Aug 13 Weekly.

Models in both groups71 only in Aug 13 Weekly
Rounds used by both3Aug 19 Weekly has 8 more
Did the order change?Hardly changedBased on models in both groups
Did the top model change?No3 of 3 stayed in the top three
Model by model

How did each model's result change?

Each group uses all of its completed rounds. Models that appear in only one group are marked clearly.

Aug 19 Weekly#1Grok 4.3xAIAug 13 Weekly#1No change
Aug 19 Weekly#3Claude Fable 5AnthropicAug 13 Weekly#2Up 1
Aug 19 Weekly#2GPT-5.6 SolOpenAIAug 13 Weekly#3Down 1
Aug 19 Weekly#5Claude Opus 5AnthropicAug 13 Weekly#4Up 1
Aug 19 Weekly#6Gemini 3.1 ProGoogleAug 13 Weekly#5Up 1
Aug 19 Weekly#4Grok 4.5xAIAug 13 Weekly#6Down 2
Aug 19 WeeklyNot includedClaude Opus 4.8Only in Aug 13 WeeklyAug 13 Weekly#7Added
Aug 19 Weekly#7Grok 4.6xAIAug 13 Weekly#8Down 1
Fair comparison

Did both groups use the same rounds?

3 completed rounds are used by both groups. Aug 19 Weekly also uses 8 other rounds. Different rounds can change scores and ranks.

3 used by both8 only in Aug 19 Weekly0 only in Aug 13 Weekly
Grok 4.3Overall score on 3 shared rounds8.6Overall score on 8 other Aug 19 Weekly rounds6.2Overall score on all 11 Aug 19 Weekly rounds7.4
Claude Fable 5Overall score on 3 shared rounds5.7Overall score on 8 other Aug 19 Weekly rounds8.6Overall score on all 11 Aug 19 Weekly rounds7.1
GPT-5.6 SolOverall score on 3 shared rounds4.8Overall score on 8 other Aug 19 Weekly rounds9.5Overall score on all 11 Aug 19 Weekly rounds7.2
Claude Opus 5Overall score on 3 shared rounds4.8Overall score on 8 other Aug 19 Weekly rounds7.3Overall score on all 11 Aug 19 Weekly rounds6.1
Gemini 3.1 ProOverall score on 3 shared rounds4.3Overall score on 8 other Aug 19 Weekly rounds3.5Overall score on all 11 Aug 19 Weekly rounds3.9
Grok 4.5Overall score on 3 shared rounds4.3Overall score on 8 other Aug 19 Weekly rounds8.6Overall score on all 11 Aug 19 Weekly rounds6.4
Grok 4.6Overall score on 3 shared rounds-2.0Overall score on 8 other Aug 19 Weekly rounds7.1Overall score on all 11 Aug 19 Weekly rounds2.6
See exactly which rounds were used

Used by both: CB-2026-08-13-1W, CB-2026-08-15-1W, CB-2026-08-18-1W

Only in Aug 19 Weekly: CB-2026-08-19-1W, CB-2026-08-20-1W, CB-2026-08-21-1W, CB-2026-08-23-1W, CB-2026-08-24-1W, CB-2026-08-25-1W, CB-2026-08-26-1W, CB-2026-08-27-1W

Only in Aug 13 Weekly: None

Why were 8 rounds left out of Aug 13 Weekly?

All 8 rounds were left out because Claude Opus 4.8 had no recorded result. To keep the ranking fair, those rounds were left out for every model in Aug 13 Weekly.

Claude Opus 4.8 had no result: Aug 19, 2026, Aug 20, 2026, Aug 21, 2026, Aug 23, 2026, Aug 24, 2026, Aug 25, 2026, Aug 26, 2026, Aug 27, 2026

What this means

Which results should you rely on?

Use Aug 19 Weekly as the more reliable ranking because it has 11 completed rounds. Aug 13 Weekly has 3 and needs 3 more before it has enough evidence to become the main ranking.

All results

What are all the numbers?

ModelIncluded inAug 19 Weekly rankAug 13 Weekly rankChangeAug 19 Weekly overall scoreAug 13 Weekly overall scoreOverall score change
Grok 4.3xAIBoth groups#1#1No change7.48.6+1.2
Claude Fable 5AnthropicBoth groups#3#2Up 17.15.7-1.5
GPT-5.6 SolOpenAIBoth groups#2#3Down 17.24.8-2.3
Claude Opus 5AnthropicBoth groups#5#4Up 16.14.8-1.3
Gemini 3.1 ProGoogleBoth groups#6#5Up 13.94.3+0.4
Grok 4.5xAIBoth groups#4#6Down 26.44.3-2.2
Claude Opus 4.8AnthropicOnly Aug 13 Weekly#7Addedn/a-0.0n/a
Grok 4.6xAIBoth groups#7#8Down 12.6-2.0-4.6