Caura.aiPeerRank · Model Evaluation · October 2026

PeerRank · AI judging AI

AIs Have Egos: What Happened When 6 Frontier Models Graded Each Other

We asked six frontier AI models to grade each other’s work. They behaved a lot like people do.

They gave themselves higher marks than their peers did. They got more generous as soon as they saw a famous name. They favored whoever spoke first. And two models from the same family graded so much alike that they could pass for one judge.

These aren’t anecdotes. They come from 3,570 peer scores collected in the October 2026 PeerRank run, and the numbers are below.

+0.76
GPT-6 Astra
grading itself
+0.25
Name bonus
for Opus 5.5
+0.13
Bonus for
going first
0.89
Opus–Fable
agreement
3,570
Peer
scores

The setup: six models, three jobs each

Every model played three roles. It wrote questions, it answered everyone’s questions, and it graded everyone’s answers.

  1. Ask. Each model wrote 20 questions across five categories (creative, current events, factual, practical, reasoning), for 120 in total.
  2. Answer. All six models answered all 120 questions. Current-events questions came with the same web search results for every model.
  3. Judge. Each model scored every answer from 1 to 10, three separate times:
  • Blind + shuffled: names hidden, order random. This is the fair baseline, called the peer score.
  • Names shown: order random, model names visible.
  • Fixed order: names hidden, answers always in the same order.

The models were GPT-6 Astra, Claude Opus 5.5, Claude Fable 5.1, Gemini 3.8 Flash, Grok 4.7 and DeepSeek V4.1 Flash. Comparing the three judging modes shows how much a model’s identity and its position changed its grade.

Ego: five of six models liked their own work best

GPT-6 Astra scored its own answers 0.76 points higher than the other five models scored them. On a 10-point scale, that moves an answer from “good” to “great” based only on who wrote it.

Self bias

Each model’s grade for its own answers, against the grade the other five models gave them.

ModelScore from peersScore from itselfSelf bias
GPT-6 Astra8.449.19+0.76
DeepSeek V4.1 Flash7.478.01+0.54
Claude Fable 5.18.328.73+0.41
Gemini 3.8 Flash7.567.87+0.32
Claude Opus 5.58.558.61+0.07
Grok 4.76.966.94−0.02

Only two models came close to judging themselves honestly. Claude Opus 5.5 inflated itself by under a tenth of a point. Grok 4.7 actually rated itself slightly lower than its peers did, and it was also the only model that did significantly worse on questions it had written itself (−0.64 points). It was writing questions it couldn’t answer.

Brand bias: a famous name is worth points

All six models scored higher once judges could see the model’s name. The answers stayed exactly the same. The only thing that changed was the label next to them.

Name bonus

The same answers, graded with names hidden and then with names shown.

ModelBlind scoreName shownName bonus
Claude Opus 5.58.558.80+0.25
GPT-6 Astra8.448.61+0.18
Claude Fable 5.18.328.50+0.18
DeepSeek V4.1 Flash7.477.61+0.14
Gemini 3.8 Flash7.567.65+0.09
Grok 4.76.967.05+0.09

The brand premium wasn’t spread evenly. The three top-ranked models got the biggest boosts, so reputation added to a lead that already existed. AI judges, like human ones, give the benefit of the doubt to names they already trust.

Position bias: it pays to go first

Only the answer shown first got a boost. With names hidden and the order fixed, the first answer gained 0.13 points over its fair score, while answers in slots 3, 4 and 5 lost between 0.16 and 0.19.

People do the same thing. Interviewers, wine tasters and talent-show judges all reward whoever shows up first. The models seem to have picked up the habit along with everything else they learned from us.

Family resemblance: the two Claudes judge like twins

Claude Opus 5.5 and Claude Fable 5.1 agreed with each other more than any other pair of judges did (r = 0.89). Every pairing across different companies landed between 0.56 and 0.80.

Judge agreement

Correlation (r) between two judges’ scores.

Most similar judgesr
Opus 5.5 and Fable 5.10.89
Grok 4.7 and DeepSeek V4.10.80
GPT-6 Astra and DeepSeek V4.10.79
Least similar judgesr
GPT-6 Astra and Fable 5.10.56
Fable 5.1 and DeepSeek V4.10.60
Fable 5.1 and Gemini 3.80.62

This isn’t loyalty, since neither Claude knew which answers were the other’s. It’s shared upbringing: models from the same lab tend to share a sense of what a good answer looks like.

The twist: the toughest critic came out on top

The model that graded others most harshly is the one its peers rated highest. Claude Opus 5.5 gave the lowest average score of the six judges (7.50), barely inflated its own score, and still finished first. DeepSeek V4.1 Flash was the most generous judge (8.27) and finished fifth.

Final standings

Ranked by peer score. Elo, from the 7,140 pairwise matches, puts Fable 5.1 just ahead of GPT-6 Astra.

RankModelPeer scoreEloAvg grade it gave
1Claude Opus 5.58.5516447.50
2GPT-6 Astra8.4415757.95
3Claude Fable 5.18.3215917.77
4Gemini 3.8 Flash7.5614507.76
5DeepSeek V4.1 Flash7.4714078.27
6Grok 4.76.9613338.03

The top three are separated by less than a quarter of a point, and the overall winner doesn’t win every category. GPT-6 Astra led in creative, practical and current-events questions, while Opus led in factual and reasoning questions. Current events were the hardest category for every model, even with web search: the average score for “who won the most recent FIFA World Cup” was just 3.9 out of 10.

What this means if you let AI judge AI

AI already grades AI everywhere: eval pipelines, agent self-review, reward models, “pick the best of five drafts.” These results show a single AI judge carries the same biases a single human judge would. Here are four practical rules:

  1. Never let a model grade itself. A self-score can run three-quarters of a point high.
  2. Hide the names. Brand alone was worth up to a quarter of a point.
  3. Shuffle the order. Coming first is an advantage that has nothing to do with quality.
  4. Use judges from different families. Two judges from the same lab act like one judge counted twice.

The bigger lesson: there’s no single best AI, only the best AI for a specific task, judged by a panel you can trust. That’s why at Caura we build to stay model-agnostic. The best model this month may not be the best next month, and your agents shouldn’t be stuck with either one.

Swap the model. Keep the memory.

Caura is governed, shared memory for agent fleets on any LLM. Change models whenever the leaderboard does, and what your agents have learned stays with them.

Start free →Read the PeerRank paper

Data: PeerRank evaluation, revision October 7, 2026. 6 models, 120 questions, 3,570 peer scores, 7,140 pairwise Elo matches. Methodology: PeerRank, arXiv:2602.02589.