PeerRank · AI judging AI
AIs Have Egos: What Happened When 6 Frontier Models Graded Each Other
We asked six frontier AI models to grade each other’s work. They behaved a lot like people do.
They gave themselves higher marks than their peers did. They got more generous as soon as they saw a famous name. They favored whoever spoke first. And two models from the same family graded so much alike that they could pass for one judge.
These aren’t anecdotes. They come from 3,570 peer scores collected in the October 2026 PeerRank run, and the numbers are below.
grading itself
for Opus 5.5
going first
agreement
scores
The setup: six models, three jobs each
Every model played three roles. It wrote questions, it answered everyone’s questions, and it graded everyone’s answers.
- Ask. Each model wrote 20 questions across five categories (creative, current events, factual, practical, reasoning), for 120 in total.
- Answer. All six models answered all 120 questions. Current-events questions came with the same web search results for every model.
- Judge. Each model scored every answer from 1 to 10, three separate times:
- Blind + shuffled: names hidden, order random. This is the fair baseline, called the peer score.
- Names shown: order random, model names visible.
- Fixed order: names hidden, answers always in the same order.
The models were GPT-6 Astra, Claude Opus 5.5, Claude Fable 5.1, Gemini 3.8 Flash, Grok 4.7 and DeepSeek V4.1 Flash. Comparing the three judging modes shows how much a model’s identity and its position changed its grade.
Ego: five of six models liked their own work best
GPT-6 Astra scored its own answers 0.76 points higher than the other five models scored them. On a 10-point scale, that moves an answer from “good” to “great” based only on who wrote it.
Self bias
Each model’s grade for its own answers, against the grade the other five models gave them.
| Model | Score from peers | Score from itself | Self bias |
|---|---|---|---|
| GPT-6 Astra | 8.44 | 9.19 | +0.76 |
| DeepSeek V4.1 Flash | 7.47 | 8.01 | +0.54 |
| Claude Fable 5.1 | 8.32 | 8.73 | +0.41 |
| Gemini 3.8 Flash | 7.56 | 7.87 | +0.32 |
| Claude Opus 5.5 | 8.55 | 8.61 | +0.07 |
| Grok 4.7 | 6.96 | 6.94 | −0.02 |
Only two models came close to judging themselves honestly. Claude Opus 5.5 inflated itself by under a tenth of a point. Grok 4.7 actually rated itself slightly lower than its peers did, and it was also the only model that did significantly worse on questions it had written itself (−0.64 points). It was writing questions it couldn’t answer.
Brand bias: a famous name is worth points
All six models scored higher once judges could see the model’s name. The answers stayed exactly the same. The only thing that changed was the label next to them.
Name bonus
The same answers, graded with names hidden and then with names shown.
| Model | Blind score | Name shown | Name bonus |
|---|---|---|---|
| Claude Opus 5.5 | 8.55 | 8.80 | +0.25 |
| GPT-6 Astra | 8.44 | 8.61 | +0.18 |
| Claude Fable 5.1 | 8.32 | 8.50 | +0.18 |
| DeepSeek V4.1 Flash | 7.47 | 7.61 | +0.14 |
| Gemini 3.8 Flash | 7.56 | 7.65 | +0.09 |
| Grok 4.7 | 6.96 | 7.05 | +0.09 |
The brand premium wasn’t spread evenly. The three top-ranked models got the biggest boosts, so reputation added to a lead that already existed. AI judges, like human ones, give the benefit of the doubt to names they already trust.
Position bias: it pays to go first
Only the answer shown first got a boost. With names hidden and the order fixed, the first answer gained 0.13 points over its fair score, while answers in slots 3, 4 and 5 lost between 0.16 and 0.19.
People do the same thing. Interviewers, wine tasters and talent-show judges all reward whoever shows up first. The models seem to have picked up the habit along with everything else they learned from us.
Family resemblance: the two Claudes judge like twins
Claude Opus 5.5 and Claude Fable 5.1 agreed with each other more than any other pair of judges did (r = 0.89). Every pairing across different companies landed between 0.56 and 0.80.
Judge agreement
Correlation (r) between two judges’ scores.
| Most similar judges | r |
|---|---|
| Opus 5.5 and Fable 5.1 | 0.89 |
| Grok 4.7 and DeepSeek V4.1 | 0.80 |
| GPT-6 Astra and DeepSeek V4.1 | 0.79 |
| Least similar judges | r |
|---|---|
| GPT-6 Astra and Fable 5.1 | 0.56 |
| Fable 5.1 and DeepSeek V4.1 | 0.60 |
| Fable 5.1 and Gemini 3.8 | 0.62 |
This isn’t loyalty, since neither Claude knew which answers were the other’s. It’s shared upbringing: models from the same lab tend to share a sense of what a good answer looks like.
The twist: the toughest critic came out on top
The model that graded others most harshly is the one its peers rated highest. Claude Opus 5.5 gave the lowest average score of the six judges (7.50), barely inflated its own score, and still finished first. DeepSeek V4.1 Flash was the most generous judge (8.27) and finished fifth.
Final standings
Ranked by peer score. Elo, from the 7,140 pairwise matches, puts Fable 5.1 just ahead of GPT-6 Astra.
| Rank | Model | Peer score | Elo | Avg grade it gave |
|---|---|---|---|---|
| 1 | Claude Opus 5.5 | 8.55 | 1644 | 7.50 |
| 2 | GPT-6 Astra | 8.44 | 1575 | 7.95 |
| 3 | Claude Fable 5.1 | 8.32 | 1591 | 7.77 |
| 4 | Gemini 3.8 Flash | 7.56 | 1450 | 7.76 |
| 5 | DeepSeek V4.1 Flash | 7.47 | 1407 | 8.27 |
| 6 | Grok 4.7 | 6.96 | 1333 | 8.03 |
The top three are separated by less than a quarter of a point, and the overall winner doesn’t win every category. GPT-6 Astra led in creative, practical and current-events questions, while Opus led in factual and reasoning questions. Current events were the hardest category for every model, even with web search: the average score for “who won the most recent FIFA World Cup” was just 3.9 out of 10.
What this means if you let AI judge AI
AI already grades AI everywhere: eval pipelines, agent self-review, reward models, “pick the best of five drafts.” These results show a single AI judge carries the same biases a single human judge would. Here are four practical rules:
- Never let a model grade itself. A self-score can run three-quarters of a point high.
- Hide the names. Brand alone was worth up to a quarter of a point.
- Shuffle the order. Coming first is an advantage that has nothing to do with quality.
- Use judges from different families. Two judges from the same lab act like one judge counted twice.
The bigger lesson: there’s no single best AI, only the best AI for a specific task, judged by a panel you can trust. That’s why at Caura we build to stay model-agnostic. The best model this month may not be the best next month, and your agents shouldn’t be stuck with either one.
Swap the model. Keep the memory.
Caura is governed, shared memory for agent fleets on any LLM. Change models whenever the leaderboard does, and what your agents have learned stays with them.
Data: PeerRank evaluation, revision October 7, 2026. 6 models, 120 questions, 3,570 peer scores, 7,140 pairwise Elo matches. Methodology: PeerRank, arXiv:2602.02589.