Skip to main content
Full record · Run 1

Comms AI Pentathlon 2026

Eight AI assistants answered the same five comms briefs on Wednesday 30 September 2026, in one sitting. Then each of them judged all 40 entries blind: 320 scores of 40 entries, from one run, each with a written justification. This is the full record, with every brief, entry and score, and a human judge’s blind ranking alongside.

It is a snapshot of eight assistants as they were served on one day, not a buying verdict. For a decision about your own team’s tools, test them on your own work with the Bench Test.

The write-up, on what the results mean for comms teams, is on Applied / Comms With AI.

Medal table

Each event score is the mean of the eight AI judges’ scores, to one decimal place. The total adds the five event scores. The last column is the human judge’s overall order from blind ranks, which sits outside the medal table.

Medal table: event scores out of 10, totals out of 50, medals, and the human judge’s order
Rank Assistant E1 E2 E3 E4 E5 Total Gold Silver Bronze Human judge
1 Claude 8.68.69.18.48.1 42.8 3 2 0 3rd
2 GLM 8.68.48.88.28.3 42.3 2 2 1 2nd
3 GPT 7.78.18.57.87.3 39.4 0 0 2 7th
4 Kimi 8.17.38.47.48.0 39.2 0 0 1 1st
5 DeepSeek 8.46.68.18.77.3 39.1 1 0 1 6th
6 Copilot 7.88.06.37.26.0 35.3 0 0 0 4th
7 Gemini 7.97.15.87.36.9 35.0 0 0 0 5th
8 Siri 7.00.45.37.14.3 24.1 0 0 0 8th

Event score is the mean of the eight judges' scores, to one decimal place. Medals by rank on that mean; tied means share the medal. Overall order is the total of the five event means; tie-break most golds, then head-to-head (not needed: no totals tied). E1 to E5 link to each event’s entries and scores.

The five events

Judging the judges

Every assistant judged its own entry without knowing which it was. The own-entry scoring gap, the measure set before the run, is the score a judge gave its own blind entry minus the mean of the other seven judges’ scores for that entry. A positive gap means it marked its own entry higher than the others did.

Own-entry scoring gap per judge and event
Judge E1E2E3E4E5 Own-entry scoring gap, average
GPT +1.04+1.03+0.64+0.19+0.26 +0.63
Claude −0.17+0.73−0.36−0.69+0.04 −0.09
Gemini +0.03+1.64+0.79+1.06+0.61 +0.83
DeepSeek −0.40+0.61+0.33−0.51−0.77 −0.15
GLM +0.26+0.31+0.59+0.56+0.53 +0.45
Kimi −1.03+0.80+0.01+1.26+0.94 +0.40
Copilot +1.76+1.66+1.51+1.69+0.57 +1.44
Siri +0.63+0.71−0.86+1.66+0.21 +0.47

Adjusted for how harsh each judge was added 1 October 2026

The own-entry gap also picks up whether a judge marks everyone hard or soft. Claude and GPT marked everyone else’s entries down by about a point, so the raw gap flatters them; a raw gap near zero does not show an absence of self-preference. The adjusted figure is each judge’s average own-entry gap minus its average gap on the other 35 entries it judged, each measured the same way.

Own-entry scoring gap, gap on other entrants’ entries, and the harshness-adjusted figure, per judge
Judge Own-entry scoring gap Gap on the other 35 entries Adjusted
GPT +0.63 −1.02 +1.65
Claude −0.09 −1.03 +0.94
Gemini +0.83 +0.28 +0.55
DeepSeek −0.15 +0.13 −0.28
GLM +0.45 +0.20 +0.25
Kimi +0.40 +0.14 +0.26
Copilot +1.44 +0.50 +0.93
Siri +0.47 +0.23 +0.24
  • Each judge had five own entries, one per event, so each adjusted figure rests on five data points. Treat it as a pointer worth investigating, not a verdict.
  • It cannot show that any judge recognised its own entry or favoured it deliberately. Entries were blind; a preference for its own house style would produce the same pattern.
  • A single average does not capture a judge that is harsh in different ways in different events.
  • It was added after the run, on 1 October 2026, because the original measure on its own flatters the harshest judges. The original measure and its figures above are unchanged. Every row is in the supplementary CSV.

Scores each judge gave

Mean, lowest and highest of the 40 scores each judge gave, and its mean per event.

Scores given by each judge
Judge Mean Lowest Highest E1E2E3E4E5
GPT 6.70 0.0 9.3 8.466.156.547.235.14
Claude 6.62 0.5 9.2 6.846.586.736.416.55
Gemini 7.72 0.0 9.5 7.517.457.968.017.66
DeepSeek 7.50 1.0 9.2 8.007.047.397.687.39
GLM 7.62 0.0 9.4 8.166.458.407.747.33
Kimi 7.57 0.5 9.2 7.806.907.658.097.40
Copilot 7.96 0.0 9.6 8.817.037.938.807.23
Siri 7.64 1.0 9.5 8.266.757.698.107.41

A human judge

Michael ranked each event 1 to 8 (he did not score out of 10), from a separate blind pack: entries renumbered 1 to 8 in a fresh random order per event, text identical to the AI judges' packs, same brief and criteria. His ranks sit outside the eight-judge means and medals. He knew the overall results before judging; the only entry he recalled was Siri's Event 2 refusal. The key was opened only after all five sheets were complete. Notes are his rough notes, with spelling corrected and wording unchanged. His rank and note for every entry are on each event’s page.

The human judge’s ranks per event and overall order, beside the AI judges’ overall rank
Human judge’s order Assistant E1E2E3E4E5 Sum of ranks AI judges’ rank
1 Kimi 23311 10 4
2 GLM 31234 13 2
3 Claude 77122 19 1
4 Copilot 44645 23 6
5 Gemini 62583 24 7
6 DeepSeek 16776 27 5
7 GPT 85467 30 3
8 Siri 58858 34 8

Order is the sum of the five event ranks, lowest first (no ties). 1 is the best rank in each event.

Agreement with the consensus

Spearman's rank correlation of each judge's order with the consensus order, per event. For an AI judge, the consensus is the mean of the other seven judges' scores; for Michael, the mean of all eight. Tied scores take averaged ranks. 1 means the same order as the consensus; 0 means no relationship. Event 2 figures are lifted for every judge by Siri’s refusal, which almost every judge placed last.

Rank correlation of each judge with the consensus
Judge E1E2E3E4E5 Average
Claude 0.810.930.980.950.95 0.92
GLM 0.810.950.760.830.98 0.87
DeepSeek 0.780.950.900.640.81 0.82
Kimi 0.520.930.860.930.83 0.82
Siri 0.430.930.830.600.81 0.72
Gemini 0.450.600.830.760.93 0.71
Copilot 0.480.500.860.740.76 0.67
GPT 0.810.790.830.400.29 0.62
Michael (human) 0.380.260.880.070.59 0.44

How it was run

The heats

  • Wednesday 30 September 2026, one sitting, about 07:30 to 09:10 BST. Five briefs for fictional organisations, each pasted word for word into a fresh chat in every app.
  • Each app’s default settings and model. Memory, custom instructions and personalisation off wherever the app offered the choice.
  • One shot: the first answer stands. No regenerations, no follow-ups. A refusal is a result and was judged as the entry.
  • Lanes followed three rules set before the heats: the entry paid tier where an app had one, and the best chat model it offered on the day; what the app served on the day competes; the field is frozen at eight.

The judging

  • Wednesday 30 September 2026, about 09:15 to 12:45 BST. Every assistant judged all 40 entries, on the same surface and settings it competed on.
  • Entries were labelled A to H (the same letters for every judge within an event, different letters between events) and shown to each judge in its own random order.
  • One identical prompt for every judge: a score out of 10 to one decimal place, weighing the event’s three criteria equally, and a justification of no more than two sentences.
  • All eight judges scored all 40 entries: 320 scores of 40 entries, from one run, each parsed from the judge’s saved reply and checked against the scoring workbook.

Lanes, as each app showed them

Each lane’s assistant, model and surface as shown, and monthly cost
Lane Assistant Model and surface, as shown at the heats Monthly cost logged
1 GPT ChatGPT Plus, Mac desktop app, Work tab: "GPT-6 Sol High" (GPT-6.1 Sol announced that morning, not offered in the picker); no folder, no plugins £20 (Plus)
2 Claude claude.ai incognito chat: "Opus 5.5 Medium"; Michael's own custom skills off, Anthropic-provided skills on £15 (Pro, the entry tier serving Opus 5.5; Michael's account is on Max)
3 Gemini gemini.google.com, Workspace "Work · Pro": "Thinking" (3.6 Thinking, the default) Included in Google Workspace Business Standard, £14 (July invoice)
4 DeepSeek chat.deepseek.com (branded "DSeek" in the UK): model not named; Deep thinking off, Smart Search on (defaults). Replaced Sakana Fugu, region-blocked in the UK (403) Free
5 GLM chat.z.ai: "GLM-5.3" (default, "Flagship model"), Deep Think Max Free
6 Kimi kimi.com Plus: "K3 High", context Standard (defaults) $19, charged £14.94
7 Copilot copilot.microsoft.com, work account "M365 Copilot (Premium)" on a fresh default Microsoft 365 Business Premium with Copilot tenant: Auto, Work IQ on, web search on (defaults); custom instructions and saved memories off One-month trial (list price not logged)
8 Siri Siri (iOS 27.0 beta, "Powered by Apple Intelligence"): iPhone 16 Pro for Event 1; Siri app on the Mac (same Apple account) for Events 2 to 5 and all judging No charge logged

Where the method bent

  • Sakana Fugu, the planned lane 4, was region-blocked in the UK on prep day. DeepSeek took the lane.
  • Event 3’s judging packs ran to about 118,000 characters. Six judges passed a 125,188-character input test and judged it in one chat. GLM (input ceiling about 50,000 characters) judged it across four fresh chats and Siri across two, with the same prompt in each part and the presentation order kept.
  • Pasting a pack into the Siri app on the Mac did not work, so all six Siri packs were uploaded as files. Kimi, DeepSeek and GLM turned pasted packs into text attachments themselves.
  • Some replies listed the entries in a different order from the one shown (usually alphabetical), and some departed from the format: longer justifications, an unrequested opening line, one justification cut off mid-sentence, and one score given without its letter, recorded from the single entry in that chat. Scores were unaffected.
  • Before judging, entries were transcribed to plain text. Interface artefacts and one assistant’s closing self-reference were removed, one subject line and one sign-off were restored from the screen. Nothing else was edited; errors stay in. Every edit is in the collation log in the download.
  • Claude helped iterate the test design, competed, judged and won, and this page was built with Claude. None of that needs deliberate interference to matter: a model can shape briefs, criteria and framing in ways that suit its own style. The formula protects the arithmetic, not those choices, which is why every file is published here.

Screenshots from prep and the heats

Sakana Fugu from the UK on prep day: 403, region restricted. DeepSeek took the lane.
Sakana Fugu from the UK on prep day: 403, region restricted. DeepSeek took the lane.
DeepSeek’s UK name-change notice (“DSeek”) at login, prep day.
DeepSeek’s UK name-change notice (“DSeek”) at login, prep day.
ChatGPT Plus desktop app, Work tab: GPT-6 Sol High, which competed.
ChatGPT Plus desktop app, Work tab: GPT-6 Sol High, which competed.
Gemini in Google Workspace at prep: 3.6 Thinking, the default, competed. 3.1 Pro was also offered.
Gemini in Google Workspace at prep: 3.6 Thinking, the default, competed. 3.1 Pro was also offered.
Z.ai’s model menu at prep, with Flash selected for an early length test. GLM-5.3, the default “Flagship model”, competed.
Z.ai’s model menu at prep, with Flash selected for an early length test. GLM-5.3, the default “Flagship model”, competed.
Microsoft 365 Copilot work chat at prep: mode left on Auto, the default. The GPT submenu offered GPT 5.6 Sol and GPT 6.0 Sol.
Microsoft 365 Copilot work chat at prep: mode left on Auto, the default. The GPT submenu offered GPT 5.6 Sol and GPT 6.0 Sol.
Siri’s Event 2 reply, the one refusal in the heats.
Siri’s Event 2 reply, the one refusal in the heats.

Cropped where needed to remove account details. The surfaces are listed in short form here: GPT, ChatGPT Plus, Work tab, GPT-6 Sol High; Claude, claude.ai incognito chat, Opus 5.5 Medium; Gemini, Gemini in Google Workspace, 3.6 Thinking; DeepSeek, chat.deepseek.com, default settings; GLM, chat.z.ai, GLM-5.3, Deep Think Max; Kimi, kimi.com Plus, K3 High; Copilot, M365 Copilot (Premium), fresh Business Premium trial, Auto; Siri, Siri app, iOS 27 beta.

Download the record

Free to reuse under CC BY 4.0: credit Comms With AI. The entries are AI outputs, published as test evidence; the organisations and people in the briefs are fictional.

  • Scoring workbook

    All 320 scores; means, ranks, medal table and self-preference calculate from them. Includes the human judge’s ranks.

  • Scores (CSV)

    320 rows: event, judge, blind letter, assistant, score, position in that judge’s order, own-entry flag, justification.

  • Human judge (CSV)

    40 rows: Michael MacLennan’s blind rank and note for every entry, beside the AI judges’ rank and mean.

  • Full record (JSON)

    Everything in one file: briefs, criteria, blind letters, presentation orders, entries, scores, justifications and results. These pages are built from it.

  • Everything (ZIP)

    The above, plus the exact message each judge was sent, each reply as saved, the human judge’s blind packs and key, the length tests, the collation log and the screenshots.

  • Supplementary: self-preference adjusted for harshness (CSV)

    Added 1 October 2026; the original pre-test measure is unchanged. 320 rows: every judge’s gap on every entry, marked own or other.

  • Supplementary: the adjustment and its limits (note)

    Added 1 October 2026; the original pre-test measure is unchanged. Why it was added, how it is calculated, and what it cannot show.

Run it yourself

  1. Add your assistant as a ninth entrant. Paste one of the five briefs into a fresh chat, keep the first answer, add it as Entry I to that event’s judging pack from the download, and judge all nine with the same prompt.
  2. Re-run the whole thing. The download has the briefs, the judging prompt, the blind letters and the order each judge saw the entries in.
  3. For a real decision, use your own work. The Bench Test runs the same blind method on your team’s recurring tasks, with your colleagues scoring.