Comms AI Pentathlon 2026
Eight AI assistants answered the same five comms briefs on Wednesday 30 September 2026, in one sitting. Then each of them judged all 40 entries blind: 320 scores of 40 entries, from one run, each with a written justification. This is the full record, with every brief, entry and score, and a human judge’s blind ranking alongside.
It is a snapshot of eight assistants as they were served on one day, not a buying verdict. For a decision about your own team’s tools, test them on your own work with the Bench Test.
The write-up, on what the results mean for comms teams, is on Applied / Comms With AI.
Medal table
Each event score is the mean of the eight AI judges’ scores, to one decimal place. The total adds the five event scores. The last column is the human judge’s overall order from blind ranks, which sits outside the medal table.
| Rank | Assistant | E1 | E2 | E3 | E4 | E5 | Total | Gold | Silver | Bronze | Human judge |
|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | Claude | 8.6 | 8.6 | 9.1 | 8.4 | 8.1 | 42.8 | 3 | 2 | 0 | 3rd |
| 2 | GLM | 8.6 | 8.4 | 8.8 | 8.2 | 8.3 | 42.3 | 2 | 2 | 1 | 2nd |
| 3 | GPT | 7.7 | 8.1 | 8.5 | 7.8 | 7.3 | 39.4 | 0 | 0 | 2 | 7th |
| 4 | Kimi | 8.1 | 7.3 | 8.4 | 7.4 | 8.0 | 39.2 | 0 | 0 | 1 | 1st |
| 5 | DeepSeek | 8.4 | 6.6 | 8.1 | 8.7 | 7.3 | 39.1 | 1 | 0 | 1 | 6th |
| 6 | Copilot | 7.8 | 8.0 | 6.3 | 7.2 | 6.0 | 35.3 | 0 | 0 | 0 | 4th |
| 7 | Gemini | 7.9 | 7.1 | 5.8 | 7.3 | 6.9 | 35.0 | 0 | 0 | 0 | 5th |
| 8 | Siri | 7.0 | 0.4 | 5.3 | 7.1 | 4.3 | 24.1 | 0 | 0 | 0 | 8th |
Event score is the mean of the eight judges' scores, to one decimal place. Medals by rank on that mean; tied means share the medal. Overall order is the total of the five event means; tie-break most golds, then head-to-head (not needed: no totals tied). E1 to E5 link to each event’s entries and scores.
The five events
The 100m: headline sprint
- Gold: GLM (8.6)
- Gold: Claude (8.6)
- Bronze: DeepSeek (8.4)
- Human judge’s first: DeepSeek
The 110m hurdles: crisis statement
- Gold: Claude (8.6)
- Silver: GLM (8.4)
- Bronze: GPT (8.1)
- Human judge’s first: GLM
The marathon: comms strategy
- Gold: Claude (9.1)
- Silver: GLM (8.8)
- Bronze: GPT (8.5)
- Human judge’s first: Claude
The archery: media pitch
- Gold: DeepSeek (8.7)
- Silver: Claude (8.4)
- Bronze: GLM (8.2)
- Human judge’s first: Kimi
The artistic gymnastics: creative campaign
- Gold: GLM (8.3)
- Silver: Claude (8.1)
- Bronze: Kimi (8.0)
- Human judge’s first: Kimi
Judging the judges
Every assistant judged its own entry without knowing which it was. The own-entry scoring gap, the measure set before the run, is the score a judge gave its own blind entry minus the mean of the other seven judges’ scores for that entry. A positive gap means it marked its own entry higher than the others did.
| Judge | E1 | E2 | E3 | E4 | E5 | Own-entry scoring gap, average |
|---|---|---|---|---|---|---|
| GPT | +1.04 | +1.03 | +0.64 | +0.19 | +0.26 | +0.63 |
| Claude | −0.17 | +0.73 | −0.36 | −0.69 | +0.04 | −0.09 |
| Gemini | +0.03 | +1.64 | +0.79 | +1.06 | +0.61 | +0.83 |
| DeepSeek | −0.40 | +0.61 | +0.33 | −0.51 | −0.77 | −0.15 |
| GLM | +0.26 | +0.31 | +0.59 | +0.56 | +0.53 | +0.45 |
| Kimi | −1.03 | +0.80 | +0.01 | +1.26 | +0.94 | +0.40 |
| Copilot | +1.76 | +1.66 | +1.51 | +1.69 | +0.57 | +1.44 |
| Siri | +0.63 | +0.71 | −0.86 | +1.66 | +0.21 | +0.47 |
Adjusted for how harsh each judge was added 1 October 2026
The own-entry gap also picks up whether a judge marks everyone hard or soft. Claude and GPT marked everyone else’s entries down by about a point, so the raw gap flatters them; a raw gap near zero does not show an absence of self-preference. The adjusted figure is each judge’s average own-entry gap minus its average gap on the other 35 entries it judged, each measured the same way.
| Judge | Own-entry scoring gap | Gap on the other 35 entries | Adjusted |
|---|---|---|---|
| GPT | +0.63 | −1.02 | +1.65 |
| Claude | −0.09 | −1.03 | +0.94 |
| Gemini | +0.83 | +0.28 | +0.55 |
| DeepSeek | −0.15 | +0.13 | −0.28 |
| GLM | +0.45 | +0.20 | +0.25 |
| Kimi | +0.40 | +0.14 | +0.26 |
| Copilot | +1.44 | +0.50 | +0.93 |
| Siri | +0.47 | +0.23 | +0.24 |
- Each judge had five own entries, one per event, so each adjusted figure rests on five data points. Treat it as a pointer worth investigating, not a verdict.
- It cannot show that any judge recognised its own entry or favoured it deliberately. Entries were blind; a preference for its own house style would produce the same pattern.
- A single average does not capture a judge that is harsh in different ways in different events.
- It was added after the run, on 1 October 2026, because the original measure on its own flatters the harshest judges. The original measure and its figures above are unchanged. Every row is in the supplementary CSV.
Scores each judge gave
Mean, lowest and highest of the 40 scores each judge gave, and its mean per event.
| Judge | Mean | Lowest | Highest | E1 | E2 | E3 | E4 | E5 |
|---|---|---|---|---|---|---|---|---|
| GPT | 6.70 | 0.0 | 9.3 | 8.46 | 6.15 | 6.54 | 7.23 | 5.14 |
| Claude | 6.62 | 0.5 | 9.2 | 6.84 | 6.58 | 6.73 | 6.41 | 6.55 |
| Gemini | 7.72 | 0.0 | 9.5 | 7.51 | 7.45 | 7.96 | 8.01 | 7.66 |
| DeepSeek | 7.50 | 1.0 | 9.2 | 8.00 | 7.04 | 7.39 | 7.68 | 7.39 |
| GLM | 7.62 | 0.0 | 9.4 | 8.16 | 6.45 | 8.40 | 7.74 | 7.33 |
| Kimi | 7.57 | 0.5 | 9.2 | 7.80 | 6.90 | 7.65 | 8.09 | 7.40 |
| Copilot | 7.96 | 0.0 | 9.6 | 8.81 | 7.03 | 7.93 | 8.80 | 7.23 |
| Siri | 7.64 | 1.0 | 9.5 | 8.26 | 6.75 | 7.69 | 8.10 | 7.41 |
A human judge
Michael ranked each event 1 to 8 (he did not score out of 10), from a separate blind pack: entries renumbered 1 to 8 in a fresh random order per event, text identical to the AI judges' packs, same brief and criteria. His ranks sit outside the eight-judge means and medals. He knew the overall results before judging; the only entry he recalled was Siri's Event 2 refusal. The key was opened only after all five sheets were complete. Notes are his rough notes, with spelling corrected and wording unchanged. His rank and note for every entry are on each event’s page.
| Human judge’s order | Assistant | E1 | E2 | E3 | E4 | E5 | Sum of ranks | AI judges’ rank |
|---|---|---|---|---|---|---|---|---|
| 1 | Kimi | 2 | 3 | 3 | 1 | 1 | 10 | 4 |
| 2 | GLM | 3 | 1 | 2 | 3 | 4 | 13 | 2 |
| 3 | Claude | 7 | 7 | 1 | 2 | 2 | 19 | 1 |
| 4 | Copilot | 4 | 4 | 6 | 4 | 5 | 23 | 6 |
| 5 | Gemini | 6 | 2 | 5 | 8 | 3 | 24 | 7 |
| 6 | DeepSeek | 1 | 6 | 7 | 7 | 6 | 27 | 5 |
| 7 | GPT | 8 | 5 | 4 | 6 | 7 | 30 | 3 |
| 8 | Siri | 5 | 8 | 8 | 5 | 8 | 34 | 8 |
Order is the sum of the five event ranks, lowest first (no ties). 1 is the best rank in each event.
Agreement with the consensus
Spearman's rank correlation of each judge's order with the consensus order, per event. For an AI judge, the consensus is the mean of the other seven judges' scores; for Michael, the mean of all eight. Tied scores take averaged ranks. 1 means the same order as the consensus; 0 means no relationship. Event 2 figures are lifted for every judge by Siri’s refusal, which almost every judge placed last.
| Judge | E1 | E2 | E3 | E4 | E5 | Average |
|---|---|---|---|---|---|---|
| Claude | 0.81 | 0.93 | 0.98 | 0.95 | 0.95 | 0.92 |
| GLM | 0.81 | 0.95 | 0.76 | 0.83 | 0.98 | 0.87 |
| DeepSeek | 0.78 | 0.95 | 0.90 | 0.64 | 0.81 | 0.82 |
| Kimi | 0.52 | 0.93 | 0.86 | 0.93 | 0.83 | 0.82 |
| Siri | 0.43 | 0.93 | 0.83 | 0.60 | 0.81 | 0.72 |
| Gemini | 0.45 | 0.60 | 0.83 | 0.76 | 0.93 | 0.71 |
| Copilot | 0.48 | 0.50 | 0.86 | 0.74 | 0.76 | 0.67 |
| GPT | 0.81 | 0.79 | 0.83 | 0.40 | 0.29 | 0.62 |
| Michael (human) | 0.38 | 0.26 | 0.88 | 0.07 | 0.59 | 0.44 |
How it was run
The heats
- Wednesday 30 September 2026, one sitting, about 07:30 to 09:10 BST. Five briefs for fictional organisations, each pasted word for word into a fresh chat in every app.
- Each app’s default settings and model. Memory, custom instructions and personalisation off wherever the app offered the choice.
- One shot: the first answer stands. No regenerations, no follow-ups. A refusal is a result and was judged as the entry.
- Lanes followed three rules set before the heats: the entry paid tier where an app had one, and the best chat model it offered on the day; what the app served on the day competes; the field is frozen at eight.
The judging
- Wednesday 30 September 2026, about 09:15 to 12:45 BST. Every assistant judged all 40 entries, on the same surface and settings it competed on.
- Entries were labelled A to H (the same letters for every judge within an event, different letters between events) and shown to each judge in its own random order.
- One identical prompt for every judge: a score out of 10 to one decimal place, weighing the event’s three criteria equally, and a justification of no more than two sentences.
- All eight judges scored all 40 entries: 320 scores of 40 entries, from one run, each parsed from the judge’s saved reply and checked against the scoring workbook.
Lanes, as each app showed them
| Lane | Assistant | Model and surface, as shown at the heats | Monthly cost logged |
|---|---|---|---|
| 1 | GPT | ChatGPT Plus, Mac desktop app, Work tab: "GPT-6 Sol High" (GPT-6.1 Sol announced that morning, not offered in the picker); no folder, no plugins | £20 (Plus) |
| 2 | Claude | claude.ai incognito chat: "Opus 5.5 Medium"; Michael's own custom skills off, Anthropic-provided skills on | £15 (Pro, the entry tier serving Opus 5.5; Michael's account is on Max) |
| 3 | Gemini | gemini.google.com, Workspace "Work · Pro": "Thinking" (3.6 Thinking, the default) | Included in Google Workspace Business Standard, £14 (July invoice) |
| 4 | DeepSeek | chat.deepseek.com (branded "DSeek" in the UK): model not named; Deep thinking off, Smart Search on (defaults). Replaced Sakana Fugu, region-blocked in the UK (403) | Free |
| 5 | GLM | chat.z.ai: "GLM-5.3" (default, "Flagship model"), Deep Think Max | Free |
| 6 | Kimi | kimi.com Plus: "K3 High", context Standard (defaults) | $19, charged £14.94 |
| 7 | Copilot | copilot.microsoft.com, work account "M365 Copilot (Premium)" on a fresh default Microsoft 365 Business Premium with Copilot tenant: Auto, Work IQ on, web search on (defaults); custom instructions and saved memories off | One-month trial (list price not logged) |
| 8 | Siri | Siri (iOS 27.0 beta, "Powered by Apple Intelligence"): iPhone 16 Pro for Event 1; Siri app on the Mac (same Apple account) for Events 2 to 5 and all judging | No charge logged |
Where the method bent
- Sakana Fugu, the planned lane 4, was region-blocked in the UK on prep day. DeepSeek took the lane.
- Event 3’s judging packs ran to about 118,000 characters. Six judges passed a 125,188-character input test and judged it in one chat. GLM (input ceiling about 50,000 characters) judged it across four fresh chats and Siri across two, with the same prompt in each part and the presentation order kept.
- Pasting a pack into the Siri app on the Mac did not work, so all six Siri packs were uploaded as files. Kimi, DeepSeek and GLM turned pasted packs into text attachments themselves.
- Some replies listed the entries in a different order from the one shown (usually alphabetical), and some departed from the format: longer justifications, an unrequested opening line, one justification cut off mid-sentence, and one score given without its letter, recorded from the single entry in that chat. Scores were unaffected.
- Before judging, entries were transcribed to plain text. Interface artefacts and one assistant’s closing self-reference were removed, one subject line and one sign-off were restored from the screen. Nothing else was edited; errors stay in. Every edit is in the collation log in the download.
- Claude helped iterate the test design, competed, judged and won, and this page was built with Claude. None of that needs deliberate interference to matter: a model can shape briefs, criteria and framing in ways that suit its own style. The formula protects the arithmetic, not those choices, which is why every file is published here.
Screenshots from prep and the heats
Cropped where needed to remove account details. The surfaces are listed in short form here: GPT, ChatGPT Plus, Work tab, GPT-6 Sol High; Claude, claude.ai incognito chat, Opus 5.5 Medium; Gemini, Gemini in Google Workspace, 3.6 Thinking; DeepSeek, chat.deepseek.com, default settings; GLM, chat.z.ai, GLM-5.3, Deep Think Max; Kimi, kimi.com Plus, K3 High; Copilot, M365 Copilot (Premium), fresh Business Premium trial, Auto; Siri, Siri app, iOS 27 beta.
Download the record
Free to reuse under CC BY 4.0: credit Comms With AI. The entries are AI outputs, published as test evidence; the organisations and people in the briefs are fictional.
- Scoring workbook
All 320 scores; means, ranks, medal table and self-preference calculate from them. Includes the human judge’s ranks.
- Scores (CSV)
320 rows: event, judge, blind letter, assistant, score, position in that judge’s order, own-entry flag, justification.
- Human judge (CSV)
40 rows: Michael MacLennan’s blind rank and note for every entry, beside the AI judges’ rank and mean.
- Full record (JSON)
Everything in one file: briefs, criteria, blind letters, presentation orders, entries, scores, justifications and results. These pages are built from it.
- Everything (ZIP)
The above, plus the exact message each judge was sent, each reply as saved, the human judge’s blind packs and key, the length tests, the collation log and the screenshots.
- Supplementary: self-preference adjusted for harshness (CSV)
Added 1 October 2026; the original pre-test measure is unchanged. 320 rows: every judge’s gap on every entry, marked own or other.
- Supplementary: the adjustment and its limits (note)
Added 1 October 2026; the original pre-test measure is unchanged. Why it was added, how it is calculated, and what it cannot show.
Run it yourself
- Add your assistant as a ninth entrant. Paste one of the five briefs into a fresh chat, keep the first answer, add it as Entry I to that event’s judging pack from the download, and judge all nine with the same prompt.
- Re-run the whole thing. The download has the briefs, the judging prompt, the blind letters and the order each judge saw the entries in.
- For a real decision, use your own work. The Bench Test runs the same blind method on your team’s recurring tasks, with your colleagues scoring.