AI vs Human Media

AVHM-0008

Writing

Can AI write a better customer complaint response than a human?

The investigation is complete. The result, judging, evidence, limitations and final verdict are now part of the public record.

PublishedCreated 2026-08-08

Result At A Glance

Human vs AI.
Here's what happened.

The headline numbers are simple. How strongly we can interpret them is not.

Human

Human

Score

35 / 50

Timing

Approximately 9 minutes

Submission Status

Valid with minor deviations

System

Model

Provider

AI

AI

Observed Winner

Score

38 / 50

Timing

Not precisely measured

Submission Status

Valid with independence limitation

System

ChatGPT

Model

GPT-5.6 Sol

Provider

OpenAI

AVHM Verdict

AI ADVANTAGE OBSERVED - CLEAN VICTORY NOT ESTABLISHED

The observed result favoured AI, but the strength of that result depends on the evidence and experimental controls.

Evidence Confidence

Moderate

Independence

Material AI prior-familiarity limitation

What Happened

What the evidence showed.

The important findings first. The full evidence interpretation remains available below.

35 / 50

Human judging score

The Human submission's recorded judging score.

38 / 50

AI judging score

The AI submission's recorded judging score.

Read the full evidence finding+
The AI won this specific blind judging round. The two responses received identical scores for: - Accuracy: 8-8 - Clarity: 7-7 - Practicality: 7-7 The entire three-point difference came from: - Empathy: AI 9, Human 7 - Trust: AI 7, Human 6 The result therefore shows that Blind Judge 01 rated the AI response more highly for Empathy and Trust in this particular comparison. It does not establish that AI is generally more empathetic or trustworthy than humans. The judge's incorrect identity guess also shows that this judge did not correctly identify which of these two responses was AI-generated. It does not establish that AI writing is generally indistinguishable from Human writing.

Full finding preserved from the investigation record

Blind Judging

The submissions were scored before the identities were revealed.

The judge scored the anonymous submissions using the categories fixed before the result was known.

Human

35 / 50

Blind judging score

AI

Higher Score

38 / 50

Blind judging score

Judging Result

AI received the higher overall blind judging score.

The full judging record remains available below, including the scoring process, judge comments and known limitations.

Read the full blind judging record+
The two responses were anonymised as: **Submission A** and **Submission B** The judging pack withheld participant identities and completion times. Blind Judge 01 scored both responses using the five locked criteria. The judge was the Human participant's father. That relationship is disclosed as a limitation. Only one blind judge was used. The preserved scores, preference, qualitative comments and identity guess were reported by the AVHM founder following the judging session. The current evidence record does not independently establish that those results were entered directly by the judge onto a preserved original completed score sheet. The original blind judging pack itself is preserved in the evidence record.

Full judging record preserved from the investigation file

The Reveal

Human or AI?

The anonymous submissions were revealed only after judging had been completed.

Human Submission

Human

Score

35 / 50

Timing

Approximately 9 minutes

Submission Status

Valid with minor deviations

System

Model

Provider

AI Submission

Observed Winner

AI

Score

38 / 50

Timing

Not precisely measured

Submission Status

Valid with independence limitation

System

ChatGPT

Model

GPT-5.6 Sol

Provider

OpenAI

Reveal Result

The observed result favoured AI.

The full reveal record below preserves the investigation's detailed interpretation and any limitations.

Read the full reveal record+
The locked mapping was: **Submission A = AI** **Submission B = Human** The blind judge had therefore preferred the AI response. The judge's AI identity guess was incorrect. Final score: **AI: 38 / 50** **Human: 35 / 50** Observed competition winner: **AI** Winning margin: **3 points**

Full reveal record preserved from the investigation file

What Surprised Us

The result was not quite what the headline suggests.

The most interesting part was not simply who won. It was where the differences appeared — and where they did not.

Read the full surprise analysis+
The most interesting part of the result was where the AI advantage appeared. The entire winning margin came from Empathy and Trust. In this blind comparison, the categories that might intuitively be expected to favour Human communication instead produced the AI's advantage. The identity guess added another unexpected result. The judge preferred the AI response while incorrectly believing that the Human response was AI-generated. These findings are interesting precisely because they challenge an intuitive expectation. They must still be interpreted within the limits of a single judge and a single scenario.

Full analysis preserved from the investigation record

What We Learned

Every investigation should improve the next one.

The result matters. So does what the investigation taught us about running better tests next time.

01

Lesson 01

Blindness Must Include Contextual Clues

02

Lesson 02

The Competing AI Must Be Separated From Experiment Design

03

Lesson 03

Timing Needs Independent Measurement

04

Lesson 04

Subjective Tests Benefit From More Judges

05

Lesson 05

Audience Results Must Remain Separate

Read the full lessons record+
AVHM-0008 exposed several useful methodological lessons. ## Blindness Must Include Contextual Clues Hiding participant names is not always enough. Completion times and other contextual information can reveal participant identity. Future blind judging should remove unnecessary identity clues before scoring. ## The Competing AI Must Be Separated From Experiment Design The most important procedural weakness was that the competing AI had previously participated in developing and auditing the brief. Future comparable AVHM investigations should separate these roles. Once a brief is locked, the competing AI should receive only the final locked competition materials in a fresh context. ## Timing Needs Independent Measurement The AI completion time was not independently measured with sufficient precision. Future timed investigations should establish an independent timing method before either attempt begins. ## Subjective Tests Benefit From More Judges One blind judge can determine the winner of a locked competition. One judge cannot establish broad customer preference. Future subjective investigations should consider multiple independent blind judges or a separate audience study. ## Audience Results Must Remain Separate If AVHM later allows the audience to judge Submission A and Submission B, those votes should form a separate dataset. They must not retroactively alter the original locked Blind Judge 01 result.

Full lessons record preserved from the investigation file

AVHM Recommendation

What should someone actually do with this result?

A winner is only useful if the evidence tells us what the result means in practice.

Observed Result

AI advantage

The recommendation starts with what the investigation actually observed.

Evidence Confidence

Moderate

This rating describes confidence in the investigation evidence itself. It is separate from any confidence rating attached to the practical recommendation.

Editorial Recommendation

The practical recommendation depends on the evidence, not just the winner.

The full recommendation below preserves the investigation-specific advice, limitations and practical interpretation exactly where they belong: in the investigation record itself.

Decision Basis

Observed result plus evidence strength

AVHM Principle

Never recommend more than the evidence supports

The expanded recommendation may include its own recommendation-confidence rating. That is separate from the Evidence Confidence shown above.

Read the full recommendation+

The evidence does not justify recommending that AI universally replace Human customer-service writers.

It does demonstrate that an AI-generated complaint response can compete strongly with Human-written customer-service communication and, in this test, receive the higher blind score.

A reasonable practical interpretation is:

AI can be useful for drafting customer complaint responses, with appropriate Human oversight and organisational controls.

This recommendation is narrower than claiming AI is generally superior.

The investigation tested one scenario, one Human participant, one AI system and one blind judge.

Full recommendation preserved from the investigation record

Final AVHM Interpretation

The result matters.
The limitations matter too.

The observed result tells us what happened. The evidence determines how confidently we can interpret it.

Final Verdict

AI ADVANTAGE OBSERVED - CLEAN VICTORY NOT ESTABLISHED

The observed result favoured AI. AVHM does not treat that observed advantage as stronger than the evidence allows.

Evidence Confidence

Moderate

Independence

Material AI prior-familiarity limitation

Read the full final interpretation+

AVHM-0008 produced a clear result within the locked blind judging process: the AI response received the higher score.

Blind Judge 01 scored the AI response 38 / 50 and the Human response 35 / 50.

The observed advantage came entirely from the subjective categories of Empathy and Trust. Accuracy, Clarity and Practicality were scored equally.

The result is meaningful for this specific comparison, but it does not establish that AI is generally better than humans at writing customer complaint responses.

The competing AI had prior familiarity with the competition brief because it had participated in developing and auditing the experiment before its own attempt.

AI completion time was not independently measured with sufficient precision, so no reliable speed winner is claimed.

Evidence confidence: MODERATE

Full interpretation preserved from the investigation record

The Method Behind The Result

Now inspect how the test was designed.

AVHM publishes the protocol alongside the result so the question, rules and limitations can be examined rather than simply trusted.

Why This Investigation Matters

Why are we asking this question?

Customer complaints are a useful test of more than writing ability. A good response needs to understand the problem, communicate clearly, show appropriate empathy and offer a practical way forward.

AI systems can produce polished text extremely quickly, but polished writing is not necessarily effective communication. A human writer may be better at recognising tone, nuance and what an unhappy customer actually needs to hear.

This investigation will test whether AI can turn the same complaint information into a better response than a human when both work from the same brief.

Compiled directly from the investigation record

Objective

What are we trying to determine?

Determine whether an AI system can produce a higher-quality customer complaint response than a skilled human writer when both receive the same brief, the same information and the same time limit.

Scope

What does this investigation cover?

This investigation compares only the quality of the final written response. It does not measure customer satisfaction, legal compliance or long-term business outcomes.

Investigation Protocol

The rules are set before we know the result.

The success criteria, competition rules and judging method are recorded in advance so the test cannot quietly move the goalposts after either side performs.

Success Criteria

What does a good result look like?

A successful response should address the customer's complaint accurately, communicate with an appropriate level of empathy and provide a clear and practical way forward.

Both submissions will be assessed against the same five criteria:

1.

Accuracy — Does the response correctly understand and address the facts of the complaint without inventing information or making unsupported claims?

2.

Empathy — Does the response recognise the customer's frustration appropriately and sound sincere without becoming exaggerated, defensive or dismissive?

3.

Clarity — Is the response easy to understand, well structured and free from unnecessary wording or ambiguity?

4.

Practicality — Does the response provide a useful, realistic next step or resolution based only on the information available in the brief?

5.

Trust — Would the response increase the customer's confidence that their complaint has been understood and is being handled fairly?

Each criterion will be scored out of 10 using the same judging standard for both submissions.

The higher combined judging score will determine which response performed better overall.

Completion time will also be recorded and reported as part of the investigation, but speed alone will not determine the winner unless explicitly incorporated into the locked judging method before either attempt begins.

Competition Rules

A fair test needs fixed rules.

1.

Same complaint — The AI and Human participants must respond to exactly the same customer complaint.

2.

Same brief — Both participants must receive the same instructions, customer information, company information and permitted resolution options.

3.

Same information — Neither participant may be given additional facts, context or guidance that is unavailable to the other.

4.

Independent attempts — Neither participant may see, review or receive information about the other participant's response before their own final response has been submitted.

5.

No outside research — Neither participant may search the web, consult another person or use external reference material unless the locked brief explicitly permits it.

6.

Human tools — The Human participant may use a standard text editor and normal spelling or grammar checking, but may not use generative AI or AI-assisted rewriting.

7.

AI tools — The AI participant may use only the designated AI system and the information contained in the locked brief. No additional human editing of the AI response is permitted before submission.

8.

Timing — Timing begins when the participant receives the complete locked brief and ends when they declare their response final.

9.

Final submission — Once a participant declares their response final, it is locked. No rewriting, correction or improvement is permitted before judging.

10.

Response format — Both participants must produce a customer-facing written response suitable for sending directly to the customer. Any required length or formatting constraints must be defined in the locked brief before either attempt begins.

11.

Blind judging — The final responses will be presented as Submission A and Submission B. Information that directly identifies which response was produced by AI or Human will be removed where practical.

12.

Judging independence — Judges must score the written responses using the locked judging criteria before the identities of the participants are revealed.

13.

Completion time — Completion time will be recorded and disclosed, but will not be shown to judges before scoring because it could reveal participant identity and influence assessment of writing quality.

14.

No silent rule changes — Any ambiguity, deviation or unexpected issue that occurs during the investigation must be recorded. The rules must not be retrospectively changed to favour either participant.

15.

Evidence preservation — The brief, prompts or instructions, final submissions, timing information, judging records and any material interventions must be preserved as part of the investigation evidence.

16.

Core principle — The published result must reflect the evidence produced by the test, including any weaknesses or limitations in the experiment.

Judging Criteria

How were AI and Human judged?

Both sides faced the same five categories and the same scoring scale.

The Human and AI responses will be judged against the five criteria defined in the Success Criteria:

Accuracy

Empathy

Clarity

Practicality

Trust

Each criterion will be scored independently from 1 to 10, giving each submission a maximum possible score of 50 points.

Judging Procedure

1.

The Human and AI responses will be labelled Submission A and Submission B before being shown to the judge.

2.

The order of the submissions will not indicate which response was produced by the Human or AI participant.

3.

The judge will receive the original customer complaint, the relevant locked brief and both final responses.

4.

The judge will not be shown participant identities, completion times or process information before scoring.

5.

Each submission will be scored independently against all five criteria.

6.

The judge must record a score for every criterion before participant identities are revealed.

7.

The five scores for each submission will then be added to produce a total score out of 50.

8.

The submission with the higher total judging score will be the judged winner.

9.

If both submissions receive the same total score, the judged result will be recorded as a draw.

10.

After all scores are locked, the Human and AI identities will be revealed.

Judge Commentary

The judge will also be invited to explain the reasoning behind the scores and identify the strongest and weakest aspects of each response.

Any qualitative comments will be preserved alongside the numerical scores so the published result shows not only who scored higher, but why.

Identity Guess

Before the reveal, the judge will be asked which submission they believe was produced by AI and which was produced by a Human, together with their confidence in that judgement.

This identity guess will not affect the competition score. It will be recorded separately as an additional Turing-test-style observation.

Speed

Completion time will be revealed only after judging has been completed and the scores have been locked.

Speed will be reported as a separate performance measure and will not contribute points to the judging score.

Limitations

What could this test not prove?

This investigation is designed to compare the quality of two written complaint responses under controlled conditions. It cannot prove that one participant would always perform better in every real customer-service situation.

The main limitations are:

1.

Single complaint scenario — The result will be based on one complaint brief. Different complaint types, industries or customer circumstances may produce different results.

2.

Single Human participant — The Human response represents one individual's writing ability and judgement. It should not be treated as representative of all human customer-service writers.

3.

Single AI system — The AI response represents the designated system and model used for this investigation at the time of testing. Other systems, models or later versions may perform differently.

4.

Written response only — The test evaluates the final written reply. It does not measure live conversation ability, telephone handling, negotiation, follow-up performance or long-term customer satisfaction.

5.

Simulated customer context — The judge can assess the quality of the response, but the investigation does not measure how an actual customer would react after receiving it.

6.

Judging subjectivity — Accuracy can be assessed relatively objectively, but empathy, clarity, practicality and trust involve human judgement. Different judges may reasonably score the same response differently.

7.

Blindness may be imperfect — Even when identities are hidden, writing style or other characteristics may cause a judge to suspect whether a response was produced by AI or Human.

8.

Time is reported separately — Completion time may show an efficiency difference, but it will not determine which response is judged better.

9.

Brief-dependent result — Both participants are restricted to the information contained in the locked brief. The quality of the brief therefore affects the quality and realism of both responses.

10.

No claim beyond the evidence — The final result will apply only to the conditions of this investigation. AVHM will not generalise the finding beyond what the evidence supports.

Timeline

How did the investigation unfold?

The investigation will progress through the following stages. Dates and times will be recorded in the investigation evidence as each stage is completed.

1.

Brief development — Define the investigation question, scope, success criteria, competition rules, judging criteria and known limitations.

2.

Brief lock — Freeze the final competition brief before either participant begins. Any later amendment must be recorded rather than silently incorporated.

3.

Human attempt — Provide the Human participant with the locked brief, record the start time and preserve the final submitted response and completion time.

4.

AI attempt — Provide the designated AI system with the same locked brief, record the start time and preserve the prompts, final submitted response and completion time.

5.

Submission lock — Confirm that both final responses are preserved and cannot be edited before judging.

6.

Blind preparation — Remove identifying information where practical and prepare the responses as Submission A and Submission B.

7.

Blind judging — Give the judge the complaint, relevant brief and both anonymous responses. Record all five criterion scores and qualitative commentary before revealing participant identities.

8.

Identity guess — Ask the judge which submission they believe was produced by AI and which by a Human, together with their confidence in that judgement.

9.

Reveal — Reveal the identities only after the judging scores and identity guess have been locked.

10.

Evidence review — Check the preserved brief, submissions, timings, judging records, interventions, deviations and known limitations before drawing the final conclusion.

11.

Editorial conclusion — Determine what the evidence supports, what it does not support and the appropriate level of confidence in the result.

12.

Publication — Publish the result, judging, evidence summary, limitations and underlying protocol as part of the AVHM public record.

Core Rule

The published result must reflect the evidence produced by the test, including any weaknesses or limitations in the experiment. The rules must not be retrospectively changed to favour either participant. Test Everything. Hype Nothing.

Investigation Record

Published investigation

The investigation result, judging, evidence, interpretation, limitations and underlying protocol are now part of the public record.

Evidence Principle

Test Everything. Hype Nothing.

AVHM records the result and the weaknesses in the experiment rather than hiding inconvenient evidence.