Human
Human
Score
35 / 50
Timing
Approximately 9 minutes
Submission Status
Valid with minor deviations
System
—
Model
—
Provider
—
AVHM-0008
WritingThe investigation is complete. The result, judging, evidence, limitations and final verdict are now part of the public record.
Result At A Glance
The headline numbers are simple. How strongly we can interpret them is not.
Human
Human
Score
35 / 50
Timing
Approximately 9 minutes
Submission Status
Valid with minor deviations
System
—
Model
—
Provider
—
AI
AI
Score
38 / 50
Timing
Not precisely measured
Submission Status
Valid with independence limitation
System
ChatGPT
Model
GPT-5.6 Sol
Provider
OpenAI
Comparison
Human
Human
AI
AI
Score
35 / 50
38 / 50
Timing
Approximately 9 minutes
Not precisely measured
Submission Status
Valid with minor deviations
Valid with independence limitation
System
—
ChatGPT
Model
—
GPT-5.6 Sol
Provider
—
OpenAI
AVHM Verdict
The observed result favoured AI, but the strength of that result depends on the evidence and experimental controls.
Evidence Confidence
Moderate
Independence
Material AI prior-familiarity limitation
What Happened
The important findings first. The full evidence interpretation remains available below.
35 / 50
Human judging score
The Human submission's recorded judging score.
38 / 50
AI judging score
The AI submission's recorded judging score.
Full finding preserved from the investigation record
Blind Judging
The judge scored the anonymous submissions using the categories fixed before the result was known.
Human
35 / 50
Blind judging score
AI
Higher Score38 / 50
Blind judging score
Judging Result
AI received the higher overall blind judging score.
The full judging record remains available below, including the scoring process, judge comments and known limitations.
Full judging record preserved from the investigation file
The Reveal
The anonymous submissions were revealed only after judging had been completed.
Human Submission
Human
Score
35 / 50
Timing
Approximately 9 minutes
Submission Status
Valid with minor deviations
System
—
Model
—
Provider
—
AI Submission
Observed WinnerAI
Score
38 / 50
Timing
Not precisely measured
Submission Status
Valid with independence limitation
System
ChatGPT
Model
GPT-5.6 Sol
Provider
OpenAI
Reveal Result
The observed result favoured AI.
The full reveal record below preserves the investigation's detailed interpretation and any limitations.
Full reveal record preserved from the investigation file
What Surprised Us
The most interesting part was not simply who won. It was where the differences appeared — and where they did not.
Full analysis preserved from the investigation record
What We Learned
The result matters. So does what the investigation taught us about running better tests next time.
Lesson 01
Lesson 02
Lesson 03
Lesson 04
Lesson 05
Full lessons record preserved from the investigation file
AVHM Recommendation
A winner is only useful if the evidence tells us what the result means in practice.
Observed Result
AI advantage
The recommendation starts with what the investigation actually observed.
Evidence Confidence
Moderate
This rating describes confidence in the investigation evidence itself. It is separate from any confidence rating attached to the practical recommendation.
Editorial Recommendation
The full recommendation below preserves the investigation-specific advice, limitations and practical interpretation exactly where they belong: in the investigation record itself.
Decision Basis
Observed result plus evidence strength
AVHM Principle
Never recommend more than the evidence supports
The expanded recommendation may include its own recommendation-confidence rating. That is separate from the Evidence Confidence shown above.
The evidence does not justify recommending that AI universally replace Human customer-service writers.
It does demonstrate that an AI-generated complaint response can compete strongly with Human-written customer-service communication and, in this test, receive the higher blind score.
A reasonable practical interpretation is:
AI can be useful for drafting customer complaint responses, with appropriate Human oversight and organisational controls.
This recommendation is narrower than claiming AI is generally superior.
The investigation tested one scenario, one Human participant, one AI system and one blind judge.
Full recommendation preserved from the investigation record
Final AVHM Interpretation
The observed result tells us what happened. The evidence determines how confidently we can interpret it.
Final Verdict
The observed result favoured AI. AVHM does not treat that observed advantage as stronger than the evidence allows.
Evidence Confidence
Moderate
Independence
Material AI prior-familiarity limitation
AVHM-0008 produced a clear result within the locked blind judging process: the AI response received the higher score.
Blind Judge 01 scored the AI response 38 / 50 and the Human response 35 / 50.
The observed advantage came entirely from the subjective categories of Empathy and Trust. Accuracy, Clarity and Practicality were scored equally.
The result is meaningful for this specific comparison, but it does not establish that AI is generally better than humans at writing customer complaint responses.
The competing AI had prior familiarity with the competition brief because it had participated in developing and auditing the experiment before its own attempt.
AI completion time was not independently measured with sufficient precision, so no reliable speed winner is claimed.
Evidence confidence: MODERATE
Full interpretation preserved from the investigation record
The Method Behind The Result
AVHM publishes the protocol alongside the result so the question, rules and limitations can be examined rather than simply trusted.
Why This Investigation Matters
Customer complaints are a useful test of more than writing ability. A good response needs to understand the problem, communicate clearly, show appropriate empathy and offer a practical way forward.
AI systems can produce polished text extremely quickly, but polished writing is not necessarily effective communication. A human writer may be better at recognising tone, nuance and what an unhappy customer actually needs to hear.
This investigation will test whether AI can turn the same complaint information into a better response than a human when both work from the same brief.
Compiled directly from the investigation record
Objective
Determine whether an AI system can produce a higher-quality customer complaint response than a skilled human writer when both receive the same brief, the same information and the same time limit.
Scope
This investigation compares only the quality of the final written response. It does not measure customer satisfaction, legal compliance or long-term business outcomes.
Investigation Protocol
The success criteria, competition rules and judging method are recorded in advance so the test cannot quietly move the goalposts after either side performs.
Success Criteria
A successful response should address the customer's complaint accurately, communicate with an appropriate level of empathy and provide a clear and practical way forward.
Both submissions will be assessed against the same five criteria:
Accuracy — Does the response correctly understand and address the facts of the complaint without inventing information or making unsupported claims?
Empathy — Does the response recognise the customer's frustration appropriately and sound sincere without becoming exaggerated, defensive or dismissive?
Clarity — Is the response easy to understand, well structured and free from unnecessary wording or ambiguity?
Practicality — Does the response provide a useful, realistic next step or resolution based only on the information available in the brief?
Trust — Would the response increase the customer's confidence that their complaint has been understood and is being handled fairly?
Each criterion will be scored out of 10 using the same judging standard for both submissions.
The higher combined judging score will determine which response performed better overall.
Completion time will also be recorded and reported as part of the investigation, but speed alone will not determine the winner unless explicitly incorporated into the locked judging method before either attempt begins.
Competition Rules
Same complaint — The AI and Human participants must respond to exactly the same customer complaint.
Same brief — Both participants must receive the same instructions, customer information, company information and permitted resolution options.
Same information — Neither participant may be given additional facts, context or guidance that is unavailable to the other.
Independent attempts — Neither participant may see, review or receive information about the other participant's response before their own final response has been submitted.
No outside research — Neither participant may search the web, consult another person or use external reference material unless the locked brief explicitly permits it.
Human tools — The Human participant may use a standard text editor and normal spelling or grammar checking, but may not use generative AI or AI-assisted rewriting.
AI tools — The AI participant may use only the designated AI system and the information contained in the locked brief. No additional human editing of the AI response is permitted before submission.
Timing — Timing begins when the participant receives the complete locked brief and ends when they declare their response final.
Final submission — Once a participant declares their response final, it is locked. No rewriting, correction or improvement is permitted before judging.
Response format — Both participants must produce a customer-facing written response suitable for sending directly to the customer. Any required length or formatting constraints must be defined in the locked brief before either attempt begins.
Blind judging — The final responses will be presented as Submission A and Submission B. Information that directly identifies which response was produced by AI or Human will be removed where practical.
Judging independence — Judges must score the written responses using the locked judging criteria before the identities of the participants are revealed.
Completion time — Completion time will be recorded and disclosed, but will not be shown to judges before scoring because it could reveal participant identity and influence assessment of writing quality.
No silent rule changes — Any ambiguity, deviation or unexpected issue that occurs during the investigation must be recorded. The rules must not be retrospectively changed to favour either participant.
Evidence preservation — The brief, prompts or instructions, final submissions, timing information, judging records and any material interventions must be preserved as part of the investigation evidence.
Core principle — The published result must reflect the evidence produced by the test, including any weaknesses or limitations in the experiment.
Judging Criteria
Both sides faced the same five categories and the same scoring scale.
The Human and AI responses will be judged against the five criteria defined in the Success Criteria:
Accuracy
Empathy
Clarity
Practicality
Trust
Each criterion will be scored independently from 1 to 10, giving each submission a maximum possible score of 50 points.
The Human and AI responses will be labelled Submission A and Submission B before being shown to the judge.
The order of the submissions will not indicate which response was produced by the Human or AI participant.
The judge will receive the original customer complaint, the relevant locked brief and both final responses.
The judge will not be shown participant identities, completion times or process information before scoring.
Each submission will be scored independently against all five criteria.
The judge must record a score for every criterion before participant identities are revealed.
The five scores for each submission will then be added to produce a total score out of 50.
The submission with the higher total judging score will be the judged winner.
If both submissions receive the same total score, the judged result will be recorded as a draw.
After all scores are locked, the Human and AI identities will be revealed.
The judge will also be invited to explain the reasoning behind the scores and identify the strongest and weakest aspects of each response.
Any qualitative comments will be preserved alongside the numerical scores so the published result shows not only who scored higher, but why.
Before the reveal, the judge will be asked which submission they believe was produced by AI and which was produced by a Human, together with their confidence in that judgement.
This identity guess will not affect the competition score. It will be recorded separately as an additional Turing-test-style observation.
Completion time will be revealed only after judging has been completed and the scores have been locked.
Speed will be reported as a separate performance measure and will not contribute points to the judging score.
Limitations
This investigation is designed to compare the quality of two written complaint responses under controlled conditions. It cannot prove that one participant would always perform better in every real customer-service situation.
The main limitations are:
Single complaint scenario — The result will be based on one complaint brief. Different complaint types, industries or customer circumstances may produce different results.
Single Human participant — The Human response represents one individual's writing ability and judgement. It should not be treated as representative of all human customer-service writers.
Single AI system — The AI response represents the designated system and model used for this investigation at the time of testing. Other systems, models or later versions may perform differently.
Written response only — The test evaluates the final written reply. It does not measure live conversation ability, telephone handling, negotiation, follow-up performance or long-term customer satisfaction.
Simulated customer context — The judge can assess the quality of the response, but the investigation does not measure how an actual customer would react after receiving it.
Judging subjectivity — Accuracy can be assessed relatively objectively, but empathy, clarity, practicality and trust involve human judgement. Different judges may reasonably score the same response differently.
Blindness may be imperfect — Even when identities are hidden, writing style or other characteristics may cause a judge to suspect whether a response was produced by AI or Human.
Time is reported separately — Completion time may show an efficiency difference, but it will not determine which response is judged better.
Brief-dependent result — Both participants are restricted to the information contained in the locked brief. The quality of the brief therefore affects the quality and realism of both responses.
No claim beyond the evidence — The final result will apply only to the conditions of this investigation. AVHM will not generalise the finding beyond what the evidence supports.
Timeline
The investigation will progress through the following stages. Dates and times will be recorded in the investigation evidence as each stage is completed.
Brief development — Define the investigation question, scope, success criteria, competition rules, judging criteria and known limitations.
Brief lock — Freeze the final competition brief before either participant begins. Any later amendment must be recorded rather than silently incorporated.
Human attempt — Provide the Human participant with the locked brief, record the start time and preserve the final submitted response and completion time.
AI attempt — Provide the designated AI system with the same locked brief, record the start time and preserve the prompts, final submitted response and completion time.
Submission lock — Confirm that both final responses are preserved and cannot be edited before judging.
Blind preparation — Remove identifying information where practical and prepare the responses as Submission A and Submission B.
Blind judging — Give the judge the complaint, relevant brief and both anonymous responses. Record all five criterion scores and qualitative commentary before revealing participant identities.
Identity guess — Ask the judge which submission they believe was produced by AI and which by a Human, together with their confidence in that judgement.
Reveal — Reveal the identities only after the judging scores and identity guess have been locked.
Evidence review — Check the preserved brief, submissions, timings, judging records, interventions, deviations and known limitations before drawing the final conclusion.
Editorial conclusion — Determine what the evidence supports, what it does not support and the appropriate level of confidence in the result.
Publication — Publish the result, judging, evidence summary, limitations and underlying protocol as part of the AVHM public record.
Core Rule
The published result must reflect the evidence produced by the test, including any weaknesses or limitations in the experiment. The rules must not be retrospectively changed to favour either participant. Test Everything. Hype Nothing.
Investigation Record
The investigation result, judging, evidence, interpretation, limitations and underlying protocol are now part of the public record.
Evidence Principle
Test Everything. Hype Nothing.
AVHM records the result and the weaknesses in the experiment rather than hiding inconvenient evidence.