AI vs Human Media

AVHM-0012

Shopping/Research

Can AI find a cheaper supermarket basket than a human?

The investigation is complete. The result, judging, evidence, limitations and final verdict are now part of the public record.

PublishedCreated 2026-08-17

Result At A Glance

Human vs AI.
Here's what happened.

The headline numbers are simple. How strongly we can interpret them is not.

Human

Human

Supermarket

Asda

Total

£32.05

Time

25m 23s

Score

40 / 50

AI

AI

Observed Winner

Supermarket

Morrisons

Total

£31.69

Time

3m 43s

Score

45 / 50

Price

£0.36

Advantage: AI

Time

21m 40s

Advantage: AI

AVHM Verdict

AI ADVANTAGE DEMONSTRATED — CLEAN VICTORY NOT ESTABLISHED

The observed result favoured AI, but the strength of that result depends on the evidence and experimental controls.

Evidence Confidence

Qualified

Independence

Material limitation

What Happened

What the evidence showed.

The important findings first. The full evidence interpretation remains available below.

£0.36

AI price advantage

AI held the observed advantage for price.

21m 40s

AI time advantage

AI held the observed advantage for time.

40 / 50

Human judging score

The Human submission's recorded judging score.

45 / 50

AI judging score

The AI submission's recorded judging score.

Read the full evidence finding+
The evidence showed that both participants produced supermarket baskets that passed the current review of all 20 locked product and minimum-quantity requirements. The Human participant submitted an Asda basket costing: **£32.05** The AI participant submitted a Morrisons basket costing: **£31.69** Both submitted totals were independently recalculated successfully. On the locked submitted totals, the AI basket was: **£0.36 cheaper** The difference was small, at approximately **1.1%** of the Human basket total. The clearest measured difference was speed. The Human participant completed the task in: **25 minutes 23 seconds** The AI participant completed it in: **3 minutes 43 seconds** The AI was therefore **21 minutes 40 seconds faster**. Blind Judge 01 scored the submissions: **Human: 40 / 50** **AI: 45 / 50** The entire five-point difference came from Speed. The judge awarded identical scores to both participants for: - Accuracy - Creativity - Practicality - Trust This means the judging evidence did not show that the AI produced a more accurate, creative, practical or trustworthy basket than the Human participant. It showed a substantial advantage in completion speed alongside a small advantage in submitted price. The evidence also showed important limitations. Neither submission had perfect contemporaneous retailer evidence preservation, and original Morecambe-specific availability could not be independently established for every product. More importantly, the AI attempt was not fully independent. Before its timed attempt began, the AI system had already been exposed to the Human participant's supermarket and final basket total. It is not possible to establish retrospectively whether or how much that prior knowledge influenced the AI result. The evidence therefore supports the conclusion that: **AI demonstrated an advantage in this test, particularly in speed, but a clean independent victory was not established.**

Full finding preserved from the investigation record

Blind Judging

The submissions were scored before the identities were revealed.

The judge scored the anonymous submissions using the categories fixed before the result was known.

Human

40 / 50

Blind judging score

AI

Higher Score

45 / 50

Blind judging score

Judging Result

AI received the higher overall blind judging score.

The full judging record remains available below, including the scoring process, judge comments and known limitations.

Read the full blind judging record+
Judging was conducted after both submissions had been locked and independently reviewed against the investigation brief. An independent third-party judge, recorded as **Blind Judge 01**, was used. The judge was the Human participant's father. He had no prior knowledge of the AI vs Human Media project before taking part in the judging. Before judging, the two submissions were assigned anonymous labels: - Submission A - Submission B The Human/AI identity mapping was fixed and recorded separately before the judging materials were prepared. The judge was not told which submission had been produced by the Human participant and which had been produced by the AI participant. Both submissions were presented using the same judging structure. The judge received the locked shopping requirements, the two anonymous submissions and a scorecard. The judge was asked to score each submission out of 10 in the five categories fixed in the investigation brief: - Accuracy - Speed - Creativity - Practicality - Trust The judge was instructed to assess only the information supplied and not conduct additional research. The recorded completion times were included because Speed was one of the original locked judging categories. The locked scores were: | Category | Submission A | Submission B | |---|---:|---:| | Accuracy | 10 / 10 | 10 / 10 | | Speed | 5 / 10 | 10 / 10 | | Creativity | 8 / 10 | 8 / 10 | | Practicality | 8 / 10 | 8 / 10 | | Trust | 9 / 10 | 9 / 10 | | **Total** | **40 / 50** | **45 / 50** | Submission B therefore won the blind judging by **45 points to 40**. The entire five-point difference came from the Speed category. Accuracy, Creativity, Practicality and Trust received identical scores for both submissions. The judge summarised the result as: > "b won on speed, everything else is the same standard food" After completing the scoring, but before the Human/AI identities were formally revealed, the judge was asked which submission he believed had been produced by AI. He selected **Submission B** and stated that he was **very confident**. The substantially faster recorded completion time was the reason for that inference. This exposed a limitation in the blind-judging design. Although the formal participant identities had been withheld, the difference in completion times made participant identity inferable to the judge. The scores were preserved without alteration. The judging procedure was not retrospectively changed to compensate for this limitation. Details of a material independence limitation affecting the AI attempt were withheld during blind scoring because disclosing them would have revealed participant identity. That limitation remained preserved elsewhere in the investigation record for consideration after the blind scores were locked.

Full judging record preserved from the investigation file

The Reveal

Human or AI?

The anonymous submissions were revealed only after judging had been completed.

Human Submission

Human

Supermarket

Asda

Total

£32.05

Time

25m 23s

Score

40 / 50

AI Submission

Observed Winner

AI

Supermarket

Morrisons

Total

£31.69

Time

3m 43s

Score

45 / 50

Reveal Result

The observed result favoured AI.

The full reveal record below preserves the investigation's detailed interpretation and any limitations.

Read the full reveal record+
The Human/AI identities were revealed only after Blind Judge 01 had completed the scoring and the blind judgement had been preserved. The confidential mapping fixed before judging was: - **Submission A = Human** - **Submission B = AI** The mapping had been recorded before preparation of the final judging materials and was not changed after the scores were known. Before the identities were formally revealed, Blind Judge 01 was asked which submission he believed had been produced by AI. The judge predicted: **Submission B = AI** Confidence: **Very confident** The prediction was correct. The judge explained that the substantially faster completion time led him to identify Submission B as the AI submission. After the reveal, the anonymous results became: | | Human | AI | |---|---:|---:| | Basket total | £32.05 | £31.69 | | Completion time | 25m 23s | 3m 43s | | Blind judging score | 40 / 50 | 45 / 50 | The AI submission was therefore **£0.36 cheaper** on the locked submitted totals. It was also **21 minutes 40 seconds faster**. The AI received the higher blind judging score by **45 points to 40**. However, the entire five-point judging difference came from Speed. The judge scored the Human and AI submissions identically for Accuracy, Creativity, Practicality and Trust. The identity reveal did not alter the locked judging scores. After the reveal, a material limitation affecting the AI attempt returned to the final interpretation. Before the timed AI attempt began, the AI system had already been exposed to the Human participant's selected supermarket and final basket total. That information should not have been available under the intended independent-testing condition. The AI had been instructed not to use the Human participant's individual product selections or research to optimise its basket, but the intended independence condition was nevertheless not fully achieved. This information had been withheld from the blind judge because revealing it during scoring would itself have disclosed which submission was AI. The limitation was restored after the blind judgement had been locked and was carried forward into the final editorial conclusion. The reveal therefore established the observed result: **AI produced the cheaper submitted basket, completed the task substantially faster and received the higher blind judging score.** It did not establish an unconditional or fully independent AI victory. The strength of that claim remained subject to the evidence and independence limitations documented elsewhere in the investigation.

Full reveal record preserved from the investigation file

What Surprised Us

The result was not quite what the headline suggests.

The most interesting part was not simply who won. It was where the differences appeared — and where they did not.

Surprise 01

£0.36

Price

This was one of the clearest observed differences between the Human and AI submissions.

Surprise 02

21m 40s

Time

This was one of the clearest observed differences between the Human and AI submissions.

Read the full surprise analysis+
Several parts of the result were genuinely unexpected. The first was how small the price difference was. The AI submission cost **£31.69** and the Human submission cost **£32.05** — a difference of only **36p**. Despite using different supermarkets and different product selections, the final basket totals were remarkably close. The much larger difference was speed. The Human participant took **25 minutes 23 seconds**. The AI participant took **3 minutes 43 seconds**. The AI therefore completed the task **21 minutes 40 seconds faster**, while finding a basket only 36p cheaper. That contrast was more striking than the price result itself. Another surprise came from the blind judging. Blind Judge 01 scored both submissions identically for: - Accuracy - Creativity - Practicality - Trust The only category separating them was Speed. The judge summarised his assessment as: > "b won on speed, everything else is the same standard food" This meant that, despite the large difference in completion time, the judge did not regard the resulting baskets as meaningfully different in the other four scoring categories. The blind test also produced an unexpected methodological lesson. Although the Human/AI identities were withheld, the difference in completion times made the AI submission obvious to the judge. He correctly identified Submission B as AI and was very confident in that prediction. This showed that anonymising participant names alone is not always enough to create an effectively blind Human-versus-AI comparison. A measurement such as completion time can itself reveal participant identity. Finally, the investigation exposed how important evidence preservation is. Both baskets appeared straightforward when they were originally produced. Only during the later evidence review did issues such as missing contemporaneous screenshots, changing prices, later product availability and retrospective retailer verification become significant. The investigation therefore produced a lesson beyond supermarket shopping: **running the test was easier than proving exactly what happened afterward.**

Full analysis preserved from the investigation record

What We Learned

Every investigation should improve the next one.

The result matters. So does what the investigation taught us about running better tests next time.

01

Lesson 01

Capture evidence during the attempt

02

Lesson 02

Independence must begin before the timed test

03

Lesson 03

Lock the judging procedure before the attempts

04

Lesson 04

Blind does not necessarily mean unidentifiable

05

Lesson 05

Separate the headline result from the strength of the evidence

06

Lesson 06

Small differences require stronger evidence

07

Lesson 07

Record problems instead of repairing them

08

Lesson 08

Build the evidence process into the experiment itself

Read the full lessons record+
AVHM-0012 produced practical lessons that should improve future investigations. ## 1. Capture evidence during the attempt The biggest operational lesson was that evidence should be preserved as the experiment happens. Trying to reconstruct retailer evidence afterward introduced unnecessary uncertainty. Future shopping investigations should capture, where practical: - retailer pages - product names - prices - quantities - basket totals - promotions - availability information - start and finish times before the participant's attempt is considered complete. Retrospective verification can still be useful, but it is weaker than contemporaneous evidence. ## 2. Independence must begin before the timed test The AI attempt demonstrated that independence cannot be created simply by instructing a participant to ignore information it has already received. The AI already knew the Human participant's supermarket and final basket total before its timed attempt began. Future investigations should isolate competing submissions before either participant receives information about the other's attempt. If information has already been disclosed, that should be treated as a design problem rather than something that instructions alone can fully undo. ## 3. Lock the judging procedure before the attempts The judging categories were fixed in advance, which helped prevent the scoring system from being changed to suit the eventual result. That principle should be extended further. Future investigations should define before testing: - what information the judge receives - what information is withheld - how participant identities are anonymised - when identities are revealed - how objective measurements such as Speed are incorporated - how evidence limitations are presented ## 4. Blind does not necessarily mean unidentifiable Submission A and Submission B were formally anonymised, but the large difference in completion times allowed Blind Judge 01 to infer correctly which submission was AI. Future blind judging should consider whether information necessary for scoring can itself reveal participant identity. If it can, AVHM should decide before the experiment whether that is acceptable or whether objective categories should be assessed separately. ## 5. Separate the headline result from the strength of the evidence The headline numbers favoured AI: - 36p cheaper - 21 minutes 40 seconds faster - 45/50 versus 40/50 in blind judging But those numbers alone did not describe the quality of the experiment. Evidence preservation and independence limitations materially affected how strongly the result could be stated. Future investigations should continue separating: - observed result - validity - evidence quality - methodological limitations - final editorial verdict ## 6. Small differences require stronger evidence The price difference was only 36p. When the winning margin is small, relatively minor price changes, promotions or evidence uncertainties can become important. The closer the result, the more important precise evidence preservation becomes. ## 7. Record problems instead of repairing them AVHM-0012 contained imperfections. The Human contemporaneous trolley screenshots were unavailable. The AI retailer evidence was not preserved adequately during the timed attempt. Some prices changed or were promotional. Local availability could not always be reconstructed. The AI independence condition was compromised. None of these issues was silently removed from the record. That approach should remain part of AVHM's methodology: **a documented imperfect experiment is more trustworthy than a retrospectively perfect-looking one.** ## 8. Build the evidence process into the experiment itself The most important overall operational lesson is that evidence collection should not be a separate job performed after an investigation. It should be part of the investigation procedure from the beginning. For future AVHM tests, the experiment should not be considered complete until the required evidence has been captured, labelled and stored. That will make later verification faster, reduce ambiguity and allow stronger conclusions. AVHM-0012 therefore improved not only our understanding of the question being tested, but also the way future AI vs Human Media investigations should be designed.

Full lessons record preserved from the investigation file

AVHM Recommendation

What should someone actually do with this result?

A winner is only useful if the evidence tells us what the result means in practice.

Observed Result

AI advantage

The recommendation starts with what the investigation actually observed.

Evidence Confidence

Qualified

This rating describes confidence in the investigation evidence itself. It is separate from any confidence rating attached to the practical recommendation.

Editorial Recommendation

The practical recommendation depends on the evidence, not just the winner.

The full recommendation below preserves the investigation-specific advice, limitations and practical interpretation exactly where they belong: in the investigation record itself.

Decision Basis

Observed result plus evidence strength

AVHM Principle

Never recommend more than the evidence supports

The expanded recommendation may include its own recommendation-confidence rating. That is separate from the Evidence Confidence shown above.

Read the full recommendation+

Based on the evidence from AVHM-0012, our recommendation is:

USE AI FOR THE SEARCH — VERIFY BEFORE YOU BUY

The strongest demonstrated AI advantage in this investigation was speed.

The AI produced its submitted basket in 3 minutes 43 seconds, compared with 25 minutes 23 seconds for the Human participant.

Its submitted basket was also 36p cheaper.

However, the experiment does not provide enough evidence to recommend relying unquestioningly on an AI-generated supermarket basket.

The AI submission had limitations around contemporaneous retailer evidence, local availability and time-sensitive pricing.

It also did not satisfy the intended independence condition fully in this investigation.

For an ordinary shopper, the most practical use of AI suggested by this test is therefore:

1.

Give AI a precise shopping list and requirements.

2.

Ask it to identify a low-cost basket quickly.

3.

Use that result as a shortlist or starting point.

4.

Verify the important prices, promotions and availability with the retailer before purchasing.

This approach uses the strongest capability demonstrated in the experiment — rapid research — while retaining Human verification where current retailer information matters.

The evidence does not show that AI selected meaningfully better food.

Blind Judge 01 scored the Human and AI submissions identically for Accuracy, Creativity, Practicality and Trust.

The main difference was how quickly the result was produced.

For this task, AI therefore appears most useful as a fast research assistant rather than an unquestioned final authority.

AVHM Recommendation

AI-assisted shopping research with Human verification.

Confidence

Moderate

The recommendation is supported by a substantial measured speed advantage and a small observed price advantage, but confidence is limited by the evidence-preservation and independence issues documented in this investigation.

A properly controlled repeat investigation would be required before making a stronger recommendation.

Full recommendation preserved from the investigation record

Final AVHM Interpretation

The result matters.
The limitations matter too.

The observed result tells us what happened. The evidence determines how confidently we can interpret it.

Final Verdict

AI ADVANTAGE DEMONSTRATED — CLEAN VICTORY NOT ESTABLISHED

The observed result favoured AI. AVHM does not treat that observed advantage as stronger than the evidence allows.

Evidence Confidence

Qualified

Independence

Material limitation

Read the full final interpretation+

If AVHM considered only the headline numbers, the result would be simple:

AI won.

It submitted the cheaper basket.

It completed the task much faster.

It received the higher blind judging score.

But AVHM does not consider only the headline numbers.

The AI knew information about the Human result before beginning its timed attempt.

That weakens the causal claim that AI independently outperformed the Human participant.

The blind judging also revealed something important: outside Speed, the judge saw no scoring difference between the two submissions.

The strongest finding from AVHM-0012 is therefore not that AI was dramatically better at supermarket shopping.

It is that AI was able to produce a reviewed, competitive basket far faster, while achieving a submitted total marginally below the Human participant.

Whether AI would retain that price advantage under genuinely independent conditions remains unanswered.

Full interpretation preserved from the investigation record

The Method Behind The Result

Now inspect how the test was designed.

AVHM publishes the protocol alongside the result so the question, rules and limitations can be examined rather than simply trusted.

Why This Investigation Matters

Why are we asking this question?

Grocery prices affect almost everyone, and shoppers increasingly use AI tools to help compare products, plan purchases and reduce costs.

This investigation tests whether an AI can use publicly available supermarket information to assemble a cheaper valid basket than a human working from the same shopping brief and the same retailer constraints.

The question matters because a lower headline price is only useful if the products are actually comparable, available, correctly priced and suitable for the brief. The investigation therefore tests not just price finding, but whether AI can produce a result that is accurate, practical and verifiable.

It also provides a useful real-world test of how well AI handles changing online information such as supermarket prices, offers, product sizes and availability.

The aim is not to prove that AI or humans are universally better at shopping. It is to produce evidence about how each performs on one clearly defined supermarket-basket task under the same conditions.

Compiled directly from the investigation record

Objective

What are we trying to determine?

Determine whether AI or a human can identify the lowest-cost valid supermarket basket from publicly available online supermarket information when both are given the same fixed shopping list, comparison rules and research conditions.

The investigation will compare the final valid basket price achieved by each participant while also recording the accuracy, completeness and verifiability of the selections.

A cheaper basket will only count as better if every selected product satisfies the locked requirements of the investigation.

Scope

What does this investigation cover?

This investigation compares one AI participant and one human participant attempting to find the cheapest valid supermarket basket using the same locked shopping list and comparison rules.

Included:

Products available through the supermarkets and online sources permitted by the locked brief.

Publicly displayed prices and qualifying promotions available during the defined research period.

Product size, quantity and specification checks needed to determine whether an item satisfies the brief.

The final total price of each valid basket.

The research process used by both participants, including sources consulted and relevant decisions made.

Evidence required to verify the selected products and prices.

Excluded:

Loyalty-only prices or personalised discounts unless the locked brief explicitly permits them.

Delivery charges, travel costs and other costs outside the product basket unless specifically included in the brief.

Products that do not meet the locked quantity, size or specification requirements.

Substituting an unavailable product after the research period has ended merely to improve a participant's result.

Claims about which participant would perform better across all supermarkets, shopping lists or future price conditions.

The investigation is a comparison of performance on one controlled supermarket-basket task, not a general test of whether AI or humans are universally better shoppers.

Investigation Protocol

The rules are set before we know the result.

The success criteria, competition rules and judging method are recorded in advance so the test cannot quietly move the goalposts after either side performs.

Success Criteria

What does a good result look like?

A participant's final basket will be considered valid only if every required item satisfies the locked shopping brief and can be supported by verifiable evidence from an approved source.

The investigation will assess success using the following criteria:

Basket validity: Every required item must be included and meet the locked product, quantity and size requirements.

Total cost: The final valid basket price will be calculated using the recorded prices available during the investigation.

Accuracy: Product names, sizes, quantities, prices and qualifying offers must be recorded correctly.

Completeness: No required item may be omitted from the final basket.

Verifiability: Each final selection must have sufficient evidence for its price and eligibility to be independently checked.

Speed: The time taken by each participant to produce the final basket will be recorded.

Practicality: The final basket must represent a purchase that could reasonably have been made under the locked conditions.

The cheaper valid basket will win the price comparison.

A lower total will not automatically constitute a better result if the basket contains an invalid item, an unsupported price, a missing product or otherwise fails the locked brief.

If the evidence does not establish a fair and reliable winner, the investigation may result in a draw or an inconclusive finding.

Competition Rules

A fair test needs fixed rules.

The following rules must be locked before either participant begins researching the basket.

1.

Same brief: The AI and human must receive the same locked shopping list, product requirements, retailer rules and comparison conditions.

2.

Same research window: Both participants must complete their research within the defined investigation period so that price changes do not knowingly favour either side.

3.

Same permitted retailers: Both participants may search only the supermarkets or online sources specified in the locked brief.

4.

Same product requirements: Products must satisfy the same minimum quantity, size, type and other requirements. A cheaper product that fails the brief is not a valid selection.

5.

Publicly available information: Participants may use publicly accessible supermarket information permitted by the brief. Account-specific, personalised or otherwise unavailable prices must not be used unless explicitly allowed before testing begins.

6.

Offers and discounts: Promotions may be used only when they are available under the locked conditions and the evidence clearly shows that the basket qualifies for them.

7.

Availability: A selected product must be shown as available through the permitted source when it is recorded. If availability cannot reasonably be verified, that limitation must be documented.

8.

Evidence capture: Each final product selection must record enough information to verify the item, price, quantity or size, retailer and source at the time of research.

9.

No retrospective optimisation: Once a participant's final basket has been submitted, products may not be replaced simply because a cheaper option is discovered later.

10.

Independent attempts: Neither participant may see the other participant's research, selections or final basket before both submissions are locked.

11.

Interventions recorded: Any clarification, technical problem, assistance or change affecting either attempt must be recorded in the evidence log.

12.

Invalid selections: Any item that fails the locked brief must be identified during judging and cannot contribute to a claimed cheaper valid basket without the discrepancy being disclosed.

13.

Same scoring method: Both submissions will be assessed using the judging criteria locked before the competition begins.

14.

Evidence over outcome: Rules, evidence or scores must not be changed to manufacture a Human Win, AI Win or Draw.

Judging Criteria

How were AI and Human judged?

Both sides faced the same five categories and the same scoring scale.

Both participants will be scored using the same five AI vs Human Media judging categories. Each category will be scored out of 10, producing a maximum total score of 50 for each participant.

Accuracy

Measures whether the final basket correctly satisfies the locked shopping brief.

Judging will consider:

Whether every required product is included.

Whether selected products meet the required type, size and quantity.

Whether recorded prices and offers are correct.

Whether basket totals have been calculated correctly.

Whether invalid or non-comparable products have been included.

Speed

Measures the time taken to produce a completed final submission.

Timing will begin when the participant receives the locked brief and end when the final basket is submitted.

Any pauses, interruptions or technical problems that materially affect the comparison must be recorded.

Creativity

Measures the participant's ability to identify legitimate approaches that reduce the basket cost without breaking the locked rules.

This may include:

Finding less obvious valid alternatives.

Making effective use of permitted promotions.

Identifying equivalent products that satisfy the brief at a lower cost.

Using an efficient research strategy.

Creativity does not reward loopholes, invalid substitutions or rule-breaking.

Practicality

Measures whether the proposed basket represents a realistic and usable supermarket purchase.

Judging will consider:

Product availability.

Whether quantities and pack sizes make practical sense.

Whether the basket could reasonably be purchased under the recorded conditions.

Whether savings depend on unrealistic assumptions.

Whether the final result would be useful to an ordinary shopper following the same brief.

Trust

Measures how well the participant's result is supported by clear, traceable and verifiable evidence.

Judging will consider:

Quality of source evidence.

Accuracy of product and price records.

Transparency about uncertainty or missing information.

Whether important assumptions have been disclosed.

Whether another person could reasonably verify the final basket from the recorded evidence.

Final Result

The five category scores will be added to produce a total out of 50 for each participant.

The higher overall score will determine the investigation winner.

The raw basket price will also be reported separately so the audience can see which participant found the cheaper valid basket.

If the total scores are equal, the result will be recorded as a Draw.

If evidence is insufficient to support a reliable comparison, the investigation may be declared inconclusive rather than forcing a winner.

Limitations

What could this test not prove?

This investigation is designed to produce a fair comparison under controlled conditions, but several limitations must be recognised when interpreting the result.

Prices can change: Supermarket prices and promotions may change during or shortly after the investigation. The recorded result represents the prices that could be verified during the defined research period.

Availability can change: A product shown online may become unavailable, go out of stock or differ by location. Availability will be recorded where reasonably possible, but cannot be guaranteed beyond the investigation period.

Online and in-store prices may differ: The investigation relies on the sources permitted by the locked brief. A different price may exist in a physical store or through another sales channel.

Location may affect results: Product ranges, prices and availability can vary between stores and geographic areas. The result therefore applies only to the conditions defined for this investigation.

Promotions may have conditions: Multi-buy offers, loyalty schemes, minimum spends and other promotions can affect the apparent price. Only promotions permitted by the locked rules will be included.

Product comparison involves judgement: Different brands, pack sizes and product descriptions may not always be perfectly equivalent. Eligibility decisions will follow the locked brief and any uncertainty will be documented.

Search coverage cannot be guaranteed: Neither participant can prove that every possible valid product or offer was discovered. The investigation compares the results they actually achieved within the agreed conditions.

AI access may differ from human access: The AI participant may have different capabilities or restrictions when accessing current supermarket websites and online information. Any material access limitation or intervention will be recorded.

Timing can influence the comparison: Website performance, technical problems, interruptions and research methods may affect completion time. Material issues will be documented in the evidence log.

One investigation cannot establish universal superiority: The result applies to this specific basket, research period, participants and testing conditions. It does not demonstrate that AI or humans will always find cheaper supermarket baskets.

Any limitation discovered during the investigation that could materially affect the result will be added to the evidence record and disclosed when the findings are published.

Timeline

How did the investigation unfold?

The investigation will proceed through the following stages.

1. Protocol Preparation

Define and record the objective, scope, success criteria, competition rules, judging criteria, limitations and investigation method.

2. Brief Lock

Create the final shopping list and lock the permitted retailers, product requirements, research conditions, promotion rules and evidence requirements before either participant begins.

3. Human Attempt

Provide the human participant with the locked brief.

Record the start time, research process, relevant decisions, evidence sources and final basket. Lock the human submission before the AI attempt is revealed or compared.

4. AI Attempt

Provide the AI participant with the same locked brief.

Record the AI system used, start time, prompts or instructions, research process, interventions, evidence sources and final basket. Lock the AI submission before comparison.

5. Evidence Review

Verify the recorded products, prices, quantities, offers and sources for both submissions.

Identify any invalid selections, missing evidence, discrepancies or limitations before scoring begins.

6. Blind Presentation and Judging

Where practical, prepare anonymised versions of the two valid submissions so that judging does not unnecessarily reveal which participant produced each basket.

Score both submissions using the judging criteria locked before testing.

7. Result and Verdict

Calculate the final scores and compare the valid basket totals.

Record the Human Win, AI Win, Draw or inconclusive result supported by the evidence.

8. Publication

Publish the investigation record, relevant evidence, scores, limitations and final conclusion.

Any material correction discovered after publication will be recorded rather than silently altering the historical record.

Core Rule

Trust the evidence. Not us.

Investigation Record

Published investigation

The investigation result, judging, evidence, interpretation, limitations and underlying protocol are now part of the public record.

Evidence Principle

Test Everything. Hype Nothing.

AVHM records the result and the weaknesses in the experiment rather than hiding inconvenient evidence.