Human
Human
Supermarket
Asda
Total
£32.05
Time
25m 23s
Score
40 / 50
AVHM-0012
Shopping/ResearchThe investigation is complete. The result, judging, evidence, limitations and final verdict are now part of the public record.
Result At A Glance
The headline numbers are simple. How strongly we can interpret them is not.
Human
Human
Supermarket
Asda
Total
£32.05
Time
25m 23s
Score
40 / 50
AI
AI
Supermarket
Morrisons
Total
£31.69
Time
3m 43s
Score
45 / 50
Comparison
Human
Human
AI
AI
Supermarket
Asda
Morrisons
Total
£32.05
£31.69
Time
25m 23s
3m 43s
Score
40 / 50
45 / 50
Price
£0.36
Advantage: AI
Time
21m 40s
Advantage: AI
AVHM Verdict
The observed result favoured AI, but the strength of that result depends on the evidence and experimental controls.
Evidence Confidence
Qualified
Independence
Material limitation
What Happened
The important findings first. The full evidence interpretation remains available below.
£0.36
AI price advantage
AI held the observed advantage for price.
21m 40s
AI time advantage
AI held the observed advantage for time.
40 / 50
Human judging score
The Human submission's recorded judging score.
45 / 50
AI judging score
The AI submission's recorded judging score.
Full finding preserved from the investigation record
Blind Judging
The judge scored the anonymous submissions using the categories fixed before the result was known.
Human
40 / 50
Blind judging score
AI
Higher Score45 / 50
Blind judging score
Judging Result
AI received the higher overall blind judging score.
The full judging record remains available below, including the scoring process, judge comments and known limitations.
Full judging record preserved from the investigation file
The Reveal
The anonymous submissions were revealed only after judging had been completed.
Human Submission
Human
Supermarket
Asda
Total
£32.05
Time
25m 23s
Score
40 / 50
AI Submission
Observed WinnerAI
Supermarket
Morrisons
Total
£31.69
Time
3m 43s
Score
45 / 50
Reveal Result
The observed result favoured AI.
The full reveal record below preserves the investigation's detailed interpretation and any limitations.
Full reveal record preserved from the investigation file
What Surprised Us
The most interesting part was not simply who won. It was where the differences appeared — and where they did not.
Surprise 01
£0.36
This was one of the clearest observed differences between the Human and AI submissions.
Surprise 02
21m 40s
This was one of the clearest observed differences between the Human and AI submissions.
Full analysis preserved from the investigation record
What We Learned
The result matters. So does what the investigation taught us about running better tests next time.
Lesson 01
Lesson 02
Lesson 03
Lesson 04
Lesson 05
Lesson 06
Lesson 07
Lesson 08
Full lessons record preserved from the investigation file
AVHM Recommendation
A winner is only useful if the evidence tells us what the result means in practice.
Observed Result
AI advantage
The recommendation starts with what the investigation actually observed.
Evidence Confidence
Qualified
This rating describes confidence in the investigation evidence itself. It is separate from any confidence rating attached to the practical recommendation.
Editorial Recommendation
The full recommendation below preserves the investigation-specific advice, limitations and practical interpretation exactly where they belong: in the investigation record itself.
Decision Basis
Observed result plus evidence strength
AVHM Principle
Never recommend more than the evidence supports
The expanded recommendation may include its own recommendation-confidence rating. That is separate from the Evidence Confidence shown above.
Based on the evidence from AVHM-0012, our recommendation is:
The strongest demonstrated AI advantage in this investigation was speed.
The AI produced its submitted basket in 3 minutes 43 seconds, compared with 25 minutes 23 seconds for the Human participant.
Its submitted basket was also 36p cheaper.
However, the experiment does not provide enough evidence to recommend relying unquestioningly on an AI-generated supermarket basket.
The AI submission had limitations around contemporaneous retailer evidence, local availability and time-sensitive pricing.
It also did not satisfy the intended independence condition fully in this investigation.
For an ordinary shopper, the most practical use of AI suggested by this test is therefore:
Give AI a precise shopping list and requirements.
Ask it to identify a low-cost basket quickly.
Use that result as a shortlist or starting point.
Verify the important prices, promotions and availability with the retailer before purchasing.
This approach uses the strongest capability demonstrated in the experiment — rapid research — while retaining Human verification where current retailer information matters.
The evidence does not show that AI selected meaningfully better food.
Blind Judge 01 scored the Human and AI submissions identically for Accuracy, Creativity, Practicality and Trust.
The main difference was how quickly the result was produced.
For this task, AI therefore appears most useful as a fast research assistant rather than an unquestioned final authority.
AI-assisted shopping research with Human verification.
Moderate
The recommendation is supported by a substantial measured speed advantage and a small observed price advantage, but confidence is limited by the evidence-preservation and independence issues documented in this investigation.
A properly controlled repeat investigation would be required before making a stronger recommendation.
Full recommendation preserved from the investigation record
Final AVHM Interpretation
The observed result tells us what happened. The evidence determines how confidently we can interpret it.
Final Verdict
The observed result favoured AI. AVHM does not treat that observed advantage as stronger than the evidence allows.
Evidence Confidence
Qualified
Independence
Material limitation
If AVHM considered only the headline numbers, the result would be simple:
AI won.
It submitted the cheaper basket.
It completed the task much faster.
It received the higher blind judging score.
But AVHM does not consider only the headline numbers.
The AI knew information about the Human result before beginning its timed attempt.
That weakens the causal claim that AI independently outperformed the Human participant.
The blind judging also revealed something important: outside Speed, the judge saw no scoring difference between the two submissions.
The strongest finding from AVHM-0012 is therefore not that AI was dramatically better at supermarket shopping.
It is that AI was able to produce a reviewed, competitive basket far faster, while achieving a submitted total marginally below the Human participant.
Whether AI would retain that price advantage under genuinely independent conditions remains unanswered.
Full interpretation preserved from the investigation record
The Method Behind The Result
AVHM publishes the protocol alongside the result so the question, rules and limitations can be examined rather than simply trusted.
Why This Investigation Matters
Grocery prices affect almost everyone, and shoppers increasingly use AI tools to help compare products, plan purchases and reduce costs.
This investigation tests whether an AI can use publicly available supermarket information to assemble a cheaper valid basket than a human working from the same shopping brief and the same retailer constraints.
The question matters because a lower headline price is only useful if the products are actually comparable, available, correctly priced and suitable for the brief. The investigation therefore tests not just price finding, but whether AI can produce a result that is accurate, practical and verifiable.
It also provides a useful real-world test of how well AI handles changing online information such as supermarket prices, offers, product sizes and availability.
The aim is not to prove that AI or humans are universally better at shopping. It is to produce evidence about how each performs on one clearly defined supermarket-basket task under the same conditions.
Compiled directly from the investigation record
Objective
Determine whether AI or a human can identify the lowest-cost valid supermarket basket from publicly available online supermarket information when both are given the same fixed shopping list, comparison rules and research conditions.
The investigation will compare the final valid basket price achieved by each participant while also recording the accuracy, completeness and verifiability of the selections.
A cheaper basket will only count as better if every selected product satisfies the locked requirements of the investigation.
Scope
This investigation compares one AI participant and one human participant attempting to find the cheapest valid supermarket basket using the same locked shopping list and comparison rules.
Included:
Products available through the supermarkets and online sources permitted by the locked brief.
Publicly displayed prices and qualifying promotions available during the defined research period.
Product size, quantity and specification checks needed to determine whether an item satisfies the brief.
The final total price of each valid basket.
The research process used by both participants, including sources consulted and relevant decisions made.
Evidence required to verify the selected products and prices.
Excluded:
Loyalty-only prices or personalised discounts unless the locked brief explicitly permits them.
Delivery charges, travel costs and other costs outside the product basket unless specifically included in the brief.
Products that do not meet the locked quantity, size or specification requirements.
Substituting an unavailable product after the research period has ended merely to improve a participant's result.
Claims about which participant would perform better across all supermarkets, shopping lists or future price conditions.
The investigation is a comparison of performance on one controlled supermarket-basket task, not a general test of whether AI or humans are universally better shoppers.
Investigation Protocol
The success criteria, competition rules and judging method are recorded in advance so the test cannot quietly move the goalposts after either side performs.
Success Criteria
A participant's final basket will be considered valid only if every required item satisfies the locked shopping brief and can be supported by verifiable evidence from an approved source.
The investigation will assess success using the following criteria:
Basket validity: Every required item must be included and meet the locked product, quantity and size requirements.
Total cost: The final valid basket price will be calculated using the recorded prices available during the investigation.
Accuracy: Product names, sizes, quantities, prices and qualifying offers must be recorded correctly.
Completeness: No required item may be omitted from the final basket.
Verifiability: Each final selection must have sufficient evidence for its price and eligibility to be independently checked.
Speed: The time taken by each participant to produce the final basket will be recorded.
Practicality: The final basket must represent a purchase that could reasonably have been made under the locked conditions.
The cheaper valid basket will win the price comparison.
A lower total will not automatically constitute a better result if the basket contains an invalid item, an unsupported price, a missing product or otherwise fails the locked brief.
If the evidence does not establish a fair and reliable winner, the investigation may result in a draw or an inconclusive finding.
Competition Rules
The following rules must be locked before either participant begins researching the basket.
Same brief: The AI and human must receive the same locked shopping list, product requirements, retailer rules and comparison conditions.
Same research window: Both participants must complete their research within the defined investigation period so that price changes do not knowingly favour either side.
Same permitted retailers: Both participants may search only the supermarkets or online sources specified in the locked brief.
Same product requirements: Products must satisfy the same minimum quantity, size, type and other requirements. A cheaper product that fails the brief is not a valid selection.
Publicly available information: Participants may use publicly accessible supermarket information permitted by the brief. Account-specific, personalised or otherwise unavailable prices must not be used unless explicitly allowed before testing begins.
Offers and discounts: Promotions may be used only when they are available under the locked conditions and the evidence clearly shows that the basket qualifies for them.
Availability: A selected product must be shown as available through the permitted source when it is recorded. If availability cannot reasonably be verified, that limitation must be documented.
Evidence capture: Each final product selection must record enough information to verify the item, price, quantity or size, retailer and source at the time of research.
No retrospective optimisation: Once a participant's final basket has been submitted, products may not be replaced simply because a cheaper option is discovered later.
Independent attempts: Neither participant may see the other participant's research, selections or final basket before both submissions are locked.
Interventions recorded: Any clarification, technical problem, assistance or change affecting either attempt must be recorded in the evidence log.
Invalid selections: Any item that fails the locked brief must be identified during judging and cannot contribute to a claimed cheaper valid basket without the discrepancy being disclosed.
Same scoring method: Both submissions will be assessed using the judging criteria locked before the competition begins.
Evidence over outcome: Rules, evidence or scores must not be changed to manufacture a Human Win, AI Win or Draw.
Judging Criteria
Both sides faced the same five categories and the same scoring scale.
Both participants will be scored using the same five AI vs Human Media judging categories. Each category will be scored out of 10, producing a maximum total score of 50 for each participant.
Measures whether the final basket correctly satisfies the locked shopping brief.
Judging will consider:
Whether every required product is included.
Whether selected products meet the required type, size and quantity.
Whether recorded prices and offers are correct.
Whether basket totals have been calculated correctly.
Whether invalid or non-comparable products have been included.
Measures the time taken to produce a completed final submission.
Timing will begin when the participant receives the locked brief and end when the final basket is submitted.
Any pauses, interruptions or technical problems that materially affect the comparison must be recorded.
Measures the participant's ability to identify legitimate approaches that reduce the basket cost without breaking the locked rules.
This may include:
Finding less obvious valid alternatives.
Making effective use of permitted promotions.
Identifying equivalent products that satisfy the brief at a lower cost.
Using an efficient research strategy.
Creativity does not reward loopholes, invalid substitutions or rule-breaking.
Measures whether the proposed basket represents a realistic and usable supermarket purchase.
Judging will consider:
Product availability.
Whether quantities and pack sizes make practical sense.
Whether the basket could reasonably be purchased under the recorded conditions.
Whether savings depend on unrealistic assumptions.
Whether the final result would be useful to an ordinary shopper following the same brief.
Measures how well the participant's result is supported by clear, traceable and verifiable evidence.
Judging will consider:
Quality of source evidence.
Accuracy of product and price records.
Transparency about uncertainty or missing information.
Whether important assumptions have been disclosed.
Whether another person could reasonably verify the final basket from the recorded evidence.
The five category scores will be added to produce a total out of 50 for each participant.
The higher overall score will determine the investigation winner.
The raw basket price will also be reported separately so the audience can see which participant found the cheaper valid basket.
If the total scores are equal, the result will be recorded as a Draw.
If evidence is insufficient to support a reliable comparison, the investigation may be declared inconclusive rather than forcing a winner.
Limitations
This investigation is designed to produce a fair comparison under controlled conditions, but several limitations must be recognised when interpreting the result.
Prices can change: Supermarket prices and promotions may change during or shortly after the investigation. The recorded result represents the prices that could be verified during the defined research period.
Availability can change: A product shown online may become unavailable, go out of stock or differ by location. Availability will be recorded where reasonably possible, but cannot be guaranteed beyond the investigation period.
Online and in-store prices may differ: The investigation relies on the sources permitted by the locked brief. A different price may exist in a physical store or through another sales channel.
Location may affect results: Product ranges, prices and availability can vary between stores and geographic areas. The result therefore applies only to the conditions defined for this investigation.
Promotions may have conditions: Multi-buy offers, loyalty schemes, minimum spends and other promotions can affect the apparent price. Only promotions permitted by the locked rules will be included.
Product comparison involves judgement: Different brands, pack sizes and product descriptions may not always be perfectly equivalent. Eligibility decisions will follow the locked brief and any uncertainty will be documented.
Search coverage cannot be guaranteed: Neither participant can prove that every possible valid product or offer was discovered. The investigation compares the results they actually achieved within the agreed conditions.
AI access may differ from human access: The AI participant may have different capabilities or restrictions when accessing current supermarket websites and online information. Any material access limitation or intervention will be recorded.
Timing can influence the comparison: Website performance, technical problems, interruptions and research methods may affect completion time. Material issues will be documented in the evidence log.
One investigation cannot establish universal superiority: The result applies to this specific basket, research period, participants and testing conditions. It does not demonstrate that AI or humans will always find cheaper supermarket baskets.
Any limitation discovered during the investigation that could materially affect the result will be added to the evidence record and disclosed when the findings are published.
Timeline
The investigation will proceed through the following stages.
Define and record the objective, scope, success criteria, competition rules, judging criteria, limitations and investigation method.
Create the final shopping list and lock the permitted retailers, product requirements, research conditions, promotion rules and evidence requirements before either participant begins.
Provide the human participant with the locked brief.
Record the start time, research process, relevant decisions, evidence sources and final basket. Lock the human submission before the AI attempt is revealed or compared.
Provide the AI participant with the same locked brief.
Record the AI system used, start time, prompts or instructions, research process, interventions, evidence sources and final basket. Lock the AI submission before comparison.
Verify the recorded products, prices, quantities, offers and sources for both submissions.
Identify any invalid selections, missing evidence, discrepancies or limitations before scoring begins.
Where practical, prepare anonymised versions of the two valid submissions so that judging does not unnecessarily reveal which participant produced each basket.
Score both submissions using the judging criteria locked before testing.
Calculate the final scores and compare the valid basket totals.
Record the Human Win, AI Win, Draw or inconclusive result supported by the evidence.
Publish the investigation record, relevant evidence, scores, limitations and final conclusion.
Any material correction discovered after publication will be recorded rather than silently altering the historical record.
Core Rule
Trust the evidence. Not us.
Investigation Record
The investigation result, judging, evidence, interpretation, limitations and underlying protocol are now part of the public record.
Evidence Principle
Test Everything. Hype Nothing.
AVHM records the result and the weaknesses in the experiment rather than hiding inconvenient evidence.