Whiskey Knowledge Databases › Literature Notes
Complete analytical Literature Note produced from a full-source review. Claims below are bounded by the recorded evidence and limitations.
Scope and core summary
Complete synthesis of a 14-page analysis of 136 real-world tastings, each presenting seven whiskies.
Author argument
The authors find that the last whisky in a sequence received the highest ratings and favorite votes more often, even after accounting for the tendency to place higher-ABV whiskies later.
Researcher synthesis
The blend program should not evaluate many candidates in one fixed order. Use smaller flights, randomized or counterbalanced order, coded samples, palate recovery, and independently prepared repeats so a finalist does not advance because it was last.
Evidence assessment
Primary field-data study with direct practical relevance and controls for a plausible ABV confound.
Limitations and open questions
Limitations: Archival tastings were not fully randomized experiments; social setting, product selection, order practices, and participant experience may contribute. The study demonstrates bias risk, not a guaranteed effect in every session.
Open questions: What flight size and counterbalancing scheme best fits a single-person repeated blend protocol?
Connected records
- Contributors: 5
- Verified Excerpts: 3
- Citations: 1
- Zettels: 2
Completion record
Full-source pass completed: 14/14 local PDF sheets and 8,772 extracted words reviewed, including model specification, robustness checks, discussion, and references. Evidence locators retained at local sheets 1, 8–12. No completion hold remains.
September 27 Full-Source Audit
Scope and Identity
Quigley-McBride, Franco, McLaren, Mantonakis and Garry (2018), PLOS ONE 13(8), e0202732, DOI 10.1371/journal.pone.0202732. All 14 supplied PDF pages were read, including the entire introduction, methods, results, discussion, credits and 25 references. Both tables and both figures on pp. 9–10 were visually inspected. The two-page Regional Wines protocol recovered from the linked OSF archive was also read in full. This is a complete article-and-protocol review; the OSF spreadsheets and whiteboard archive remain separate, unreviewed supporting material. No model reproduction is claimed.
Research Question and Design
Does the familiar laboratory serial-position pattern persist in ordinary commercial whisky tastings? The authors retrospectively coded 136 seven-whisky sessions held at one Wellington retailer from 2002 to 2013, all hosted by the same person. The 952 observations are whisky appearances across sessions, not 952 distinct whiskies or individual participants. There were 216 distinct whisky types, and the average attendance was 29.73 people per session (SD 15.11). Participants may recur across sessions; multiplying attendance by sessions would not establish a count of unique individuals.
The dataset was 98.7% Scotch. Average whisky age was 16.07 years and average strength 50.44% ABV (range 40–66.8%). This gives useful field relevance but little direct evidence for American whiskey flights.
What Blinding Did and Did Not Mean
Attendees knew the lineup names and could see prices, except for one mystery whisky, but did not know each sample's position. The host was usually blind; repeated versions of a tasting sometimes kept the same sequence. A shop manager selected the order and tended to put stronger whiskies later. There was no random assignment of sequence.
The tasting included group nosing, group tasting with publicly voiced descriptors and votes, then unrestricted revisiting and final integer ratings. Beer and water were used between samples. Scores below five required handing the whisky to someone who passed it. This unusual social consequence can discourage low scores and makes the rating scale different from a neutral private liking scale. Retain it as a documented feature, not an Academy recommendation.
One author coded all records; a second coded 20%, with reported high agreement. Missing bottle information was sometimes recovered from tasting sheets and internet sources. Enthusiast and first-timer classifications were too incomplete and inconsistent for the intended comparison, so the study does not demonstrate an expert–novice difference.
Findings with Locators
Main pp. 8–10: multilevel models accounted for whisky identity and, for overall ratings, tasting identity, together with age, ABV, serial position and attendance. Later positions were positively associated with both favorite counts and final mean ratings after adjustment.
Table 1 reports a serial-position coefficient of 0.635, SE .0941, p<.001 for favorite counts. Table 2 reports .0997, SE .0132, p<.01 for mean ratings. These are the authors' reported coefficients. The table notes call them standardized; without the analysis code and scaling definitions, do not translate them into a guaranteed number of votes or rating points gained by moving a bottle one place.
Main pp. 10–11 report mean ratings of 8.59 for the last position versus 8.16 for the middle position, a descriptive difference of .43 on the stated scale. Figure 2's vertical axis starts at five; inspect the numeric difference rather than judging effect size from bar heights. Both graphs are descriptive position averages with 95% confidence intervals, not plots of adjusted causal effects.
The full overall-rating model has conditional R²=.5294 and marginal R²=.3211. Removing position reduced explained variance by roughly four percentage points; removing ABV reduced it by about 18 points. This does not mean order causes four percent of every person's preference, or that the 53% belongs to order alone. Favorite-count models report conditional R²=.4507 and marginal R²=.3281.
Statistical and Causal Limits
The favorite-count outcome had 22.1% missing data and a nonnormal distribution. The authors acknowledge assumption concerns. Mixed-model software does not itself resolve nonrandom missingness, a bounded/count outcome, or the dependence between favorite counts competing within the same session. A count/proportion model with an appropriate attendance denominator and session structure would be a useful robustness analysis, not an analysis completed here.
There is a reporting inconsistency: p. 8 says the favorite-count model omitted tasting as a higher-level variable, while Table 1's heading describes both tasting and whisky levels. Preserve this uncertainty until the archived analysis can resolve it.
Adjustment for ABV and age is valuable but does not randomize bottle quality, smoke intensity, style, expected finale, carryover, repeated attendance, or host behavior. The authors suggest unmeasured influences often add random noise; in a deliberately sequenced commercial event some may systematically track position. A positive position coefficient is therefore evidence of an association and a design concern, not a clean estimate of recency caused solely by order.
The discussion's argument against intoxication is plausible but not conclusive: alcohol exposure and timing were not independently measured or randomized, and the “first sip” outcome still follows a sequence. The study also cannot isolate memory recency from sensory adaptation, contrast or social influence.
All sessions use seven whiskies. The introduction's theories about shorter flights, expertise, pairwise comparison and choice inertia are background hypotheses from cited research, not factors experimentally tested here. A shorter flight may exchange recency for primacy; it is not a complete correction by itself.
Contribution and Proposed Academy Uses
This is a strong practical reminder to treat tasting order as part of the measurement. Its distinctive contribution is an archival field pattern across many sessions rather than a tightly controlled lab demonstration.
For Academy comparison exercises, collect private ratings before group discussion, use coded samples, vary or counterbalance order across people/sessions, and repeat finalists in another position. Record ABV, water additions, serving temperature, interval and sample identities. These are proposed design improvements; this paper did not experimentally validate every element or determine an ideal flight length.
For guided educational tastings, a deliberate narrative sequence can be engaging. Describe it as the teaching sequence, and avoid presenting an end-of-flight favorite vote as an independent ranking of inherent bottle quality. Do not promise a marketing “last position wins” formula.
For research literacy, contrast raw bar graphs, adjusted observational models and randomized trials. A useful classroom task is to identify how the below-five penalty and public voting might alter scores, then redesign that part of the protocol.
Connections
The dilution reviews show that preparation changes the stimulus; this study shows that sequence and social procedure may also change the response. Together they support recording both liquid preparation and presentation conditions.
Whisky dilution changes both headspace chemistry and discrimination — Ashmore et al.
Dilution of Whisky – The Molecular Perspective — Full-Source Literature Note
Existing order-effect idea: Tasting order can manufacture apparent preference.
Located variance evidence: Serial position alone explained a modest but consequential share of variation — p. 12.
Access and Verification
The original attachment and all existing relations are preserved. The article is open access under the displayed Creative Commons attribution license. It reports no specific funding or competing interests; the retailer host's author role and commercial setting remain relevant provenance. Publisher metadata and OSF access were checked live. The main mirror file has not been hash-matched to the Proton placeholder; browser rendering of the Notion pages remains unverified.
The OSF root contains whiteboards, a protocol, two data spreadsheets and an empty applied-value folder in the returned listing. The two-page protocol corroborates the tasting procedure. Spreadsheets and whiteboards are now explicit follow-up resources; their discovery is not full examination.