Logo
  • Home
  • Learn
  • Explore
  • Resources
  • About
Whisky vocabulary can be mined, but context still governs meaning — Miller, Hamilton, and Lahne
Whisky vocabulary can be mined, but context still governs meaning — Miller, Hamilton, and Lahne

Whisky vocabulary can be mined, but context still governs meaning — Miller, Hamilton, and Lahne

icon
Tags
Sensory VocabularySensory VocabularyDescriptive Sensory ProfilingDescriptive Sensory ProfilingSensory EvaluationSensory Evaluation
icon
Legacy Single Excerpt
icon
Legacy Associated Permanent Note
Content
Citation(s) & Footnote(s)
Book(s)

Whiskey Knowledge Databases › Literature Notes

✅

Complete analytical Literature Note produced from a full-source review. Claims below are bounded by the recorded evidence and limitations.

Scope and core summary

Complete synthesis of a 16-page study training a deep-learning classifier to distinguish flavor descriptors from ordinary language in online whisky reviews.

Author argument

The authors show that existing review corpora can support automated sensory-language extraction, while emphasizing the need for human labeling, contextual disambiguation, and future support for multiword descriptors.

Researcher synthesis

This can accelerate vocabulary discovery for profiles and retrieval, but extracted terms must be normalized, defined, and separated from brand, place, evaluation, and metaphor. A frequency-rich lexicon is not automatically a scientifically validated wheel.

Evidence assessment

Primary computational sensory study with labeled data, held-out evaluation, and explicit limitations.

Limitations and open questions

Limitations: Unigram model, narrow context window, source-corpus bias, and no guarantee that detected words are consistently used or perceptually anchored. Online reviews mix expert, enthusiast, and marketing language.

Open questions: How should multiword whiskey descriptors be normalized without erasing useful distinctions among literal, associative, and evaluative language?

Connected records

  • Contributors: 3
  • Verified Excerpts: 3
  • Citations: 1
  • Zettels: 2

Completion record

Full-source pass completed: 16/16 local PDF sheets and 7,579 extracted words reviewed, including corpus construction, model evaluation, limitations, conclusions, and references. Evidence locators retained at local sheets 1, 14, and 15. No completion hold remains.

Entire supplied source reviewed
Author argument separated from researcher synthesis
Evidence quality and limitations recorded
Page- or section-located excerpts connected
Citation connected
Zettel synthesis connected

Independent complete article review — September27,2026

All16PDFpages,28referenceentries,14figures and3tables were read. Every figure and table was inspected in rendered sheets5–13. OriginalPDF and existing Source/LiteratureNote identities remain preserved. No separate supplement is cited. The underlying review corpus is available only on author request because of privacy/copyright; it was not obtained, read or used to rerun the model. Article reading is complete; independent computational replication is not.

Contribution

Miller,Hamilton and Lahne (2021), Foods10,1633, DOI10.3390/foods10071633, demonstrate a practical workflow for proposing sensory vocabulary from existing English whisky reviews. Its strongest Academy contribution is separating candidate discovery from human validation: review language can reveal useful words without becoming a calibrated sensory measurement, a chemical assay or a universal flavor wheel.

The corpus contains8,036reviews: WhiskyAdvocate4,288, WhiskyCast2,309, WhiskeyJug1,095 andBreakingBourbon344. These are four editorial sources with unequal volumes, geographic/product emphases and author counts, not a representative sample of drinkers. Professional-versus-hobbyist style is the authors' explanation for transfer difficulty, not an experimentally isolated cause. Product mix, annotation tiebreaker and vocabulary also differ.

Data and modeling pipeline — pages3–7

Tokenization, part-of-speech filtering, lemmatization and frequency ranking precede annotation. Four sensory-science annotators contributed: A/B across datasets, C as tiebreaker forWA/WC andD forBB/WJ. Majority labels yielded1,794lemmas (499descriptive,1,295other) mapping to2,638unique token forms. A lemma's label was propagated to its instances; the tagging interface did not show occurrence context. This is a label-generation limitation even though the neural network later receives context. Majority agreement is not an independent physical standard, and no inter-rater agreement statistic is reported.

Each candidate and its three-word context on either side use300-dimensional GloVe vectors withzero-vectorpadding. Figure3 has separate256-unitLSTMs,128-unitdensebranches,concatenation and128/64/32dense layers beforebinarysoftmax. Adamlearningrate0.0001,batch32,binarycrossentropy andthreeepochs are reported. These describe a historical prototype; no installed Academy model or reproduced performance is implied.

The two evaluations must remain separate

Table1: training onWA/WC and testing onBB/WJ yielded90%accuracy butprecision0.779,recall0.422,F1=0.547. The high overall accuracy masks missing many positive descriptors. POSbaselineprecision0.209/recall0.946/F1=0.3422 demonstrates the precision–recall tradeoff.

Tables2–3: after pooling all four sites and randomly splitting token occurrences80/20, the model achieved99.910%testaccuracy andprecision0.99883,recall0.99912,F1=0.99898. Another20%oftrainingdata was reservedforvalidation. These scores are credible as reported results for that split, but do not establish performance on independent reviews, unseen words, future authors or new domains.

My methodological assessment: random token splitting permits words and overlapping contexts from the same review in both training and testing. Context-free lemma labels repeated across occurrences also allow word-identity learning. Exact token-position tracking prevents assigning one occurrence twice; it does not ensure independent reviews or distinct lemmas. A lemma-lookup baseline and word-only/context-ablation tests would help distinguish memorization from context-sensitive inference. The source reports neither. This is a risk inherent in the described design, not a claim that its unavailable code has been audited.

Loss-curve monitoring and early stopping address optimization overfitting within this split; they do not remedy review overlap or validate new-domain transfer. Dataset mixing and split unit changed together, so the score increase cannot be attributed solely to more varied writing styles. Report unseen-review,unseen-author,unseen-source andunseen-lemma tests separately before adopting the tool.

Figure and numerical audit

Figures1–3 clarify the annotation and architecture workflow. Figure1's word cloud makes common terms easy to tag but gives rare/context-dependent terms less attention. Figure10 visibly misses copper/heavy/coating and parts of multiword phrases; the displayed prose is a useful failure case, not a performance estimate. The text mentions a false-positive Redbreast probability, while that behavior is not clearly marked in the displayed orange/blue example; retain the distinction between prose report and visual demonstration.

Figures4–6 support a train/validation gap, but validation minima do not uniformly coincide with the curves crossing at the third epoch. Crossing alone is not a general stopping rule. Figures7–8 show diminishing,nonmonotonic gains, not a guarantee that every additional5%improves performance. The prose calls0.00238 at100%an improvement over0.00231 at95%; the quoted loss is actually higher. Figure9 ends atabout124.9k samples whereas the text discusses200ktraining/50ktest plusvalidation; the plotted count's basis is insufficiently reconciled. Its narrow accuracy axis emphasizes small changes.

Page10 prints recall=TP/(TP+TN), which is wrong: recall usesTP/(TP+FN). The reported metric tables need not have been computed with that typographical formula; without code/data, do not claim the results themselves were recalculated incorrectly.

Figures11–14 are t-SNE visualizations of embeddings/labels/predictions. Cluster separation is not independent proof that annotations are correct or descriptors are perceptually equivalent; model-generatedlabels naturallyreflectmodeldecisionpatterns. The discussion conflates labeledtestdata andunannotatedpredictions in places. t-SNE distances andclusterareas cannot certify semantic or sensorytruth.

The balanced56/44occurrence split differs from the499/1,295unique-lemma distribution. Do not mix type counts with token counts, or infer the prevalence of descriptors in all unfiltered review language from the selected labeled instances.

Academy applications and cross-source synthesis

Propose an internal vocabulary-discovery aid that retains the full source sentence,reviewauthor,date,product and locator with each candidate; groups potential synonyms for review; and preserves compound expressions,negation,appearance,mouthfeel,evaluation and metaphor as distinguishable facets. The paper's early definition includes appearance, while its later copper example excludes color from flavor: the Academy should explicitly choose its scope.

Preserve expressive comparison and storytelling. A metaphor can be valuable tasting communication even when it is not a literal ingredient or compound claim. Human reviewers should check brand/place names and expressions such as wetdog orredfruits before normalization erases meaning. Do not claim that two bottles with similar extracted words provide the same experience. The paper's$50/$300comparison is a proposed use, not a blinded substitution experiment; price/liking prediction is proposed, not causal evidence that certain flavors drive either.

Connect the existing vocabulary ideas Sensory panels and instruments are complementary measurement systemsSensory panels and instruments are complementary measurement systems and Shared sensory vocabulary does not establish assessor calibrationShared sensory vocabulary does not establish assessor calibration to sensory-method comparison Choosing among QDA, Napping, and GC-MS for whisky development — Daute et al.Choosing among QDA, Napping, and GC-MS for whisky development — Daute et al.: extracted language,descriptivepanels andconsumerliking answer different questions. Corn-study lexicon Literature Note — Corn Variety, Texas Terroir, and New-Make BourbonLiterature Note — Corn Variety, Texas Terroir, and New-Make Bourbon likewise requires explicit definitions/references rather than frequency alone.

Before any implementation,validate on a genuinely withheld Academy corpus using occurrence-levelcontextualannotation, reportdescriptorrecall/precision andphraseerrors, andcompare againstsimplebaselines. No scraping,externalmessaging,modeltraining orpubliccourseediting has been done as part of this review. Corpus/code availability,methodreplication,Protonbyteidentity andbrowserrendering remain separate limitations; none is represented asverified.

Date
September 5, 2026
icon
Contributors
Chreston MillerChreston MillerLeah HamiltonLeah HamiltonJacob LahneJacob Lahne
icon
Source
No access
icon
Excerpts
Miller et al. — lexicon constructionMiller et al. — lexicon constructionMiller et al. — descriptor organizationMiller et al. — descriptor organizationMiller et al. — lexicon requires calibrationMiller et al. — lexicon requires calibration
icon
Zettels
Sensory panels and instruments are complementary measurement systemsSensory panels and instruments are complementary measurement systemsShared sensory vocabulary does not establish assessor calibrationShared sensory vocabulary does not establish assessor calibration
icon
Citations
Miller et al. 2021 — whisky sensory lexicons — ChicagoMiller et al. 2021 — whisky sensory lexicons — Chicago
Logo

Policies

Home

Learn

Whiskey History

From Grain to Glass

Evaluating Whiskey

The Blending Lab

Explore

Whiskey Directory

Distillery Profiles

Brand Profiles

Regional Profiles

People of American Whiskey

Resources

About

About the Academy

Contact the Academy

© 2026 US Whiskey Academy LLC. All rights reserved.