Data methodology

How we use the wine dataset

In short

Our price tiers, critic-score ranges, and common-region figures are aggregate statistics from a historical wine-review dataset. They are meant for relative comparison (which grapes and regions tend to cost more or score higher), not as current retail prices. Factual claims are grounded in reference sources like Wikipedia. We never reproduce the original review text.

What we compute

We analyze a public dataset of roughly 150,000 wine reviews. From it we calculate only aggregates: counts, price tiers and historical medians, critic-score ranges, and which regions and grapes appear together most often.

  • Price tier: a relative band (value, mid-priced, premium, ultra-premium) plus the historical median, computed only from the entries that list a price. These are not current prices (see below).
  • Critic scores: the minimum, median, and maximum on a 100-point scale, as recorded in the dataset.
  • Common regions and grapes: the most frequent, by number of wines. For a region+grape page (say Napa Cabernet), figures describe that grape in that region, not the whole region.

How we handle pricing

The dataset is a historical snapshot, and wine prices move with vintage, scarcity, tariffs, and demand, not simple inflation. So we do not present dataset prices as current, and we do not adjust them and call them today's prices. Instead we use them for durable, relative signals: a wine's price tier and how categories compare (for example, Cabernet usually sits above Malbec). For a current price, check a retailer; our numbers are for orientation, not quotes.

How we keep facts accurate

Statistics come from the dataset; everything else (a grape's parentage, a region's soils, a wine law) is written from reference sources such as Wikipedia and checked by a person. Every sentence on our guides is original writing, not copied text.

What we never do

We do not copy, quote, or paraphrase the individual review descriptions in the dataset. Those belong to their authors.

Caveats worth knowing

Any dataset reflects its own coverage. This one skews toward wines reviewed in the United States and is a historical snapshot, so we treat every figure as directional context rather than a live quote. We refresh the aggregates when the underlying data changes.