How Flavor Scoring Works
Every pairing score on this site is calculated from real molecular data. Here's what the numbers mean and how we get them.
What this engine does: it ranks ingredient pairings by shared volatile chemistry, weighted toward the ~226 key food odorants identified by Dunkel et al. (2014) as the perceptually-active subset of food volatiles, and cross-checked against 610K+ curated recipes.
What it does not do: predict perceived aroma intensity in a specific food matrix. Matrix effects, mixture interactions at the olfactory-receptor level, and concentration-ratio-driven percepts are well-documented limits of compound-level analysis. We surface candidate pairings; tasting still belongs in the kitchen.
The Short Version
When you search for "garlic," we look at garlic's molecular fingerprint — the volatile compounds that create its smell and taste — and compare it to every other ingredient in our database. Ingredients that share more of these compounds score higher.
But chemistry isn't everything. We also check 610K+ curated recipes (OpenRecipes + public-domain cookbooks from Project Gutenberg and the Internet Archive) to see which ingredients chefs actually use together. The final score blends both signals: 50% molecular overlap + 50% recipe co-occurrence, with bonuses for matching aroma profiles.
Compound Score (50%)
Every ingredient has a set of volatile flavor compounds identified through GC-MS analysis (gas chromatography-mass spectrometry). We calculate a Jaccard similarity between two ingredients' compound sets:
Compound Score = IDF(Shared Compounds) / IDF(Total Unique Compounds)
Not all compounds are weighted equally. We use IDF (Inverse Document Frequency) weighting: rare compounds found in few ingredients count more than ubiquitous ones. Compounds you can actually smell — those with low odor detection thresholds — get an extra boost. A shared compound detectable at 0.01 parts per billion matters more than one you can't smell until 100,000 ppb.
Three further refinements sharpen the per-pair weight where the data supports it:
- Concentration. When a shared compound has measured concentrations in both ingredients, its weight is scaled by how abundant it is — a compound present only in trace amounts counts for less than one present in quantity. This is magnitude only; we do not model concentration ratios or matrix effects (see Known Limitations).
- Chirality. Many odorants smell different in their mirror-image forms (R- vs S-carvone is spearmint vs caraway). Where enantiomer data exists, a compound shared as different enantiomers between the two ingredients is down-weighted rather than counted as a clean match.
- Cooking state. For proteins you can select a preparation (raw, roasted, grilled, …); the engine then scores against that state's volatile profile, since cooking transforms which compounds are present.
Recipe Score (50%)
We analyze 610K+ curated recipes (OpenRecipes CC-BY, Fandom Recipes Wiki CC-BY-SA, public-domain cookbook scans from Project Gutenberg and the Internet Archive) to measure how often two ingredients appear together. We use NPMI (Normalized Pointwise Mutual Information), a statistical measure that captures whether two ingredients co-occur more than you'd expect by chance.
NPMI is better than simple co-occurrence counting because it accounts for ingredient popularity. Salt appears in almost every recipe, so "salt + chicken" isn't surprising. But "cardamom + coffee" appearing together is statistically significant — chefs are making a deliberate flavor choice.
Bonus Signals
These are an enrichment layer on top of the measured base above (compound chemistry + recipe co-occurrence). Two of them — the aroma-profile match and the chef-canonical bonus — draw on labels we generate in-house with a large language model and cross-validate against curated reference databases (94% agreement with ChemTastesDB for aroma/chemical-class labels; held-out validation for chef pairings). We name them explicitly so you can weigh measured chemistry against derived enrichment rather than take a single opaque number.
- Synergy Bonus (+20%)
- When both compound AND recipe signals are strong, they reinforce each other. This square-root bonus rewards pairings where the science and the culinary tradition agree.
- Aroma Profile Match (+10%)
- Each ingredient gets a 14-dimensional aroma profile (citrus, floral, smoky, etc.). When two ingredients have similar aroma shapes, they get a cosine-similarity bonus.
- Chef-style Pairing Bonus (+8%)
- Pairings recognized as chef-canonical (874 ingredients with curated counterpart lists, generated in-house with a large language model and cross-validated against a held-out set of established chef pairings) get an additional boost.
The "+20%" / "+10%" / "+8%" figures are the maximum each bonus can contribute. Bonuses are clamped together with the base 50/50 blend so the final pairing score always stays in [0, 1]; high-scoring matches saturate near 1.0 rather than exceeding it.
OAV markers in compound results
Pro users see compound-level OAV (Odor Activity Value — concentration ÷ detection threshold) on every shared compound. Each value is tagged with a confidence marker so you can see the data provenance at a glance:
-
Matrix-validated
Cross-referenced against the 1,251-pair OAV reference database with matrix matching, uncertainty bands, and paper provenance. Citable in publication.
-
Paper-cited
Computed from primary literature: concentration data extracted from peer-reviewed papers, divided by published threshold values. Each entry has a paper-level citation.
-
Significant aroma contributor (OAV > 10)
Concentration is at least 10× the detection threshold — the compound is reliably perceived in the food. Combine with ✓ or 📄 to read confidence: a ⚡✓ compound is a high-impact aroma you can cite.
-
SIDA-quantified threshold
At least one of the threshold sources used in the median was determined by Stable Isotope Dilution Analysis (SIDA) — a chromatographically-coupled MS method that uses a deuterated or 13C-labeled twin of the target compound as the internal standard. SIDA is the gold standard for absolute quantitation of trace odorants, with typical precision < 10%. Non-isotope internal-standard methods (the majority of the GC-MS literature) can be off by 2–10× due to matrix-dependent recovery differences. Thresholds in our database without the SIDA badge are still useful for ranking, but should be treated as approximate, not as absolute citable values.
Compounds without a confidence marker have their OAV value suppressed in the UI to avoid surfacing values without provenance. The other chemistry — threshold, aroma class, CAS — is always shown.
Why this matters: a typical OAV-weighted scoring system mixes SIDA-quantified values with non-isotope values without distinguishing them. We surface the SIDA · n/total badge on compound detail pages so you can see, per compound, how much of the published threshold data is from the gold-standard method. Current coverage: 230 compounds with at least one SIDA-tagged source. Backbone is the Rychlik/Schieberle/Grosch 1998 compilation (491 rows, 190 compounds); subsequent additions came from Dunkel et al. 2014 main-text thresholds, plus a hybrid identification pipeline over our 40K-paper corpus: BigQuery PMC OA full-text scan, abstract regex matching, Anthropic Claude refinement of borderline cases, author-based heuristic for the canonical SIDA practitioners (Schieberle, Grosch, Hofmann, Steinhaus, Dunkel, Granvogl, Buettner, Blank, et al.), and PMID-level source tracking on every threshold row so a SIDA paper's contribution flows through to every value it informed.
Key food odorants
Of the ~10,000 volatiles documented in food, fewer than 3% actually drive perceived aroma. The remainder are below detection threshold or masked by mixture effects. Dunkel et al. (2014) identified 226 key food odorants (KFOs) across 227 food samples in a meta-analysis of the SIDA-quantified literature, and showed that each food's perceived smell is encoded by just 3–40 of those KFOs in specific concentration ratios.
We flag every compound in our database that matches a KFO and surface them on the ingredient and compound detail pages with one of three labels:
- Generalist
- Appears in >25% of foods (16 compounds). Generic Maillard, lipid-oxidation, and fermentation products — methional, diacetyl, vanillin, hexanal, butyric acid, etc.
- Intermediary
- Appears in 5–25% of foods (57 compounds). Linalool, ethyl hexanoate, 2-phenylethanol, etc.
- Individualist
- Appears in <5% of foods (151 compounds). Carry diagnostic signatures: 1-p-menthene-8-thiol (grapefruit), wine lactone, diallyl disulfide (garlic).
Rare KFOs (individualists) carry more diagnostic information than common ones — the same intuition behind our IDF weighting. We have mapped 203 of the 226 KFOs from the paper's Supporting Information Table S1 (the remainder are entries without a unique CAS or with stereochemistry that the SI lists under one CAS).
Novel Pairings
A pairing is flagged as novel when it has high compound overlap but low recipe co-occurrence. These are scientifically compatible ingredients that chefs haven't widely explored yet — the frontier of flavor.
Compound Coverage Distribution
Coverage depth varies by ingredient. Well-studied foods like coffee, chocolate, and tomato have hundreds of documented compounds from decades of GC-MS research. Less-studied ingredients may have fewer data points, which we account for in our scoring confidence.
Distribution across 857 profiled ingredients. Median coverage: ~150 compounds per ingredient.
Scoring Confidence
A pairing can score 75% off rich, well-measured chemistry — or off a handful of compounds with estimated thresholds. Those aren't equally trustworthy, so every result carries a confidence band alongside the score. It is composed from data quality, not a second model, and it flags how much to trust a score without changing the ranking:
- Coverage — how richly the ingredient is characterized (the distribution above) and how many compounds the pair shares. A Jaccard/NPMI score over a thin compound set is statistically noisy.
- Measurement quality — what fraction of the shared compounds carry a measured odor threshold, and how many are SIDA / matrix-validated rather than estimated.
Results built on a thin profile or very few shared compounds are badged low confidence (the chip's tooltip names the reason); a well-covered pair backed by measured thresholds carries no badge. It's the honest counterpart to the limitations below — shared molecules suggest compatibility, and the band tells you how firmly the data backs each suggestion.
Data Sources & Pipeline
Compound profiles are aggregated from 10+ peer-reviewed food chemistry databases, deduplicated by CAS number, and enriched with odor thresholds and aroma classifications. The pipeline runs continuously — new data sources are integrated as they become available.
| Source | Coverage | Contribution |
|---|---|---|
| FlavorDB 2.0 | 857 ingredients | Base compound profiles, ingredient categories |
| FooDB | 28,000+ compounds | Deep chemical profiles, concentrations, CAS numbers |
| ChemTastesDB | 2,944 compounds | Curated taste classifications (sweet, bitter, umami, etc.) |
| PubChem | Molecular enrichment | Chemical properties, molecular weights, CAS validation |
| Flavornet | 738 compounds | GC-olfactometry aroma data |
| Rychlik M, Schieberle P, Grosch W (1998) | 224 key odorants | Licensed compilation of odor thresholds, qualities, and retention indices (Institut für Lebensmittelchemie der Technischen Universität München and Deutsche Forschungsanstalt für Lebensmittelchemie, Garching) — provides orthonasal-water detection thresholds used in OAV calculations |
| Dunkel et al. (2014) | 226 key food odorants | Defines the perceptually-active subset of ~10,000 food volatiles. Drives the “Key food odorant” classification on compound and ingredient pages. |
| Ahn et al. Flavor Network | 1,530 ingredients | Seminal flavor network dataset (Scientific Reports, 2011), CC BY 4.0 |
| USDA Dr. Duke's | Phytochemicals | Plant phytochemistry — Agricultural Research Service, public domain |
| Wikidata | 91% of ingredient pages | Taxonomy, common names, regional varieties — CC0 |
| LLM-derived chef pairings | 874 ingredients | Chef-style counterpart lists (Claude Sonnet, +8% bonus, cross-validated against a held-out chef set) |
| LLM aroma classification | 5,240 compounds | Aroma family + chemical class labels via Sonnet, validated 94% against ChemTastesDB |
| PubMed + Semantic Scholar | 44,648 papers | GC-MS concentrations, OAV, retention indices |
| Curated Recipe Corpus | 610K+ recipes | Ingredient co-occurrence (NPMI scoring). OpenRecipes + public-domain cookbooks. |
Data pipeline runs continuously. Compounds are deduplicated by CAS number and cross-validated across sources.
Our methodology papers (preprints)
“Unsupervised Recovery of Food Taxonomy from Volatile Compound Profiles: A UMAP Analysis of 450 Ingredients”
We show that ingredient identity can be recovered from volatile compound profiles alone, using UMAP dimensionality reduction. The resulting clusters reproduce conventional culinary categories without supervision — evidence that the compound graph encodes real flavor structure, not artifacts of how data was collected.
Read the paper on Zenodo (DOI: 10.5281/zenodo.19719459)“Integrating Licensed Odor Threshold Data into a Flavor Pairing Engine: Pair-Level Score Shifts and Unit Errors Caught”
Wiring licensed odor-threshold data (Rychlik / Schieberle / Grosch) into the perceptual-weighting scorer — quantifying how much it shifts per-pair scores, and the unit-conversion errors it surfaced in the published threshold literature along the way.
Read the paper on Zenodo (DOI: 10.5281/zenodo.20322645)Johnson, M. C. (2026). Unsupervised Recovery of Food Taxonomy from Volatile Compound Profiles: A UMAP Analysis of 450 Ingredients. Zenodo. https://doi.org/10.5281/zenodo.19719459
@misc{johnson2026compkitchen,
author = {Johnson, Mark C.},
title = {Unsupervised Recovery of Food Taxonomy from Volatile Compound Profiles: A {UMAP} Analysis of 450 Ingredients},
year = {2026},
publisher = {Zenodo},
doi = {10.5281/zenodo.19719459},
url = {https://doi.org/10.5281/zenodo.19719459}
}
TY - DATA AU - Johnson, Mark C. PY - 2026 TI - Unsupervised Recovery of Food Taxonomy from Volatile Compound Profiles: A UMAP Analysis of 450 Ingredients PB - Zenodo DO - 10.5281/zenodo.19719459 UR - https://doi.org/10.5281/zenodo.19719459 ER -
Methodology · Zenodo · CC BY 4.0
Data Sources & Attribution
CompKitchen is built on data from the following sources:
- Ahn et al. (2011) — Flavor network and the principles of food pairing. Scientific Reports 1, 196. Data available at Zenodo under CC BY 4.0.
- ChemTastesDB — A curated database of molecular tastants (Rojas et al., 2022). Available at Zenodo under CC BY 4.0.
- FlavorDB 2.0 — IIIT-Delhi Computational Systems Lab. cosylab.iiitd.edu.in/flavordb2.
- FooDB — The Food Database. foodb.ca.
- Dr. Duke's Phytochemical Database — USDA Agricultural Research Service. phytochem.nal.usda.gov. Public domain.
- Flavornet — Cornell University. flavornet.org.
- Rychlik M, Schieberle P, Grosch W (1998) Compilation of Odor Thresholds, Odor Qualities and Retention Indices of Key Food Odorants. Institut für Lebensmittelchemie der Technischen Universität München and Deutsche Forschungsanstalt für Lebensmittelchemie, Garching. Licensed directly from the author for use in Compound Kitchen.
- Dunkel, A., Steinhaus, M., Kotthoff, M., Nowak, B., Krautwurst, D., Schieberle, P. & Hofmann, T. (2014) — Nature's Chemical Signatures in Human Olfaction: A Foodborne Perspective for Future Biotechnology. Angew. Chem. Int. Ed. 53, 7124–7143. DOI: 10.1002/anie.201309508. Identifies the 226 key food odorants that define the perceptually-active subset of food volatiles. Used here to flag KFOs across our compound database and to frame the engine's scope.
- PubChem — National Library of Medicine. pubchem.ncbi.nlm.nih.gov. Public domain.
- Wikidata — Wikimedia Foundation. wikidata.org. CC0.
- LLM-derived layers (Claude Sonnet) — aroma family + chemical class labels (5,240 compounds, 94% agreement with ChemTastesDB), chef-style pairing bonuses (874 ingredients, cross-validated against a held-out chef set), and ingredient cooking-state inferences. Generated by us; output cross-validated against curated databases where possible.
- OpenRecipes (fictive-kin) — CC BY 3.0. 168K modern recipes with schema.org provenance.
- Fandom Recipes Wiki — CC BY-SA 4.0. ~29K community-edited recipes.
- Project Gutenberg + Internet Archive cookbook scans — public domain (expired copyright). ~27K historical recipes with parsed ingredient blocks.
Known Limitations
- Matrix effects are not modeled. A compound's perceived intensity depends on the food it sits in — fat content, pH, salt, competing volatiles. Threshold values shown on the site are typically in clean water; in a real food matrix they can shift by orders of magnitude (Dunkel et al. 2014).
- Mixture interactions are not modeled. Above ~4 components, odorant mixtures form new percepts that aren't a simple sum of parts. Two foods sharing many compounds can smell very different if the concentration ratios differ.
- Quantitative data is heterogeneous. The gold standard for absolute odorant quantitation is stable-isotope dilution analysis (SIDA). Most of the GC-MS literature in our corpus uses internal-standard methods that can be off by 2–10×. Each compound's odor threshold now carries a SIDA · n/total badge showing how many of the underlying sources used SIDA — see the markers section above.
- Texture, temperature, and presentation are not modeled. Flavor pairing is only one dimension of good cooking.
- Compound data varies by source. 59% of ingredients have 100+ compounds; 8% have fewer than 20. Ingredients with sparse data may score differently than expected.
- Cooking transforms compounds. Raw garlic and roasted garlic have different profiles. Our data represents typical preparations but can't capture every cooking method.
- Recipe data reflects Western bias. Our curated corpus is predominantly English-language and US-centric (OpenRecipes, Fandom, Project Gutenberg cookbooks), which skews co-occurrence toward Western cuisines. We partially compensate with cuisine-aware scoring modes.
- Nutrition estimates are rough. Based on typical serving sizes and a ~300-ingredient lookup table, not precise recipe analysis.
- Dietary tags are auto-generated from ingredient keywords and may contain errors. Always verify ingredients if you have allergies.