Flavor pairing, with the chemistry attached
Find what pairs with anything on your menu — ranked by shared volatiles + recipe co-occurrence. Built for menu development, recipe testing, and the “wait, would that work?” moments.
Selected for the Georgia Tech Food & Beverage Incubator · Profiled by Hypepotamus · Every score is explained
Free · Pro $12/mo · Academic $49/yr · Enterprise API
No spam · Unsubscribe anytime
Lost an ingredient to a supply, cost or regulatory shock? Screen replacements by shared compound profile — before you commit bench time.
OAV with confidence bands, key-aroma flags, peer-reviewed citations per compound. JSON API, bulk exports, open methodology.
Pro $12/mo per scientist · Enterprise — team licenses + bulk data
No spam · Unsubscribe anytime
Every pairing comes down to molecules. These volatile compounds are what give food its smell, taste, and the surprising affinities between ingredients you’d never expect.
Educator $49/yr · Classroom $199/yr · Pro $12/mo
Or pick an ingredient and trace its pairings:
Why cookies smell like cookies
Vanillin is the single compound responsible for — you guessed it — the smell of vanilla. It’s also the dominant aroma in chocolate, bourbon, brown sugar, and almost every baked good.
Each of our 5,240 compound pages has the structure, descriptors, source foods, and the peer-reviewed papers it was identified in. Built for graduate-level food science coursework — classroom plans →
How parts per billion add up to flavor
Odor Activity Value (OAV) = a compound’s concentration divided by its odor threshold. OAV > 1 means it’s above the detection floor; OAV > 100 means it’s a major contributor.
- Concentration
- ~10,000 µg/kg
- Threshold (water)
- ~25 µg/kg
- OAV
- ≈ 400
- Concentration
- ~3,800 µg/kg
- Threshold (water)
- ~64 µg/kg
- OAV
- ≈ 60
- Concentration
- ~5,000 µg/kg
- Threshold (water)
- ~6 µg/kg
- OAV
- ≈ 830
The same compound can dominate one food and be invisible in another — vanillin shows up in chocolate at ~0.5 µg/kg (OAV ≈ 0.02, undetectable), but at ~10,000 µg/kg in vanilla (OAV ≈ 400, character-defining). Full methodology →
One molecule, one food
For most foods, a single compound is responsible for the “that smells like X” signature. Strip it out, the food becomes unrecognizable.
Why tomato and basil work together
The chemistry behind one of the oldest pairings in the kitchen. Three compounds, both ingredients, doing different jobs.
Both fresh tomato and basil emit this compound when cut. It’s the “leafy, cut-grass” smell — an instant signal of just-picked. In tomato it accounts for the fresh-from-the-vine character; in basil it’s the green base layer underneath the spicier top notes.
Basil’s dominant aroma compound — soft, floral, slightly citrusy. Present in tomato too, but in much smaller amounts. Linalool is what stops basil from smelling purely grassy and adds the “perfumed” quality the pairing leans on.
Trace amounts in both. Same molecule that gives clove its signature warmth — in basil it’s subtle, in tomato barely detectable, but it threads them together with a low spicy hum the palate registers even when the nose doesn’t name it.
44,648 papers under the hood
Every compound profile traces back to peer-reviewed GC-MS work. A few examples from the corpus:
Why R&D teams pick this over typical pairing tools
Three differences that matter when the chemistry has to be defensible.
Each entry includes OAV with explicit measured or estimated confidence, odor thresholds, chemical class, taste labels, and PubChem CID. Roughly 4× the molecular coverage of typical pairing tools.
Every compound page links to the peer-reviewed GC-MS papers it was extracted from — full citation chain, with author and institution affiliations. Critical for regulatory submissions, patent prior art, and academic publication.
Scoring formula and corpus are public. ARI 0.222 ± 0.003 against a random baseline (permutation z = 3.95). Read the methodology →
What’s actually in the corpus
The depth claim, quantified. Per-compound coverage of the fields R&D actually queries on.
Bulk JSON dumps of every field above ship with Enterprise.
This is what your code gets back
A real /api/pair response for cinnamon. Compound-level detail with confidence bands and paper counts.
GET /api/pair?ingredients=cinnamon&top_n=3
{
"results": [
{
"ingredient": "clove",
"final_score": 0.892,
"compound_score": 0.781,
"recipe_score": 0.943,
"shared_compounds": 28,
"compound_details": [
{
"name": "caryophyllene",
"max_oav": 280.0,
"oav_confidence": "estimated",
"key_aroma": true,
"papers_cited": 47,
"pubchem_cid": 5281515
},
{
"name": "eugenol",
"max_oav": 1450.0,
"oav_confidence": "measured",
"key_aroma": true,
"papers_cited": 112
}
]
}
]
}
Odor Activity Value with explicit measured vs estimated tag. No opaque “compatibility score.”
True when published GC-O studies identify this compound as character-impact in this food.
Number of peer-reviewed papers in our corpus that mention this compound in this food. Cite-ready.
PubChem Compound ID for direct cross-reference with structure databases and toxicology tools.
Make a real API call
Live request against /api/pair from your browser. No signup, no key — same endpoint your code will hit.
// Click "Run" to fetch a live response. Try changing the ingredient — basil, miso, saffron…
Anon requests return 3 unlocked results + locked stubs. Pro returns all 5 + full compound_details.
Three things R&D teams use this for
Concrete situations — with the exact endpoint and the shape of what comes back.
Identifying substitutes for a discontinued aroma chemical
A flavor house loses access to an ingredient (regulatory, supply, cost). Engine returns ingredients with the closest compound profile — ranked by molecular overlap, not popularity.
The workflow: Reformulating around a supply, cost, or regulatory shock — screen candidates by compound overlap before committing bench time.
curl "https://compkitchen.com/api/sub?ingredient=butter&top_n=5"
# IDF-weighted Jaccard over shared VOLATILES only:
# → milk (55) cream cheese (32) cheese (52)
# cheddar cheese (46) swiss cheese (41)
# Each result carries shared_compound_count + the compounds themselves.
Patent landscape for a target compound
Researching prior art for a novel formulation. Our SureChEMBL cross-reference exposes which patents cite a given compound — useful surface scan before deeper attorney review.
The workflow: Prior-art surface scan — narrow the compound list before a deeper attorney review.
GET /api/compound/vanillin
{
"compound": "vanillin",
"pubchem_cid": 1183,
"surechembl_patents": 2476, // patents citing this CID
"key_aroma_in": ["vanilla", "chocolate", ...],
"papers_cited": 47
}
OAV-driven sensory panel design
Building a descriptor panel for a product? Pull the top compounds by OAV in the target ingredient — those are the ones panelists actually perceive. Filter out the “present but below threshold” noise.
The workflow: Descriptor-panel design — rank by OAV so the panel tests what is actually perceivable.
import requests
r = requests.get("https://compkitchen.com/api/pair",
params={"ingredients": "cinnamon", "top_n": 1})
compounds = r.json()["results"][0]["compound_details"]
key_aromas = [c for c in compounds if c.get("max_oav", 0) > 1]
# → cinnamaldehyde, eugenol, linalool, caryophyllene
Two tiers, no surprises
Start with Pro for individual exploration. Move to Enterprise when your team needs API integration and bulk data.
Cancel anytime
- ✓ Full site access — 300 req/min signed in (programmatic
/api/*access is a separate developer plan) - ✓ OAV, thresholds, key-aroma flags, CAS, PubChem CIDs
- ✓ CSV exports from every tool
- ✓ Citation blocks (Plain + BibTeX)
- ✓ Per-compound institution affiliations
Scoped to team size and needs
- ✓ Everything in Pro, team-wide
- ✓ Bulk JSON dumps of all corpora
- ✓ Custom endpoints for IP & regulatory work
- ✓ SLA, dedicated support, data alerts
- ✓ Academic discount for research groups
Both tiers include the full /api/* surface — only rate limits and bulk-export access differ. See the full API docs →
Three pairings worth trying tonight
Curated combinations the engine surfaced — with the molecular reason they work.
Steep a few threads of saffron in warm cream, fold into ganache for tart or truffle. The hay-floral safranal lifts chocolate’s bitter edge.
Whisk a spoonful of white miso into melted butter, toss with pasta or roasted vegetables. Glutamate + diacetyl amplifies umami without needing a stock.
Glaze for grilled chicken, roasted carrots, or salmon. The smoke phenols and honey’s phenylacetaldehyde bridge savory and sweet on a single brush.
Five ways to use the engine
Same molecular data, five different questions. Click any tile to jump in.
Type any ingredient, get its top 20 molecular matches ranked by shared volatiles + recipe co-occurrence.
e.g. cinnamon → vanilla, clove, apple, allspice, brown sugar…
Drop 4–10 ingredients you have on hand. Engine finds the best next addition plus possible trios.
e.g. chicken + lemon + garlic + thyme → add white wine, capers, or olive…
Type a flavor mood — smoky, bright, umami, funky — get the ingredients whose compound profiles best match.
e.g. smoky → bacon, smoked paprika, mezcal, bonito, lapsang…
Find ingredients with the closest compound profile to one you’re out of. Includes a ratio hint when recipe evidence is strong.
e.g. butter → milk, cream cheese, cheese, cheddar cheese
Two ingredients that don’t obviously pair. Engine finds the third ingredient that’s a great match to both.
e.g. blue cheese + chocolate → bridge with roquefort, walnut, or fig…
Drill into every shared compound between A and B, plus the recipes using both. The chemistry of why a pairing works.
e.g. tomato × basil → (Z)-3-hexenal, linalool, eugenol …
“What do I have in the fridge?”
Drop everything you have, see what the engine wants you to add next.
-
1
White wineReinforces lemon’s linalool + garlic’s sulfur compounds — classic deglaze chemistry.
-
2
CapersBridges chicken & lemon via shared sulfur volatiles + brings briny depth.
-
3
ParsleyShares (Z)-3-hexenal with thyme — rounds the green-herbal profile.
Cooking by feel, not by ingredient
Have a mood in mind but no ingredient yet? Start with the descriptor and see what fits.
Type any descriptor on the mood tool — works with multi-word phrases like “bright + smoky.”
When you search for an ingredient, we look at its molecular fingerprint — the volatile compounds that create its smell and taste — and compare it to every other ingredient in our database. Ingredients that share more compounds score higher.
But chemistry isn’t everything. We also check 612,557 curated recipes for what cooks actually use together. The final score blends both: 50% molecular overlap + 50% recipe co-occurrence, with bonuses for matching aroma profiles and Odor Activity Value alignment.
Full methodology
Flavor Tools You Won't Find Anywhere Else
Find pairings, discover novel combinations, swap ingredients, and explore by flavor mood — all powered by molecular compound analysis.
Explore All ToolsBuilt on Peer-Reviewed Food Science
Every pairing score is backed by real molecular data from published research databases
Molecular Compound Data
Volatile compound profiles from FlavorDB, FooDB, and PubChem — 5,240 unique flavor compounds with 130,000+ compound-ingredient links sourced from GC-MS analysis.
612,557 Curated Recipes
Recipe co-occurrence data from a curated corpus of OpenRecipes (CC BY), Fandom Recipes Wiki (CC BY-SA), and public-domain cookbook scans from Project Gutenberg and the Internet Archive. Used to validate compound pairings against real culinary practice.
Published Research
Aroma and taste data from the Ahn et al. Flavor Network (Nature Scientific Reports, 2011) and ChemTastesDB, with licensed odor-threshold data from Rychlik, Schieberle & Grosch (1998, TU München). Continuously supplemented by literature mining of GC-MS studies from PubMed (44,648 papers).
Read our methodology paper (preprint)The 44,648 peer-reviewed papers in our corpus were authored at 13,863 institutions. The most frequent by affiliation count are Harvard University, CNRS, TU München, University of Copenhagen, Inserm and Jiangnan University.
Affiliation counts via OpenAlex. These institutions are cited, not affiliated — their appearance here is not an endorsement. Methodology DOI 10.5281/zenodo.19719459 · CC BY 4.0.
Data from FlavorDB (IIIT Delhi) · FooDB · OpenRecipes + PD cookbooks · 44,648 peer-reviewed papers · Ahn et al. 2011
Methodology published — “Unsupervised Recovery of Food Taxonomy from Volatile Compound Profiles: A UMAP Analysis of 450 Ingredients”
Popular Flavor Pairings
Classic and surprising combinations backed by shared molecular compounds
Cilantro
Blueberry
Pineapple
Mushroom
Cherry
Prosciutto
Basil
Blackberry
Cite this dataset
Methodology published under CC BY 4.0. Three formats — APA for prose, BibTeX for LaTeX, RIS for reference managers (Zotero, Mendeley, EndNote).
Johnson, M. C. (2026). Unsupervised Recovery of Food Taxonomy from Volatile Compound Profiles: A UMAP Analysis of 450 Ingredients. Zenodo. https://doi.org/10.5281/zenodo.19719459
@misc{johnson2026compkitchen,
author = {Johnson, Mark C.},
title = {Unsupervised Recovery of Food Taxonomy from Volatile Compound Profiles: A {UMAP} Analysis of 450 Ingredients},
year = {2026},
publisher = {Zenodo},
doi = {10.5281/zenodo.19719459},
url = {https://doi.org/10.5281/zenodo.19719459}
}
TY - DATA AU - Johnson, Mark C. PY - 2026 TI - Unsupervised Recovery of Food Taxonomy from Volatile Compound Profiles: A UMAP Analysis of 450 Ingredients PB - Zenodo DO - 10.5281/zenodo.19719459 UR - https://doi.org/10.5281/zenodo.19719459 ER -
Methodology · Zenodo · CC BY 4.0