Introduction

Food is inherently complex, consisting of thousands of chemical components that interact and ultimately shape its flavor, texture, nutrition, and consumer acceptance1,2. Despite advances in analytical food chemistry and sensory science, we still lack a clear understanding of how molecular composition, component interactions, and human perception are interconnected across scales3,4.

At the chemical level, food contains thousands of compounds with diverse structures and physicochemical properties, including volatility, solubility, bioactivity, and bioavailability. Even small structural differences, such as isomerism, can lead to pronounced changes in functional or sensory behavior5. Although more than 77,000 food-related chemicals have been reported, fewer than one-tenth have annotated flavor or bioactivity information6,7. This sparse knowledge base makes manual interpretation difficult and highlights the value of machine learning (ML), which can predict molecular properties directly from structure and reveal how specific features shape a chemical’s sensory or functional role in food.

Beyond individual chemicals, food properties also arise from interactions among components. These interactions can create emergent behaviors that are not predictable by examining single compounds in isolation. Well-known examples include the balance created jointly by sugars, acids, and volatiles in fruits8, or the enhancement of flavor richness through lipid-aroma interactions in dairy matrices9. Such interaction patterns are further modified by ingredient composition, processing conditions, and storage environments, contributing to the variability observed across foods10. Analytical tools such as gas chromatography-mass spectrometry (GC-MS) and liquid chromatography-mass spectrometry (LC-MS) help identify and quantify complex molecular mixtures11, whereas sensor systems like e-noses and e-tongues mimic human chemosensation to rapidly profile collective molecular signals12,13. When combined with such approaches and ML, these multimodal data support rapid assessment of quality, authenticity, and origin and provide scalable tools for interpreting interaction-driven complexity14.

Human perception adds yet another dimension. The same chemical stimulus can elicit distinct sensory responses across individuals, influenced by genetic variation, cultural background, and personal experience15. Advances in neuroinformatics technologies, including electroencephalography (EEG), functional near-infrared spectroscopy (fNIRS), and functional magnetic resonance imaging (fMRI), provide richer psychological and physiological markers for studying these perceptual differences16. For example, fMRI and fNIRS can be used to identify specific brain regions activated by different flavor stimuli, and EEG can capture, in real time, the rapid changes in brain electrical activity associated with perception17. Integrating ML with these neurophysiological datasets offers a powerful means to model emotions and preferences elicited by food18,19, and to identify the neural features most relevant to perception20. Such approaches deepen our understanding of how humans experience food beyond its chemical and physical properties.

Although previous reviews have examined ML applications in food science, most have focused on a single data type or addressed only one layer of food complexity11,21,22. As a result, the integration of multimodal data within unified ML frameworks remains limited. In this review, we conceptualize food complexity across three interconnected dimensions: (1) the chemical diversity that defines the building blocks of food; (2) the interactions among components that give rise to the emergent properties of foods; and (3) the perceptual processes through which humans experience food (Fig. 1). Unveiling these dimensions demands large, heterogeneous datasets, motivating growing interest in data fusion strategies that integrate information across sources to tackle high-dimensional or sparse data and enhance model accuracy23,24,25. Building on this motivation, the review aims to summarize how ML has been applied to unveil each dimension of food complexity; examine data fusion approaches that unify disparate datasets; and highlight the opportunity of incorporating neuroinformatics data to deepen our understanding of food perception and improve predictive performance.

Fig. 1: Three levels of food complexity.
Fig. 1: Three levels of food complexity.
Full size image

Level 1 captures the chemical complexity of food, where diverse chemical structures give rise to distinct bioactivities and flavor characteristics. Level 2 represents the interaction level, encompassing the dynamic physical, chemical, and matrix-dependent interactions among food components that shape quality, safety, and stability. Level 3 reflects perceptual complexity, arising from the multisensory integration of taste, smell, trigeminal sensations, and higher cognitive processes that together determine the eating experience.

Representative machine learning algorithms for unveiling food complexity

ML provides a diverse set of algorithms that can extract patterns, uncover relationships, and make predictions from the heterogeneous datasets that characterize food complexity. Depending on the data structure and objective, different algorithmic families, such as supervised and unsupervised learning, offer complementary strengths for classification, regression, and clustering.

As illustrated in Fig. 2, ML workflows in food research typically involve data collection, preprocessing, feature engineering, and model training. Within this pipeline, supervised algorithms like support vector machines (SVM), decision trees (DT), random forests (RF), and k-nearest neighbors (KNN) have been widely adopted for tasks such as quality prediction, authenticity verification, and sensory property modeling26. Unsupervised techniques, including clustering and dimensionality reduction, are often used to explore data structure, extract latent features, and guide early-stage data processing27. This section introduces the core ML algorithms most relevant to decoding food complexity and their representative applications.

Fig. 2: Machine learning modeling pipeline for unveiling food complexity.
Fig. 2: Machine learning modeling pipeline for unveiling food complexity.
Full size image

The process begins with data collection, integrating information from databases, sensory evaluations, and instrumental analyses. Next, data cleaning and preprocessing standardize formats, resolve inconsistencies, and harmonize multi-source inputs. Feature engineering follows, involving either manual descriptor construction or automated representation learning through deep models. In model development, algorithms are trained, optimized, and validated to capture structure-function relationships. Finally, the resulting models are applied to predict properties across the molecular, interaction, and perception levels, enabling a multiscale understanding of food complexity.

SVM is a margin-based classifier that identifies the decision boundary maximizing the separation between classes in a transformed feature space28. It addresses nonlinear classification issues using kernel functions, including linear, radial basis, and sigmoid kernels. For example, Gerhardt et al. (2019) developed an SVM model to effectively differentiate among various types of virgin olive oil, including extra-virgin, virgin, and lampante olive oil29. Similarly, Sun et al. (2022) developed an SVM model that integrates fruit metabolomic data to predict consumer preferences, thereby facilitating consumer fruit-flavor selection and aiding breeding programs in developing more popular varieties30.

DTs use hierarchical, rule-based splits to partition data into homogeneous groups for classification or regression31. RFs and XGBoost extend DTs by constructing large ensembles of trees and aggregating their predictions to improve robustness and reduce overfitting32. Wang et al. (2021) developed an RF model to predict the retention index and flavor attributes (aromatic, bitter, sulfury, and other) of molecules in beer, achieving satisfactory performance with an R2 of 0.9633. Hu et al. (2023) developed an RF model to determine the relationship between key aroma components and the sensory properties of fragrant peanut oils34.

KNN classifies or predicts samples based on the labels or values of the most similar instances in the feature space35. Wu et al. (2019) developed a KNN model using electronic nose (e-nose) signals preprocessed with fuzzy discriminant principal component analysis (PCA) as input to differentiate between various Chinese liquor types and identify inferior or counterfeit liquids based on flavor, achieving a classification accuracy of 98.3%36. In addition to the previously mentioned ML algorithms, partial least-squares (PLS) regression, K-means clustering, and naïve Bayes have also been used to investigate food properties7.

In recent years, advances in computational capacity and the availability of large-scale datasets have enabled the application of deep learning7. Deep learning refers to a family of neural-network-based models capable of learning hierarchical representations from large, high-dimensional datasets26,37. In contrast to traditional ML techniques that require manual feature selection, deep learning models can autonomously learn hierarchical representations, thereby facilitating their deployment in complex data analyses.

Artificial neural networks (ANNs) consist of interconnected computational units that transform input features through weighted connections and nonlinear activation functions38. ANNs have been used to predict the taste attributes of chemicals based on molecular descriptors, thereby facilitating the rapid screening of novel flavorings39. In recent years, significant advances have been made in the development of ANN architectures, including convolutional neural networks (CNNs), graph neural networks (GNNs), and recurrent neural networks (RNNs), which have been proposed to address more complex tasks in food analysis.

CNNs use convolutional filters to extract spatially localized features and are widely applied to grid-like data such as images40. In food flavor research, CNNs are frequently used in the analysis of e-nose and electronic tongue (e-tongue) related tasks41,42. Wu et al. (2019) developed a CNN model to predict the perceptual pleasantness of odors based on e-nose data, achieving an accuracy of > 90% on the test dataset43. Zhang et al. (2019) integrated CNNs with neuroinformatics data to enhance flavor recognition by analysing brain electrical signals in response to various aroma stimuli, including coffee, lemon, and vanilla odors44. Xia et al. (2024) developed a CNN model that effectively decodes EEG signals corresponding to the taste of sour, sweet, bitter, and salty foods, thereby facilitating the prediction of food taste perception45.

GNNs operate on graph-structured data by iteratively aggregating information from neighboring nodes and edges46. In contrast to CNNs, which operate on a fixed set of relationships between pixels in image data, GNNs are designed to handle dynamic relationships between elements, aggregating neighboring information to update node representations. When predicting the taste of molecules, Song et al. (2023) observed that whereas most models perform reasonably well in predicting the umami taste, GNNs stand out for their superior accuracy in predicting bitter and sweet tastes47.

RNNs model sequential data by maintaining hidden states that capture temporal dependencies across time steps48. Qi et al. (2023) used an RNN combined with a multilayer perceptron to enable rapid identification of umami peptides49. However, RNNs face limitations like vanishing and exploding gradients, which hinder performance on long sequences50. To address this, advanced variants like long short-term memory (LSTM) networks were developed. LSTMs enhance RNNs by introducing gating mechanisms that regulate information flow and preserve long-range dependencies51. For instance, Jiang et al. (2023) applied an LSTM encoder to extract features from peptide sequences for umami recognition52. LSTMs have also been used in natural language processing tasks in food research, including flavor term extraction from whisky reviews and personalized recipe generation based on user preferences53,54.

Transformer architecture leverages self-attention mechanisms to model global dependencies in sequences while enabling highly parallelizable training55. This design markedly enhances model parallelism, thereby improving the training efficiency on large-scale datasets. Transformer-based models typically generate hidden state representations for each input token, thereby capturing the range of features learned from the input text. For example, Chew et al. (2022) developed a variant of the transformer model to analyse Instagram posts, and this approach exhibited superior precision and recall relative to baseline models in identifying brands and flavors of electronic nicotine delivery systems56. Additionally, large language models (LLMs) built on Transformer architectures excel at processing multimodal inputs (from text to images) to identify dish types and flavor profiles, while generating customized recipes that meet specific requirements57.

Generative models learn the underlying probability distribution of data to create new samples that resemble those in the training set. By training on extensive datasets containing both flavor chemometric data and human sensory panel evaluations, these models learn to construct latent mathematical representations that encode fundamental patterns2. The trained model can generate synthetic but chemically plausible profiles to augment limited experimental datasets or predict emergent characteristics from novel ingredient combinations by interpolating in the learned space. For example, Queiroz et al. (2023) proposed deep generative models to design new flavor molecules based on their molecular structures58. Aleixandre et al. (2025) presented a generative diffusion network to create new aromas with specified characteristics, based on mass spectrometry data of essential oils59. Zhang et al. (2026) developed a conditional variational autoencoder based on KFO-Atlas to produce category-targeted aroma formulations and validate its outputs by blinded human sensory evaluation60. These aspirational directions highlight the capacity of generative modeling to accelerate discovery in food, although most of them remain contingent on further validation with experimental and sensory data.

Overall, the aforementioned algorithms differ in complexity, data requirements, and representational power, each offering distinct advantages for food research (Table 1). Traditional ML methods like SVM, RF, and KNN are well-suited for structured, small-to-medium datasets and are often used for classification, regression, and feature selection tasks. In contrast, deep learning models excel at capturing complex, nonlinear relationships in large, high-dimensional, and multimodal datasets. They can autonomously learn hierarchical representations, making them ideal for analyzing unstructured data such as sensor signals, molecular graphs, peptide sequences, and textual descriptions. Generative models go a step further by enabling the design of novel molecules and the generation of synthetic data.

Table 1 Comparison of representative machine learning algorithms in exploring food complexity

Machine learning unveils food chemical complexity

Mapping the uncharted chemical space of foods with machine learning

While databases such as KFO-Atlas, FooDB, FlavorDB, and AdditiveChem collectively catalog tens of thousands of food-related compounds as summarized in Table 2, much of the chemical diversity of food remains uncharted, particularly secondary metabolites, processing-induced compounds, and minor constituents with sensory or health relevance. Early efforts such as FoodMine demonstrated the value of systematic literature mining by extracting over 7000 quantified measurements for garlic and cocoa; however, its reliance on manual curation highlighted the limitations of scalability61.

Table 2 Representative databases supporting food chemical and interaction complexity studies

Recent advances in natural language processing, especially large language models (LLMs), now offer automated routes to expanding the molecular inventory. Domain-adapted LLMs can extract chemical entities, normalize molecular identifiers, and associate compounds with functional annotations from large-scale scientific corpora. Models such as GPT-4 have been shown to identify food-related molecules, metabolic intermediates, and bioactive compounds from a large set of publications at speeds impossible for manual curation62.

Beyond literature, ML is also broadening the search for food-relevant metabolites through genome mining. ML-driven tools such as antiSMASH63 and PRISM64 can analyze biosynthetic gene clusters to infer the structures of secondary metabolites that may have nutritional, flavor, or preservative functions. This is particularly relevant for identifying previously uncharacterized molecules from fermented foods, edible fungi, and underutilized plant species.

Predicting multidimensional properties of food-related chemicals

Among the tens of thousands of known food-related molecules, fewer than 10% have annotated flavor or bioactivity information. This sparse coverage has motivated growing interest in ML as a means to infer molecular properties from structural information.

Traditional approaches, such as quantitative structure-activity relationship (QSAR) models, originally established the link between molecular structure and functional behavior using handcrafted physicochemical descriptors65,66. Fingerprints and descriptor-based models remain effective for small to medium datasets and have been widely applied to tasks such as odor classification, as demonstrated by Shang et al.’s (2017) work using over 1000 descriptors from Dragon software to predict odor characteristics67.

However, descriptor-based strategies inherently depend on manually engineered features, limiting their ability to capture higher-order structural patterns. This has driven a shift toward deep learning approaches that learn representations directly from raw molecular structures (Table 3). GNNs, for example, treat atoms and bonds as nodes and edges, allowing the model to capture the topological and geometric properties essential for molecular function. Pred-O3, developed by Ollitrault et al. (2024), predicted 23 odor notes and 109 human olfactory receptor interactions from 5802 food-derived odorants, revealing structure-odor relationships missing from traditional fingerprints68. Similarly, Lee et al. (2023) created a principal odor map that positioned over 500,000 potential odorants, of which only ~5000 have been characterised, highlighting a vast unexplored chemical-sensory landscape69. Emerging work has also explored the use of large language models for molecular property prediction. Song et al. (2024) fine-tuned GPT-3.5 and GPT-4 on SMILES strings for taste classification, with GPT-4 achieving 86% accuracy, illustrating LLMs’ potential to complement graph-based modeling70.

Table 3 Representative machine learning studies in exploring food complexity

As these methods advance, ML is increasingly positioned to accelerate the early-stage prediction of molecular sensory and functional properties, reducing reliance on lengthy experimental screening. Equally important, interpretable AI frameworks such as SHAP47 offer new opportunities to illuminate the structural features most responsible for molecular behavior, bridging predictive modeling with mechanistic understanding.

Integrating machine learning with instrumental analysis to decode food component interactions

Examining food properties, including quality grading, geographical origin, and freshness, facilitates food product development and quality control (Fig. 3a). Traditionally, these tasks rely on labor-intensive, time-consuming experiments, such as sensory evaluation panels and metabonomics analyses, which are often limited by human subjectivity, high costs, and low throughput. Moreover, the complexity of food matrices and the interactions among various components pose challenges for manual interpretation. In this context, ML offers a powerful tool by learning complex, nonlinear relationships from instrumental data.

Fig. 3: Machine learning with instrumental analysis and electroencephalogram (EEG) data.
Fig. 3: Machine learning with instrumental analysis and electroencephalogram (EEG) data.
Full size image

a Application of machine learning using instrumental analysis data and the representative artificial neural network (ANN) algorithm for predicting food characteristics. b Machine learning modeling with EEG data. EEG data can be manually preprocessed to extract temporal, spatial, or spectral features, which can then be used in conventional machine learning frameworks, such as support vector machines (SVMs) and random forests (RFs). Alternatively, deep learning algorithms, such as convolutional neural networks, enable automatic feature extraction and prediction, offering an advanced approach to decoding sensory responses to food.

Chromatography-based modeling

Chromatographic techniques such as GC-MS, LC-MS, and GC-IMS spectrometry provide detailed molecular fingerprints of foods and are widely used to profile volatiles, semi-volatiles, and other key constituents. When coupled with ML, these datasets enable rapid and accurate analysis of product origin, quality, and sensory attributes.

Recent studies demonstrate that statistical classifiers and modern ML models can decode subtle variations in chromatographic profiles for tasks such as geographical origin discrimination in honey71, regional classification of Atractylodes lancea72, and prediction of α-acid content in hops73. Similar approaches have been used to assess rancidity in walnuts74 and sensory grade in Sauvignon Blanc wine75. Together, these examples illustrate how ML enhances the interpretive power of chromatographic datasets by uncovering latent relationships between volatile patterns and food properties.

Chromatography-based ML models are also increasingly used to complement sensory evaluation. While sensory panels remain essential, they suffer from fatigue, cultural variability, and limited throughput15. ML can serve as an initial screening tool to map chemical profiles to sensory descriptors, as in studies linking GC-MS data to aroma attributes in peanut oils76 or predicting beer flavor preferences by integrating chemical, sensory, and consumer data77. These approaches help form a hybrid workflow in which ML provides rapid, objective screening, while human assessors calibrate and validate final outputs.

Sensor-based modeling

E-noses and e-tongues extend the capability of instrumental analysis by capturing holistic responses to odorants and tastants through sensor arrays. When integrated with ML, these systems offer high-throughput, reproducible alternatives to traditional sensory evaluation (Fig. 3a).

E-nose signals together with ML have been used to classify products such as Chinese baijiu36, discriminate brewing stages78, recognize beer odors79, and detect spoilage volatiles80. Advanced deep learning methods further improve e-nose performance. Shi et al. (2019) proposed a CNN-SVM hybrid for precise beer odor recognition, replacing the fully connected layer of a CNN with an SVM to enhance predictive capability and better capture complex sensory patterns79. Xiong et al. (2021) developed a convolutional spiking neural network to identify spoilage odors, integrating residual networks with spiking neurons to convert continuous sensor signals into discrete pulses. This design reduces redundant computations, requires fewer parameters than a one-dimensional-CNN, occupies less memory, and achieves >84% accuracy, making it well-suited for e-nose devices with limited computational resources80.

Although these approaches effectively model overall odor profiles, ML models that can predict odor activity values (OAVs) or quantify multi-compound synergistic, masking, or enhancing effects remain scarce. Traditional methods for studying aroma interactions, such as σ-τ diagrams, OAV calculations, S-curves, or distribution-based metrics81,82, require intensive mathematical formulation and often fail to capture nonlinear behaviors in real food matrices. ML offers new opportunities to learn these complex relationships directly from large datasets of known interactions, yet such datasets remain limited and models for multi-component interaction prediction are still in their infancy.

E-tongue systems show similar potential. ML-enabled classifiers have been developed for tea authentication83 and for predicting Pu’er tea storage time using transfer-learning approaches that leverage pre-trained models on large temporal signal datasets84. Data augmentation strategies, such as introducing controlled noise, also help mitigate small sample sizes frequently encountered in sensor-based studies84.

Despite many applications, it is worth noting that taste stimuli dissolve in saliva before reaching their respective receptors. Once dissolved, the components of saliva can interact with these stimuli and their receptors, thereby influencing taste perception85. However, solvents commonly used in food research with e-tongues often exhibit notable differences in their physicochemical properties compared to those of human saliva. Such information has been less considered in ML modeling. It is recommended that future research using the e-tongue account for the effects of saliva and consider using artificial saliva as a solvent to more accurately replicate oral environmental conditions.

Data fusion modeling

Insights from chromatography- and sensor-based modeling show that individual analytical platforms capture only fragments of the complex interactions within food matrices. Chromatography provides detailed molecular fingerprints, particularly for volatiles, whereas e-noses, e-tongues, and other sensor systems capture holistic olfactory and gustatory responses. Integrating these complementary data streams can substantially improve food property prediction by representing a broader spectrum of chemical and sensory information. Yet this richness comes with challenges, including high dimensionality, redundancy, and noise.

Data fusion addresses these limitations by combining multimodal information to enhance model robustness and predictive performance86. Fusion strategies typically operate at three levels: low, mid, and high (Fig. 4a)87,88. Low-level fusion merges raw or preprocessed signals into a unified input matrix. Mid-level fusion extracts and concatenates informative features across modalities to preserve complementary structure while reducing dimensionality. High-level fusion integrates predictions from separate models trained on different data types.

Fig. 4: Data fusion strategies with multimodal input.
Fig. 4: Data fusion strategies with multimodal input.
Full size image

a Combining multi-source heterogeneous data at low, mid, and high levels. b Data fusion pipeline involving electroencephalogram (EEG) and functional near-infrared spectroscopy (fNIRS) data, utilizing both early fusion and late fusion strategies to improve predictive accuracy in understanding food perception. Note: E-nose electronic nose, E-tongue electronic tongue, GC-MS gas chromatography-mobility spectrometry, LC-MS liquid chromatography-mass spectrometry, VIS visible spectroscopy, NIR near-infrared spectroscopy, HIS hyperspectral imaging spectroscopy.

Applications across food quality and flavor research illustrate the advantages of this approach. For example, combining nuclear magnetic resonance, LC-MS, and GC-MS data has revealed synergistic insights into compositional changes in heat-treated apple juice89. Studies fusing MS and NIR signals demonstrate that mid-level fusion often yields higher accuracy than low-level approaches when predicting sensory attributes such as bitterness and grassiness in olive oil90. Multisensor fusion of e-nose and e-tongue signals similarly enhances flavor modeling by jointly capturing olfactory and taste information, enabling accurate classification, quality monitoring, and regression tasks91,92. Incorporating visual data from electronic eye systems further broadens sensory representation, as shown in Longjing tea evaluation, integrating aroma, taste, and color93.

While multimodal fusion outperforms unimodal methods, practical implementation remains challenging. Differences in data formats, resolutions, scales, and naming conventions can hinder integration and reduce model generalizability. Food matrices also exhibit heterogeneity and temporal variability, complicating alignment across platforms. Preprocessing strategies, including wavelet filtering, Savitzky-Golay smoothing, independent component analysis, dynamic time warping, correlation-optimized warping, and retention-time correction, help mitigate noise and reconcile instrumental shifts94,95. Batch effects arising from platform discrepancies can be reduced using standard normal variate scaling, multiplicative scatter correction, or empirical-Bayes adjustments. At a broader level, adherence to the Findable, Accessible, Interoperable, Reusable (FAIR) data principles enhances metadata harmonization and reproducibility96. Incorporating these procedures is critical for enabling accurate, scalable, and comprehensive food property profiling.

Integrating machine learning with neuroinformatics to decode perception complexity

Human flavor perception emerges from the integration of taste, smell, texture, visual cues, and trigeminal sensations, processed across distributed neural circuits. Recent advances in neuroinformatics, particularly EEG, fNIRS, and fMRI, provide non-invasive ways to probe these processes and link subjective experiences to objective neural responses16. Each modality offers complementary strengths: EEG captures rapid neural dynamics, fNIRS enables naturalistic monitoring of cortical haemodynamics, and fMRI provides fine-grained spatial localization. Together, these tools create an opportunity to build ML frameworks that decode perceptual complexity directly from brain signals and move toward individualized food prediction.

Machine learning for decoding EEG-based flavor responses

EEG is a neurophysiological technique that records brain electrical activity by measuring the electrical potential difference between the scalp and skull using electrodes placed on the scalp. These signals represent temporal variations in electrical activity across distinct brain regions. Owing to its high temporal resolution and relatively low experimental cost, EEG has been extensively used in food research to enhance understanding of the relationship between food flavor perception and individuals’ cortical processing97.

Using EEG, Yang et al. (2023) identified neurophysiological indicators associated with four tastes (sour, sweet, bitter, and salty) and their respective intensities98. They discovered that different taste qualities could be distinguished within 250-1,500 ms of stimulation, and the alpha and theta frequency bands show greater sensitivity to different tastes than do the delta, beta, and gamma bands98. These responses reflect not only sensory processing but also reward valuation and cognitive control mechanisms.

Nevertheless, analyzing extensive EEG data to discern distinctive patterns and correlations with specific tastes presents considerable challenges. The application of ML techniques could prove instrumental in addressing this challenge. Figure 3b shows the ML model for food-flavor-induced EEG signals, with an emphasis on feature extraction, model construction, and prediction. Classical ML models, including decision trees, SVMs, and k-nearest neighbors, have been successfully applied to tasks such as classifying sweet vs. non-sweet stimuli99 and distinguishing fresh vs. non-fresh foods based on spectral features20.

Deep learning further advances EEG analysis by automatically extracting multiscale temporal-spatial representations. Multi-scale CNNs enable robust taste-category recognition across sour, sweet, bitter, salty, and umami stimuli19, and data-augmentation techniques such as spatiotemporal reconstruction mitigate issues of limited sample size45. ML models have also been used to predict olfactory pleasantness from EEG signatures, with graph-based features showing particularly strong performance18,100. Collectively, these studies highlight the emerging potential of EEG-based ML systems to decode affective and perceptual dimensions of food experience.

Machine learning for decoding fNIRS and fMRI signals

Besides EEG, fNIRS and fMRI are widely used non-invasive techniques that help to evaluate food perception by providing physiological indicators of brain activity. These technologies offer higher spatial resolution than EEG, enabling the observation of brain activity through hemodynamic changes. In particular, fNIRS uses near-infrared light to penetrate the skull and quantify the alterations in cerebral cortical blood oxygenation. Although its spatial resolution is inferior to that of fMRI, fNIRS is relatively portable and places fewer constraints on participants during use, making it well-suited for practical applications101. In contrast, fMRI measures brain activity by detecting blood-oxygen-level-dependent signals, offering exceptional spatial resolution that allows precise localisation of cognitive functions within the brain cortex101. Both fNIRS and fMRI facilitate the examination of brain responses to flavor stimuli during sensory evaluation, thereby providing valuable insights into the neural processing of sensory information102,103,104,105.

Integrating ML with these imaging modalities improves classification performance and reduces reliance on manual feature engineering. CNNs trained on resting-state fMRI data can discriminate between participant groups106, and SVM models have identified brain networks distinguishing responses to high-calorie foods (potato chips) and low-calorie foods (zucchini)107. These approaches demonstrate how data-driven methods can extract perceptual information that is not readily apparent in univariate analyses.

Multimodal neuroinformatics and data fusion for food perception prediction

While various neuroinformatics technologies can extract numerous features, including time, frequency, time-frequency, and spatial features, the characteristics derived from different neuroinformatics modalities inherently possess unique strengths and limitations. For instance, EEG provides rapid responsiveness but lacks spatial localisation. In contrast, fMRI provides high spatial resolution but is highly susceptible to head and body movements. Meanwhile, fNIRS enables monitoring of the cerebral cortex surface with greater tolerance to body movements, though it has slower responsiveness and lower spatial resolution108,109. A meta-LDA classifier that combines EEG band power changes with fNIRS hemoglobin concentration data can markedly enhance the accuracy of motor imagery classification relative to the use of EEG data alone110. This suggests that combining neuroinformatics technologies through multimodal data integration, supported by advanced ML techniques, could leverage their complementary strengths.

Multimodal fusion methods in neuroinformatics can be categorized as symmetric or asymmetric. Symmetric fusion treats all modalities equally, while asymmetric fusion prioritizes one modality as a reference or constraint for another111. Feature fusion can also be classified as early or late, depending on the timing of integration. Fig. 4b illustrates the data fusion strategies for EEG and fNIRS data. Early fusion combines data from multiple sources into a unified format before processing, suitable for strongly correlated data, whereas late fusion integrates results after independent analyses, ideal for independent data sources112. For example, Sun et al. (2020) extracted features from fNIRS and EEG signals, concatenated them to create enhanced feature vectors, and classified participants’ emotional states while watching videos using an SVM algorithm113. Furthermore, an increasing number of studies have examined the influence of auditory and visual cues on food perception114,115.

Challenges and opportunities

ML at the molecular layer benefits from growing databases that support the prediction of sensory attributes, receptor activities, and toxicological profiles. However, progress is constrained by three major gaps. First, molecular databases contain substantial redundancy and inconsistent annotations due to overlapping curation sources116, complicating data integration and reducing modeling reliability. Second, sensory annotations remain narrowly focused on olfaction and taste, while trigeminal and chemesthetic dimensions are largely undocumented, with PungentDB being a rare resource linking molecules to TRP channels117 (Table 2). Third, generative models produce promising molecular suggestions but lack interpretability and are rarely validated experimentally, limiting real-world applicability58,59.

Modeling interactions among food ingredients is far more complex than predicting single-molecule properties. First, most ML studies still focus on small volatile compounds because they have well-defined structures, richer databases, and established analytical workflows. In contrast, interactions involving macromolecules, such as lipids, proteins, and polysaccharides, remain largely unexplored, even though they play critical roles in flavor release, stability, and texture. This gap arises from the structural heterogeneity of macromolecules, the strong dependence of interactions on matrix conditions (pH, temperature, ionic strength), the scarcity of high-resolution in situ datasets, and the difficulty of representing molecular and supramolecular features within unified ML frameworks. Second, although databases such as KFO-Atlas, Open Food Facts, Recipe1M, FoodRepo, and FoodData Central map food-ingredient relationships, current computational approaches offer limited chemical interpretability and struggle to incorporate reaction dynamics and spatial constraints60 (Table 2). Third, empirical data generation remains a major bottleneck: matrix effects, chromatographic variability, and isomer ambiguity compromise measurement stability and hinder reproducible quantification5. Inter-laboratory differences and inconsistent labeling further reduce dataset comparability118. Moreover, food systems evolve continuously through oxidation, enzymatic transformations, and Maillard reactions29,89, yet available datasets are predominantly static snapshots that fail to capture these temporal dynamics.

Despite growing interest in neuroinformatics, its application to food perception remains limited due to multiple structural and methodological barriers. First, large-scale EEG, fNIRS, and fMRI datasets specifically focused on flavor are scarce, and each modality has inherent limitations: EEG lacks spatial precision; fNIRS and fMRI are expensive, sensitive to motion artifacts, and difficult to deploy in naturalistic settings110. Small sample sizes, inconsistent acquisition protocols, heterogeneous hardware configurations, and culturally narrow participant pools further undermine reproducibility and model generalizability. Neuroinformatics signals are inherently noisy and highly susceptible to environmental and physiological interference, making the collection of high-quality data outside controlled laboratories challenging. Second, the field lacks standardized preprocessing pipelines and evaluation criteria. These inconsistencies reduce cross-study comparability and limit the transferability of models across populations and contexts. At the modeling level, explainability remains insufficient: existing frameworks mainly serve methodological inspection rather than providing actionable insights for product development, sensory evaluation, or consumer applications. Third, most current studies examine olfactory or gustatory pathways in isolation. However, real-world flavor perception is inherently multisensory, shaped by trigeminal stimulation as well as visual, auditory, and oral tactile cues.

Looking ahead, progress in understanding food complexity requires coordinated advances across the molecular, interaction, and perceptual levels. At the chemical level, databases should expand beyond olfactory and gustatory data to include trigeminal responses and receptor-level information. High-throughput screening, combined with carefully curated, standardized datasets, will be essential for mapping the vast space of uncharacterized food molecules and improving structure-function understanding. At the interaction level, standardization across labs, open data sharing, and unified ontologies are crucial for making data comparable and enabling reproducible models. These efforts should go hand in hand with multi-scale analytical approaches capable of capturing dynamic changes during processing and storage. At the perceptual level, there is a pressing need for large, culturally diverse neuroimaging and sensory datasets. Federated learning offers a promising solution by enabling multi-center model training without exchanging raw neurophysiological data, thereby safeguarding participant privacy while improving model generalizability across populations and experimental settings; in parallel, synthetic data generation, using generative models such as variational autoencoders, GANs, or diffusion-based approaches, can help alleviate small-sample limitations by augmenting training data with statistically realistic neural signals, reducing overfitting to site-specific noise119,120. There also remains considerable scope to develop end-user-centered, explainable AI systems that tailor model outputs to the needs of different stakeholders. For example, developing interfaces that map neural features onto sensory attributes familiar to product developers, or consumer-oriented visual summaries that communicate how flavor cues influence predicted affective responses121. Furthermore, the development of industry-wide standards, rigorous validation frameworks, and systematic cost-benefit and return-on-investment assessments will be essential for narrowing the gap between laboratory findings and real-world applications. Finally, incorporating modalities including olfactory, gustatory, trigeminal, visual, and auditory sensations through multimodal data fusion and ML frameworks will enable more comprehensive decoding of flavor responses and support future developments in personalized nutrition, virtual tasting, and immersive multisensory applications.

Advances in ML will support progress across all three layers. Ongoing development of flexible deep learning models is essential to combine diverse data types, including chemical, sensory, and neural signals, while automated feature engineering and domain-specific explainable AI can lessen dependence on manual processing and enhance interpretability. Creating multisensory models that identify nonlinear interactions among sensory pathways will be vital. Ultimately, close collaboration among food scientists, chemists, data scientists, neuroscientists, and sensory experts will accelerate the development of shared resources, standardized workflows, and integrated modeling tools, thereby advancing the field’s understanding of food complexity.

Conclusions

This review outlines the transformative potential of ML in decoding the three interrelated layers of food complexity: (1) the structural and physicochemical properties of individual food molecules; (2) the food properties arising from interactions among food components; and (3) the neurophysiological processes underlying human food perception. By harnessing a growing ecosystem of molecular, instrumental, and neuroinformatics data, ML provides a powerful means to bridge micro-level chemical data with macro-level sensory experiences.

We summarised the application of advanced ML and deep learning architectures, including CNNs, GNNs, and transformers, which are particularly well-suited to modeling the complex, nonlinear, and multimodal nature of food-related data. We also examined the growing role of neuroinformatics technologies such as EEG, fNIRS, and fMRI in decoding individual perceptual responses, and how their integration with ML can advance personalised food experience modeling. Moreover, we emphasised that integrating diverse data modalities through data fusion across chromatographic, sensor-based, imaging, and neurophysiological platforms offers a critical solution to challenges such as data sparsity, heterogeneity, and noise. These strategies enable the development of more accurate and generalisable predictive models for food quality assessment, flavor profiling, consumer preference prediction, and product authentication.

Looking forward, progress will depend on three priorities: (1) building comprehensive and standardized multimodal databases, including underrepresented dimensions such as trigeminal perception; (2) developing interpretable and hybrid ML models that balance predictive accuracy with mechanistic insight; and (3) fostering interdisciplinary collaboration across food science, chemistry, neuroscience, and data science to establish shared platforms and validation pipelines. By aligning technical development with rigorous standards and collaborative practices, we expect ML to not only advance the scientific understanding of food complexity but also accelerate its translation into and consumer-relevant applications.