Frontiers | Systematic Multi-Omics Integration (... https://www.frontiersin.org/articles/10.3389/fpls....
About us All journals All articles Submit your research Search Login
Frontiers in Plant Science Sections Articles Research Topics Editorial Board About journal
Download Article 13,430 total views (https://www.altmetric.com
View Article Impact
SHARE
/details.php?domain=www.frontiersin.org& ON
citation_id=84836336)
REVIEW article
Front. Plant Sci., 26 June 2020
Sec. Plant Systems and Synthetic Biology
https://doi.org/10.3389/fpls.2020.00944 (https://doi.org/10.3389/fpls.2020.00944)
Systematic Multi-Omics Integration (MOI) Approach in Plant Systems Biology
Ili Nadhirah Jamil (https://www.frontiersin.org/people/u/971527)1, Juwairiah Remali (https://www.frontiersin.org/people/u/1009029)1,
Kamalrul Azlan Azizan (https://www.frontiersin.org/people/u/1009037)1,
Nor Azlan Nor Muhammad (https://www.frontiersin.org/people/u/795994)1, Masanori Arita (https://www.frontiersin.org/people/u/13070)2,3,
Hoe-Han Goh (https://www.frontiersin.org/people/u/113859)1 and Wan Mohd Aizat (https://www.frontiersin.org/people/u/387348)1*
1
Institute of Systems Biology (INBIOSIS), Universiti Kebangsaan Malaysia (UKM), Bangi, Malaysia
2 Bioinformation & DDBJ Center, National Institute of Genetics (NIG), Mishima, Japan
3 Metabolome Informatics Team, RIKEN Center for Sustainable Resource Science, Yokohama, Japan
Across all facets of biology, the rapid progress in high-throughput data generation has enabled us to perform multi-omics systems biology
research. Transcriptomics, proteomics, and metabolomics data can answer targeted biological questions regarding the expression of
transcripts, proteins, and metabolites, independently, but a systematic multi-omics integration (MOI) can comprehensively assimilate,
annotate, and model these large data sets. Previous MOI studies and reviews have detailed its usage and practicality on various organisms
including human, animals, microbes, and plants. Plants are especially challenging due to large poorly annotated genomes, multi-
organelles, and diverse secondary metabolites. Hence, constructive and methodological guidelines on how to perform MOI for plants are
needed, particularly for researchers newly embarking on this topic. In this review, we thoroughly classify multi-omics studies on plants and
verify work�ows to ensure successful omics integration with accurate data representation. We also propose three levels of MOI, namely
element-based (level 1), pathway-based (level 2), and mathematical-based integration (level 3). These MOI levels are described in relation
to recent publications and tools, to highlight their practicality and function. The drawbacks and limitations of these MOI are also discussed
for future improvement toward more amenable strategies in plant systems biology.
Introduction
The acquisition of multi-omics data sets has become an integral component of modern molecular biology and biotechnology. This is due to
technological advancements, such as the next-generation sequencing technology (Illumina, PacBio, and Nanopore) and mass spectrometry
coupled with gas- and liquid chromatography, which o�er high-throughput data generation (Fondi and Liò, 2015). The core data sets of
systems biology are transcriptomics, proteomics, and metabolomics, providing the expression levels of transcripts, proteins, and metabolites,
respectively (Aizat et al., 2018a). The data generated from these platforms can be massive, often without clear connections between them.
For instance, it is nearly impossible to manually associate hundred thousands of transcripts to their respective proteins or metabolic
pathways. In fact, the bottleneck of omics research is considered as the biological/machine/human resource allocation for data processing
and integration (Palsson and Zengler, 2010). What we need is a well-de�ned methodological scheme for multi-omics integration (MOI) to
extract, combine, and critically associate di�erent data sets to allow researchers to decipher the seemingly complex biological results at hand
(Fondi and Liò, 2015; Hughes, 2015; Wang et al., 2018).
MOI approach has been extensively studied and reviewed in studies on human (Chen and Chen, 2019; Cho et al., 2019; Shetty et al., 2019),
animals (García-Sevillano et al., 2014), microbes (Denman et al., 2018; Gutleben et al., 2018; Wang et al., 2019), and their combinations
(Wanichthanarak et al., 2015; Cavill et al., 2016; Pinu et al., 2019). In comparison, MOI in plants has been more di�cult due to their metabolic
diversity, poorly annotated large genomes (particularly for non-model species), and the presence of numerous symbionts with complex
interaction networks. Several comprehensive reviews are available speci�cally on plant MOI (Fukushima et al., 2009; Fukushima et al., 2014;
Rajasundaram and Selbig, 2016; Rai et al., 2017; Rai et al., 2019) and its practical usage in green systems biology, precision plant breeding,
and other biotechnological applications (Weckwerth, 2011; Weckwerth, 2019; Weckwerth et al., 2020). However, the advancement of high-
throughput technologies and large omics data sets leading to big data biology can be overwhelming, and perhaps an “Achilles' heel” for
inexperience researchers. Omics data from poorly characterized species are often feed into software without proper manual curation and
oblivious of the limitation of each technology, which could result in incorrect interpretations. Further, there are also a large collection of
software platforms, statistical rigor, and modeling (Weckwerth, 2011; Pinu et al., 2019) which must be selected appropriately by users, yet
these can be viewed as extraneous to untrained researchers. Hence, suitable methodological work�ow for MOI must be identi�ed to ensure
accurate large-scale data analysis and representation. Previously, the di�erent levels of MOI have been summarized as “conceptual,”
“statistical,” and “model-based” integration (Cavill et al., 2016; Rai et al., 2017); instances of such integration have been detailed elsewhere (de
Oliveira Dal'Molin and Nielsen, 2013; de Oliveira Dal'Molin and Nielsen, 2018; Seaver et al., 2018). Let us start from a critical review on the
1 of 19 10/18/22, 6:35 PM
,Frontiers | Systematic Multi-Omics Integration (... https://www.frontiersin.org/articles/10.3389/fpls....
previous three-level classi�cation.
About us All journals All articles Submit your research Search Login
The “conceptual” integration refers to multiple omics data sets being analyzed separately and are matched without further statistical analysis.
Even though this approach can produce valuable insights, it may miss reproducible associations when multiple omics data sets are analyzed
Frontiers in Plant
together Science
(Cavill et al., 2016;Sections
Rai et al.,
2017). We therefore
Articles argueTopics
Research that the conceptual integration
Editorial Board Aboutisjournal
an arbitrary
connection without proper
analysis and should not be considered a part of MOI approach. Instead, we re-classify the “statistical” integration where statistical
associations are sought between elements from di�erent data sets (Cavill et al., 2016; Rai et al., 2017). The e�ective use of prior knowledge is
separated as the pathway-based integration from unbiased, element-based integration. Finally, “model-based” integration is also re-classi�ed
so that qualitative reconstruction of biological pathways or systematic regulatory pathways is separated from their quantitative, mathematical
evaluation to generate working models for hypothesis testing (Thiele and Palsson, 2010; Rai et al., 2017).
Thus, we re-de�ne the MOI work�ow into three main integration levels (levels 1 to 3) with increasing complexity (Figure 1). Level 1 is the
unbiased, “element-based” integration with three subclasses: correlation, clustering, and multivariate analyses. Level 2 is the knowledge-
based “pathway” integration, which includes co-expression and mapping-based approaches. Finally, level 3 is the “mathematical” integration
with two subclasses, namely, di�erential and genome-scale analyses. The three levels are discussed in relation to recent omics reports from
the year 2014 to 2020 (Table 1) gathered from comprehensive literature searches, including Web of Science, Scopus, and Google Scholar
databases providing an updated comprehensive overview of MOI applications in various plant systems. Furthermore, this review focuses on
expression-based omics from transcriptomics, proteomics, and metabolomics to further clarify strategies taken to integrate such large-scale
expression data (transcript, protein, and metabolite).
Figure 1
www.frontiersin.org FIGURE 1 Current approaches in multi-omics integration (MOI) of plant systems biology. This MOI strategy is
(https://www.frontiersin.org/�les classi�ed into three main levels with increasing degrees of complexity.
/Articles/540561/fpls-11-00944-
HTML-r1/image_m/fpls-11-00944-
g001.jpg)
Table 1
www.frontiersin.org TABLE 1 Summary of recent publications and their tools and methods used in the three di�erent levels of
(https://www.frontiersin.org/�les multi-omics integration (MOI).
/Articles/540561/fpls-11-00944-
HTML-r1/image_m/fpls-11-00944-
t001.jpg)
Level 1 MOI: Element-Based Approach
Correlation Analysis
The �rst level of MOI is an element-based integration approach speci�cally using correlation analysis (Figure 1). The advantage of this
integration is its simplicity and intuitiveness. The standard approach is correlative association between two or more di�erent omics data sets
(i.e., transcriptomics, proteomics, and metabolomics data sets). Such analysis is performed using Pearson's (Benesty et al., 2009) and
Spearman's correlation coe�cients (Myers and Sirois, 2004), which assess linear and ranked relationships, respectively. Other studies also
analyzed their omics data sets using Fisher's transformation to transform skewed data sets to normally distributed data for calculating
corresponding correlation coe�cients (Mata et al., 2018). In general, signi�cant positive or negative coe�cient suggests strong direct or
inverse relationship between data sets, respectively.
Correlation-based MOI has been performed between transcripts and their cognate proteins. Such analysis is straightforward, assuming that
di�erential expression of transcripts will also be observed at their translational (protein) level. However, this is often not the case, most studies
reported weak correlations (di�erent patterns) between transcript-protein levels. For instance, salt treatment on salt-tolerant Earlistaple 7 and
salt-sensitive Nan Dan Ba Di Da Hua cotton revealed scarce correlation (r=0.03) between transcript and corresponding protein patterns,
regardless of genotypic background (Peng et al., 2018). Another example includes methyl jasmonate (MeJA) stress hormone treatment on
Persicaria minor Huds. herbal plants, with poor overall proteome-transcriptome correlation (r=0.341) (Aizat et al., 2018b). Similarly, transcripts
and proteins related to ethylene pathway (ethylene receptors [ETRs] and downstream signaling proteins, constitutive triple response-like
proteins [CTRs], and ethylene insensitive 2 [EIN2]) were not well correlated during the ripening process of tomato (Solanum lycopersicum)
(Mata et al., 2018). This suggests the existence of post-transcriptional and post-translational regulation (such as proteasomal degradation) for
the majority components of stress and ripening pathways. Despite transcriptome can be weakly correlated to proteome, it serves as an
excellent database for protein identi�cation in proteomics informed by transcriptomics approach for non-model plants (Aizat et al., 2018b;
Wan Zakaria et al., 2019) as well as studying allele-speci�c expression (van Wesemael et al., 2018).
On the other hand, an interesting emerging pattern arises when transcript-protein is compared between speci�c protein groups. For
example, signi�cantly upregulated proteins were positively correlated with their cognate transcripts in the stress response of various plants
(Ye et al., 2017; Aizat et al., 2018b). Speci�cally, proteins related to defense such as proteases and peroxidases in MeJA-treated P. minor (Aizat
et al., 2018b) and secondary metabolite biosynthesis such as �avonoid in phytoplasma-infected Ziziphus jujuba Mill. leaf (Ye et al., 2017) were
upregulated coherently with their transcripts. This may suggest the concerted molecular upregulation of defense-related proteins to
overcome stress signals and infection. Meanwhile, proteins related to growth such as photosynthetic and structural proteins were
signi�cantly suppressed in these studies, perhaps as a response toward the stress signal to conserve energy and recycling molecular
2 of 19 10/18/22, 6:35 PM
, Frontiers | Systematic Multi-Omics Integration (... https://www.frontiersin.org/articles/10.3389/fpls....
resources (Ye et al., 2017; Aizat et al., 2018b). Interestingly, such downregulation was mainly observed at the protein level, but not the
transcript levelAbout
(Aizatusetal., 2018b), perhaps asAll
All journals a mechanism
articles toSubmit
quicklyyour
resume protein synthesis, when the stress is relieved. However,
research Searchwe Login
could not rule out the possibility that changes at both transcript and protein levels are not simultaneous. Even when sampling of both is done
at the same time, the translational and post-translational degradation and modi�cation rates may di�er among proteins. However, this is
Frontiers in Plant
often Sciencefrom the
unpredictable Sections Articles
genome sequences aloneResearch Topics
(Weckwerth, 2019; Editorial Board About journal
Weckwerth et al., 2020), convoluting meaningful and direct
interpretation between expression data.
While comparisons between transcripts and corresponding proteins are generally performed in multi-omics studies, correlation of these two
with metabolites are relatively fewer. Perhaps, one such recent example is the transcriptomics and metabolomics investigation of Ginkgo
biloba during leaf maturity process (Guo et al., 2020). Correlation analysis was performed in this study between all di�erentially expressed
transcript (DET) and metabolites (DEM), however with no regard for their biochemical pathway relationships. While this study may be
interested only in the pattern consistency between DET and DEM, the corresponding biochemical pathway should be considered before
such correlation is performed. Importantly, metabolites should be classi�ed as either being substrates or products of certain enzymatic
pathways to be accurately correlated with their corresponding transcripts/proteins. This has been performed by Silva et al. (2017) for
transcripts and associated metabolites to elucidate primary metabolism in Arabidopsis seed germination and growth.
Clustering Analysis
Clustering analysis allows grouping of omics data sets with similar attribute such as expression levels to deduce underlying associations and
patterns. There are two main approaches in clustering, either hierarchical such as HCA (hierarchical cluster analysis) or non-hierarchical
methods. However, the latter approach (non-hierarchical) is more applicable in the integration of multiple omics especially using machine
learning algorithms, such as the k-means clustering and random forest (Ma et al., 2014; Silva et al., 2019). k-means clustering groups
available data points (in this case from the omics expression data) such that clear, distinctive groupings emerge to di�erentiate expression
patterns. Meanwhile, random forest classi�es a group of genes/proteins/metabolites based on prior training data sets (from omics
experiments) to associate them to a particular characteristic/trait of interest (Ma et al., 2014). These techniques have been used widely in
plant multi-omics research.
For instance, Keller and Simm (2018) reported two modes of protein translation when comparing transcriptome and proteome of tomato
pollen development under either control or heat stress condition. The study employed the k-means clustering approach (Table 1), which
clustered expressed transcripts and proteins to di�erent clusters according to developmental stages. This has revealed the underlying
mechanism for protein translation; one that signi�cantly correlated between transcript-protein pair at one particular stage (direct translation)
or if certain proteins were only di�erentially expressed in the next stage after their corresponding DET at one stage (delayed translation). The
latter phenomenon may explain the weak correlation between transcript-protein pairs at certain stages of pollen development, primarily
those proteins related to carbohydrate and energy metabolism. Furthermore, upon heat stress, heat-shock proteins were regulated mostly at
the translational level (synthesis and degradation) rather than transcription, suggesting immediate plant response toward stresses (Keller and
Simm, 2018).
In addition, k-means clustering approach has also been used to integrate proteomics and metabolomics (Table 1) from developing cacao
seeds (Wang et al., 2016) and grape fruits (Wang et al., 2017). Such integration successfully identi�ed stage-speci�c clusters, whereby
secondary metabolites such as �avonoids were found concomitantly increased with the upregulation of corresponding biosynthetic
enzymes (Wang et al., 2016; Wang et al., 2017). Interestingly, Granger causality network analysis performed by Wang et al. (2017) on co-
regulated clusters further revealed signi�cant time-shift correlation between protein and metabolite pairs in grapes. This suggests that
protein abundance may be directly responsible for metabolic modulation during fruit development and ripening (Wang et al., 2017)
highlighting the importance of systematic MOI in elucidating key regulatory elements in plants.
In another study by Acharjee et al. (2016), a random forest approach was utilized to cluster and correlate transcriptomics, metabolomics, and
proteomics data sets against certain potato tuber phenotypic traits (�esh color, shape, starch gelatinization, and discoloration after peeling).
Interestingly, this study revealed that the di�erent omics was associated strongly with the di�erent tuber traits (Acharjee et al., 2016). For
example, traits related to color were more likely to be correlated to metabolite data (such as carotenoids) whereas tuber shape was
in�uenced strongly by transcripts related to size. This implies that certain omics are more suited to reveal the underlying mechanism of a
certain phenotypic or experimental condition. Hence, it is important to choose the most suitable omics platform for any investigation,
especially those related to phenotypic changes for relatable and descriptive results.
Multivariate Analysis
Multivariate analysis can handle more complex omics data sets, while allowing greater �exibility in experimental design and metadata analysis
(Rai et al., 2017). This approach enables the user to predict di�erent aspects or trends of data sets, including the discovery of variance or
covariance associations (Meng et al., 2014) as well as investigating the dynamic relationships and topological networks between
transcript/protein/metabolite elements (Weckwerth, 2019). Among the most common multivariate techniques are principal component
analysis (PCA), partial least squares (PLS) and orthogonal projection to latent structures discriminant analysis (OPLS-DA) (Mamat et al., 2018;
Mazlan et al., 2018; Reinke et al., 2018). Selecting di�erent multivariate techniques, optimal parameters, and model validation can be
overwhelming for new users, and hence several reading materials on this topic provide excellent learning resources (Tabachnick et al., 2007;
Meng et al., 2014; Saccenti and Timmerman, 2016).
Recently, Obudulu et al. (2018) performed OnPLS (multiple-block orthogonal projections to latent structures), an extension of the OPLS
technique, to integrate transcriptomics, proteomics, and metabolomics of poplar transgenic plants lacking PttSCAMP3 (Populus tremula x
tremuloides Secretory Carrier-Associated Membrane Protein3) gene, potentially important for wood development. Evidently, several
biomarkers related to the wood formation and secondary cell wall components have been successfully documented using this approach.
Other forms of multivariate analyses, such as MCIA (multiple co-inertia analysis) and GFLASSO (graph-guided fused least absolute shrinkage
and selection operator) have also been applied in multi-omics plant studies. For instance, MCIA was used to integrate metabolome and
proteome of a near-isogenic maize line (control) and its transgenic counterpart (glyphosate-tolerant maize, NK603) (Mesnage et al., 2016).
The study successfully identi�ed metabolic di�erences between the two, in particular, sugar metabolism and polyamine biosynthesis
(Mesnage et al., 2016). Another study in maize further illustrates the use of multivariate analysis such as GFLASSO to integrate transcriptome
and metabolome in deciphering its lipid biosynthesis (De Abreu E Lima et al., 2018).
3 of 19 10/18/22, 6:35 PM
About us All journals All articles Submit your research Search Login
Frontiers in Plant Science Sections Articles Research Topics Editorial Board About journal
Download Article 13,430 total views (https://www.altmetric.com
View Article Impact
SHARE
/details.php?domain=www.frontiersin.org& ON
citation_id=84836336)
REVIEW article
Front. Plant Sci., 26 June 2020
Sec. Plant Systems and Synthetic Biology
https://doi.org/10.3389/fpls.2020.00944 (https://doi.org/10.3389/fpls.2020.00944)
Systematic Multi-Omics Integration (MOI) Approach in Plant Systems Biology
Ili Nadhirah Jamil (https://www.frontiersin.org/people/u/971527)1, Juwairiah Remali (https://www.frontiersin.org/people/u/1009029)1,
Kamalrul Azlan Azizan (https://www.frontiersin.org/people/u/1009037)1,
Nor Azlan Nor Muhammad (https://www.frontiersin.org/people/u/795994)1, Masanori Arita (https://www.frontiersin.org/people/u/13070)2,3,
Hoe-Han Goh (https://www.frontiersin.org/people/u/113859)1 and Wan Mohd Aizat (https://www.frontiersin.org/people/u/387348)1*
1
Institute of Systems Biology (INBIOSIS), Universiti Kebangsaan Malaysia (UKM), Bangi, Malaysia
2 Bioinformation & DDBJ Center, National Institute of Genetics (NIG), Mishima, Japan
3 Metabolome Informatics Team, RIKEN Center for Sustainable Resource Science, Yokohama, Japan
Across all facets of biology, the rapid progress in high-throughput data generation has enabled us to perform multi-omics systems biology
research. Transcriptomics, proteomics, and metabolomics data can answer targeted biological questions regarding the expression of
transcripts, proteins, and metabolites, independently, but a systematic multi-omics integration (MOI) can comprehensively assimilate,
annotate, and model these large data sets. Previous MOI studies and reviews have detailed its usage and practicality on various organisms
including human, animals, microbes, and plants. Plants are especially challenging due to large poorly annotated genomes, multi-
organelles, and diverse secondary metabolites. Hence, constructive and methodological guidelines on how to perform MOI for plants are
needed, particularly for researchers newly embarking on this topic. In this review, we thoroughly classify multi-omics studies on plants and
verify work�ows to ensure successful omics integration with accurate data representation. We also propose three levels of MOI, namely
element-based (level 1), pathway-based (level 2), and mathematical-based integration (level 3). These MOI levels are described in relation
to recent publications and tools, to highlight their practicality and function. The drawbacks and limitations of these MOI are also discussed
for future improvement toward more amenable strategies in plant systems biology.
Introduction
The acquisition of multi-omics data sets has become an integral component of modern molecular biology and biotechnology. This is due to
technological advancements, such as the next-generation sequencing technology (Illumina, PacBio, and Nanopore) and mass spectrometry
coupled with gas- and liquid chromatography, which o�er high-throughput data generation (Fondi and Liò, 2015). The core data sets of
systems biology are transcriptomics, proteomics, and metabolomics, providing the expression levels of transcripts, proteins, and metabolites,
respectively (Aizat et al., 2018a). The data generated from these platforms can be massive, often without clear connections between them.
For instance, it is nearly impossible to manually associate hundred thousands of transcripts to their respective proteins or metabolic
pathways. In fact, the bottleneck of omics research is considered as the biological/machine/human resource allocation for data processing
and integration (Palsson and Zengler, 2010). What we need is a well-de�ned methodological scheme for multi-omics integration (MOI) to
extract, combine, and critically associate di�erent data sets to allow researchers to decipher the seemingly complex biological results at hand
(Fondi and Liò, 2015; Hughes, 2015; Wang et al., 2018).
MOI approach has been extensively studied and reviewed in studies on human (Chen and Chen, 2019; Cho et al., 2019; Shetty et al., 2019),
animals (García-Sevillano et al., 2014), microbes (Denman et al., 2018; Gutleben et al., 2018; Wang et al., 2019), and their combinations
(Wanichthanarak et al., 2015; Cavill et al., 2016; Pinu et al., 2019). In comparison, MOI in plants has been more di�cult due to their metabolic
diversity, poorly annotated large genomes (particularly for non-model species), and the presence of numerous symbionts with complex
interaction networks. Several comprehensive reviews are available speci�cally on plant MOI (Fukushima et al., 2009; Fukushima et al., 2014;
Rajasundaram and Selbig, 2016; Rai et al., 2017; Rai et al., 2019) and its practical usage in green systems biology, precision plant breeding,
and other biotechnological applications (Weckwerth, 2011; Weckwerth, 2019; Weckwerth et al., 2020). However, the advancement of high-
throughput technologies and large omics data sets leading to big data biology can be overwhelming, and perhaps an “Achilles' heel” for
inexperience researchers. Omics data from poorly characterized species are often feed into software without proper manual curation and
oblivious of the limitation of each technology, which could result in incorrect interpretations. Further, there are also a large collection of
software platforms, statistical rigor, and modeling (Weckwerth, 2011; Pinu et al., 2019) which must be selected appropriately by users, yet
these can be viewed as extraneous to untrained researchers. Hence, suitable methodological work�ow for MOI must be identi�ed to ensure
accurate large-scale data analysis and representation. Previously, the di�erent levels of MOI have been summarized as “conceptual,”
“statistical,” and “model-based” integration (Cavill et al., 2016; Rai et al., 2017); instances of such integration have been detailed elsewhere (de
Oliveira Dal'Molin and Nielsen, 2013; de Oliveira Dal'Molin and Nielsen, 2018; Seaver et al., 2018). Let us start from a critical review on the
1 of 19 10/18/22, 6:35 PM
,Frontiers | Systematic Multi-Omics Integration (... https://www.frontiersin.org/articles/10.3389/fpls....
previous three-level classi�cation.
About us All journals All articles Submit your research Search Login
The “conceptual” integration refers to multiple omics data sets being analyzed separately and are matched without further statistical analysis.
Even though this approach can produce valuable insights, it may miss reproducible associations when multiple omics data sets are analyzed
Frontiers in Plant
together Science
(Cavill et al., 2016;Sections
Rai et al.,
2017). We therefore
Articles argueTopics
Research that the conceptual integration
Editorial Board Aboutisjournal
an arbitrary
connection without proper
analysis and should not be considered a part of MOI approach. Instead, we re-classify the “statistical” integration where statistical
associations are sought between elements from di�erent data sets (Cavill et al., 2016; Rai et al., 2017). The e�ective use of prior knowledge is
separated as the pathway-based integration from unbiased, element-based integration. Finally, “model-based” integration is also re-classi�ed
so that qualitative reconstruction of biological pathways or systematic regulatory pathways is separated from their quantitative, mathematical
evaluation to generate working models for hypothesis testing (Thiele and Palsson, 2010; Rai et al., 2017).
Thus, we re-de�ne the MOI work�ow into three main integration levels (levels 1 to 3) with increasing complexity (Figure 1). Level 1 is the
unbiased, “element-based” integration with three subclasses: correlation, clustering, and multivariate analyses. Level 2 is the knowledge-
based “pathway” integration, which includes co-expression and mapping-based approaches. Finally, level 3 is the “mathematical” integration
with two subclasses, namely, di�erential and genome-scale analyses. The three levels are discussed in relation to recent omics reports from
the year 2014 to 2020 (Table 1) gathered from comprehensive literature searches, including Web of Science, Scopus, and Google Scholar
databases providing an updated comprehensive overview of MOI applications in various plant systems. Furthermore, this review focuses on
expression-based omics from transcriptomics, proteomics, and metabolomics to further clarify strategies taken to integrate such large-scale
expression data (transcript, protein, and metabolite).
Figure 1
www.frontiersin.org FIGURE 1 Current approaches in multi-omics integration (MOI) of plant systems biology. This MOI strategy is
(https://www.frontiersin.org/�les classi�ed into three main levels with increasing degrees of complexity.
/Articles/540561/fpls-11-00944-
HTML-r1/image_m/fpls-11-00944-
g001.jpg)
Table 1
www.frontiersin.org TABLE 1 Summary of recent publications and their tools and methods used in the three di�erent levels of
(https://www.frontiersin.org/�les multi-omics integration (MOI).
/Articles/540561/fpls-11-00944-
HTML-r1/image_m/fpls-11-00944-
t001.jpg)
Level 1 MOI: Element-Based Approach
Correlation Analysis
The �rst level of MOI is an element-based integration approach speci�cally using correlation analysis (Figure 1). The advantage of this
integration is its simplicity and intuitiveness. The standard approach is correlative association between two or more di�erent omics data sets
(i.e., transcriptomics, proteomics, and metabolomics data sets). Such analysis is performed using Pearson's (Benesty et al., 2009) and
Spearman's correlation coe�cients (Myers and Sirois, 2004), which assess linear and ranked relationships, respectively. Other studies also
analyzed their omics data sets using Fisher's transformation to transform skewed data sets to normally distributed data for calculating
corresponding correlation coe�cients (Mata et al., 2018). In general, signi�cant positive or negative coe�cient suggests strong direct or
inverse relationship between data sets, respectively.
Correlation-based MOI has been performed between transcripts and their cognate proteins. Such analysis is straightforward, assuming that
di�erential expression of transcripts will also be observed at their translational (protein) level. However, this is often not the case, most studies
reported weak correlations (di�erent patterns) between transcript-protein levels. For instance, salt treatment on salt-tolerant Earlistaple 7 and
salt-sensitive Nan Dan Ba Di Da Hua cotton revealed scarce correlation (r=0.03) between transcript and corresponding protein patterns,
regardless of genotypic background (Peng et al., 2018). Another example includes methyl jasmonate (MeJA) stress hormone treatment on
Persicaria minor Huds. herbal plants, with poor overall proteome-transcriptome correlation (r=0.341) (Aizat et al., 2018b). Similarly, transcripts
and proteins related to ethylene pathway (ethylene receptors [ETRs] and downstream signaling proteins, constitutive triple response-like
proteins [CTRs], and ethylene insensitive 2 [EIN2]) were not well correlated during the ripening process of tomato (Solanum lycopersicum)
(Mata et al., 2018). This suggests the existence of post-transcriptional and post-translational regulation (such as proteasomal degradation) for
the majority components of stress and ripening pathways. Despite transcriptome can be weakly correlated to proteome, it serves as an
excellent database for protein identi�cation in proteomics informed by transcriptomics approach for non-model plants (Aizat et al., 2018b;
Wan Zakaria et al., 2019) as well as studying allele-speci�c expression (van Wesemael et al., 2018).
On the other hand, an interesting emerging pattern arises when transcript-protein is compared between speci�c protein groups. For
example, signi�cantly upregulated proteins were positively correlated with their cognate transcripts in the stress response of various plants
(Ye et al., 2017; Aizat et al., 2018b). Speci�cally, proteins related to defense such as proteases and peroxidases in MeJA-treated P. minor (Aizat
et al., 2018b) and secondary metabolite biosynthesis such as �avonoid in phytoplasma-infected Ziziphus jujuba Mill. leaf (Ye et al., 2017) were
upregulated coherently with their transcripts. This may suggest the concerted molecular upregulation of defense-related proteins to
overcome stress signals and infection. Meanwhile, proteins related to growth such as photosynthetic and structural proteins were
signi�cantly suppressed in these studies, perhaps as a response toward the stress signal to conserve energy and recycling molecular
2 of 19 10/18/22, 6:35 PM
, Frontiers | Systematic Multi-Omics Integration (... https://www.frontiersin.org/articles/10.3389/fpls....
resources (Ye et al., 2017; Aizat et al., 2018b). Interestingly, such downregulation was mainly observed at the protein level, but not the
transcript levelAbout
(Aizatusetal., 2018b), perhaps asAll
All journals a mechanism
articles toSubmit
quicklyyour
resume protein synthesis, when the stress is relieved. However,
research Searchwe Login
could not rule out the possibility that changes at both transcript and protein levels are not simultaneous. Even when sampling of both is done
at the same time, the translational and post-translational degradation and modi�cation rates may di�er among proteins. However, this is
Frontiers in Plant
often Sciencefrom the
unpredictable Sections Articles
genome sequences aloneResearch Topics
(Weckwerth, 2019; Editorial Board About journal
Weckwerth et al., 2020), convoluting meaningful and direct
interpretation between expression data.
While comparisons between transcripts and corresponding proteins are generally performed in multi-omics studies, correlation of these two
with metabolites are relatively fewer. Perhaps, one such recent example is the transcriptomics and metabolomics investigation of Ginkgo
biloba during leaf maturity process (Guo et al., 2020). Correlation analysis was performed in this study between all di�erentially expressed
transcript (DET) and metabolites (DEM), however with no regard for their biochemical pathway relationships. While this study may be
interested only in the pattern consistency between DET and DEM, the corresponding biochemical pathway should be considered before
such correlation is performed. Importantly, metabolites should be classi�ed as either being substrates or products of certain enzymatic
pathways to be accurately correlated with their corresponding transcripts/proteins. This has been performed by Silva et al. (2017) for
transcripts and associated metabolites to elucidate primary metabolism in Arabidopsis seed germination and growth.
Clustering Analysis
Clustering analysis allows grouping of omics data sets with similar attribute such as expression levels to deduce underlying associations and
patterns. There are two main approaches in clustering, either hierarchical such as HCA (hierarchical cluster analysis) or non-hierarchical
methods. However, the latter approach (non-hierarchical) is more applicable in the integration of multiple omics especially using machine
learning algorithms, such as the k-means clustering and random forest (Ma et al., 2014; Silva et al., 2019). k-means clustering groups
available data points (in this case from the omics expression data) such that clear, distinctive groupings emerge to di�erentiate expression
patterns. Meanwhile, random forest classi�es a group of genes/proteins/metabolites based on prior training data sets (from omics
experiments) to associate them to a particular characteristic/trait of interest (Ma et al., 2014). These techniques have been used widely in
plant multi-omics research.
For instance, Keller and Simm (2018) reported two modes of protein translation when comparing transcriptome and proteome of tomato
pollen development under either control or heat stress condition. The study employed the k-means clustering approach (Table 1), which
clustered expressed transcripts and proteins to di�erent clusters according to developmental stages. This has revealed the underlying
mechanism for protein translation; one that signi�cantly correlated between transcript-protein pair at one particular stage (direct translation)
or if certain proteins were only di�erentially expressed in the next stage after their corresponding DET at one stage (delayed translation). The
latter phenomenon may explain the weak correlation between transcript-protein pairs at certain stages of pollen development, primarily
those proteins related to carbohydrate and energy metabolism. Furthermore, upon heat stress, heat-shock proteins were regulated mostly at
the translational level (synthesis and degradation) rather than transcription, suggesting immediate plant response toward stresses (Keller and
Simm, 2018).
In addition, k-means clustering approach has also been used to integrate proteomics and metabolomics (Table 1) from developing cacao
seeds (Wang et al., 2016) and grape fruits (Wang et al., 2017). Such integration successfully identi�ed stage-speci�c clusters, whereby
secondary metabolites such as �avonoids were found concomitantly increased with the upregulation of corresponding biosynthetic
enzymes (Wang et al., 2016; Wang et al., 2017). Interestingly, Granger causality network analysis performed by Wang et al. (2017) on co-
regulated clusters further revealed signi�cant time-shift correlation between protein and metabolite pairs in grapes. This suggests that
protein abundance may be directly responsible for metabolic modulation during fruit development and ripening (Wang et al., 2017)
highlighting the importance of systematic MOI in elucidating key regulatory elements in plants.
In another study by Acharjee et al. (2016), a random forest approach was utilized to cluster and correlate transcriptomics, metabolomics, and
proteomics data sets against certain potato tuber phenotypic traits (�esh color, shape, starch gelatinization, and discoloration after peeling).
Interestingly, this study revealed that the di�erent omics was associated strongly with the di�erent tuber traits (Acharjee et al., 2016). For
example, traits related to color were more likely to be correlated to metabolite data (such as carotenoids) whereas tuber shape was
in�uenced strongly by transcripts related to size. This implies that certain omics are more suited to reveal the underlying mechanism of a
certain phenotypic or experimental condition. Hence, it is important to choose the most suitable omics platform for any investigation,
especially those related to phenotypic changes for relatable and descriptive results.
Multivariate Analysis
Multivariate analysis can handle more complex omics data sets, while allowing greater �exibility in experimental design and metadata analysis
(Rai et al., 2017). This approach enables the user to predict di�erent aspects or trends of data sets, including the discovery of variance or
covariance associations (Meng et al., 2014) as well as investigating the dynamic relationships and topological networks between
transcript/protein/metabolite elements (Weckwerth, 2019). Among the most common multivariate techniques are principal component
analysis (PCA), partial least squares (PLS) and orthogonal projection to latent structures discriminant analysis (OPLS-DA) (Mamat et al., 2018;
Mazlan et al., 2018; Reinke et al., 2018). Selecting di�erent multivariate techniques, optimal parameters, and model validation can be
overwhelming for new users, and hence several reading materials on this topic provide excellent learning resources (Tabachnick et al., 2007;
Meng et al., 2014; Saccenti and Timmerman, 2016).
Recently, Obudulu et al. (2018) performed OnPLS (multiple-block orthogonal projections to latent structures), an extension of the OPLS
technique, to integrate transcriptomics, proteomics, and metabolomics of poplar transgenic plants lacking PttSCAMP3 (Populus tremula x
tremuloides Secretory Carrier-Associated Membrane Protein3) gene, potentially important for wood development. Evidently, several
biomarkers related to the wood formation and secondary cell wall components have been successfully documented using this approach.
Other forms of multivariate analyses, such as MCIA (multiple co-inertia analysis) and GFLASSO (graph-guided fused least absolute shrinkage
and selection operator) have also been applied in multi-omics plant studies. For instance, MCIA was used to integrate metabolome and
proteome of a near-isogenic maize line (control) and its transgenic counterpart (glyphosate-tolerant maize, NK603) (Mesnage et al., 2016).
The study successfully identi�ed metabolic di�erences between the two, in particular, sugar metabolism and polyamine biosynthesis
(Mesnage et al., 2016). Another study in maize further illustrates the use of multivariate analysis such as GFLASSO to integrate transcriptome
and metabolome in deciphering its lipid biosynthesis (De Abreu E Lima et al., 2018).
3 of 19 10/18/22, 6:35 PM