Novogene AMEA
  • Novogene AMEA
  • Genomics
    • Human Whole Genome Sequencing
    • Plant and Animal Whole Genome Sequencing
    • Microbial Whole Genome Sequencing
    • Plant and Animal De novo Sequencing
    • Microbial De novo Sequencing
    • Shotgun Metagenomics Sequencing
    • Amplicon Sequencing
    • Whole Exome Sequencing
    Transcriptomics
    • mRNA Sequencing
    • Total RNA Sequencing
    • Full-Length Transcriptome Sequencing
    • Whole Transcriptome Sequencing
    • Small RNA Sequencing
    • Circular RNA Sequencing
    • Metatranscriptome Sequencing
    • Prokaryotic RNA Sequencing
    Single Cell & Spatial Omics
    • Single Cell Gene Expression
    • Single Cell Immune Profiling Sequencing
    • Single Cell Long Read Transcriptome
    • Visium HD Spatial Gene Expression
    • Stereo-Seq Spatial Gene Expression
    • Xenium In Situ Spatial Transcriptome
    Epigenomics
    • Whole Genome Bisulfite Sequencing (WGBS)
    • Directed DNA Methylation Sequencing (DM-Seq) NEW
    • Reduced Representation Bisulfite Sequencing (RRBS)
    • Chromatin Immunoprecipitation Sequencing (ChIP-seq)
    • RNA Immunoprecipitation Sequencing (RIP-seq)
    • Assay for Transposase-Accessible Chromatin with Sequencing (ATAC-seq)

    Premade Library

    • Sequencing Only on Illumina Sequencer
    • Sequencing Only on PacBio Sequencer
    Proteomics and Metabolomics
    • Olink Proteomics
    • Quantitative Proteomics
    • Untargeted Metabolomics
  • PromotionsPromotions
    • Platforms
    • Automated Delivery Platform (Falcon)
    • Bioinformatics Analysis Tool (NovoMagic)
    • Customer Service System (CSS)
    • Brochures
    • Case Studies
    • Webinar
    • Blog
    • Sample Guidelines
    • Community
    • Cancer Research
    • Immuno-oncology
    • Agrigenomics
    • Environment
    • Food Science
    • Human Microbiome
    • Plant and Animal Microbiome
    • Drug Discovery and Development
    • Rare and Complex Diseases
    • About Us
    • Our Locations
    • News
    • Careers
  • Contact UsContact Us
  1. Home
  2. Resources
  3. Blog
  4. Advancing Untargeted Metabolomics with NovoMetDB-UM V4.0

Advancing Untargeted Metabolomics with NovoMetDB-UM V4.0

Expanded database coverage to support more confident metabolite annotation

Jiaxin Geng, 26 August 2026

Untargeted metabolomics can detect thousands of molecular features in a biological sample, but determining what those signals represent remains a significant challenge. Limited reference spectra, incomplete chemical coverage and retention-time variability across analytical conditions can leave many detected features without confident annotation. For researchers, this can constrain pathway interpretation, biomarker discovery and the development of a clear biological picture.

To help address this challenge, Novogene continues to expand its metabolite reference resources. NovoMetDB-UM V4.0 is the latest version of Novogene’s proprietary database for untargeted metabolomics analysis. It brings together a high-quality secondary spectral database including an in-house standard database (20,000+ standards) secondary spectrum library, and an AI-predicted spectral library across more than 500,000 total libraries, ensuring comprehensive and high-quality search results.

For researchers, this expanded coverage can support more confident metabolite annotation and provide a stronger foundation for translating detected molecular features into biologically meaningful findings.

What is NovoMetDB-UM?

NovoMetDB-UM is Novogene’s proprietary database used within its untargeted metabolomics analysis workflow. Detected molecular features are compared with database records using available evidence such as retention time, MS1 and MS2 information. Under Novogene’s annotation framework, matches are classified into four levels:

  • Level 1: Retention time, MS1 and MS2 match to an authentic reference standard—the highest annotation level
  • Level 2: MS1 and MS2 spectral match
  • Level 3.1: MS1-only match
  • Level 3.2: MS1 and MS2 match to an AI-predicted spectral record

This framework allows researchers to distinguish annotations supported by authentic reference standards from those based on spectral-library or computationally predicted evidence.

Novogene Untargeted Metabolomics Annotation Framework
Fig 1. International Gold Standard for Metabolite Identification
What has changed in V4.0?
NovoMetDB-UM V4.0 brings together more than 500,000 database entries across three complementary resources:
  • More than 20,000 authentic reference standards: Novogene’s in-house library contains experimentally measured retention time, MS1 and MS2 data, supporting Level 1 annotation.
  • More than 300,000 curated public records: Data integrated from established resources, including HMDB, GNPS, MoNA and MiMeDB.
  • More than 200,000 AI-predicted spectral records: Predicted spectra that extend coverage to compounds without available experimental reference spectra.
Novogene’s in-house authentic reference-standard library has expanded from more than 5,000 compounds in 2024 to more than 20,000 in V4.0. This provides more opportunities to match detected features against experimentally measured reference data.
The curated public and AI-predicted libraries extend coverage beyond compounds represented by available authentic standards, supporting candidate annotation across less-characterised areas of the metabolome.
Authentic reference standards support the highest-confidence annotations, while computational prediction extends coverage to features not yet represented in experimental spectral libraries. Computational evidence complements, rather than replaces, experimental reference data.
Extending coverage with AI-predicted spectra
V4.0 includes more than 200,000 predicted spectral records generated using graph neural networks and the FIORA model. These records extend coverage to compounds without available experimental reference spectra and support Level 3.2 candidate annotations.
V4.0 also incorporates an AI classifier trained using more than 500 reference compounds to distinguish metabolites associated with animal and plant sources. The classifier achieved a reported accuracy of at least 90% under Novogene’s validation conditions.
What does this mean for researchers?
The value of an expanded database lies not only in the number of entries it contains, but also in the quality and diversity of the evidence supporting each annotation. NovoMetDB-UM V4.0 may provide researchers with:
  • Broader coverage across the metabolome
  • More opportunities for Level 1 matching against authentic reference standards
  • Additional candidate annotations for less-characterised compounds
  • Clearer differentiation between experimental and computational evidence
  • A stronger foundation for pathway analysis and biological interpretation
In Novogene’s internal datasets, the Untargeted Metabolomics Plus workflow produced more than 1,200 Level 1 annotations on average and more than 5,000 total annotations.
Statistics of detected features in Untargeted Metabolomics Plus
Fig 2. Untargeted Metabolomics PlusHQ achieved over 1,200 Level 1 annotations and more than 5,000 total annotations on average across diverse sample types in Novogene’s internal datasets.
Statistics of detected features in Untargeted Metabolomics Pro
Fig 3. The Untargeted Metabolomics HQ Pro workflow produced more than 1,500 Level 1 annotations on average, with selected datasets exceeding 2,000 Level 1 annotations and 9,000 total annotations.
Results will vary according to sample type, biological matrix, metabolite composition and data quality. These internal results should not be interpreted as guaranteed annotation numbers for every study.
Supporting consistency across large-cohort studies
Database coverage is only one part of a reliable metabolomics workflow. Analytical stability and batch management are particularly important when samples are processed across extended periods. Novogene’s workflow incorporates:
  • 8 quality-control dimensions, multiple QC samples monitoring.
  • 4 batch-correction algorithms are evaluated in parallel, with performance assessed across eight indicators to select the most appropriate correction method.
  • Workflow stability assessment across 20 weeks of continuous instrument operation.
Together, these measures are designed to support reproducibility and comparability in large-cohort studies.
Untargeted Metabolomics Sample and Data Quality Control
Fig 4. Sample and Data Quality Control
Untargeted Metabolomics Sample Quality Control (QC) Fig 5. Sample Quality Control (QC)
From detected features to biological understanding
Untargeted metabolomics will continue to reveal more molecular features than can be conclusively identified. Addressing this challenge requires authentic reference measurements, carefully curated spectral information and computational approaches that extend beyond currently characterised compounds.
By combining these resources while preserving the distinction between annotation levels, NovoMetDB-UM V4.0 provides researchers with broader coverage and a more transparent foundation for interpreting metabolomics data.
To discuss how NovoMetDB-UM V4.0 may support your study, contact your local Novogene representative.

ServicesServices menu

SupportSupport menu

CompanyCompany menu

Services
Whole Genome SequencingDe novo SequencingAmplicon SequencingShotgun Metagenomic SequencingDirected DNA Methylation Sequencing (DM-Seq)mRNA SequencingSingle Cell Gene ExpressionVisium HD Spatial Gene ExpressionXenium In Situ Spatial TranscriptomeOlink ProteomicsUntargeted Metabolomics
Support
NovoMagic Bioinformatics Analysis ToolCustomer Service SystemFalcon Intelligent Delivery Platform
Company
About UsOur LocationsOur PlatformsNewsCareersContact Us
LinkedInLinkedIn hoverYouTubeYouTube hoverXX hover
Copyright © 2026 Novogene Inc. All rights reserved.For Research Use Only. Not for Clinical Diagnostic Use.
Novogene AMEA
  • Novogene AMEA
  • Genomics
    • Human Whole Genome Sequencing
    • Plant and Animal Whole Genome Sequencing
    • Microbial Whole Genome Sequencing
    • Plant and Animal De novo Sequencing
    • Microbial De novo Sequencing
    • Shotgun Metagenomics Sequencing
    • Amplicon Sequencing
    • Whole Exome Sequencing
    Transcriptomics
    • mRNA Sequencing
    • Total RNA Sequencing
    • Full-Length Transcriptome Sequencing
    • Whole Transcriptome Sequencing
    • Small RNA Sequencing
    • Circular RNA Sequencing
    • Metatranscriptome Sequencing
    • Prokaryotic RNA Sequencing
    Single Cell & Spatial Omics
    • Single Cell Gene Expression
    • Single Cell Immune Profiling Sequencing
    • Single Cell Long Read Transcriptome
    • Visium HD Spatial Gene Expression
    • Stereo-Seq Spatial Gene Expression
    • Xenium In Situ Spatial Transcriptome
    Epigenomics
    • Whole Genome Bisulfite Sequencing (WGBS)
    • Directed DNA Methylation Sequencing (DM-Seq) NEW
    • Reduced Representation Bisulfite Sequencing (RRBS)
    • Chromatin Immunoprecipitation Sequencing (ChIP-seq)
    • RNA Immunoprecipitation Sequencing (RIP-seq)
    • Assay for Transposase-Accessible Chromatin with Sequencing (ATAC-seq)

    Premade Library

    • Sequencing Only on Illumina Sequencer
    • Sequencing Only on PacBio Sequencer
    Proteomics and Metabolomics
    • Olink Proteomics
    • Quantitative Proteomics
    • Untargeted Metabolomics
  • PromotionsPromotions
    • Platforms
    • Automated Delivery Platform (Falcon)
    • Bioinformatics Analysis Tool (NovoMagic)
    • Customer Service System (CSS)
    • Brochures
    • Case Studies
    • Webinar
    • Blog
    • Sample Guidelines
    • Community
    • Cancer Research
    • Immuno-oncology
    • Agrigenomics
    • Environment
    • Food Science
    • Human Microbiome
    • Plant and Animal Microbiome
    • Drug Discovery and Development
    • Rare and Complex Diseases
    • About Us
    • Our Locations
    • News
    • Careers
  • Contact UsContact Us
  1. Home
  2. Resources
  3. Blog
  4. Advancing Untargeted Metabolomics with NovoMetDB-UM V4.0

Advancing Untargeted Metabolomics with NovoMetDB-UM V4.0

Expanded database coverage to support more confident metabolite annotation

Jiaxin Geng, 26 August 2026

Untargeted metabolomics can detect thousands of molecular features in a biological sample, but determining what those signals represent remains a significant challenge. Limited reference spectra, incomplete chemical coverage and retention-time variability across analytical conditions can leave many detected features without confident annotation. For researchers, this can constrain pathway interpretation, biomarker discovery and the development of a clear biological picture.

To help address this challenge, Novogene continues to expand its metabolite reference resources. NovoMetDB-UM V4.0 is the latest version of Novogene’s proprietary database for untargeted metabolomics analysis. It brings together a high-quality secondary spectral database including an in-house standard database (20,000+ standards) secondary spectrum library, and an AI-predicted spectral library across more than 500,000 total libraries, ensuring comprehensive and high-quality search results.

For researchers, this expanded coverage can support more confident metabolite annotation and provide a stronger foundation for translating detected molecular features into biologically meaningful findings.

What is NovoMetDB-UM?

NovoMetDB-UM is Novogene’s proprietary database used within its untargeted metabolomics analysis workflow. Detected molecular features are compared with database records using available evidence such as retention time, MS1 and MS2 information. Under Novogene’s annotation framework, matches are classified into four levels:

  • Level 1: Retention time, MS1 and MS2 match to an authentic reference standard—the highest annotation level
  • Level 2: MS1 and MS2 spectral match
  • Level 3.1: MS1-only match
  • Level 3.2: MS1 and MS2 match to an AI-predicted spectral record

This framework allows researchers to distinguish annotations supported by authentic reference standards from those based on spectral-library or computationally predicted evidence.

Novogene Untargeted Metabolomics Annotation Framework
Fig 1. International Gold Standard for Metabolite Identification
What has changed in V4.0?
NovoMetDB-UM V4.0 brings together more than 500,000 database entries across three complementary resources:
  • More than 20,000 authentic reference standards: Novogene’s in-house library contains experimentally measured retention time, MS1 and MS2 data, supporting Level 1 annotation.
  • More than 300,000 curated public records: Data integrated from established resources, including HMDB, GNPS, MoNA and MiMeDB.
  • More than 200,000 AI-predicted spectral records: Predicted spectra that extend coverage to compounds without available experimental reference spectra.
Novogene’s in-house authentic reference-standard library has expanded from more than 5,000 compounds in 2024 to more than 20,000 in V4.0. This provides more opportunities to match detected features against experimentally measured reference data.
The curated public and AI-predicted libraries extend coverage beyond compounds represented by available authentic standards, supporting candidate annotation across less-characterised areas of the metabolome.
Authentic reference standards support the highest-confidence annotations, while computational prediction extends coverage to features not yet represented in experimental spectral libraries. Computational evidence complements, rather than replaces, experimental reference data.
Extending coverage with AI-predicted spectra
V4.0 includes more than 200,000 predicted spectral records generated using graph neural networks and the FIORA model. These records extend coverage to compounds without available experimental reference spectra and support Level 3.2 candidate annotations.
V4.0 also incorporates an AI classifier trained using more than 500 reference compounds to distinguish metabolites associated with animal and plant sources. The classifier achieved a reported accuracy of at least 90% under Novogene’s validation conditions.
What does this mean for researchers?
The value of an expanded database lies not only in the number of entries it contains, but also in the quality and diversity of the evidence supporting each annotation. NovoMetDB-UM V4.0 may provide researchers with:
  • Broader coverage across the metabolome
  • More opportunities for Level 1 matching against authentic reference standards
  • Additional candidate annotations for less-characterised compounds
  • Clearer differentiation between experimental and computational evidence
  • A stronger foundation for pathway analysis and biological interpretation
In Novogene’s internal datasets, the Untargeted Metabolomics Plus workflow produced more than 1,200 Level 1 annotations on average and more than 5,000 total annotations.
Statistics of detected features in Untargeted Metabolomics Plus
Fig 2. Untargeted Metabolomics PlusHQ achieved over 1,200 Level 1 annotations and more than 5,000 total annotations on average across diverse sample types in Novogene’s internal datasets.
Statistics of detected features in Untargeted Metabolomics Pro
Fig 3. The Untargeted Metabolomics HQ Pro workflow produced more than 1,500 Level 1 annotations on average, with selected datasets exceeding 2,000 Level 1 annotations and 9,000 total annotations.
Results will vary according to sample type, biological matrix, metabolite composition and data quality. These internal results should not be interpreted as guaranteed annotation numbers for every study.
Supporting consistency across large-cohort studies
Database coverage is only one part of a reliable metabolomics workflow. Analytical stability and batch management are particularly important when samples are processed across extended periods. Novogene’s workflow incorporates:
  • 8 quality-control dimensions, multiple QC samples monitoring.
  • 4 batch-correction algorithms are evaluated in parallel, with performance assessed across eight indicators to select the most appropriate correction method.
  • Workflow stability assessment across 20 weeks of continuous instrument operation.
Together, these measures are designed to support reproducibility and comparability in large-cohort studies.
Untargeted Metabolomics Sample and Data Quality Control
Fig 4. Sample and Data Quality Control
Untargeted Metabolomics Sample Quality Control (QC) Fig 5. Sample Quality Control (QC)
From detected features to biological understanding
Untargeted metabolomics will continue to reveal more molecular features than can be conclusively identified. Addressing this challenge requires authentic reference measurements, carefully curated spectral information and computational approaches that extend beyond currently characterised compounds.
By combining these resources while preserving the distinction between annotation levels, NovoMetDB-UM V4.0 provides researchers with broader coverage and a more transparent foundation for interpreting metabolomics data.
To discuss how NovoMetDB-UM V4.0 may support your study, contact your local Novogene representative.

ServicesServices menu

SupportSupport menu

CompanyCompany menu

Services
Whole Genome SequencingDe novo SequencingAmplicon SequencingShotgun Metagenomic SequencingDirected DNA Methylation Sequencing (DM-Seq)mRNA SequencingSingle Cell Gene ExpressionVisium HD Spatial Gene ExpressionXenium In Situ Spatial TranscriptomeOlink ProteomicsUntargeted Metabolomics
Support
NovoMagic Bioinformatics Analysis ToolCustomer Service SystemFalcon Intelligent Delivery Platform
Company
About UsOur LocationsOur PlatformsNewsCareersContact Us
LinkedInLinkedIn hoverYouTubeYouTube hoverXX hover
Copyright © 2026 Novogene Inc. All rights reserved.For Research Use Only. Not for Clinical Diagnostic Use.
Privacy PolicyCookie PolicyCareers
Privacy PolicyCookie PolicyCareers