Advancing Untargeted Metabolomics with NovoMetDB-UM V4.0
By Jiaxin Geng,
———-
Expanded database coverage to support more confident metabolite annotation
Untargeted metabolomics can detect thousands of molecular features in a biological sample, but determining what those signals represent remains a significant challenge. Limited reference spectra, incomplete chemical coverage and retention-time variability across analytical conditions can leave many detected features without confident annotation. For researchers, this can constrain pathway interpretation, biomarker discovery and the development of a clear biological picture.
To help address this challenge, Novogene continues to expand its metabolite reference resources. NovoMetDB-UM V4.0 is the latest version of Novogene’s proprietary database for untargeted metabolomics analysis. It integrates an in-house library containing more than 20,000 authentic reference standards, curated public spectral records and an AI-predicted spectral library, providing access to more than 500,000 database entries in total.
For researchers, this expanded coverage can support more confident metabolite annotation and provide a stronger foundation for translating detected molecular features into biologically meaningful findings.
What is NovoMetDB-UM?
NovoMetDB-UM is Novogene’s proprietary database used within its untargeted metabolomics analysis workflow. Detected molecular features are compared with database records using available evidence such as retention time, MS1 and MS2 information. Under Novogene’s annotation framework, matches are classified into four levels:
- Level 1: Retention time, MS1 and MS2 match to an authentic reference standard—the highest annotation level
- Level 2: MS1 and MS2 spectral match
- Level 3.1: MS1-only match
- Level 3.2: MS1 and MS2 match to an AI-predicted spectral record
This framework allows researchers to distinguish annotations supported by authentic reference standards from those based on spectral-library or computationally predicted evidence.

What’s new in NovoMetDB-UM V4.0?
NovoMetDB-UM V4.0 brings together more than 500,000 database entries across three complementary resources:
- More than 20,000 authentic reference standards:Novogene’s in-house library contains experimentally measured retention time, MS1 and MS2 data, supporting Level 1 annotation.
- More than 300,000 curated public records:Data integrated from established resources, including HMDB, GNPS, MoNA and MiMeDB.
- More than 200,000 AI-predicted spectral records:Predicted spectra that extend coverage to compounds without available experimental reference spectra.
Novogene’s in-house authentic reference-standard library has expanded from more than 5,000 compounds in 2024 to more than 20,000 in V4.0. This provides more opportunities to match detected features against experimentally measured reference data.
The curated public and AI-predicted libraries extend coverage beyond compounds represented by available authentic standards, supporting candidate annotation across less-characterised areas of the metabolome.
Authentic reference standards support the highest-confidence annotations, while computational prediction extends coverage to features not yet represented in experimental spectral libraries. Computational evidence complements, rather than replaces, experimental reference data.
Extending metabolite coverage with AI-predicted spectra
V4.0 includes more than 200,000 predicted spectral records generated using graph neural networks and the FIORA model. These records extend coverage to compounds without available experimental reference spectra and support Level 3.2 candidate annotations.
V4.0 also incorporates an AI classifier trained using more than 500 reference compounds to distinguish metabolites associated with animal and plant sources. The classifier achieved a reported accuracy of at least 90% under Novogene’s validation conditions.
Benefits of NovoMetDB-UM V4.0
The value of an expanded database lies not only in the number of entries it contains, but also in the quality and diversity of the evidence supporting each annotation. NovoMetDB-UM V4.0 may provide researchers with:
- Broader coverage across the metabolome
- More opportunities for Level 1 matching against authentic reference standards
- Additional candidate annotations for less-characterised compounds
- Clearer differentiation between experimental and computational evidence
- A stronger foundation for pathway analysis and biological interpretation
In Novogene’s internal datasets, the Untargeted Metabolomics Plus workflow produced more than 1,200 Level 1 annotations on average and more than 5,000 total annotations.


Results will vary according to sample type, biological matrix, metabolite composition and data quality. These internal results should not be interpreted as guaranteed annotation numbers for every study.
Quality control for large-cohort untargeted metabolomics studies
Database coverage is only one part of a reliable metabolomics workflow. Analytical stability and batch management are particularly important when samples are processed across extended periods. Novogene’s workflow incorporates:
- Eight quality-control dimensions and monitoring using multiple QC samples
- Four batch-correction algorithms evaluated in parallel, with performance assessed across eight indicators to select the most appropriate correction method
- Workflow stability assessment across 20 weeks of continuous instrument operation
Together, these measures are designed to support reproducibility and comparability in large-cohort studies.


From metabolite detection to biological interpretation
Untargeted metabolomics will continue to reveal more molecular features than can be conclusively identified. Addressing this challenge requires authentic reference measurements, carefully curated spectral information and computational approaches that extend beyond currently characterised compounds.
By combining these resources while preserving the distinction between annotation levels, NovoMetDB-UM V4.0 provides researchers with broader coverage and a more transparent foundation for interpreting metabolomics data.
Learn more
Interested in applying untargeted metabolomics to your research? Explore Novogene’s untargeted metabolomics services or contact our team to discuss your project.
———- Jiaxin Geng is a Product Manager at Novogene with expertise in mass spectrometry and metabolomics workflows. She shares practical insights to help researchers apply emerging metabolomics technologies to real-world scientific challenges.