Advancing Untargeted Metabolomics with NovoMetDB-UM V4.0
Expanded database coverage to support more confident metabolite annotation
Jiaxin Geng, 26 August 2026Untargeted metabolomics can detect thousands of molecular features in a biological sample, but determining what those signals represent remains a significant challenge. Limited reference spectra, incomplete chemical coverage and retention-time variability across analytical conditions can leave many detected features without confident annotation. For researchers, this can constrain pathway interpretation, biomarker discovery and the development of a clear biological picture.
To help address this challenge, Novogene continues to expand its metabolite reference resources. NovoMetDB-UM V4.0 is the latest version of Novogene’s proprietary database for untargeted metabolomics analysis. It brings together a high-quality secondary spectral database including an in-house standard database (20,000+ standards) secondary spectrum library, and an AI-predicted spectral library across more than 500,000 total libraries, ensuring comprehensive and high-quality search results.
For researchers, this expanded coverage can support more confident metabolite annotation and provide a stronger foundation for translating detected molecular features into biologically meaningful findings.
What is NovoMetDB-UM?
NovoMetDB-UM is Novogene’s proprietary database used within its untargeted metabolomics analysis workflow. Detected molecular features are compared with database records using available evidence such as retention time, MS1 and MS2 information. Under Novogene’s annotation framework, matches are classified into four levels:
- Level 1: Retention time, MS1 and MS2 match to an authentic reference standard—the highest annotation level
- Level 2: MS1 and MS2 spectral match
- Level 3.1: MS1-only match
- Level 3.2: MS1 and MS2 match to an AI-predicted spectral record
This framework allows researchers to distinguish annotations supported by authentic reference standards from those based on spectral-library or computationally predicted evidence.
Fig 1. International Gold Standard for Metabolite Identification
-
More than 20,000 authentic reference standards: Novogene’s in-house library contains experimentally measured retention time, MS1 and MS2 data, supporting Level 1 annotation.
-
More than 300,000 curated public records: Data integrated from established resources, including HMDB, GNPS, MoNA and MiMeDB.
-
More than 200,000 AI-predicted spectral records: Predicted spectra that extend coverage to compounds without available experimental reference spectra.
-
Broader coverage across the metabolome
-
More opportunities for Level 1 matching against authentic reference standards
-
Additional candidate annotations for less-characterised compounds
-
Clearer differentiation between experimental and computational evidence
-
A stronger foundation for pathway analysis and biological interpretation
Fig 2. Untargeted Metabolomics PlusHQ achieved over 1,200 Level 1 annotations and more than 5,000 total annotations on average across diverse sample types in Novogene’s internal datasets.
Fig 3. The Untargeted Metabolomics HQ Pro workflow produced more than 1,500 Level 1 annotations on average, with selected datasets exceeding 2,000 Level 1 annotations and 9,000 total annotations.
-
8 quality-control dimensions, multiple QC samples monitoring.
-
4 batch-correction algorithms are evaluated in parallel, with performance assessed across eight indicators to select the most appropriate correction method.
-
Workflow stability assessment across 20 weeks of continuous instrument operation.
Fig 4. Sample and Data Quality Control
Fig 5. Sample Quality Control (QC)