Jesper Løve Hinrich, Pia Susan Mayer, Bekzod Khakimov +2q-bio.QM cs.LG
Overlapping peaks and sample-dependent chemical shift variability prevent reliable metabolite recovery from complex biological spectra. This problem is critical in one-dimensional proton (1D 1H) NMR which has become the standard method providing fast acquisition and information-rich spectra in metabolomics and foodomics. This study demonstrates how chemical shifts can be utilised as a strength in 1D 1H NMR, when suitably modeled through the proposed Bayesian Shift-Invariant Non-negative Matrix Factorization (BSI-NMF) procedure. We find that BSI-NMF accurately recovers the underlying chemical signals in 1D 1H NMR spectra missed by existing analyses approaches across simulations, laboratory created datasets, and a large urine dataset obtained from 2439 people across Europe. Our study highlights how shifts in the chemical signatures - until now perceived as a nuisance - can in fact when suitably modelled be instrumental for unique recovery of metabolites. This creates an opportunity to experimentally induce chemical shifts changes to facilitate unique recovery of spectra.
Zheqi Jin, Grace Armitage, Richard Cox +3physics.chem-ph cs.LG
Here, we present a platform built on our inverted Graph Transformer Network, IMPRESSION-G2, which can accurately and rapidly reconstruct molecular bonding directly from experimental nuclear magnetic resonance (NMR) spectroscopic information. It comprises three interconnected stages: a one-shot model that predicts bond connectivity between atoms; a structure-correction stage that corrects the predicted structures by removing uncertain bonds and iteratively reassigning them; noise-augmented multi-shot prediction, generating an ensemble of candidate structures, which are ranked to identify the best-fit structure. By integrating a range of $^{1}$H and $^{13}$C NMR data, including two-dimensional (2D) experiments such as COSY, HSQC, and HMBC, the inverse-IMPRESSION platform correctly identifies the structures of 77.8% of molecules with up to 30 heavy atoms (H, C, N, O and F) using simulated NMR data, and 10 of 19 (53%) molecules using experimental NMR data. The experimental structures solved have molecular weights of up to 480 Da and are representative of the complex structures in synthetic and natural products that routinely challenge chemists. The inverse-IMPRESSION framework thus provides the first effective approach for automated molecular structure elucidation using graph-based machine learning on experimental data.
Nuclear Magnetic Resonance (NMR) spectroscopy is the gold standard for molecular structure elucidation, yet interpreting complex spectra for unknown molecules remains a bottleneck reliant on human expertise. While artificial intelligence has advanced this field, current methods face a critical trade-off: database retrieval cannot identify novel scaffolds, while de novo molecular structure elucidation models operate as black boxes, lacking the atom-level interpretability required for rigorous scientific validation. Here, we present NMRAgent, an evidential reasoning agent powered by large language models (LLMs) that bridges this gap by integrating specialized spectral analysis tools with chemical knowledge graphs. Unlike previous approaches, NMRAgent mimics the deductive reasoning of human experts: it takes experimental NMR spectra and molecular formula as input, plans the elucidation process, proposes candidate structures, verifies peak-atom consistency, and refines misaligned substructure through formula-aware fragment optimization. Enabled by its evidential reasoning, NMRAgent outperforms state-of-the-art methods, improving top-1 accuracy by 46.5% and Tanimoto similarity by 0.502 on a scaffold-split benchmark with novel scaffolds in the test set. Besides, we demonstrate the agent's practical utility by elucidating the structures of two previously unknown natural products isolated from Hydrangea davidii and Vitex trifolia, and by correcting structural misassignments in established literature. By combining high-accuracy prediction with transparent and evidence-based reasoning, NMRAgent establishes a new paradigm for interpretable AI in analytical chemistry.
Chen Yang, Zheng Fang, Hanyu Sun +7physics.chem-ph cs.AI cs.LG
Nuclear Magnetic Resonance (NMR) spectroscopy is a powerful tool for molecular structure analysis, and spectral artificial intelligence offers great potential for its rapid and automated interpretation. However, the scarcity of experimental NMR datasets has constrained deep learning in this domain to narrow, task-specific applications that lack broad generalization. Here, we introduce UltraNMR, a large-scale foundation model for NMR that leverages the intrinsic properties of NMR spectra to learn generalizable spectral representations. We collected 158 million paired simulated $^{1}$H and $^{13}$C NMR spectra to train UltraNMR, employing multiple domain-specific pre-training objectives. UltraNMR captures both intra-spectral and inter-spectral dependencies, enabling seamless simulation-to-real adaptation. We demonstrate that adapting UltraNMR to a range of molecular structure analysis tasks on experimental NMR spectra consistently yields state-of-the-art performance and clearly outperforms UltraNMR variants trained directly on downstream data without simulation pre-training. We also construct a large-scale NMR spectral vector library by encoding simulated NMR spectra using UltraNMR, covering 94 million unique molecules and enabling effective structure-aware retrieval. In real-world applications, UltraNMR facilitates the structural elucidation of two previously unknown natural products from Chinese herbal medicines recorded in the Chinese Pharmacopoeia. These results suggest that large-scale simulation pre-training can effectively bridge the simulation-to-real gap, enabling robust and generalizable molecular structure analysis of real-world NMR spectra.