Researchers Develop ConfSeq to Encode 3D Molecular Conformations as Language

Artificial intelligence (AI) is reshaping the research paradigm across scientific disciplines. Central to this transformation are language models (LMs), which learn complex patterns from large-scale token sequences through self-supervision and exhibit remarkable generative capabilities. In pharmaceutical and chemical research, chemical language models (CLMs) have emerged as a pivotal focus. By encoding molecular topological structures as discrete token sequences, such as SMILES and SELFIES, CLMs enable state-of-the-art (SOTA) performance in two-dimensional (2D) molecular modeling tasks, including molecular translation, de novo generation, and representation learning. However, these notations describe only the two-dimensional topology of molecules. The three-dimensional (3D) conformation of a molecule, which ultimately governs its physical, chemical, and biological behavior, has remained largely beyond the reach of language models.

To address these challenges, a research team led by ZHENG Mingyue and ZHANG Sulin at the Shanghai Institute of Materia Medica, Chinese Academy of Sciences, published an article in Nature Machine Intelligence. This study designed a new conformation description language called ConfSeq and established ConfSeq as a robust framework for extending LLMs in 3D molecular modeling.

ConfSeq encodes molecular conformations using three key intrinsic geometric elements, dihedral angles, bond angles, and pseudo-chirality. These elements are discretized into tokens and strategically integrated into the SMILES framework. This design preserves human readability while naturally guaranteeing SE(3) invariance, and it enables LMs to learn geometry–structure relationships by associating geometric tokens with their atomic and bonding context. Leveraging ConfSeq, the team reformulated a range of 3D molecular modeling tasks, including conformation prediction, unconditional and shape-conditioned 3D molecular generation, and 3D representation learning, as a unified sequence-modeling problem. Using only standard Transformer architectures, ConfSeq achieved SOTA performance across multiple benchmarks. Compared with prevailing graph diffusion models, ConfSeq offers both substantially faster inference and an intrinsic confidence-scoring mechanism for prioritizing high-quality structures. In real-world drug discovery, the team used ConfSeq-derived 3D representations to perform ligand-based virtual screening, identifying several novel STING and ALDH1B1 inhibitors with IC₅₀ values ranging from 0.338 to 3.51 μM.

This study challenges the prevailing view that language models are inherently limited in 3D molecular tasks. By establishing a robust bridge between language models and 3D molecular structures, ConfSeq lays a foundation and opens a new direction for AI-driven molecular modeling, drug discovery, and molecular design.

An overview of ConfSeq, which encodes 3D molecular structures into discrete token sequences and unifies conformation prediction, 3D molecular generation, and 3D molecular representation learning as sequence-modeling tasks (Image by ZHENG Mingyue’s group)


DOI: 10.1038/s42256-026-01250-8

Link: https://www.nature.com/articles/s42256-026-01250-8

Keywords: Conformation description language; 3D molecular representation; Drug design


Contact:

DIAO Wentong

Shanghai Institute of Materia Medica

E-mail: diaowentong@simm.ac.cn