NTH

Graph Learning on Ensembles of Cyclic Peptides: An Investigation of Molecular Ensemble Modeling

AuthorsAaron Feller, Kris Deibler, Maxim Secor

July 29, 2026 2 min read
Watch on YouTube
The one-line take

A pretrained geometric model that treats cyclic peptides as conformational ensembles improves property prediction over sequence-only modeling.

Key results

5
Conformers retained per peptide

Top Boltzmann-weighted conformers used for each molecular ensemble

2979
Evaluation set size

Filtered CREMP-CycPeptMPDB measurements used in 5-fold cross-validation

0.005
Random-init EnsembleEGNN R2

Downstream performance without CREMP pretraining

0.477
Pretrained EnsembleEGNN R2

Performance after self-supervised geometric pretraining

0.538
Hybrid model R2

Best result from co-training EnsembleEGNN with PeptideCLM-2

0.737
Hybrid model Pearson r

Correlation achieved on held-out membrane-permeability prediction

What the paper found

Researchers at the University of Texas at Austin and Novo Nordisk introduce EnsembleEGNN, a molecular ensemble foundation model designed for cyclic peptides whose solution behavior depends on multiple Boltzmann-weighted conformers rather than one static structure. The model applies shared E(n)-Equivariant Graph Neural Network layers to each conformer, then combines representations using Boltzmann-informed attention and a Set Attention Block; experiments retain the top 5 conformers per peptide. Self-supervised pretraining on CREMP jointly performs masked atom-token recovery, noisy-coordinate denoising, and pairwise-distance reconstruction. On the 2979-example CREMP-CycPeptMPDB membrane-permeability benchmark, training EnsembleEGNN from scratch nearly fails, producing R2 = 0.005, while pretrained EnsembleEGNN reaches R2 = 0.477 and Pearson r = 0.699, exceeding the sequence-only BERT baseline at R2 = 0.439. Co-training the geometric encoder with the PeptideCLM-2 sequence encoder yields the strongest Hybrid model, with R2 = 0.538 and Pearson r = 0.737. The 11.3M-parameter Hybrid outperforms the 114M-parameter BERT-style architecture, suggesting that explicit ensemble geometry can improve data-efficient cyclic-peptide property prediction, although the study remains limited by a low-data benchmark and the absence of direct cross-conformer message passing.

Original abstract

Molecular property prediction from structure often uses a single representative conformation, even though many molecules exist as conformational ensembles in solution. We introduce EnsembleEGNN, a molecular ensemble foundation model that encodes an ensemble by first encoding each conformer with shared Equivariant Graph Neural Network (EGNN) layers, then pooling the resulting conformer representations with a Set Attention Block. We pretrain the model on CREMP, a cyclic peptide ensemble dataset, using a multi-task self-supervised objective combining masked token recovery, noisy-coordinate reconstruction, and pairwise distance reconstruction. On the CREMP-CycPeptMPDB dataset, training EnsembleEGNN from scratch fails entirely ($R^2=0.005$). However, the pretrained model reaches $R^2=0.477$ and Pearson $r=0.699$, outperforming the sequence-only BERT baseline ($R^2=0.439$, Pearson $r=0.667$). When EnsembleEGNN is co-trained end-to-end with the BERT sequence encoder, the hybrid model improves further to $R^2=0.538$ and Pearson $r=0.737$. These results demonstrate that encoding conformational ensembles into a single thermodynamically informed embedding improves cyclic-peptide property prediction.

Read the original paper

More in Graph Learning

Browse all 32 papers →
02Graph Learning

GraphWrit3R: End-to-End 3D Scene Graph Writing

Luka Milivojevic, Nikola Popovic, Sayan Deb Sarkar, Sebastian Koch, Iro Armeni, Luc Van Gool, Danda Pani Paudel

GraphWrit3R turns 3D spatial data into open-vocabulary scene graphs using multimodal encoders and an LLM, without requiring ground-truth object annotations at inference.

Read analysis