This Week In Cheminformatics: Issue #030
How useful are Cavity Prediction Algorithms, NMR-AI, (Failing at) Predicting Alternative Binding Modes via Docking and a long list of papers...and maybe the last one of these from PT.
Highlights
Can Cavity Prediction Algorithms Help in Docking Experiments?
This paper from CCDC asks if your docking tool needs a binding site as input, and you don’t have one, is cavity prediction actually good enough to bridge that gap? They ran four cavity predictors (PUResNet-2.0, LVPocket, CAVIAR, Fpocket) on 486 curated protein-ligand structures and fed the best predicted cavity into GOLD, using four different ways of translating a cavity into a binding site definition (atom-vicinity vs. residue-vicinity, at 10 and 12 Å). Fpocket and CAVIAR come out ahead on cavity centroid accuracy (means of 3.5 and 4.2 Å vs. 6.6-6.9 Å for LVPocket/PUResNet-2.0), and centroid-based cavity definitions consistently beat residue-based ones for GOLD, likely because residue expansion drags in false-positive atoms that GOLD’s scoring can reward. The more interesting result is correlation between cavity accuracy and final pose accuracy drops to 0.49 when restricted to the two best cavity methods, meaning close-enough cavities don’t reliably translate to correct poses, and conversely some docking runs succeeded even with cavity centroids >9 Å off. They also benchmark against DiffDock and Boltz-1 blind docking, where GOLD+Fpocket/CAVIAR wins on mean distance but DiffDock has a lower median with heavier-tailed outliers (quite not the usual “AI blind docking dominates” narrative you see elsewhere). Good read !
NMR-AI: An Open Platform for NMR-Enhanced Molecular Representations and Physicochemical Property Prediction
Kurczab’s group keeps pushing the “spectra as descriptors” line, and this one tests the question that actually matters: is NMR redundant if one uses ECFP4, or does it add something ? They built SpectraPRINTS which is concatenation of binned predicted ¹H/¹³C shifts (400-dim, HOSE-code/nmrshiftdb2-derived) with ECFP4. They ran the identical CNN regression pipeline across logP, logS, logD at three pH values, and both macroscopic pKa endpoints, so any difference in RMSE is due to representation and not architecture. They observed that lipophilicity/solubility endpoints get consistent gains (16–39% RMSE reduction, largest at logD pH 10.5), while pKa shows no systematic improvement which interestingly they attribute to macroscopic pKa collapsing multiple protonation microstates/tautomers into one number, which makes sense. They also shipped it as a live webapp (link here: NMR-AI).
Benchmarking Docking Protocols on Predicting Alternative Binding Modes
Docking benchmarks have long leaned on the Astex diverse set and its ~2 Å RMSD success criterion, but this paper points out that criterion quietly assumes there’s one right answer to find. Kubincová et al. instead curate 60 PDB structures with alternate locations for the same ligand. These are crystallographically resolved binding-mode pairs, split into ring flips, ring puckers, and other rearrangements. They ask four templated docking protocols (RDKit/smina, rDock, PLANTS, AutoDock-GPU) to recover the non-input pose when given the other as a starting point. Success drops to 30–50% after pose selection, versus 70–90% on Astex (!!) even though the docking is templated and should have been easier. Worth reading if you rely on docking to generate starting poses for RBFE. The paper argues (and I think reasonably) that the 2 Å criterion from virtual-screening benchmarks doesn’t transfer to the pose-curation use case, where getting the wrong binding mode can cost you several kcal/mol downstream.
Long List
Cheminformatics
EvoDiffMol: evolutionary diffusion framework for 3D molecular design with optimized properties
Multi-omics network reconstruction with collaborative graphical lasso
Cross-domain transfer learning from peptides to metabolites using a multi-property fine-tuned LLM
From data to scent: validating an ensemble AI model that predicts novel insect odourant interactions
EDEL: enhancing dense retrievers for curation of biomedical knowledge bases
Bayesian Uncertainty-Guided Fidelity Fusion for Bioactivity Prediction
Iterative Interaction Fingerprints-Guided Multiobjective Molecular Generation
Gradient-Guided Graph Contrastive Learning for Mass Spectrometry-Based Proteomics Clustering
HELM-BERT: Topology-Aware Representations for Chemically Modified Peptides
BenzDB, an Exhaustive Database of Benzenoids up to Nine 6-Membered Rings
Unveiling the KRAS Relationship between Affinity and Dynamics: A Molecular Simulations Study
Familywise Feature Importance Stability in Chemical and Materials Machine Learning
Assessing Molecular Contacts Using Atom Environments Described by Ranked Lists
HELMify: A Hybrid Rule- and LLM-Based Generator of Peptide Monomer HELM Names
CatIF-RL: Activity-Oriented Enzyme Sequence Design by Steered Inverse Protein Folding
MedChem
Other
I made an app for Garmin watches with gentle reminders to be present in the moment, might be of interest to some : https://github.com/Manas02/HereAndNow
Palate Cleanser
Signing off one last time from PT ; )
Muito Obrigado,
Manas




























































