This Week In Cheminformatics: Issue #029
M-JEPA, reaction network reduction, MolRank, how to find interesting molecules, automated parallel synthesis, and a long list of papers
Highlights
Automated Parallel Synthesis Accelerates Virtual Screening Hit Discovery
McKenna, Šícho, and colleagues built COMBINAUT, a repurposed peptide synthesizer running SPS workflows (amide/sulfonamide coupling, reductive amination, SNAr, nitro reduction, heterocyclization) across 715 in-house building blocks, enumerating a 22.9M compound virtual library and physically producing hits within 32h instead of the 4-6 week make-on-demand turnaround !! They report that not all but 78/100 selected VS hits actually got synthesized to spec, and only 9 of those were validated in radioligand binding against the CCR2 allosteric pocket (12% hit rate, pretty good !). They also try to justify these numbers in the paper carefully. They then used the same platform for rapid SAR on two hits, getting 4-6x potency improvements in a few hours of parallel synthesis per round, ending with two cell-active antagonists (IC50 256 nM and 1.3 µM in β-arrestin assays). It’s a modest library by ultra-large VS standards (22.9M vs billions), but the closed-loop docking followed by synthesis, assay, re-docking cycle all in hours rather than weeks is the actual contribution here, and it’s pretty amazing if you think about it.
M-JEPA: Predictive Self-Supervised Learning for Molecular Graphs with Scaffold-Shift Evaluation on Tox21
M-JEPA ports the JEPA recipe (context encoder, EMA teacher, latent-space prediction instead of contrastive InfoNCE) to molecular graphs, using BFS-grown connected-subgraph masks as the target region. Under matched compute, JEPA beats a matched InfoNCE baseline on ESOL proxy RMSE and trains faster, and hybrid fine-tuning from the JEPA checkpoint lifts Tox21 mean ROC-AUC only modestly (0.561 to 0.609) but drops mean ECE substantially (0.279 to 0.064) which is a useful reminder that ranking and calibration under scaffold “shift” are separable properties and gains in one don’t guarantee the other. Their connected vs noncontiguous masking ablation was statistically non-signigicant on both proxy RMSE and motif-level coherence, so the “chemically coherent” argument for BFS masking seems basically aesthetic.
Graph-Based Generation and Reduction of Complex Chemical Reaction Networks
Here, Fite & Gross add a neat methodological contribution to the reaction network reduction problem. Most reduction algorithms are validated on a handful of published mechanisms, which tells you little about generality. The authors build a preferential attachment based generator fitted to ammonia, methane, and hydrogen combustion mechanisms, producing 2,613 artificial networks with sampled thermodynamics and Arrhenius kinetics, to get a large, unbiased test set. On top of that they introduce MolRank, a PageRank style topological algorithm that estimates upper bounds on maximal species concentration and reaction rate using only thermodynamic data. MolRank is markedly better at identifying redundant species than redundant reactions. They also report Eyring-Polanyi based scores give better bounds than true rate constants but don’t preserve ranking quality.
Strategies for Identifying Molecules of Interest in Large Chemical Spaces
Klein et al. (BioSolveIT + Pfizer) benchmark SpaceMACS, SpaceLight, and FTrees across 2,917 CHEMBL queries on REAL Space, eXplore, and PfizerVL, and the headline result is that the three methods barely overlap under 1.2% shared hits across all three for a given query. MCS similarity, fingerprint (ECFP4/fCSFP) similarity, and FTrees pharmacophore similarity are sampling quite different neighborhoods of a chemical space, so which one and at what thresholds should you use them becomes the real question. This comparison was done with similarity scored on synthons rather than fully enumerated products. New bond formed at assembly for instance are invisible to standard ECFP4, which is why fCSFP was built with smaller, non-circular features to survive the synthon boundary better (recovering 90-100% of close analogs vs 70-95% for ECFP4). This paper gives you calibrated per-algorithm recommended thresholds for very similar, similar, and dissimilar compounds from search results (check Figure 16). Good read.
Long List
Cheminformatics
Transition Mechanism of the Human Dopamine Transporter via Molecular Dynamics Simulations
Enhancing Generalization in Synthesizability Prediction of Structurally Dissimilar Materials
Circumventing the Synthesizability Problem in Generative Molecular Design
A Numerical Implementation to Calculate Elastic Properties of Biological Membrane Simulations
Navigating chemical-linguistic sharing space with heterogeneous molecular encoding
The Chemical Smiler: CHEMICal Abstraction Leading to SMILEs of Reference
Mapping Evolution of Molecules across Biochemistry with Assembly Theory
Bridging between Structure-Based and Data-Driven Affinity Prediction
Tightly coupled equivariant flow matching for molecular docking with multimodal physical constraints
SpaBiT: enhancing spatial transcriptomics resolution via bidirectional attention transformers
IDR searcher: a search engine solution for public image resources
Meta2DB: curated shotgun metagenomic feature sets and metadata for health state prediction
Partial Least Squares Regression Models for Outdoor Air Pollutant Forecasting
Guiding generative models to uncover diverse and novel crystals via reinforcement learning
ScrambleBench: a workflow for comparative assessment of structure-based de novo generative models
MINERVA: a public XAI-powered platform advancing multi-target discovery in Alzheimer’s disease
Smiles-based bioactivity prediction through molecular encoder selection and data augmentation
TPS-Flow: Physics-Guided Flow-Based Generative Modeling of Protein Transition Paths
Collective Variable-Guided Engineering of the Free-Energy Surface of a Small Peptide
Can LLMs Solve Solubility Tasks? The SoluBench Benchmark for Pure and Mixed Solvent Systems
Deep Learning Models Capture Umbrella Sampling-Derived Energetic Trends: A Troponin C Case Study
RetroMPA: A Molecular Property-Aware Auxiliary Framework for Enhancing Retrosynthesis Prediction
Insights into the Indexing of Powder X-ray Diffraction from a Robust Transformer Deep Learning
Protein Language Model-Based Fitness Estimates Facilitate Resistance Mutation Identification
Protein Frustration Reveals Orthosteric and Allosteric Active Sites in GPCR:G Protein Complexes
Discovery of novel cinnamic acid derivatives with anti-Helicobacter pylori mechanism
A linear models approach to optimize carbazole-based dyes for solar cell applications
DOCKweb: a web-based GUI platform for molecular modeling with DOCK6
MedChem
From inhibition to degradation: advances in highly selective targeting strategies for HPK1
Targeted Degradation of KAT6A: Expanding the Therapeutic Frontier of Epigenetic Drug Discovery
Riding toward Selectivity: Optimization of Covalent 7-Azaindole-Based BMX Kinase Inhibitors
Palate Cleanser
Best,
Manas























