This Week In Cheminformatics: Issue #026
Strasbourg Summer School In Cheminformatics 2026, Properties of Macrocycles, De Novo A2A Hit and a long list of papers
Strasbourg Summer School in Chemoinformatics – 2026
Those who read this weekly would notice I’m very late on this post and that is because I’m currently attending the Summer School of Cheminformatics in Strasbourg. Before I write any more, I must thank all the organizers. Special note on Prof. Horvath and Prof. Varnek who started the conference with arguably the best introductory speech ever (10th edition special !!), I wish I could have half of their wit :)
We started this summer school on Friday with a hackathon. My group’s project was under Dr (& hopefully soon Prof.). Martin Šícho. We worked on integrating biased splits into QSPRPred along with some CLI bug fixes and addition of some dimensionally reduction plotting. It was a wonderful experience working with Martin and getting to know the fellow Erasmus students from different tracks.
I feel quite lucky and privileged meeting such talented fellow students working and studying amazing things and our EMJM Program and the coordinators who made it all possible (muito obrigado !!).
The best thing so far that we (the students) collectively did was to roam around the city on the Fête de la Musique, which I must say is an experience everyone should try once (hopefully in better weather condition). I won’t go into the fun part of the summer school here for the sake of brevity.
Monday started with our hackathon presentations. We then went into a brilliant tutorial on SMARTS by Hadi Vareno. This was followed by an amazing keynote by Prof. Matthias Rarey.
Today we started off with a keynote from Prof. Schneider on De novo molecular design. Prof. Volkamer followed this up with a talk on Data-Driven Exploration of Kinase Inhibitor Space and I must say it was packed with insights and the discussion really brought up fresh views and ideas about model validation, co-folding, TeachOpenCADD, and much more.
After a quick coffee break, the day just kept getting better with very interesting talks by Prof. José L. Medina-franco who presented Latin American Product Database amongst other projects, Mikhail Kabeshov from AstraZeneca presented some amazing work from MolecularAI team such as SMARTS-RX, Bonafide, etc. Arkadii Lin from Insilico medicine presented work on methods for reliable single-step retrosynthesis prediction, Yuliana Zabolotna from Eli Lilly presented SAR-Gate and how they are making it accessible to med chemists in the company, and Timur Madzhidov from Reaxys gave a wonderful recap on reaction condition prediction and what it takes to make it into a production ready utilitarian model. The last talk of the day was Prof. Marcou’s excellent tutorial on “Trustworthiness, the Key to Grid-Based Map-Driven Predictive Model Enhancement and Applicability Domain Control“ based on GTM (I never thought I’d like an applicability domain talk but his approach was quite honest and refreshing).
The day ended with a poster presentation session where I tried to “network“.
More on this next week. Anyways back to the usual stuff now…
Highlights
What Is in a Structure? Cell Permeability and Solubility of Series of Macrocycles and Linear Matched Pairs
Here, Tyagi et al. challenge the assumption that macrocyclization inherently improves cell permeability over linear structures. They evaluated series of 18 and 19-membered macrocycles against their linear matched pairs and observed that the linear analogues actually exhibited higher transcellular Caco-2 permeability. Using Monte Carlo conformational sampling and DFT optimizations, they demonstrated that the greater flexibility of the linear compounds allows them to have less polar conformations in membrane-like environments, effectively shielding their amide bonds via intramolecular hydrogen bonds far better than the structurally restricted macrocyclic rings. Additionally, within the macrocycle series itself, the formation of chameleonic NH-π interactions between specific side-chains and the backbone reduced the SASA, which increased permeability by an average of 3.5-fold. I highly recommend reading this work because it provides concrete evidence that optimizing beyond-Rule-of-5 chemical space requires us to move past static 2D property filters and use ensemble-based 3D conformational analysis to accurately capture these environment-dependent desolvation penalties.
Applying Deep-Learning-Driven De Novo Design to Hit Identification: A Case Study on A2A Adenosine Receptor Antagonists
Persico et al. use REINVENT to generate A2A adenosine receptor antagonists while higjlighting the practical limitations of purely score-driven reinforcement learning in de novo design. They started with a MPO scoring function with a QSAR model (CNS-MPO) for blood-brain barrier permeability, and GOLD docking scores. What makes this a compelling read is how they systematically troubleshooted their way to the hit. They start with a micromolar hit from the initial pool, but more importantly, they did a second generative loop that applies penalty for missing the Asn253 interaction. This worked quite well for them. It gave three structurally distinct compounds with nanomolar affinities. It is a grounded, practical case study showing something we already know and believe i.e. off-the-shelf docking scores are more often than not insufficient for generative design but knowing your target’s chemistry and hardcoding required target-specific pharmacophores directly into the optimization loop will most likely work.
Long List
Cheminformatics
EC-Dock: A Fast Equivariant Consistency Model for Molecular Docking and Virtual Screening
Finding Balance: Multiobjective Optimization in Molecular Generative Modeling
Learning High-Resolution Protein Embeddings from Multimodal Data via Self-Supervised Integration
PAIRMAP: A Unified Geometry-Aware Pairwise-Map Framework for Molecular Representation Learning
Energetics of Noncovalent Interactions of Protein–Ligand Complexes for Drug Discovery
Strategic Template Filtering Accelerates Fragment-Based Peptide Docking
Protein language visualizer: a repository for homology exploration with language model embeddings
Modular Framework for 3D Molecular Generation in Computational Chemistry Applications
MedChem
Other
Palate Cleanser
PS: It’s soooooooooooooooooooo hot here and I’m writing this at 3am so please forgive me for the typos, grammar, etc.
Best,
Manas























