# Chem(o)Info Radar article snapshot
> 27 records returned from 3050 current records.

Generated: 2026-09-11T22:25:22.065536+00:00
Filters: topic=reaction, days=30

Views are bounded by topic, source, days and limit; text search runs in the dashboard page, not on the server. Blog entries are metadata-only; use their source URL for full text.

## Enhancing Biocatalytic Retrosynthesis with a Graph-to-Graph Model
- Source: Chemical Science (journals)
- Date: 2026-09-09T08:03:49Z
- Categories: Reaction Informatics
- Authors: Binju Wang, Lina Dong, Lin Yao, Yucheng Yang, Yuxiang Gao, Zhihui Jiang
- Journal: Chemical Science
- DOI: 10.1039/d6sc04187f
- Keywords: Retrosynthesis
- Source URL: <https://doi.org/10.1039/d6sc04187f>
- Dashboard article: <https://tagirshin.com/radar/article?u=https%3A%2F%2Fdoi.org%2F10.1039%2Fd6sc04187f>

Abstract: Biocatalytic synthesis offers a green and sustainable route for chemical production, yet the rational design of biocatalytic routes remains challenging due to the need to jointly consider reaction feasibility and...

## TSBench: A physics-grounded benchmark for evaluating LLM understanding of chemical reaction mechanisms
- Source: arXiv (preprints)
- Date: 2026-09-08T09:44:37Z
- Categories: Reaction Informatics, LLMs & Agents
- Authors: Xiaohu Xu, Tong Zhu
- External ID: 2609.08503v1
- Keywords: LLM, LLMs, synthesis planning
- Source URL: <https://arxiv.org/abs/2609.08503v1>
- Dashboard article: <https://tagirshin.com/radar/article?u=https%3A%2F%2Farxiv.org%2Fabs%2F2609.08503v1>
- PDF: <https://arxiv.org/pdf/2609.08503v1>

Abstract: Understanding a chemical reaction requires mapping a symbolic reactant-product description onto the three-dimensional pathway through which atoms rearrange, yet chemistry benchmarks for large language models (LLMs) largely probe factual knowledge and text-based reasoning. Here we introduce TSBench, a benchmark in which an LLM agent uses structure-editing tools to construct three-dimensional transition-state (TS) guesses verified by an automated quantum-chemical pipeline, yielding a physics-grounded pass/fail verdict. Across 546 evaluations of seven frontier LLMs on 78 elementary reactions, the aggregate success rate rose from 50.4% to 66.8% under diagnosis-driven revision; the best models approached 90% on the simplest reactions, yet performance dropped sharply with mechanistic complexity. Most failed attempts produced locally plausible saddle points whose reaction paths led to the wrong reactant-product pair, revealing local geometric intuition without a robust grasp of the global reaction coordinate. TSBench establishes a mechanism-level yardstick for LLM agents in mechanism-sensitive tasks such as synthesis planning and autonomous experimentation.

## SynOmega: Simplifying Retrosynthesis for Efficient Synthesizability Scoring
- Source: Journal of Chemical Information and Modeling (journals)
- Date: 2026-09-08T00:00:00+00:00
- Categories: Cheminformatics, Reaction Informatics
- Authors: Baicheng Zhang, Guoqing Zhang, Jun Jiang, Yi Luo
- Journal: Journal of Chemical Information and Modeling
- DOI: 10.1021/acs.jcim.6c02529
- Keywords: AiZynthFinder, ChEMBL, Retrosynthesis
- Source URL: <https://doi.org/10.1021/acs.jcim.6c02529>
- Dashboard article: <https://tagirshin.com/radar/article?u=https%3A%2F%2Fdoi.org%2F10.1021%2Facs.jcim.6c02529>

Abstract: Retrosynthesis-based synthesizability scoring triages molecules from generative design but is expensive: every score requires a multi-step search. We present SynOmega, an open-source toolkit that couples a single-step template model, an AND–OR route search, and a route-based synthesizability score (SynScore). Its single-step model can be restricted, at the reaction-template level, to simplifying disconnections that split the target into smaller precursors. On 1000 ChEMBL drug molecules this yields two findings. (1) The simplifying constraint cuts node expansions by about 30% on jointly solved targets and lowers the median search time by about a third. This reduction in search effort is the robust, budget-independent result; the small accompanying rise in solved rate is a secondary effect between the two separately trained models, not the isolated result of toggling one model’s action space. (2) As a complete system, under matched search depth, width and iteration budget, SynOmega reaches about 1.8× the solved rate of the open-source planner AiZynthFinder while searching about 13× faster. SynOmega thus offers a cheap, data-level action-space constraint that makes route-based synthesizability scoring more efficient without sacrificing solvability.

## SynOmega: Simplifying Retrosynthesis for Efficient Synthesizability Scoring
- Source: Journal of Chemical Information and Modeling (journals)
- Date: 2026-09-08T00:00:00Z
- Categories: Reaction Informatics
- Journal: Journal of Chemical Information and Modeling
- DOI: 10.1021/acs.jcim.6c02529
- External ID: 905b8bebd010d0b02daa3843dcb82948a88b82ec
- Keywords: AiZynthFinder, ChEMBL, Retrosynthesis
- Source URL: <https://doi.org/10.1021/acs.jcim.6c02529>
- Dashboard article: <https://tagirshin.com/radar/article?u=https%3A%2F%2Fdoi.org%2F10.1021%2Facs.jcim.6c02529>

Abstract: Retrosynthesis-based synthesizability scoring triages molecules from generative design but is expensive: every score requires a multi-step search. We present SynOmega, an open-source toolkit that couples a single-step template model, an AND–OR route search, and a route-based synthesizability score (SynScore). Its single-step model can be restricted, at the reaction-template level, to simplifying disconnections that split the target into smaller precursors. On 1000 ChEMBL drug molecules this yields two findings. (1) The simplifying constraint cuts node expansions by about 30% on jointly solved targets and lowers the median search time by about a third. This reduction in search effort is the robust, budget-independent result; the small accompanying rise in solved rate is a secondary effect between the two separately trained models, not the isolated result of toggling one model’s action space. (2) As a complete system, under matched search depth, width and iteration budget, SynOmega reaches about 1.8× the solved rate of the open-source planner AiZynthFinder while searching about 13× faster. SynOmega thus offers a cheap, data-level action-space constraint that makes route-based synthesizability scoring more efficient without sacrificing solvability.

## A retrieval-augmented agent bridges virtual screening and automated synthesis
- Source: ChemRxiv (preprints)
- Date: 2026-09-07T00:00:00Z
- Categories: Design de novo, Reaction Informatics, LLMs & Agents
- Authors: Shan Zhu, Xin Chen, Pinqi Wang, Quan Jiang, Mengzhu Li, Qimeng Wang, Guoyun Zhao, Yuan Qi, Fenglei Cao, Li-Cheng Xu
- DOI: 10.26434/chemrxiv.15008409/v1
- External ID: 10.26434/chemrxiv.15008409/v1
- Keywords: virtual screening, retrosynthetic prediction, LLM, generative model, synthesis planning, retrosynthesis, retrosynthetic, reaction conditions
- Source URL: <https://doi.org/10.26434/chemrxiv.15008409/v1>
- Dashboard article: <https://tagirshin.com/radar/article?u=https%3A%2F%2Fdoi.org%2F10.26434%2Fchemrxiv.15008409%2Fv1>

Abstract: Translating virtually designed drug molecules into executable synthesis protocols is essential for connecting computational discovery with automated synthesis. However, most retrosynthesis methods focus on precursor generation, offer limited route diversity, and lack the detailed reaction conditions required for executable protocols. Here, we present Syn-RRAG, a retrieval-augmented LLM-driven synthesis agent connecting virtual screening with automated synthesis and validation. Syn-RRAG integrates initial retrosynthetic prediction, reaction-precedent retrieval and hierarchical refinement to generate structured protocols with complete experimental parameters. Its local generative model searches over complete reaction hypotheses rather than relying solely on token-level beam search, enabling exploration of feasible routes with distinct reaction classes and bond disconnections. It achieved strong retrosynthesis performance on two patent-derived reaction datasets. Syn-RRAG further generated executable plans for three virtually screened drug candidates, all of which were synthesized and validated on an automated platform. These results establish a workflow connecting virtual molecular design, synthesis planning and experimental validation.

## CampChem: Rubric-Grounded Adaptive Campaigns for Multi-Round Organic Synthesis with Residual Spectral Observation
- Source: ChemRxiv (preprints)
- Date: 2026-09-07T00:00:00Z
- Categories: Cheminformatics, Reaction Informatics, LLMs & Agents
- Authors: Zixuan Shi, Mengyao Qian, Yichao Sun, Haotian Gu, Jiarui Feng
- DOI: 10.26434/chemrxiv.15008453/v1
- External ID: 10.26434/chemrxiv.15008453/v1
- Keywords: SMILES, LLM, Mistral, reinforcement learning, retrosynthesis
- Source URL: <https://doi.org/10.26434/chemrxiv.15008453/v1>
- Dashboard article: <https://tagirshin.com/radar/article?u=https%3A%2F%2Fdoi.org%2F10.26434%2Fchemrxiv.15008453%2Fv1>

Abstract: Organic synthesis is a multi-round campaign, not a one-shot SMILES completion: a failed solvent, a poisoned catalyst, or leftover starting-material peaks in the 1H NMR force the chemist to revise conditions and try again. Tool-using chemistry agents and condition classifiers still emit a single recipe and then stop. We propose CampChem, an adaptive campaign agent built on a chemistry-instruction LLM. A longitudinal campaign adaptive planner (LCAP) keeps a short memory of failed condition slots and peak signatures and decodes the next catalyst/solvent/reagent tuple under a mask. Chemistry rubric-as-judge reinforcement learning (Chem-RaJ) scores stoichiometry, hazard compatibility, slot completeness, and spectral consistency instead of using a free-form LLM judge. A residual spectral booster (RSB) injects a permutation-invariant peak-set residual when NMR is available and gates to zero otherwise. On SMolInstruct, CampChem reaches 67.1% forward-synthesis and 36.1% retrosynthesis Top-1 exact match with a Mistral-7B LoRA backbone. On USPTO-Condition it attains 0.2940 overall Top-1, and on AutoLabs Experiment 5 it raises protocol F1 from 0.75 to 0.81. Ablations show that Chem-RaJ carries most of the spectrum-free gain, LCAP matters once a prior tuple has been rejected, and RSB contributes mainly on simulated NMR revision.

## Machine Learning for High-Throughput Reaction Yield Prediction in DNA-Encoded Library Synthesis
- Source: ChemRxiv (preprints)
- Date: 2026-09-07T00:00:00Z
- Categories: Property Prediction, Reaction Informatics, Library Design
- Authors: Tao Tang, Li Gao, Sen Gao, Hongyao Zhu, Guansai Liu, Jin Li, Xuemin Cheng
- DOI: 10.26434/chemrxiv.15008275/v2
- External ID: 10.26434/chemrxiv.15008275/v2
- Keywords: Reaction Yield Prediction, graph neural network, GNN, DNA Encoded Library, molecular representation, reaction type
- Source URL: <https://doi.org/10.26434/chemrxiv.15008275/v2>
- Dashboard article: <https://tagirshin.com/radar/article?u=https%3A%2F%2Fdoi.org%2F10.26434%2Fchemrxiv.15008275%2Fv2>

Abstract: DNA-encoded library (DEL) synthesis relies on efficient DNA-compatible reactions to ensure library fidelity and the reliability of downstream affinity selection. However, most existing reaction yield prediction models were developed for general organic synthesis and typically use full-atom representations of complete reactants and products. In DEL single-cycle synthesis validation, an appropriate model reaction is to focus on the exposed reactive handle, its local chemical environment, the incoming building block (BB), and the reaction type, rather than large constant backgrounds such as the DNA tag, linker, and pre-existing scaffold. Here, we propose a functional-group-centric machine learning framework for DEL single-cycle reaction yield prediction. The model uses only the exposed substrate functional group (FG) and its coarse-grained local environment, the incoming BB, and the reaction type as inputs. Benchmarking on a real-world high-throughput DEL single-cycle validation dataset shows that graph neural network (GNN) models based on this localized representation outperform traditional fingerprint-based baselines, with a graph attention network (GAT) achieving the best overall performance and showing strong robustness on highly imbalanced industrial data. These results indicate that foregrounding the local reaction center is not merely a simplification of molecular representation, but a task-aligned strategy that better reflects the chemistry and decision logic of DEL single-cycle transformations. This lightweight framework is readily compatible with practical DEL workflows and provides valuable supports on reagent quality control (QC), building block selection, high-fidelity library design, and retrospective analysis of anomalous results.

## Near-Complete Synthon Library via Retrosynthetic Analysis for Automated Organic Synthesis
- Source: ChemRxiv (preprints)
- Date: 2026-09-07T00:00:00Z
- Categories: Reaction Informatics
- Authors: Baicheng Zhang, Jing Chen, Xiaolong Zhang, Pieter E. S. Smith, Aoyuan Cheng, Runyang Miao, Hongping Liu, Linjiang Chen, Bin Jiang, Guoqing Zhang, Yi Luo, Jun Jiang
- DOI: 10.26434/chemrxiv.15008417/v1
- External ID: 10.26434/chemrxiv.15008417/v1
- Keywords: PubChem, Retrosynthetic, retrosynthesis
- Source URL: <https://doi.org/10.26434/chemrxiv.15008417/v1>
- Dashboard article: <https://tagirshin.com/radar/article?u=https%3A%2F%2Fdoi.org%2F10.26434%2Fchemrxiv.15008417%2Fv1>

Abstract: As idealized structural fragments for guiding bond disconnections, synthons are considered the conceptual backbone of retrosynthetic analysis. Yet retrosynthesis can be difficult because the choice of synthons requires extensive experience in organic synthesis, which prevents non-experts from accessing desired products. Here we address this knowledge gap by constructing a curated, machine interpretable database of commercially available synthon equivalents (hereafter valid synthons) distilled from large-scale retrosynthetic analysis. Mining over 40 million molecules from PubChem, we identify 1,617 essential valid synthons. These synthons, organized into (i) molecular backbones, (ii) functional groups, and (iii) linker motifs, were obtained via iterative retrosynthetic decomposition and stringent frequency based filtering tied to documented reaction precedents. The resulting set enables the in silico reconstruction of over 94% of these 40 million molecules. We demonstrate that such synthons reduce dependence on expansive reagent inventories, even though the routes themselves are non-conventional and algorithm-generated: using only seven representative valid synthons, we accessed three structurally and functionally divergent targets (melatonin, a natural sleep hormone; an azo dye; and a photoactive melatonin derivative), illustrating the breadth and practicality of a synthons to products methodology in the context of automated synthesis for non-experts in organic synthesis.

## Ranking-Based Surrogate Modeling for Bayesian Optimization under Small-Data Conditions
- Source: Journal of Chemical Information and Modeling (journals)
- Date: 2026-09-06T00:00:00+00:00
- Categories: Reaction Informatics
- Authors: Yuya Endo, Hiromasa Kaneko
- Journal: Journal of Chemical Information and Modeling
- DOI: 10.1021/acs.jcim.6c02115
- Source URL: <https://doi.org/10.1021/acs.jcim.6c02115>
- Dashboard article: <https://tagirshin.com/radar/article?u=https%3A%2F%2Fdoi.org%2F10.1021%2Facs.jcim.6c02115>

Abstract: Bayesian optimization (BO) is widely used for reaction optimization, but under small-data conditions, regression of absolute objective values may not align with the practical goal of prioritizing promising experiments. Here, we compared ranking-based BO (RankBO) combined with Thompson sampling (Rank-TS) with regression-based BO using two reaction benchmarks. Rank-TS showed a clear advantage on Direct Pd-catalyzed arylation in optimization performance, global ranking quality, and recovery of high-yielding conditions. For Suzuki–Miyaura coupling, which comprised 12 substrate combinations, its improvement was small in the pooled analysis. However, in the equal-weight macro analysis across the 12 substrate-defined search spaces, Rank-TS showed higher ranking quality and faster high-yield-condition recovery. These results indicate that Rank-TS is particularly useful for prioritizing reaction conditions within a fixed-substrate combination under small-data conditions, and that its relative performance depends on how the chemical search space and ranking task are defined.

## AbSynth: A Century of Total Syntheses as a Machine-Readable Resource for Analyzing Strategic Logic in Organic Chemistry
- Source: Journal of the American Chemical Society (journals)
- Date: 2026-09-02T00:00:00+00:00
- Categories: Reaction Informatics
- Authors: Maciej Czuba, Konrad Barnowski, Karol Molga, Zofia Pławecka, Agnieszka Wołos, Szymon Stodółkiewicz, Sebastian Baś, Wojciech Petrykowski, Wiktor Beker, Evan Saillard, Konrad Łęgowski, Rafał Roszak, Paweł Kowalczyk, Barbara Mikulak-Klucznik, Piotr Janasz, Emilia Mielke, Patrycja Pietruszewska, Maja Łazuchiewicz, Sylwia Wójcik, Jakub Zabielski, Dariusz Nostkiewicz, Gabriela Smaga, Ahmad Makkawi, Louis P. J.-L. Gadina, Jakub Skoczek, Tomasz Klucznik, Maxence Deschamps, Anna Żądło-Dobrowolska, Michał Michalak, Sara Szymkuć, Daniel T. Gryko, Yasemin Bilgi-Gadina, Bartosz A. Grzybowski
- Journal: Journal of the American Chemical Society
- DOI: 10.1021/jacs.6c08434
- Keywords: retrosynthetic algorithms, synthesis planning, retrosynthetic
- Source URL: <https://doi.org/10.1021/jacs.6c08434>
- Dashboard article: <https://tagirshin.com/radar/article?u=https%3A%2F%2Fdoi.org%2F10.1021%2Fjacs.6c08434>

Abstract: The design of synthetic routes to complex, stereochemically rich targets remains one of the most demanding challenges in organic chemistry. To interrogate the strategic logic underlying such efforts, we curated and digitized over 3,000 classical total syntheses comprising close to 60,000 individual steps, now freely available as the AbSynth collection. Analyses across this data set reveal that synthesis planning cannot be reduced to purely local, one-step-at-a-time decision-making: known complexity metrics and scoring functions fail to vary monotonically along synthetic trajectories, and conventional bond-disconnection heuristics apply only sporadically and lack the specificity required for automation. Instead, our study shows that nearly half of all steps correspond to structurally unproductive – yet strategically essential – plateaus of skeletal complexity. Navigating these plateaus requires retrosynthetic algorithms to embrace multistep reasoning rather than incremental scoring. We identify several such multistep heuristics, including tolerable lengths of nonskeletal disconnections and the “longevity” of functional groups across routes, and anticipate that AbSynth will serve as a resource for uncovering additional trends in synthesis design. By providing a century-spanning, machine-readable data set of complete routes, this work offers a foundation for advancing chemical AI toward the complexity of human expert planning.

## PKP-Diffmol: A Physicochemical Knowledge-Prompt Encoding and Latent Diffusion Framework for Molecular Property Prediction
- Source: Journal of Chemical Information and Modeling (journals)
- Date: 2026-09-01T00:00:00+00:00
- Categories: Cheminformatics, Property Prediction, Docking & Screening, ADMET & Safety, Reaction Informatics
- Authors: Ruizi Liu, Tongtong Yuan, Molin Guo
- Journal: Journal of Chemical Information and Modeling
- DOI: 10.1021/acs.jcim.6c01547
- Keywords: RDKit, virtual screening, MoleculeNet, SMILES, Transformer, Property Prediction, Molecular Property, bioactivity, Molecular representations
- Source URL: <https://doi.org/10.1021/acs.jcim.6c01547>
- Dashboard article: <https://tagirshin.com/radar/article?u=https%3A%2F%2Fdoi.org%2F10.1021%2Facs.jcim.6c01547>

Abstract: Molecular representations that capture both chemical information and property-relevant differences are essential for reliable molecular property prediction in drug discovery and toxicity assessments. Fragment-level SMILES representation learning is effective at modeling substructure semantics, but structural reconstruction objectives alone do not explicitly encode the descriptor-derived physicochemical information needed for downstream prediction, especially in small data and scaffold-split settings. To address this limitation, we propose PKP-DiffMol, a physicochemical knowledge-prompt encoding and latent diffusion framework built upon the pretrained SMI-EDITOR encoder. PKP-DiffMol improves fragment-level molecular representations by using numerical and semantic physicochemical knowledge. The numerical physicochemical prior branch encodes RDKit descriptors as continuous quantitative priors, while the physicochemical knowledge prompt encoder transforms the same descriptors into semantic physicochemical representations by using a Transformer-based text encoder. The two complementary views are then integrated through hierarchical physicochemical knowledge fusion to construct a chemically enriched, fused latent space. In this space, a quality-controlled conditional latent diffusion module learns label-conditioned distributions, generates synthetic molecular latent representations, and retains reliable samples through a Mahalanobis distance quality gate. Experiments on seven MoleculeNet classification datasets under scaffold splitting show that PKP-DiffMol increases the mean ROC-AUC from 77.80% for the SMI-EDITOR backbone to 80.53% and obtains the best result on five of the seven datasets. The largest gains are observed on BACE (+6.91%), MUV (+4.29%), and SIDER (+3.92%), showing strong predictive performance across bioactivity prediction, virtual screening, and adverse drug reaction prediction tasks. Further analyses indicate that numerical and semantic physicochemical knowledge improves class separability and the correspondence between representation distances and RDKit descriptor distances, while Mahalanobis-filtered label-conditioned latent augmentation supports prediction-head refinement.

## An Agentic Retrobiosynthesis Framework with Learned Frontier Selection
- Source: arXiv (preprints)
- Date: 2026-08-31T12:40:42Z
- Categories: Design de novo, Reaction Informatics, LLMs & Agents
- Authors: Philippe Meyer, Guillaume Gricourt, Thomas Duigou, Joan Hérisson, Jean-Loup Faulon
- External ID: 2608.30702v1
- Keywords: RL, retrosynthesis
- Source URL: <https://arxiv.org/abs/2608.30702v1>
- Dashboard article: <https://tagirshin.com/radar/article?u=https%3A%2F%2Farxiv.org%2Fabs%2F2608.30702v1>
- PDF: <https://arxiv.org/pdf/2608.30702v1>

Abstract: Large language models are increasingly used as agents for multistep retrosynthesis, raising the question of how much their search policy contributes independently of the underlying reaction model. We investigate this question in a biological setting through rule-based retrobiosynthesis: a deterministic biochemical engine generates the same validated transitions for every method, searching for routes that terminate in metabolites available to an \\emph\{Escherichia coli\} chassis, while the policy only selects which frontier molecule to expand next. Prompted and LoRA-tuned Qwen2.5-7B policies use a strict choice-only interface. The fine-tuned policy reaches $65\\pm1$\\% solve rate at 10 expansions on LASER versus 59\\% for MCTS, and at 200 expansions reaches $78\\pm1$\\% versus 75\\% on LASER, $88\\pm3$\\% versus 80\\% on the RetroPath RL Golden benchmark, and $63\\pm2$\\% versus 45\\% on the BioNavi-NP benchmark. Fine-tuning also consistently outperforms direct prompting. These results show that route-supervised frontier selection can improve budgeted search without altering biochemical generation, although performance remains dependent on frontier construction and reaction ranking.

## Large language models as uncertainty-calibrated optimizers for experimental discovery
- Source: Nature Machine Intelligence (journals)
- Date: 2026-08-28T00:00:00+00:00
- Categories: Reaction Informatics, LLMs & Agents
- Authors: Bojana Ranković, Ryan-Rhys Griffiths, Philippe Schwaller
- Journal: Nature Machine Intelligence
- DOI: 10.1038/s42256-026-01283-z
- Keywords: LLMs, LLM, Buchwald–Hartwig, chemical descriptors
- Source URL: <https://doi.org/10.1038/s42256-026-01283-z>
- Dashboard article: <https://tagirshin.com/radar/article?u=https%3A%2F%2Fdoi.org%2F10.1038%2Fs42256-026-01283-z>

Abstract: From reaction optimization to molecular design, experimental discovery poses the same expensive question: which candidate to test next under time and resource constraints. Bayesian optimization provides principled answers but depends on domain expertise that rarely transfers. Large language models (LLMs) contain rich scientific knowledge but lack the calibrated uncertainty estimates crucial for high-stakes decisions. Here we show how training language models through Bayesian objectives enables their use as reliable optimizers guided by natural language. Our approach, GOLLuM (Gaussian process Optimized LLMs), teaches LLMs from experimental outcomes under uncertainty, transforming their overconfidence from a fundamental flaw into a precise learning signal. This signal reshapes the LLM embeddings so that experiments with similar outcomes cluster together, revealing structure in the design space. Starting from only ten low-performing experiments, GOLLuM generalizes across 23 tasks in organic synthesis, materials science, process chemistry and molecular design, ranking first on average among all competing methods. It matches traditional Bayesian optimization with over 40% fewer experiments and nearly doubles the discovery of high-performing Buchwald–Hartwig reactions over expert quantum-chemical descriptors and state-of-the-art LLMs (43% versus 24–25%). More broadly, GOLLuM points to a different paradigm for specializing foundation models: not through more data but through richer, uncertainty-guided information.

## Mapping Structured Absences in Scientific Papers: An Auditable LLM-Assisted Workflow Demonstrated on Cross-Coupling Chemistry
- Source: ChemRxiv (preprints)
- Date: 2026-08-28T00:00:00Z
- Categories: Reaction Informatics, LLMs & Agents
- Authors: Yuriko Ono, Masaharu Yoshioka, Tetsuya Taketsugu
- DOI: 10.26434/chemrxiv.15007996/v1
- External ID: 10.26434/chemrxiv.15007996/v1
- Keywords: LLM, Cross Coupling, Kumada
- Source URL: <https://doi.org/10.26434/chemrxiv.15007996/v1>
- Dashboard article: <https://tagirshin.com/radar/article?u=https%3A%2F%2Fdoi.org%2F10.26434%2Fchemrxiv.15007996%2Fv1>

Abstract: Scientific papers record what was discovered and communicated, but the resolution at which evidence relations are made explicit—a paper’s documentary resolution—often falls short of what later computational reuse, benchmarking, or data-driven synthesis requires. We present UQS (Unasked Question Structuring), an auditable LLM-assisted pipeline that takes scientific papers as input and produces structured absence records with evidence chains, an Evidence Pack (EP) dictionary specifying the evidence needed to reach the next evidential level, a bounded filling-status map across later literature under pre-locked criteria, and cross-paper research-question briefs. Applied to Kumada’s 1972 Ni-catalysed cross-coupling communication over a 54-year record window (1972–2026), UQS derives ten EPs and reveals a VQ-hierarchy inversion: methodology-level documentation is addressed early, whereas system-matched mechanistic documentation (VQ-3) appears only after 49 years (Mazet 2021), and quantitative-reliability documentation (VQ-4) remains a bounded non-detection—a pattern consistent with publication-incentive structure. The paper claims auditability, configuration-pinned rerun consistency, and internal robustness (four-model cross-comparison, N=6 multi-run aggregation); external expert validation is registered as future work.

## Mechanistic Reaction Prediction via Discrete Flow Matching on Graph-Structured Electron Occupation
- Source: arXiv (preprints)
- Date: 2026-08-27T17:50:44Z
- Categories: Design de novo, Reaction Informatics
- Authors: Nguyen Xuan-Vu, Octavian Susanu, Daniel Armstrong, Philippe Schwaller
- External ID: 2608.27429v2
- Keywords: de novo generation, Reaction Prediction, reaction type
- Source URL: <https://arxiv.org/abs/2608.27429v2>
- Dashboard article: <https://tagirshin.com/radar/article?u=https%3A%2F%2Farxiv.org%2Fabs%2F2608.27429v2>
- PDF: <https://arxiv.org/pdf/2608.27429v2>

Abstract: Chemical reactions are fundamentally transformations in electron space, yet most machine learning approaches model them either through de novo generation of product molecules or through heuristic graph edits that operate directly on molecular topology. We introduce MAELLE (MechAnistic Edit fLow-matching on eLectron rEarrangements), which instead models reactions as discrete flow matching over electron occupation vectors. Concretely, we formulate the reactant-to-product mapping as a Continuous-time Markov Chain (CTMC) over the graph-structured integer-valued electron occupation space defined on all bonding, non-bonding, and hydrogen sites. To construct the interpolants between the reactants and products, we generalize the discrete flow matching mixture path to an edit-based formulation, where the electron moves are interpolated using Optimal Transport, yielding a mechanism-like set of moves without elementary step annotations. MAELLE achieves competitive performance on the USPTO-480K benchmark compared with leading reaction prediction models. Beyond in-distribution learning, we evaluate robustness across two out-of-distribution settings - structural complexity and reaction type - and find that MAELLE maintains strong performance where existing methods degrade. Finally, because the learned flow operates over the full electron redistribution, MAELLE naturally recovers mechanistic trajectories that align with known chemistry and can predict side products of a reaction.

## Data-driven Effective Modeling of Stochastic Chemical Reaction Networks
- Source: arXiv (preprints)
- Date: 2026-08-26T06:24:43Z
- Categories: Reaction Informatics
- Authors: Yuan Chen, Weize Mao, Dongbin Xiu
- External ID: 2608.25421v1
- Source URL: <https://arxiv.org/abs/2608.25421v1>
- Dashboard article: <https://tagirshin.com/radar/article?u=https%3A%2F%2Farxiv.org%2Fabs%2F2608.25421v1>
- PDF: <https://arxiv.org/pdf/2608.25421v1>

Abstract: The Stochastic Simulation Algorithm (SSA), widely considered an exact algorithm for stochastic chemical reaction networks, suffers from high computational cost. In this work, we propose a data-driven effective model that operates on a user-defined coarse time step independent of the underlying microscopic reaction-event scale. This is accomplished by directly approximating the finite-time transition kernel of the continuous-time Markov chain induced by SSA, using a generative machine learning model trained on short bursts of SSA simulation data. The trained model constructs a stochastic propagator that recursively generates statistically consistent trajectories at the constant coarse time step, with significantly reduced computational cost. In this paper, we employ conditional normalizing flow as the stochastic propagator. A comprehensive set of numerical examples is presented to demonstrate the accuracy and efficiency of the proposed method.

## Bridging CGR representations and language models for reaction property prediction
- Source: Digital Discovery (journals)
- Date: 2026-08-25T15:43:19Z
- Categories: Cheminformatics, Reaction Informatics
- Authors: Giustino Sulpizio, Charlotte Gerhaher, Esther Heid, Kjell Jorner
- Journal: Digital Discovery
- DOI: 10.1039/d6dd00120c
- Keywords: reaction SMILES, SMILES, reaction property, atom mapping, Transformer, BERT, property prediction
- Source URL: <https://doi.org/10.1039/d6dd00120c>
- Dashboard article: <https://tagirshin.com/radar/article?u=https%3A%2F%2Fdoi.org%2F10.1039%2Fd6dd00120c>

Abstract: Accurate prediction of reaction properties, such as activation energies, enthalpies, rate constants and yields, is key to the design of efficient and sustainable chemical processes. While graph-based models using Condensed Graphs of Reaction (CGR) have shown strong performance, text-based Transformer models trained on reaction SMILES have gained popularity due to their scalability and performance. In this work, we analyze the performance of an existing CGR-based string representation for reaction property prediction, the SMILES/CGR. In contrast to conventional reaction SMILES, this representation explicitly encodes atom and bond changes within a unified SMILES-like sequence. A key limitation of SMILES/CGR is that it does not encode stereochemistry. To overcome this, we introduce the superimposed reaction SMILES (sr-SMILES), a compact CGR-based representation that preserves full stereochemical information. To evaluate both representations, we pretrained a BERT transformer on two million USPTO reactions via masked language modeling, and then finetuned separate models on six benchmark datasets. We compared against models trained on standard reaction SMILES, naïve baselines, and results from the literature. Our results show that sr-SMILES consistently matches or outperforms reaction SMILES, particularly on benchmarks where mechanistic information is critical, as well as in low-data regimes. In contrast, SMILES/CGR tends to underperform, highlighting that the details of CGR encoding matter. While sr-SMILES did not close the performance gap with graph-based CGR methods, it shows that language models can benefit from explicit mechanistic information when appropriately encoded. Moreover, by comparing SMILES/CGR and sr-SMILES models with different atom mappings, we reinforce the notion that CGR-based approaches are highly sensitive to the mapping quality. A plausible explanation for the similar performance observed across string-based representations is that language models implicitly learn the atom mapping during (pre-)training, rendering explicit mapping unnecessary for many prediction tasks.

## Leveraging the Condensed Graph of Reaction for Clustering Retrosynthetic Pathways
- Source: Journal of Chemical Information and Modeling (journals)
- Date: 2026-08-21T00:00:00+00:00
- Categories: Property Prediction, Reaction Informatics
- Authors: Almaz Gilmullin, Tagir Akhmetshin, Dmitry Zankov, Olga Klimchuk, Dragos Horvath, Timur Madzhidov, Alexandre Varnek
- Journal: Journal of Chemical Information and Modeling
- DOI: 10.1021/acs.jcim.6c00489
- Keywords: retrosynthesis planning, Retrosynthetic, retrosynthesis
- Source URL: <https://doi.org/10.1021/acs.jcim.6c00489>
- Dashboard article: <https://tagirshin.com/radar/article?u=https%3A%2F%2Fdoi.org%2F10.1021%2Facs.jcim.6c00489>
- Code: <https://github.com/Laboratoire-de-Chemoinformatique/SynPlanner>

Abstract: Modern retrosynthetic tools can propose hundreds of alternative pathways for a single target, making it challenging to effectively explore and navigate the resulting route space. We present a CGR-based framework for the analysis and clustering of synthetic routes that integrates both target-centered and all-species-centered perspectives. Entire reaction pathways are encoded as single-molecule graphs (RouteCGR) or reduced representations retaining only target atoms (SB-CGR), enabling automatic identification of strategic bond patterns (SBPs). These representations can be transformed into Morgan fingerprints for a quantitative route similarity assessment. We propose a two-level clustering strategy in which routes are first grouped by shared SBPs, ensuring high interpretability based on key retrosynthetic disconnections, and then further differentiated using RouteCGR similarity to capture variations in starting materials and auxiliary transformations. The method demonstrates near-linear scalability and computational efficiency for large data sets. Application to synthetic routes for apatinib generated by multiple planning tools reveals tool-dependent diversity in strategic disconnections and highlights the benefit of combining tools to expand route space. The framework also supports cross-target analysis, enabling the identification of reusable route families that share common strategic disconnections and building blocks across related molecules. Overall, the SBP-based approach provides an interpretable and scalable solution for automated synthesis route analysis and informed decision-making. The proposed approach is implemented in SynPlanner retrosynthesis planning software and is available as a standalone module at https://github.com/Laboratoire-de-Chemoinformatique/SynPlanner.

## Using MR-chordless circuits for efficient enumeration of autocatalytic cores in large chemical reaction networks
- Source: Journal of Cheminformatics (journals)
- Date: 2026-08-20T00:00:00+00:00
- Categories: Reaction Informatics
- Authors: Richard Golnik, Nicola Vassena, Peter F. Stadler, Thomas Gatter
- Journal: Journal of Cheminformatics
- DOI: 10.1186/s13321-026-01240-3
- Source URL: <https://doi.org/10.1186/s13321-026-01240-3>
- Dashboard article: <https://tagirshin.com/radar/article?u=https%3A%2F%2Fdoi.org%2F10.1186%2Fs13321-026-01240-3>

Abstract: Autocatalysis is an important property of chemical reaction networks (CRNs) that is particularly prevalent in metabolic networks. A set of well-defined autocatalytic cores prominently features minimal subsystems that determine the autocatalytic capabilities. Recently, a graph-theoretic characterization has become available that enabled the enumeration of moderate-sized autocatalytic cores in real-life metabolic networks. Such an approach relies on enumerating and properly assembling elementary circuits in the bipartite graph associated to a CRN. Here, we improve on this approach in two ways: (1) We elaborate on algorithms for the enumeration of elementary circuits restricted to so-called MR-chordless circuits. These circuits do not have a chord from a Metabolite to a Reaction vertex, and are the only candidates to find autocatalytic cores. (2) We interleave our new algorithm with tests for autocatalysis to further limit the number of circuits that need to be stored for the construction of autocatalytic cores more complex than elementary MR-chordless circuits. Combined, these innovations achieve a performance gain of several orders of magnitude and make it possible to exhaustively enumerate all autocatalytic cores in real-life metabolic reaction networks comprising several hundred metabolites and reactions. Importantly, we find that reaction networks with irreversible reactions contain complex autocatalytic cores comprising more than a single “cycle with an ear”. Such structures exceed the established classification of autocatalytic cores for fully reversible networks into five types. Scientific contribution We developed a new graph-theoretic algorithm for enumerating autocatalytic cores that can handle large genome-scale metabolic models. Implemented in the Python program , it is up to four orders of magnitude faster than previous methods. Applications to large metabolic network models that involve both reversible and nonreversible reactions reveal that more complex autocatalytic cores exist than predicted by existing classification schemes for reversible reactions.

## PGFS++: Molecular Property Improvement under Synthesis and Diversity Constraints
- Source: arXiv (preprints)
- Date: 2026-08-19T17:17:31Z
- Categories: Property Prediction, Docking & Screening, Design de novo, Reaction Informatics
- Authors: Boqiao Zhang, Godbless James, Sai Krishna Gottipati, Andrew Fitzgibbon
- External ID: 2608.19121v1
- Keywords: drug likeness, Policy Gradient, reinforcement learning, RL, Molecular Property, binding affinity
- Source URL: <https://arxiv.org/abs/2608.19121v1>
- Dashboard article: <https://tagirshin.com/radar/article?u=https%3A%2F%2Farxiv.org%2Fabs%2F2608.19121v1>
- PDF: <https://arxiv.org/pdf/2608.19121v1>

Abstract: Improving molecular properties, such as drug-likeness or binding affinity, is a recurring task in early-stage drug discovery. However, molecules optimized in an unconstrained chemical space have limited practical value if they cannot be synthesized. Policy Gradient for Forward Synthesis (PGFS) is a synthesis-aware reinforcement learning method for molecular improvement, but its use of reactant embedding prediction makes reactant selection indirect, which, as we show, limits learning effectiveness. We first develop PGFS+, in which reaction templates and second reactants are represented by trainable embedding lookup tables. Combined with a more effective scoring function and RL algorithm, PGFS+ significantly improves the desired property. However, it exposes a reward-hacking failure mode: a powerful reactant search can map diverse input molecules to the same high-reward magnet molecule, improving the reward while collapsing the output diversity. We therefore introduce PGFS++, a synthesis-aware reinforcement learning framework for input-specific molecular improvement. Given an input molecule, PGFS++ treats it as the start of a forward-synthesis trajectory, applies learned reaction templates with compatible in-stock building blocks, and produces a molecule with improved target properties, an explicit synthesis route, and structural similarity to the input. Experiments on molecular improvement tasks show that PGFS++ improves target properties while preserving high output diversity.

## Training Chemical Plausibility-Aware Large Language Models for Single-Step Retrosynthesis
- Source: arXiv (preprints)
- Date: 2026-08-19T14:08:05Z
- Categories: Reaction Informatics, LLMs & Agents
- Authors: Bogdan Zagribelnyy, Ivan Ilin, Nikita Bondarev, Maksim Kuznetsov, Mathieu Reymond, Vladimir Aladinskiy, Alex Aliper, Alex Zhavoronkov
- External ID: 2608.18940v1
- Keywords: Single Step Retrosynthesis, LLMs, LLM, synthesis planning, Retrosynthesis
- Source URL: <https://arxiv.org/abs/2608.18940v1>
- Dashboard article: <https://tagirshin.com/radar/article?u=https%3A%2F%2Farxiv.org%2Fabs%2F2608.18940v1>
- PDF: <https://arxiv.org/pdf/2608.18940v1>
- Code: <https://github.com/insilicomedicine/ChemCensor>

Abstract: Single-step retrosynthesis is a central component of computer-aided synthesis planning, yet its intrinsically one-to-many nature is poorly captured by single-answer evaluation and benchmarking protocols. To address this, we introduce Top-K prompting as a robust training and inference paradigm to better capture diverse, plausible reaction predictions. We compile CREED-CCV-2+USPTO-XL, an ultra-large-scale dataset of ~45.6 million verified reactions to train the C3LM (Chemistry Constraint-Consistent Language Model). By integrating fine-tuning with ChemCensor-based and novelty-oriented rewards, our model achieves state-of-the-art performance on the OOD URSA-expert-2026 benchmark. Further analysis of reaction uniqueness shows that LLMs and conventional models explore complementary reaction spaces, motivating ensemble-based retrosynthesis systems. Overall, our results establish Top-K, plausibility-aware training as a practical new direction for robust future LLM-based synthesis planning.

## RetroMPA: A Molecular Property-Aware Auxiliary Framework for Enhancing Retrosynthesis Prediction
- Source: arXiv (preprints)
- Date: 2026-08-17T05:03:30Z
- Categories: Cheminformatics, Property Prediction, Reaction Informatics
- Authors: Mianzhi Liu, Fan Xiao, Zhiliang Yu, Huayang Huang, Yuke Li, Yi Yang, Wenbo Liu, Yu Wu
- DOI: 10.1021/acs.jcim.6c01506
- External ID: 2608.16111v1
- Keywords: SMILES, Retrosynthesis Prediction, retrosynthesis models, Molecular Property, Retrosynthesis, Suzuki Miyaura
- Source URL: <https://arxiv.org/abs/2608.16111v1>
- Dashboard article: <https://tagirshin.com/radar/article?u=https%3A%2F%2Farxiv.org%2Fabs%2F2608.16111v1>
- PDF: <https://arxiv.org/pdf/2608.16111v1>
- Code: <https://github.com/MengzhouLu/RetroMPA>

Abstract: Retrosynthesis is a cornerstone of drug discovery and organic synthesis. While data-driven deep learning models have shown remarkable progress, they autonomously learn reaction patterns from extensive datasets with limited integration of established chemical knowledge as priors. To address this limitation, we introduce RetroMPA, a molecular property-aware, post-hoc enhancement module that injects chemical knowledge into the retrosynthesis pipeline. Rather than functioning as an independent SMILES sequence generator, RetroMPA is a broadly applicable, model-agnostic chemical filter designed to recalibrate and optimize the predictive pathways of existing algorithms. This plug-and-play framework integrates seamlessly with a range of data-driven retrosynthesis methods, enhancing outputs without modifying model architecture or requiring resource-intensive retraining. By leveraging a property-aware latent embedding space, RetroMPA consistently improves top-1 accuracy across eight representative retrosynthesis models by an average of 5.50% on USPTO-50K. Furthermore, we validate its scalability on the large-scale USPTO-Full dataset, achieving an average improvement of about 2.03% across both template-based and template-free architectures. Wet-lab experiments provide preliminary support for the practical utility of the framework. These syntheses confirmed viable, previously unreported substrate combinations for classic reaction paradigms---specifically, Suzuki-Miyaura coupling, Bucherer reaction, and Friedel-Crafts acylation---suggesting that RetroMPA can operate beyond mere data fitting. The code is open-sourced at https://github.com/MengzhouLu/RetroMPA.

## AI-driven design of a cost-effective synthetic route to malaria drug Ganaplacide
- Source: ChemRxiv (preprints)
- Date: 2026-08-17T00:00:00Z
- Categories: Reaction Informatics
- Authors: Fernando Huerta, Erik Leonhardt, Lianyan Xu, Trevor Laird, John Dillon, David M. Parry, Thomas Galeandro-Diamant
- DOI: 10.26434/chemrxiv.15006146/v2
- External ID: 10.26434/chemrxiv.15006146/v2
- Keywords: retrosynthesis algorithm, synthetic route, retrosynthesis, retrosynthetic
- Source URL: <https://doi.org/10.26434/chemrxiv.15006146/v2>
- Dashboard article: <https://tagirshin.com/radar/article?u=https%3A%2F%2Fdoi.org%2F10.26434%2Fchemrxiv.15006146%2Fv2>

Abstract: Malaria is a deadly infectious disease that disproportionately affects low- and middle-income countries. Ganaplacide is a drug candidate that is highly effective against malaria, and the ability to produce it at low cost is essential. Here we show how we used a combination of AI-driven design and expert chemistry knowledge to develop a novel synthetic route to ganaplacide. The AI retrosynthesis algorithm suggested a non-evident two-steps synthesis of the imidazolopiperazine core of the molecule from benzyl-protected dimethylpiperazinone through a strategic thioimidate intermediate. We also present the Synthetic Confidence Score, which assesses the level of risk associated with a given retrosynthetic disconnection and allows balancing creativity and reliability. The synergy between AI creativity, metrics such as the Synthetic Confidence Score and expert chemist knowledge has the potential to unlock non-evident efficient routes for important molecules, when expert chemists alone have explored mainstream chemistry unsuccessfully.

## Bringing retrosynthesis in-house: strategies for integrating custom reactions into retrosynthetic planners
- Source: ChemRxiv (preprints)
- Date: 2026-08-17T00:00:00Z
- Categories: Reaction Informatics
- Authors: Almaz Gilmullin, Tagir Akhmetshin, Olga Klimchuk, Dragos Horvath, Timur Madzhidov, Alexandre Varnek
- DOI: 10.26434/chemrxiv.15007550/v1
- External ID: 10.26434/chemrxiv.15007550/v1
- Keywords: retrosynthetic planning, classification model, synthesis planning, retrosynthesis, retrosynthetic
- Source URL: <https://doi.org/10.26434/chemrxiv.15007550/v1>
- Dashboard article: <https://tagirshin.com/radar/article?u=https%3A%2F%2Fdoi.org%2F10.26434%2Fchemrxiv.15007550%2Fv1>

Abstract: Template-based computer-aided synthesis planning offers interpretability and explicit control over reaction application, but its behavior is strongly influenced by the training data and template-ranking policy used during search. This may limit chemists’ ability to apply a rare or project-specific transformation poorly represented in public data. Here, we introduce priority rule injection, a planning-time mechanism for incorporating curated in-house reaction knowledge into a template-based retrosynthesis planner without retraining the underlying single-step policy. During Monte Carlo Tree Search node expansion, curated priority rules are queried by substructure matching before the default Top-K model-ranked proposals; matched rule applications are inserted into the search tree as privileged expansion candidates. We demonstrate this approach using a case study of the Ugi four-component reaction, which is often ignored by retrosynthesis engines as corresponding templates are rejected at preparation. Priority rule injection is compared with retraining a template-classification model and retraining or fine-tuning a Modern Hopfield Network, using Ugi-free models as negative controls. Although generic target solvability remained high across the evaluated settings, recovery of the requested Ugi strategy differed substantially. In contrast, priority rule injection recovered at least one solved Ugi-containing route for all 10 hold-out Ugi products. These results show that the main benefit of priority rule injection is reliable recovery of chemist-intended, reaction class-constrained routes. The approach provides an interpretable and operationally simple route for incorporating proprietary or project-specific transformations into retrosynthetic planning, while leaving the validated single-step model unchanged. All calculations were performed using the SynPlanner tool for retrosynthetic planning.

## Molecular Fragment-Based Graph Isomorphism Networks for Interpretable Prediction of Synergistic Drug Combinations
- Source: Journal of Chemical Information and Modeling (journals)
- Date: 2026-08-17T00:00:00+00:00
- Categories: Cheminformatics, Property Prediction, Reaction Informatics
- Authors: Lifeng Shao, Jianqiang Sun, Hongzhan Ma, Qi Zhao
- Journal: Journal of Chemical Information and Modeling
- DOI: 10.1021/acs.jcim.6c01930
- Source URL: <https://doi.org/10.1021/acs.jcim.6c01930>
- Dashboard article: <https://tagirshin.com/radar/article?u=https%3A%2F%2Fdoi.org%2F10.1021%2Facs.jcim.6c01930>

Abstract: Drug combination therapy plays an increasingly important role in the clinical treatment of complex diseases, such as cancer, as rational drug combinations can enhance therapeutic efficacy and reduce toxic side effects. However, existing methods still exhibit limitations in the granularity of drug molecular representation, drug interaction modeling, and cell line context awareness, which restrict further improvements in predictive performance. To address these issues, we propose FragSyn, a deep graph learning framework for predicting synergistic drug combinations based on molecular fragmentations. FragSyn first decomposes drug molecules into chemically meaningful fragments according to breaks of retrosynthetically interesting chemical substructure rules and learns fragment-level molecular representations through a graph isomorphism network with edge features. It then captures nonlinear relationships between drug pairs from multiple perspectives while introducing a gating modulation mechanism conditioned on cell line features, enabling drug representations to adapt dynamically to the cell line context. Finally, multisource features are fused to perform binary classification of synergy versus antagonism. FragSyn achieves AUC, AUPR, and ACC of 0.944, 0.942, and 0.872, respectively, outperforming eight baseline models, and demonstrates optimal generalization performance in both leave-one-out cross-validation and external validation. Ablation studies and interpretability analyses further validate the rationality of FragSyn and its ability to identify key fragments. These results indicate that FragSyn, through the synergistic design of fragment-level representation and cellular context awareness, provides an effective and interpretable new approach to synergistic drug combination prediction.

## Superhuman Centaur Retrosynthesis Through Targeted LLM-enabled Decomposition in DeepRetro2
- Source: ChemRxiv (preprints)
- Date: 2026-08-17T00:00:00Z
- Categories: Reaction Informatics, LLMs & Agents
- Authors: Sharanabasava D. Hiremath, Priyanka Makkar, Pramod Kumar, Shreyas Vinaya Sathyanarayana, Riya Singh, Rishikesh Panda, Kamal Manchanella, Aryan Amit Barsainyan, Rahil Shah Kirankumar, Bharath Ramsundar
- DOI: 10.26434/chemrxiv.15007537/v1
- External ID: 10.26434/chemrxiv.15007537/v1
- Keywords: reaction prediction, LLM, synthetic route, Retrosynthesis, retrosynthetic
- Source URL: <https://doi.org/10.26434/chemrxiv.15007537/v1>
- Dashboard article: <https://tagirshin.com/radar/article?u=https%3A%2F%2Fdoi.org%2F10.26434%2Fchemrxiv.15007537%2Fv1>

Abstract: Retrosynthesis of complex molecules requires many chemically plausible transformations to be integrated into a coherent synthetic route. Although computational methods have improved reaction prediction and route identification, systematically decomposing complex molecules and coordinating multiple retrosynthetic steps remains challenging. Here we present DeepRetro2, an agentic retrosynthesis framework that extends large language model (LLM)-based retrosynthesis through recursive molecular decomposition, iterative generation, evaluation and expansion of retrosynthetic subproblems. DeepRetro2 enables autonomous retrosynthetic analysis for targets of moderate complexity, while substantially larger targets can be solved through a human-in-the-loop workflow in which expert 1 input is introduced at chemically critical decision points. We dub this mode of human-agentic collaboration “Centaur Retrosynthesis,” in homage to the tradition of human-AI “Centaur Chess.” We evaluated this framework on four chemically distinct targets: the unsolved ladder polyether maitotoxin (MTX), bryostatin 1, bryostatin 3 and luvesilocin. For MTX, DeepRetro2 generated a hierarchical decomposition into five major polyether fragments and subsequently elaborated these fragments into simpler synthons. For bryostatin 1 and bryostatin 3, it identified convergent fragment-level disconnections that preserved the stereochemically defined pyran frameworks, while luvesilocin provided a smaller-scale test of the same retrosynthetic framework. Across these targets, human-guided selection of protecting groups and functional-group transformations refined the computationally generated pathways, maintained orthogonal reactivity during fragment elaboration and facilitated chemically compatible fragment unions. DeepRetro2 therefore combines autonomous agentic retrosynthesis with expert-guided chemical reasoning, providing an iterative framework for decomposing complex molecular structures and developing chemically executable retrosynthetic pathways.

## Unraveling the Size Determination Mechanism of Nanocrystal Synthesis via Interpretable Neural Networks
- Source: arXiv (preprints)
- Date: 2026-08-13T07:43:27Z
- Categories: Reaction Informatics
- Authors: Kai Gu, Haizheng Zhong
- External ID: 2608.14734v1
- Keywords: reaction conditions
- Source URL: <https://arxiv.org/abs/2608.14734v1>
- Dashboard article: <https://tagirshin.com/radar/article?u=https%3A%2F%2Farxiv.org%2Fabs%2F2608.14734v1>
- PDF: <https://arxiv.org/pdf/2608.14734v1>
- Code: <https://github.com/ime1452/Nanocrystal-Equation-Learner>

Abstract: Deep learning models of nanocrystal synthesis enable the prediction of size and shape by encoding precursors and reaction conditions. However, their black-box nature hinders gaining deep insights into the underlying synthetic mechanisms. Here, we develop the Nanocrystal Equation Learner (NanoEQL), a fully white-box neural network to unravel the size determination mechanisms of nanocrystal synthesis. Building on the EQL architecture, eight operators are introduced to replace standard activation functions to fit the mathematical equations in nanocrystal synthesis. Among these operators, three smoothed operators address the gradient explosion of singular operators at zero. To evaluate the weights of different precursors, we develop a temperature-gated attention pooling strategy that encodes concentration-driven and reactivity-driven chemical synthesis mechanisms into the temperature gate. The NanoEQL model illustrates that the final nanocrystal size can be described by a linear equation composed of three scalars representing nanocrystallization capability (-Zp), growth capability (Zrea), and external input potential (-Zops). These interpretable scalars not only advance the rational design of nanocrystal synthesis but also establish a generalizable paradigm for deciphering chemical reaction mechanisms through white-box machine learning.
