# Chem(o)Info Radar article snapshot
> 27 records returned from 3035 current records.

Generated: 2026-09-22T18:24:41.688621+00:00
Filters: topic=reaction, days=30

Views are bounded by topic, source, days and limit; text search runs in the dashboard page, not on the server. Blog entries are metadata-only; use their source URL for full text.

## Collaborative Laboratory Edge Networks
- Source: ChemRxiv (preprints)
- Date: 2026-09-21T00:00:00Z
- Categories: Reaction Informatics
- Authors: Andre Williams, Shanice Brown, Kevin Thompson, Jamar Campbell, Alicia Morgan
- DOI: 10.26434/chemrxiv.15009129/v1
- External ID: 10.26434/chemrxiv.15009129/v1
- Keywords: retrosynthetic
- Source URL: <https://doi.org/10.26434/chemrxiv.15009129/v1>
- Dashboard article: <https://tagirshin.com/chemradar/article?u=https%3A%2F%2Fdoi.org%2F10.26434%2Fchemrxiv.15009129%2Fv1>

Abstract: Chemistry language models now predict reactions and retrosynthetic routes with useful accuracy, but they assume that proprietary reaction records can be pooled into a central cluster and that inference runs on abundant accelerators. Neither assumption holds in pharmaceutical and fine-chemical practice, where yield records cannot leave the plant network, the points of use are heterogeneous resource-limited sites, and different laboratories occupy disjoint regions of chemical space. We present ChemFedFusion, a heterogeneity-aware framework for federated chemistry-model training and offloading over collaborative laboratory edge networks. ChemUCP is a chemistry-aware utility–cost predictor that scores a task–node pair by expected chemistry gain, expected latency, and an explicit out-of-distribution risk of emitting unparsable or unreliable chemistry, using a dual-channel tokenizer and a scaffold-coverage encoder that is corrected online from execution feedback. CSAW replaces sample-count aggregation with the product of a scaffold-space complementarity term, a site reliability term, and a weakened volume guard, amplifying rare chemistry while suppressing noisy sites. ARCO couples offloading and aggregation through a shared adaptation matrix and an event-driven asynchronous aggregator with a staleness bound, so that successful offloads teach the aggregator which node is competent for which chemistry task. On Mol-Instruction and USPTO-50K, ChemFedFusion reaches 0.883 Exact and 51.36% Top-1, and on the FedChem suite it reduces regression RMSE by up to 9.99% and classification ROC-AUC improvements are consistent across all nine datasets and three heterogeneity levels. The framework reduces convergence rounds by 22.4% and raises the fraction of answers that are both on time and chemically valid by 8.8 percentage points. A failure analysis shows that the remaining errors are dominated by synthetically infeasible routes that parse correctly, which no validity-based router can detect.

## Using Data Science Tools to Explore Rate Matching in a Nickel-Catalyzed Cross-Electrophile Coupling of Alkyl and Aryl Halides (Cl, Br) with a Tridentate Monoanionic Ligand
- Source: Journal of the American Chemical Society (journals)
- Date: 2026-09-21T00:00:00+00:00
- Categories: Reaction Informatics
- Authors: Haruka Takenaka, Alexandra J. Ring, Avijit Hazra, Therese H. Wild, Samuel M. R. Powell, Tianhua Tang, Sarah E. Reisman, Matthew S. Sigman
- Journal: Journal of the American Chemical Society
- DOI: 10.1021/jacs.6c10810
- Keywords: cross coupling
- Source URL: <https://doi.org/10.1021/jacs.6c10810>
- Dashboard article: <https://tagirshin.com/chemradar/article?u=https%3A%2F%2Fdoi.org%2F10.1021%2Fjacs.6c10810>

Abstract: Nickel-catalyzed cross-electrophile coupling (XEC) between aryl and alkyl electrophiles has emerged as an enabling synthetic method, yet predicting the outcomes of specific electrophile pairs is difficult due to the complexity of the mechanistic pathways. It has been proposed that “rate matching” between the oxidative addition of the aryl halide and the rate of formation of the alkyl radical, in addition to the stability of the Ni(II)–Ar complex, dictates the yield of the cross-coupling product. To probe this quantitatively, we developed a new tridentate monoanionic ligand that activatesalkyl and aryl chlorides and exhibits well-behaved electrochemistry, allowing the measurement of relative rate constants. This provides a framework to interrogate the rate-matching concept using both aryl and alkyl bromides and chlorides in cross-electrophile coupling. By combining an electroanalytical workflow with machine learning, the prediction of activation rate constants for all aryl/alkyl chlorides and bromides was achieved and correlated with reaction yields. We found that productive coupling occurs in a specific rate regime, and the synthetic performance could be predictably modulated by choosing the “matched” halide (either chloride or bromide) for each coupling partner.

## CDGA: A Host-Based Alignment and Descriptor Analysis Tool for Cyclodextrin Inclusion Complexes
- Source: Journal of Chemical Information and Modeling (journals)
- Date: 2026-09-17T00:00:00+00:00
- Categories: Property Prediction, Reaction Informatics
- Authors: Ewa Napiórkowska, Łukasz Szeleszczuk
- Journal: Journal of Chemical Information and Modeling
- DOI: 10.1021/acs.jcim.6c02866
- Keywords: atom to atom mapping, atom mapping
- Source URL: <https://doi.org/10.1021/acs.jcim.6c02866>
- Dashboard article: <https://tagirshin.com/chemradar/article?u=https%3A%2F%2Fdoi.org%2F10.1021%2Facs.jcim.6c02866>

Abstract: Computational studies of cyclodextrin inclusion complexes are widespread; however, quantitative comparison of host–guest geometries remains difficult because native cyclodextrins are cyclic and pseudosymmetric, with chemically equivalent glucose units and no unique structural anchor for conventional atom-to-atom mapping. Consequently, RMSD values, guest orientations, and pose comparisons may depend on arbitrary atom numbering, molecular file representation, or software-specific alignment procedures rather than on genuine structural differences. To address this limitation, we developed CDGA, the Cyclodextrin Guest Analysis tool, an open-source KNIME-based tool for quantitative analysis of cyclodextrin host–guest complexes comprising CDGA Align for host-based alignment and CDGA NoAlign for analysis directly from their original coordinates; both workflows identify host and guest fragments, perform chemically meaningful atom mapping, and calculate standardized descriptors including guest depth, guest orientation, guest principal-axis angle, and heavy-atom RMSD, while an additional host module characterizes host geometry. CDGA was validated using controlled geometrical and chemical modifications and cross-docked and redocked poses generated for conformationally diverse native cyclodextrins, demonstrating robust and reproducible descriptor calculation despite differences in atom ordering, bond representation, host orientation, and molecular file generation, thereby enabling standardized comparison of cyclodextrin inclusion geometries, transparent reporting, benchmarking, and large-scale analysis of computational host–guest studies.

## CDGA: A Host-Based Alignment and Descriptor Analysis Tool for Cyclodextrin Inclusion Complexes
- Source: Journal of Chemical Information and Modeling (journals)
- Date: 2026-09-17T00:00:00Z
- Categories: Reaction Informatics
- Authors: E. Napiórkowska, Ł. Szeleszczuk
- Journal: Journal of Chemical Information and Modeling
- DOI: 10.1021/acs.jcim.6c02866
- External ID: 41f4b99282177c13bf76cd6326d5c9764dd48b59
- Keywords: atom to atom mapping, atom mapping
- Source URL: <https://doi.org/10.1021/acs.jcim.6c02866>
- Dashboard article: <https://tagirshin.com/chemradar/article?u=https%3A%2F%2Fdoi.org%2F10.1021%2Facs.jcim.6c02866>

Abstract: Computational studies of cyclodextrin inclusion complexes are widespread; however, quantitative comparison of host–guest geometries remains difficult because native cyclodextrins are cyclic and pseudosymmetric, with chemically equivalent glucose units and no unique structural anchor for conventional atom-to-atom mapping. Consequently, RMSD values, guest orientations, and pose comparisons may depend on arbitrary atom numbering, molecular file representation, or software-specific alignment procedures rather than on genuine structural differences. To address this limitation, we developed CDGA, the Cyclodextrin Guest Analysis tool, an open-source KNIME-based tool for quantitative analysis of cyclodextrin host–guest complexes comprising CDGA Align for host-based alignment and CDGA NoAlign for analysis directly from their original coordinates; both workflows identify host and guest fragments, perform chemically meaningful atom mapping, and calculate standardized descriptors including guest depth, guest orientation, guest principal-axis angle, and heavy-atom RMSD, while an additional host module characterizes host geometry. CDGA was validated using controlled geometrical and chemical modifications and cross-docked and redocked poses generated for conformationally diverse native cyclodextrins, demonstrating robust and reproducible descriptor calculation despite differences in atom ordering, bond representation, host orientation, and molecular file generation, thereby enabling standardized comparison of cyclodextrin inclusion geometries, transparent reporting, benchmarking, and large-scale analysis of computational host–guest studies.

## CUE: A Chemical Uncertainty-Aware Embedding Framework for Multimodal Drug Selectivity Prediction
- Source: Journal of Chemical Information and Modeling (journals)
- Date: 2026-09-17T00:00:00+00:00
- Categories: Cheminformatics, Property Prediction, Docking & Screening, Reaction Informatics
- Authors: Jin Hyuk Kim, Gyeong Hwan Kim, Hyeon Jun Park, Jonghwan Choi
- Journal: Journal of Chemical Information and Modeling
- DOI: 10.1021/acs.jcim.6c01761
- Keywords: virtual screening, molecular fingerprints, QSAR
- Source URL: <https://doi.org/10.1021/acs.jcim.6c01761>
- Dashboard article: <https://tagirshin.com/chemradar/article?u=https%3A%2F%2Fdoi.org%2F10.1021%2Facs.jcim.6c01761>
- Code: <https://github.com/jjjabcd/CUE>

Abstract: Accurate prediction of drug selectivity is critical for prioritizing candidate compounds with reduced off-target effects in AI-based drug discovery. Although selectivity can be inferred indirectly from drug–target affinity (DTA) prediction, cumulative errors from affinity predictions across multiple targets can reduce reliability, motivating the development of methods that directly predict compound-level selectivity. We propose a Chemical Uncertainty-aware Embedding (CUE) framework that integrates molecular fingerprints and 2D molecular structure image embeddings. The two modalities are combined via Loss Trajectory Analysis for Uncertainty (LTAU)-based weighted fusion, which adaptively reweights features according to their predictive reliability. Across eight benchmark data sets, CUE achieved RMSE values ranging from 0.163 to 1.691 and outperformed affinity-based and quantitative structure–activity relationship (QSAR) baseline models. Furthermore, a virtual screening case study for EGFR(T790M/C797S) inhibitors further demonstrated its utility for identifying selective hits. The source code is available at https://github.com/jjjabcd/CUE.

## Progress with FAK inhibitors in the patent literature (2020-present).
- Source: Expert opinion on therapeutic patents (journals)
- Date: 2026-09-17T00:00:00Z
- Categories: Reaction Informatics
- Authors: Bing-Bing Chen, Rui-Peng Feng, Jin-Bo Niu, Jian Song, Yuan-Bo Cui, Sai Zhang
- Journal: Expert opinion on therapeutic patents
- DOI: 10.1080/13543776.2026.2735861
- External ID: 358917ba7e794242c53b6ab00d5b71d3f7b47c3b
- Keywords: SciFinder, kinase, receptor
- Source URL: <https://doi.org/10.1080/13543776.2026.2735861>
- Dashboard article: <https://tagirshin.com/chemradar/article?u=https%3A%2F%2Fdoi.org%2F10.1080%2F13543776.2026.2735861>

Abstract: INTRODUCTION Focal adhesion kinase (FAK) is a non-receptor tyrosine kinase that transduces signals from integrins, receptor tyrosine kinases, and growth factors to orchestrate cell adhesion, migration, and survival. Aberrant FAK hyperactivation promotes tumor proliferation, invasion, anti-apoptosis, and chemoresistance. Consequently, the recent U.S. FDA approval of the FAK inhibitor defactinib combined with avutometinib for recurrent KRAS-mutant low‑grade serous ovarian cancer (LGSOC) validates FAK as a clinically actionable anticancer target. AREAS COVERED This review discusses recent advances in FAK inhibitor development, focusing on small-molecule patents published from January 2020 to August 2026. A systematic search of SciFinder and WIPO databases was conducted for this period. It categorizes these novel inhibitors into several classes and summarizes their structural features, biological activities, design strategies, and structure-activity relationships (SAR). EXPERT OPINION Recent advances in FAK drug discovery have driven a transition from conventional kinase inhibition toward broader FAK pathway modulation, including improved inhibitors, dual-target strategies, and targeted protein degradation. The clinical success of defactinib-based combination therapy validates the therapeutic potential of FAK targeting, while emerging approaches offer opportunities to overcome resistance. Future progress will depend on biomarker-guided patient selection, rational combination regimens, and exploration of non-catalytic FAK functions to achieve more effective and durable therapies.

## Bayesian reaction optimization with fixed MaleCNS features: performance, wiring controls, and reward–aversion diagnostics
- Source: ChemRxiv (preprints)
- Date: 2026-09-16T00:00:00Z
- Categories: Reaction Informatics
- Authors: Yuuya Nagata
- DOI: 10.26434/chemrxiv.15008959/v1
- External ID: 10.26434/chemrxiv.15008959/v1
- Keywords: Suzuki–Miyaura, Suzuki
- Source URL: <https://doi.org/10.26434/chemrxiv.15008959/v1>
- Dashboard article: <https://tagirshin.com/chemradar/article?u=https%3A%2F%2Fdoi.org%2F10.26434%2Fchemrxiv.15008959%2Fv1>

Abstract: Connectome-derived features provide one possible representation for reaction optimization, but their utility need not depend on measured neuronal pairing. We examine a Bayesian predictor using fixed mushroom-body features derived from MaleCNS connectivity, with observed prediction errors separated into engineered reward and aversion signals. An offline evaluator returns only selected recorded outcomes. We distinguish this predictor from earlier models that modify KC–MBON efficacy incrementally or reconstruct it from observed response ranks. The rank comparison changes both the learning head and acquisition rule. In the exploratory predictive-model evaluation, the mean best response over post-initial observations exceeded rank reconstruction by 0.486 percentage points on C–N coupling and 3.346 on Suzuki–Miyaura coupling. Differences from development-tuned Gaussian-process optimization were smaller, and their cluster-bootstrap intervals included zero. Native-versusrewired intervals also included zero. C–N responses were analytical yields; Suzuki responses were UV-area percentages. High/low response diagnostics separate input-dependent KC activity from outcome-dependent prediction updates: circuit activity remains fixed under an identical-cue outcome intervention. Reused datasets and input ensembles, and one rewired graph restrict inference. These results support evaluating anatomical representation, learning rule, search performance, and prediction accuracy separately; they establish neither statistical equivalence nor a unique advantage of native connectivity.

## Integrated Retrosynthesis and Large Language Modeling for Chemistry-Informed Circular Chemical Reaction Networks
- Source: ACS Sustainable Chemistry & Engineering (journals)
- Date: 2026-09-16T00:00:00Z
- Categories: Reaction Informatics, LLMs & Agents
- Authors: Sunghoon Kim, Avan Kumar, H. Kodamana, Manojkumar Ramteke, B. Bakshi
- Journal: ACS Sustainable Chemistry & Engineering
- DOI: 10.1021/acssuschemeng.6c06648
- External ID: 0a9eede00ab590d34aa536d6bf946005f1aca161
- Keywords: LLM, LLMs, Retrosynthesis, retrosynthetic, chemicals
- Source URL: <https://doi.org/10.1021/acssuschemeng.6c06648>
- Dashboard article: <https://tagirshin.com/chemradar/article?u=https%3A%2F%2Fdoi.org%2F10.1021%2Facssuschemeng.6c06648>

Abstract: Discovery and development of manufacturing routes that explicitly account for the entire product life cycle are essential for the transformation to a sustainable and circular chemical industry. Retrosynthesis is a promising approach, but operates within a gate-to-gate paradigm, limiting its ability to explicitly integrate end-of-life waste streams into upstream production pathways. We evaluated four reaction discovery strategies that combine human intervention, pattern recognition, retrosynthesis, and large language models (LLM) within a common comparative framework. When evaluated with methanol as a benchmark system, the strategies reveal distinct trade-offs between pathway discovery and technological maturity. Combining LLMs with retrosynthesis achieves the broadest expansion of reaction space with 165 reactions involving 12 unique chemicals and the highest novelty relative to the conventional business-as-usual reference network. The resulting circular chemical reaction network (CCRN) for a more complex molecule, polyethylene, contains more than 1,000 reactions, and human-guided extraction increased the total number of identified reactions by 34–35%, depending on the keyword-search strategy, while recovering additional end-of-life reactions from experimental results. The proposed framework demonstrates that combining language-based knowledge extraction with structure-informed retrosynthetic reasoning enables scalable construction of CCRNs and supports the systematic exploration of candidate circular chemical pathways for subsequent economic and environmental evaluation.

## Discovering Kinetically Significant Reaction Mechanisms Beyond Chemical Intuition in Condensed-Phase Radiolysis
- Source: arXiv (preprints)
- Date: 2026-09-15T02:03:36Z
- Categories: Cheminformatics, Reaction Informatics
- Authors: Nitesh Kumar, Jacob R. Milton, Eric Sivonxay, Brett A. Helms, Frances A. Houle, Samuel M. Blau
- External ID: 2609.16512v1
- Keywords: DFT
- Source URL: <https://arxiv.org/abs/2609.16512v1>
- Dashboard article: <https://tagirshin.com/chemradar/article?u=https%3A%2F%2Farxiv.org%2Fabs%2F2609.16512v1>
- PDF: <https://arxiv.org/pdf/2609.16512v1>

Abstract: Many important chemical systems, from radiation-driven processes to condensed-phase photochemistry, involve reaction mechanisms that are so complex it is a challenge to characterize them experimentally or predict them from chemical intuition. Existing computational approaches to mechanism discovery typically assess pathway importance through thermodynamic favorability alone, which does not provide time-dependent kinetics or inform how to model spatial inhomogeneities. Here we describe an integrated workflow that discovers complex reaction mechanisms without prescribing them and connects molecular-scale reactivity to spatiotemporal observables. The workflow combines high-throughput DFT, automated reaction network construction with chemical plausibility filtering, stochastic pathway sampling to identify reactions which are likely to occur, and spatially resolved reaction-diffusion kinetics simulations with explicit tracking of species in space and time. To demonstrate the workflow on a system of high complexity, we apply it to radiolytic chemistry in an extreme ultraviolet (EUV) organic polymer thin film photoresist, where a single 92 eV photon initiates cascades of radical ions, fragments, and low-energy electrons across a nanoscale radiolytic spur. Starting from over 3,300 species and millions of candidate reactions, the workflow identifies the most likely reaction pathways and produces spatiotemporal maps that resolve product formation on femtosecond-to-nanosecond timescales across a 15.5-nm domain. The simulations predict products detected experimentally and reveal that the identity of the initially photoionized species profoundly shapes the downstream product distribution through multi-step pathways governing the balance between deprotection and crosslinking reactions. The methodology is broadly applicable to complex condensed-phase reactive systems.

## VizChemoton: Visualization of Quantum Chemical Reaction Networks and Their Cheminformatics Features
- Source: Journal of Chemical Information and Modeling (journals)
- Date: 2026-09-15T00:00:00+00:00
- Categories: Cheminformatics, Reaction Informatics
- Authors: Enric Petrus, Diego Garay-Ruiz, Thomas Weymuth, Markus Reiher, Thomas B. Hofstetter
- Journal: Journal of Chemical Information and Modeling
- DOI: 10.1021/acs.jcim.6c02689
- Keywords: RDKit, Cheminformatics, density functional theory, quantum chemistry
- Source URL: <https://doi.org/10.1021/acs.jcim.6c02689>
- Dashboard article: <https://tagirshin.com/chemradar/article?u=https%3A%2F%2Fdoi.org%2F10.1021%2Facs.jcim.6c02689>

Abstract: Automated explorations of chemical reaction networks (CRNs) guided by quantum chemistry (QC) calculations are undergoing rapid development that enables numerous applications. However, CRNs are accessible primarily through specialized quantum chemistry software without interface for cheminformatic exploitation of the generated data. Here, we present VizChemoton, an open-source module within the reaction exploration framework Software for Chemical Interaction Networks (SCINE). VizChemoton enables visualization of CRNs and serves as an interface for cheminformatics applications based on conversion of molecular structures that accounts for radicals, zwitterions, and molecular complexes into string-based representations via RDKit mol objects. The conversion performance exceeded 90% of QC-generated compounds based on the evaluation of four CRNs constructed with density functional theory and extended tight-binding methods. We find that compounds absent from public structural databases correspond primarily to molecular complexes and unstable reaction intermediates. All essential QC and associated cheminformatics data are distributed in interoperable CSV and JSON formats to support downstream machine learning applications whereas reaction networks are readily visualized through a browser-based standalone HTML interface.

## Enhancing biocatalytic retrosynthesis with a graph-to-graph model
- Source: Chemical Science (journals)
- Date: 2026-09-09T08:03:49Z
- Categories: Reaction Informatics
- Authors: Lina Dong, Lin Yao, Yucheng Yang, Yuxiang Gao, Zhihui Jiang, Binju Wang
- Journal: Chemical Science
- DOI: 10.1039/d6sc04187f
- Keywords: retrosynthesis, enzyme
- Source URL: <https://doi.org/10.1039/d6sc04187f>
- Dashboard article: <https://tagirshin.com/chemradar/article?u=https%3A%2F%2Fdoi.org%2F10.1039%2Fd6sc04187f>

Abstract: BioG2G\_ESR is a unified graph-to-graph framework for one-step biocatalytic retrosynthesis, with enzyme sequence recommendation. It translates retrosynthesis into graph-to-graph, predicting reactants from products, preserving topology/stereochemistry.

## TSBench: A physics-grounded benchmark for evaluating LLM understanding of chemical reaction mechanisms
- Source: arXiv (preprints)
- Date: 2026-09-08T09:44:37Z
- Categories: Reaction Informatics, LLMs & Agents
- Authors: Xiaohu Xu, Tong Zhu
- External ID: 2609.08503v1
- Keywords: LLM, LLMs, synthesis planning
- Source URL: <https://arxiv.org/abs/2609.08503v1>
- Dashboard article: <https://tagirshin.com/chemradar/article?u=https%3A%2F%2Farxiv.org%2Fabs%2F2609.08503v1>
- PDF: <https://arxiv.org/pdf/2609.08503v1>

Abstract: Understanding a chemical reaction requires mapping a symbolic reactant-product description onto the three-dimensional pathway through which atoms rearrange, yet chemistry benchmarks for large language models (LLMs) largely probe factual knowledge and text-based reasoning. Here we introduce TSBench, a benchmark in which an LLM agent uses structure-editing tools to construct three-dimensional transition-state (TS) guesses verified by an automated quantum-chemical pipeline, yielding a physics-grounded pass/fail verdict. Across 546 evaluations of seven frontier LLMs on 78 elementary reactions, the aggregate success rate rose from 50.4% to 66.8% under diagnosis-driven revision; the best models approached 90% on the simplest reactions, yet performance dropped sharply with mechanistic complexity. Most failed attempts produced locally plausible saddle points whose reaction paths led to the wrong reactant-product pair, revealing local geometric intuition without a robust grasp of the global reaction coordinate. TSBench establishes a mechanism-level yardstick for LLM agents in mechanism-sensitive tasks such as synthesis planning and autonomous experimentation.

## SynOmega: Simplifying Retrosynthesis for Efficient Synthesizability Scoring
- Source: Journal of Chemical Information and Modeling (journals)
- Date: 2026-09-08T00:00:00+00:00
- Categories: Cheminformatics, Reaction Informatics
- Authors: Baicheng Zhang, Guoqing Zhang, Jun Jiang, Yi Luo
- Journal: Journal of Chemical Information and Modeling
- DOI: 10.1021/acs.jcim.6c02529
- Keywords: ChEMBL, AiZynthFinder, Retrosynthesis
- Source URL: <https://doi.org/10.1021/acs.jcim.6c02529>
- Dashboard article: <https://tagirshin.com/chemradar/article?u=https%3A%2F%2Fdoi.org%2F10.1021%2Facs.jcim.6c02529>

Abstract: Retrosynthesis-based synthesizability scoring triages molecules from generative design but is expensive: every score requires a multi-step search. We present SynOmega, an open-source toolkit that couples a single-step template model, an AND–OR route search, and a route-based synthesizability score (SynScore). Its single-step model can be restricted, at the reaction-template level, to simplifying disconnections that split the target into smaller precursors. On 1000 ChEMBL drug molecules this yields two findings. (1) The simplifying constraint cuts node expansions by about 30% on jointly solved targets and lowers the median search time by about a third. This reduction in search effort is the robust, budget-independent result; the small accompanying rise in solved rate is a secondary effect between the two separately trained models, not the isolated result of toggling one model’s action space. (2) As a complete system, under matched search depth, width and iteration budget, SynOmega reaches about 1.8× the solved rate of the open-source planner AiZynthFinder while searching about 13× faster. SynOmega thus offers a cheap, data-level action-space constraint that makes route-based synthesizability scoring more efficient without sacrificing solvability.

## A retrieval-augmented agent bridges virtual screening and automated synthesis
- Source: ChemRxiv (preprints)
- Date: 2026-09-07T00:00:00Z
- Categories: Design de novo, Reaction Informatics, LLMs & Agents
- Authors: Shan Zhu, Xin Chen, Pinqi Wang, Quan Jiang, Mengzhu Li, Qimeng Wang, Guoyun Zhao, Yuan Qi, Fenglei Cao, Li-Cheng Xu
- DOI: 10.26434/chemrxiv.15008409/v1
- External ID: 10.26434/chemrxiv.15008409/v1
- Keywords: retrosynthetic prediction, virtual screening, LLM, generative model, reaction conditions, retrosynthesis, retrosynthetic, synthesis planning
- Source URL: <https://doi.org/10.26434/chemrxiv.15008409/v1>
- Dashboard article: <https://tagirshin.com/chemradar/article?u=https%3A%2F%2Fdoi.org%2F10.26434%2Fchemrxiv.15008409%2Fv1>

Abstract: Translating virtually designed drug molecules into executable synthesis protocols is essential for connecting computational discovery with automated synthesis. However, most retrosynthesis methods focus on precursor generation, offer limited route diversity, and lack the detailed reaction conditions required for executable protocols. Here, we present Syn-RRAG, a retrieval-augmented LLM-driven synthesis agent connecting virtual screening with automated synthesis and validation. Syn-RRAG integrates initial retrosynthetic prediction, reaction-precedent retrieval and hierarchical refinement to generate structured protocols with complete experimental parameters. Its local generative model searches over complete reaction hypotheses rather than relying solely on token-level beam search, enabling exploration of feasible routes with distinct reaction classes and bond disconnections. It achieved strong retrosynthesis performance on two patent-derived reaction datasets. Syn-RRAG further generated executable plans for three virtually screened drug candidates, all of which were synthesized and validated on an automated platform. These results establish a workflow connecting virtual molecular design, synthesis planning and experimental validation.

## CampChem: Rubric-Grounded Adaptive Campaigns for Multi-Round Organic Synthesis with Residual Spectral Observation
- Source: ChemRxiv (preprints)
- Date: 2026-09-07T00:00:00Z
- Categories: Cheminformatics, Reaction Informatics, LLMs & Agents
- Authors: Zixuan Shi, Mengyao Qian, Yichao Sun, Haotian Gu, Jiarui Feng
- DOI: 10.26434/chemrxiv.15008453/v1
- External ID: 10.26434/chemrxiv.15008453/v1
- Keywords: SMILES, LLM, Mistral, reinforcement learning, retrosynthesis
- Source URL: <https://doi.org/10.26434/chemrxiv.15008453/v1>
- Dashboard article: <https://tagirshin.com/chemradar/article?u=https%3A%2F%2Fdoi.org%2F10.26434%2Fchemrxiv.15008453%2Fv1>

Abstract: Organic synthesis is a multi-round campaign, not a one-shot SMILES completion: a failed solvent, a poisoned catalyst, or leftover starting-material peaks in the 1H NMR force the chemist to revise conditions and try again. Tool-using chemistry agents and condition classifiers still emit a single recipe and then stop. We propose CampChem, an adaptive campaign agent built on a chemistry-instruction LLM. A longitudinal campaign adaptive planner (LCAP) keeps a short memory of failed condition slots and peak signatures and decodes the next catalyst/solvent/reagent tuple under a mask. Chemistry rubric-as-judge reinforcement learning (Chem-RaJ) scores stoichiometry, hazard compatibility, slot completeness, and spectral consistency instead of using a free-form LLM judge. A residual spectral booster (RSB) injects a permutation-invariant peak-set residual when NMR is available and gates to zero otherwise. On SMolInstruct, CampChem reaches 67.1% forward-synthesis and 36.1% retrosynthesis Top-1 exact match with a Mistral-7B LoRA backbone. On USPTO-Condition it attains 0.2940 overall Top-1, and on AutoLabs Experiment 5 it raises protocol F1 from 0.75 to 0.81. Ablations show that Chem-RaJ carries most of the spectrum-free gain, LCAP matters once a prior tuple has been rejected, and RSB contributes mainly on simulated NMR revision.

## Machine Learning for High-Throughput Reaction Yield Prediction in DNA-Encoded Library Synthesis
- Source: ChemRxiv (preprints)
- Date: 2026-09-07T00:00:00Z
- Categories: Property Prediction, Reaction Informatics, Library Design
- Authors: Tao Tang, Li Gao, Sen Gao, Hongyao Zhu, Guansai Liu, Jin Li, Xuemin Cheng
- DOI: 10.26434/chemrxiv.15008275/v2
- External ID: 10.26434/chemrxiv.15008275/v2
- Keywords: Reaction Yield Prediction, graph neural network, GNN, reaction type, molecular representation, DNA Encoded Library
- Source URL: <https://doi.org/10.26434/chemrxiv.15008275/v2>
- Dashboard article: <https://tagirshin.com/chemradar/article?u=https%3A%2F%2Fdoi.org%2F10.26434%2Fchemrxiv.15008275%2Fv2>

Abstract: DNA-encoded library (DEL) synthesis relies on efficient DNA-compatible reactions to ensure library fidelity and the reliability of downstream affinity selection. However, most existing reaction yield prediction models were developed for general organic synthesis and typically use full-atom representations of complete reactants and products. In DEL single-cycle synthesis validation, an appropriate model reaction is to focus on the exposed reactive handle, its local chemical environment, the incoming building block (BB), and the reaction type, rather than large constant backgrounds such as the DNA tag, linker, and pre-existing scaffold. Here, we propose a functional-group-centric machine learning framework for DEL single-cycle reaction yield prediction. The model uses only the exposed substrate functional group (FG) and its coarse-grained local environment, the incoming BB, and the reaction type as inputs. Benchmarking on a real-world high-throughput DEL single-cycle validation dataset shows that graph neural network (GNN) models based on this localized representation outperform traditional fingerprint-based baselines, with a graph attention network (GAT) achieving the best overall performance and showing strong robustness on highly imbalanced industrial data. These results indicate that foregrounding the local reaction center is not merely a simplification of molecular representation, but a task-aligned strategy that better reflects the chemistry and decision logic of DEL single-cycle transformations. This lightweight framework is readily compatible with practical DEL workflows and provides valuable supports on reagent quality control (QC), building block selection, high-fidelity library design, and retrospective analysis of anomalous results.

## Near-Complete Synthon Library via Retrosynthetic Analysis for Automated Organic Synthesis
- Source: ChemRxiv (preprints)
- Date: 2026-09-07T00:00:00Z
- Categories: Reaction Informatics
- Authors: Baicheng Zhang, Jing Chen, Xiaolong Zhang, Pieter E. S. Smith, Aoyuan Cheng, Runyang Miao, Hongping Liu, Linjiang Chen, Bin Jiang, Guoqing Zhang, Yi Luo, Jun Jiang
- DOI: 10.26434/chemrxiv.15008417/v1
- External ID: 10.26434/chemrxiv.15008417/v1
- Keywords: PubChem, Retrosynthetic, retrosynthesis
- Source URL: <https://doi.org/10.26434/chemrxiv.15008417/v1>
- Dashboard article: <https://tagirshin.com/chemradar/article?u=https%3A%2F%2Fdoi.org%2F10.26434%2Fchemrxiv.15008417%2Fv1>

Abstract: As idealized structural fragments for guiding bond disconnections, synthons are considered the conceptual backbone of retrosynthetic analysis. Yet retrosynthesis can be difficult because the choice of synthons requires extensive experience in organic synthesis, which prevents non-experts from accessing desired products. Here we address this knowledge gap by constructing a curated, machine interpretable database of commercially available synthon equivalents (hereafter valid synthons) distilled from large-scale retrosynthetic analysis. Mining over 40 million molecules from PubChem, we identify 1,617 essential valid synthons. These synthons, organized into (i) molecular backbones, (ii) functional groups, and (iii) linker motifs, were obtained via iterative retrosynthetic decomposition and stringent frequency based filtering tied to documented reaction precedents. The resulting set enables the in silico reconstruction of over 94% of these 40 million molecules. We demonstrate that such synthons reduce dependence on expansive reagent inventories, even though the routes themselves are non-conventional and algorithm-generated: using only seven representative valid synthons, we accessed three structurally and functionally divergent targets (melatonin, a natural sleep hormone; an azo dye; and a photoactive melatonin derivative), illustrating the breadth and practicality of a synthons to products methodology in the context of automated synthesis for non-experts in organic synthesis.

## Ranking-Based Surrogate Modeling for Bayesian Optimization under Small-Data Conditions
- Source: Journal of Chemical Information and Modeling (journals)
- Date: 2026-09-06T00:00:00+00:00
- Categories: Reaction Informatics
- Authors: Yuya Endo, Hiromasa Kaneko
- Journal: Journal of Chemical Information and Modeling
- DOI: 10.1021/acs.jcim.6c02115
- Source URL: <https://doi.org/10.1021/acs.jcim.6c02115>
- Dashboard article: <https://tagirshin.com/chemradar/article?u=https%3A%2F%2Fdoi.org%2F10.1021%2Facs.jcim.6c02115>

Abstract: Bayesian optimization (BO) is widely used for reaction optimization, but under small-data conditions, regression of absolute objective values may not align with the practical goal of prioritizing promising experiments. Here, we compared ranking-based BO (RankBO) combined with Thompson sampling (Rank-TS) with regression-based BO using two reaction benchmarks. Rank-TS showed a clear advantage on Direct Pd-catalyzed arylation in optimization performance, global ranking quality, and recovery of high-yielding conditions. For Suzuki–Miyaura coupling, which comprised 12 substrate combinations, its improvement was small in the pooled analysis. However, in the equal-weight macro analysis across the 12 substrate-defined search spaces, Rank-TS showed higher ranking quality and faster high-yield-condition recovery. These results indicate that Rank-TS is particularly useful for prioritizing reaction conditions within a fixed-substrate combination under small-data conditions, and that its relative performance depends on how the chemical search space and ranking task are defined.

## AbSynth: A Century of Total Syntheses as a Machine-Readable Resource for Analyzing Strategic Logic in Organic Chemistry
- Source: Journal of the American Chemical Society (journals)
- Date: 2026-09-01T00:00:00+00:00
- Categories: Reaction Informatics
- Authors: Maciej Czuba, Konrad Barnowski, Karol Molga, Zofia Pławecka, Agnieszka Wołos, Szymon Stodółkiewicz, Sebastian Baś, Wojciech Petrykowski, Wiktor Beker, Evan Saillard, Konrad Łęgowski, Rafał Roszak, Paweł Kowalczyk, Barbara Mikulak-Klucznik, Piotr Janasz, Emilia Mielke, Patrycja Pietruszewska, Maja Łazuchiewicz, Sylwia Wójcik, Jakub Zabielski, Dariusz Nostkiewicz, Gabriela Smaga, Ahmad Makkawi, Louis P. J.-L. Gadina, Jakub Skoczek, Tomasz Klucznik, Maxence Deschamps, Anna Żądło-Dobrowolska, Michał Michalak, Sara Szymkuć, Daniel T. Gryko, Yasemin Bilgi-Gadina, Bartosz A. Grzybowski
- Journal: Journal of the American Chemical Society
- DOI: 10.1021/jacs.6c08434
- Keywords: retrosynthetic algorithms, retrosynthetic, synthesis planning
- Source URL: <https://doi.org/10.1021/jacs.6c08434>
- Dashboard article: <https://tagirshin.com/chemradar/article?u=https%3A%2F%2Fdoi.org%2F10.1021%2Fjacs.6c08434>

Abstract: The design of synthetic routes to complex, stereochemically rich targets remains one of the most demanding challenges in organic chemistry. To interrogate the strategic logic underlying such efforts, we curated and digitized over 3,000 classical total syntheses comprising close to 60,000 individual steps, now freely available as the AbSynth collection. Analyses across this data set reveal that synthesis planning cannot be reduced to purely local, one-step-at-a-time decision-making: known complexity metrics and scoring functions fail to vary monotonically along synthetic trajectories, and conventional bond-disconnection heuristics apply only sporadically and lack the specificity required for automation. Instead, our study shows that nearly half of all steps correspond to structurally unproductive – yet strategically essential – plateaus of skeletal complexity. Navigating these plateaus requires retrosynthetic algorithms to embrace multistep reasoning rather than incremental scoring. We identify several such multistep heuristics, including tolerable lengths of nonskeletal disconnections and the “longevity” of functional groups across routes, and anticipate that AbSynth will serve as a resource for uncovering additional trends in synthesis design. By providing a century-spanning, machine-readable data set of complete routes, this work offers a foundation for advancing chemical AI toward the complexity of human expert planning.

## PKP-Diffmol: A Physicochemical Knowledge-Prompt Encoding and Latent Diffusion Framework for Molecular Property Prediction
- Source: Journal of Chemical Information and Modeling (journals)
- Date: 2026-09-01T00:00:00+00:00
- Categories: Cheminformatics, Property Prediction, Docking & Screening, ADMET & Safety, Reaction Informatics
- Authors: Ruizi Liu, Tongtong Yuan, Molin Guo
- Journal: Journal of Chemical Information and Modeling
- DOI: 10.1021/acs.jcim.6c01547
- Keywords: MoleculeNet, RDKit, SMILES, virtual screening, Transformer, Property Prediction, Molecular Property, bioactivity, Molecular representations
- Source URL: <https://doi.org/10.1021/acs.jcim.6c01547>
- Dashboard article: <https://tagirshin.com/chemradar/article?u=https%3A%2F%2Fdoi.org%2F10.1021%2Facs.jcim.6c01547>

Abstract: Molecular representations that capture both chemical information and property-relevant differences are essential for reliable molecular property prediction in drug discovery and toxicity assessments. Fragment-level SMILES representation learning is effective at modeling substructure semantics, but structural reconstruction objectives alone do not explicitly encode the descriptor-derived physicochemical information needed for downstream prediction, especially in small data and scaffold-split settings. To address this limitation, we propose PKP-DiffMol, a physicochemical knowledge-prompt encoding and latent diffusion framework built upon the pretrained SMI-EDITOR encoder. PKP-DiffMol improves fragment-level molecular representations by using numerical and semantic physicochemical knowledge. The numerical physicochemical prior branch encodes RDKit descriptors as continuous quantitative priors, while the physicochemical knowledge prompt encoder transforms the same descriptors into semantic physicochemical representations by using a Transformer-based text encoder. The two complementary views are then integrated through hierarchical physicochemical knowledge fusion to construct a chemically enriched, fused latent space. In this space, a quality-controlled conditional latent diffusion module learns label-conditioned distributions, generates synthetic molecular latent representations, and retains reliable samples through a Mahalanobis distance quality gate. Experiments on seven MoleculeNet classification datasets under scaffold splitting show that PKP-DiffMol increases the mean ROC-AUC from 77.80% for the SMI-EDITOR backbone to 80.53% and obtains the best result on five of the seven datasets. The largest gains are observed on BACE (+6.91%), MUV (+4.29%), and SIDER (+3.92%), showing strong predictive performance across bioactivity prediction, virtual screening, and adverse drug reaction prediction tasks. Further analyses indicate that numerical and semantic physicochemical knowledge improves class separability and the correspondence between representation distances and RDKit descriptor distances, while Mahalanobis-filtered label-conditioned latent augmentation supports prediction-head refinement.

## Swarm intelligence for chemical reaction optimization
- Source: Chem (journals)
- Date: 2026-09-01T00:00:00+00:00
- Categories: Reaction Informatics
- Authors: Rémi Schlama, Joshua W. Sin, Ryan P. Burwood, Kurt Püntener, Raphael Bigler, Philippe Schwaller
- Journal: Chem
- DOI: 10.1016/j.chempr.2026.103035
- Keywords: reaction conditions, Buchwald Hartwig, Suzuki
- Source URL: <https://doi.org/10.1016/j.chempr.2026.103035>
- Dashboard article: <https://tagirshin.com/chemradar/article?u=https%3A%2F%2Fdoi.org%2F10.1016%2Fj.chempr.2026.103035>

Abstract: We report ⍺-particle swarm optimization (PSO), a nature-inspired metaheuristic algorithm that augments canonical PSO with machine learning (ML) for highly parallel reaction optimization. Unlike black-box ML approaches that obscure decision-making processes, ⍺-PSO uses simple, physically intuitive swarm dynamics directly connected to experimental observables, enabling practitioners to understand each component driving optimization. We establish a theoretical framework for reaction landscape analysis using local Lipschitz constants to quantify reaction space "roughness," from smoothly varying reaction surfaces to landscapes with many reactivity cliffs. This analysis guides adaptive ⍺-PSO parameter selection, optimizing performance for different reaction topologies. Evaluation of ⍺-PSO across pharmaceutically relevant reaction benchmarks demonstrates competitive performance with state-of-the-art Bayesian optimization (BO) methods, whereas two prospective high-throughput experimentation (HTE) campaigns showed that ⍺-PSO identified optimal reaction conditions more rapidly than BO. Alongside our open-source ⍺-PSO implementation, we release 989 new high-quality Pd-catalyzed Buchwald-Hartwig and Suzuki reactions generated in these campaigns.

## An Agentic Retrobiosynthesis Framework with Learned Frontier Selection
- Source: arXiv (preprints)
- Date: 2026-08-31T12:40:42Z
- Categories: Design de novo, Reaction Informatics, LLMs & Agents
- Authors: Philippe Meyer, Guillaume Gricourt, Thomas Duigou, Joan Hérisson, Jean-Loup Faulon
- External ID: 2608.30702v1
- Keywords: RL, retrosynthesis
- Source URL: <https://arxiv.org/abs/2608.30702v1>
- Dashboard article: <https://tagirshin.com/chemradar/article?u=https%3A%2F%2Farxiv.org%2Fabs%2F2608.30702v1>
- PDF: <https://arxiv.org/pdf/2608.30702v1>

Abstract: Large language models are increasingly used as agents for multistep retrosynthesis, raising the question of how much their search policy contributes independently of the underlying reaction model. We investigate this question in a biological setting through rule-based retrobiosynthesis: a deterministic biochemical engine generates the same validated transitions for every method, searching for routes that terminate in metabolites available to an \\emph\{Escherichia coli\} chassis, while the policy only selects which frontier molecule to expand next. Prompted and LoRA-tuned Qwen2.5-7B policies use a strict choice-only interface. The fine-tuned policy reaches $65\\pm1$\\% solve rate at 10 expansions on LASER versus 59\\% for MCTS, and at 200 expansions reaches $78\\pm1$\\% versus 75\\% on LASER, $88\\pm3$\\% versus 80\\% on the RetroPath RL Golden benchmark, and $63\\pm2$\\% versus 45\\% on the BioNavi-NP benchmark. Fine-tuning also consistently outperforms direct prompting. These results show that route-supervised frontier selection can improve budgeted search without altering biochemical generation, although performance remains dependent on frontier construction and reaction ranking.

## Large language models as uncertainty-calibrated optimizers for experimental discovery
- Source: Nature Machine Intelligence (journals)
- Date: 2026-08-28T00:00:00+00:00
- Categories: Reaction Informatics, LLMs & Agents
- Authors: Bojana Ranković, Ryan-Rhys Griffiths, Philippe Schwaller
- Journal: Nature Machine Intelligence
- DOI: 10.1038/s42256-026-01283-z
- Keywords: LLMs, LLM, chemical descriptors, Buchwald–Hartwig
- Source URL: <https://doi.org/10.1038/s42256-026-01283-z>
- Dashboard article: <https://tagirshin.com/chemradar/article?u=https%3A%2F%2Fdoi.org%2F10.1038%2Fs42256-026-01283-z>

Abstract: From reaction optimization to molecular design, experimental discovery poses the same expensive question: which candidate to test next under time and resource constraints. Bayesian optimization provides principled answers but depends on domain expertise that rarely transfers. Large language models (LLMs) contain rich scientific knowledge but lack the calibrated uncertainty estimates crucial for high-stakes decisions. Here we show how training language models through Bayesian objectives enables their use as reliable optimizers guided by natural language. Our approach, GOLLuM (Gaussian process Optimized LLMs), teaches LLMs from experimental outcomes under uncertainty, transforming their overconfidence from a fundamental flaw into a precise learning signal. This signal reshapes the LLM embeddings so that experiments with similar outcomes cluster together, revealing structure in the design space. Starting from only ten low-performing experiments, GOLLuM generalizes across 23 tasks in organic synthesis, materials science, process chemistry and molecular design, ranking first on average among all competing methods. It matches traditional Bayesian optimization with over 40% fewer experiments and nearly doubles the discovery of high-performing Buchwald–Hartwig reactions over expert quantum-chemical descriptors and state-of-the-art LLMs (43% versus 24–25%). More broadly, GOLLuM points to a different paradigm for specializing foundation models: not through more data but through richer, uncertainty-guided information.

## Mapping Structured Absences in Scientific Papers: An Auditable LLM-Assisted Workflow Demonstrated on Cross-Coupling Chemistry
- Source: ChemRxiv (preprints)
- Date: 2026-08-28T00:00:00Z
- Categories: Reaction Informatics, LLMs & Agents
- Authors: Yuriko Ono, Masaharu Yoshioka, Tetsuya Taketsugu
- DOI: 10.26434/chemrxiv.15007996/v1
- External ID: 10.26434/chemrxiv.15007996/v1
- Keywords: LLM, Cross Coupling, Kumada
- Source URL: <https://doi.org/10.26434/chemrxiv.15007996/v1>
- Dashboard article: <https://tagirshin.com/chemradar/article?u=https%3A%2F%2Fdoi.org%2F10.26434%2Fchemrxiv.15007996%2Fv1>

Abstract: Scientific papers record what was discovered and communicated, but the resolution at which evidence relations are made explicit—a paper’s documentary resolution—often falls short of what later computational reuse, benchmarking, or data-driven synthesis requires. We present UQS (Unasked Question Structuring), an auditable LLM-assisted pipeline that takes scientific papers as input and produces structured absence records with evidence chains, an Evidence Pack (EP) dictionary specifying the evidence needed to reach the next evidential level, a bounded filling-status map across later literature under pre-locked criteria, and cross-paper research-question briefs. Applied to Kumada’s 1972 Ni-catalysed cross-coupling communication over a 54-year record window (1972–2026), UQS derives ten EPs and reveals a VQ-hierarchy inversion: methodology-level documentation is addressed early, whereas system-matched mechanistic documentation (VQ-3) appears only after 49 years (Mazet 2021), and quantitative-reliability documentation (VQ-4) remains a bounded non-detection—a pattern consistent with publication-incentive structure. The paper claims auditability, configuration-pinned rerun consistency, and internal robustness (four-model cross-comparison, N=6 multi-run aggregation); external expert validation is registered as future work.

## Mechanistic Reaction Prediction via Discrete Flow Matching on Graph-Structured Electron Occupation
- Source: arXiv (preprints)
- Date: 2026-08-27T17:50:44Z
- Categories: Design de novo, Reaction Informatics
- Authors: Nguyen Xuan-Vu, Octavian Susanu, Daniel Armstrong, Philippe Schwaller
- External ID: 2608.27429v2
- Keywords: de novo generation, Reaction Prediction, reaction type
- Source URL: <https://arxiv.org/abs/2608.27429v2>
- Dashboard article: <https://tagirshin.com/chemradar/article?u=https%3A%2F%2Farxiv.org%2Fabs%2F2608.27429v2>
- PDF: <https://arxiv.org/pdf/2608.27429v2>

Abstract: Chemical reactions are fundamentally transformations in electron space, yet most machine learning approaches model them either through de novo generation of product molecules or through heuristic graph edits that operate directly on molecular topology. We introduce MAELLE (MechAnistic Edit fLow-matching on eLectron rEarrangements), which instead models reactions as discrete flow matching over electron occupation vectors. Concretely, we formulate the reactant-to-product mapping as a Continuous-time Markov Chain (CTMC) over the graph-structured integer-valued electron occupation space defined on all bonding, non-bonding, and hydrogen sites. To construct the interpolants between the reactants and products, we generalize the discrete flow matching mixture path to an edit-based formulation, where the electron moves are interpolated using Optimal Transport, yielding a mechanism-like set of moves without elementary step annotations. MAELLE achieves competitive performance on the USPTO-480K benchmark compared with leading reaction prediction models. Beyond in-distribution learning, we evaluate robustness across two out-of-distribution settings - structural complexity and reaction type - and find that MAELLE maintains strong performance where existing methods degrade. Finally, because the learned flow operates over the full electron redistribution, MAELLE naturally recovers mechanistic trajectories that align with known chemistry and can predict side products of a reaction.

## Data-driven Effective Modeling of Stochastic Chemical Reaction Networks
- Source: arXiv (preprints)
- Date: 2026-08-26T06:24:43Z
- Categories: Reaction Informatics
- Authors: Yuan Chen, Weize Mao, Dongbin Xiu
- External ID: 2608.25421v1
- Source URL: <https://arxiv.org/abs/2608.25421v1>
- Dashboard article: <https://tagirshin.com/chemradar/article?u=https%3A%2F%2Farxiv.org%2Fabs%2F2608.25421v1>
- PDF: <https://arxiv.org/pdf/2608.25421v1>

Abstract: The Stochastic Simulation Algorithm (SSA), widely considered an exact algorithm for stochastic chemical reaction networks, suffers from high computational cost. In this work, we propose a data-driven effective model that operates on a user-defined coarse time step independent of the underlying microscopic reaction-event scale. This is accomplished by directly approximating the finite-time transition kernel of the continuous-time Markov chain induced by SSA, using a generative machine learning model trained on short bursts of SSA simulation data. The trained model constructs a stochastic propagator that recursively generates statistically consistent trajectories at the constant coarse time step, with significantly reduced computational cost. In this paper, we employ conditional normalizing flow as the stochastic propagator. A comprehensive set of numerical examples is presented to demonstrate the accuracy and efficiency of the proposed method.

## Bridging CGR representations and language models for reaction property prediction
- Source: Digital Discovery (journals)
- Date: 2026-08-25T15:43:19Z
- Categories: Cheminformatics, Reaction Informatics
- Authors: Giustino Sulpizio, Charlotte Gerhaher, Esther Heid, Kjell Jorner
- Journal: Digital Discovery
- DOI: 10.1039/d6dd00120c
- Keywords: atom mapping, SMILES, reaction SMILES, reaction property, Transformer, BERT, property prediction
- Source URL: <https://doi.org/10.1039/d6dd00120c>
- Dashboard article: <https://tagirshin.com/chemradar/article?u=https%3A%2F%2Fdoi.org%2F10.1039%2Fd6dd00120c>

Abstract: Accurate prediction of reaction properties, such as activation energies, enthalpies, rate constants and yields, is key to the design of efficient and sustainable chemical processes. While graph-based models using Condensed Graphs of Reaction (CGR) have shown strong performance, text-based Transformer models trained on reaction SMILES have gained popularity due to their scalability and performance. In this work, we analyze the performance of an existing CGR-based string representation for reaction property prediction, the SMILES/CGR. In contrast to conventional reaction SMILES, this representation explicitly encodes atom and bond changes within a unified SMILES-like sequence. A key limitation of SMILES/CGR is that it does not encode stereochemistry. To overcome this, we introduce the superimposed reaction SMILES (sr-SMILES), a compact CGR-based representation that preserves full stereochemical information. To evaluate both representations, we pretrained a BERT transformer on two million USPTO reactions via masked language modeling, and then finetuned separate models on six benchmark datasets. We compared against models trained on standard reaction SMILES, naïve baselines, and results from the literature. Our results show that sr-SMILES consistently matches or outperforms reaction SMILES, particularly on benchmarks where mechanistic information is critical, as well as in low-data regimes. In contrast, SMILES/CGR tends to underperform, highlighting that the details of CGR encoding matter. While sr-SMILES did not close the performance gap with graph-based CGR methods, it shows that language models can benefit from explicit mechanistic information when appropriately encoded. Moreover, by comparing SMILES/CGR and sr-SMILES models with different atom mappings, we reinforce the notion that CGR-based approaches are highly sensitive to the mapping quality. A plausible explanation for the similar performance observed across string-based representations is that language models implicitly learn the atom mapping during (pre-)training, rendering explicit mapping unnecessary for many prediction tasks.
