# Bioinfo Radar article snapshot
> 570 records returned from 9971 current records.

Generated: 2026-09-22T17:06:57.374284+00:00
Filters: days=7

Views are bounded by topic, source, days and limit; text search runs in the dashboard page, not on the server. Blog entries are metadata-only; use their source URL for full text.

## Deep Learning Protocols for Predicting Drug Mechanism of Action and Drug-Target Interactions.
- Source: Methods in molecular biology (Clifton, N.J.) (journals)
- Date: 2027-01-01
- Categories: Proteins & structural biology
- Authors: Yan Sun, Chengyou Liu, Zihao Jing, Yan Yi Li, Pingzhao Hu
- Journal: Methods in molecular biology (Clifton, N.J.)
- DOI: 10.1007/978-1-0716-5539-9\_3
- External ID: 42734743
- Source URL: <https://doi.org/10.1007/978-1-0716-5539-9_3>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1007%2F978-1-0716-5539-9_3>

Abstract: Understanding drug mechanisms of action (MOA) and predicting drug-target interactions (DTIs) are fundamental challenges in modern drug discovery and development, hindered by high costs, long development timelines, and limited knowledge of compound activity and molecular targets. Here, we present two deep learning-based computational protocols designed to address these challenges. The first framework employs directed message passing neural networks (D-MPNN) to predict drug MOA from chemical-genetic interaction profiles (CGIPs), by learning how molecular structures perturb biological pathways through systematic profiling across genetically sensitized strains. The second framework, iNGNN-DTI, utilizes interpretable nested graph neural networks combined with pretrained molecule models to predict DTIs, leveraging cross-attention mechanisms to provide insights into binding determinants. We highlight the application of these methods to key therapeutic areas, including antibacterial drug discovery and drug repurposing for COVID-19 therapeutics. Each protocol provides comprehensive guidance on data preparation, model implementation, validation strategies, and result analysis. These computational approaches offer scalable, cost-effective tools for accelerating therapeutic development by bridging chemical structure, molecular interactions, and systems-level biological responses.

## Identification of Genome-Wide Chromatin Structural Aberration in Cancer by Hi-C Analysis.
- Source: Methods in molecular biology (Clifton, N.J.) (journals)
- Date: 2027-01-01
- Authors: Shuntaro Isogai, Atsushi Okabe, Atsushi Kaneda
- Journal: Methods in molecular biology (Clifton, N.J.)
- DOI: 10.1007/978-1-0716-5539-9\_17
- External ID: 42734757
- Source URL: <https://doi.org/10.1007/978-1-0716-5539-9_17>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1007%2F978-1-0716-5539-9_17>

Abstract: Aberrant three-dimensional genome organization is a hallmark of cancer, often driving oncogene activation through mechanisms such as enhancer hijacking. High-throughput chromosome conformation capture (Hi-C) maps these interactions on a genome-wide scale. Unlike earlier dilution-based methods, in situ Hi-C performs proximity ligation within intact nuclei, minimizing random ligation noise and enabling fine-scale structure detection. This chapter describes an optimized in situ Hi-C protocol tailored for cancer cell lines using MboI digestion and biotin-mediated pull-down to generate high-complexity libraries. We further outline a computational workflow that extends beyond standard topological mapping of compartments and topologically associating domains to identify cancer-specific aberrations. Specifically, we focus on detecting chromosomal rearrangements (structural variants) and characterizing the distinct circular topology of extrachromosomal DNA. This integrated experimental and analytical framework provides the necessary tools to dissect the spatial dysregulation underlying tumor evolution.

## Improving Image Quality in 10× Visium Spatial Transcriptomics Using Vispro.
- Source: Methods in molecular biology (Clifton, N.J.) (journals)
- Date: 2027-01-01
- Categories: Biological imaging
- Authors: Huifang Ma, Zhicheng Ji
- Journal: Methods in molecular biology (Clifton, N.J.)
- DOI: 10.1007/978-1-0716-5539-9\_20
- External ID: 42734760
- Source URL: <https://doi.org/10.1007/978-1-0716-5539-9_20>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1007%2F978-1-0716-5539-9_20>

Abstract: 10× Visium is a widely used spatial transcriptomics platform that enables joint profiling of gene expression and the spatial locations of cells. However, the histology images generated by the 10× Visium platform often contain technical artifacts, including fiducial markers and background noise, which degrade image quality. Here, we describe how a computational method, Vispro, can be applied to process and enhance these images. The resulting high-quality images lead to improved performance across a range of downstream analyses.

## In Silico Single-Cell Frame work for Modeling Intestinal Stem and Transit-Amplifying Progenitor Cells Dynamics.
- Source: Methods in molecular biology (Clifton, N.J.) (journals)
- Date: 2027-01-01
- Categories: Genomics & sequence analysis, Single-cell & spatial
- Authors: Brinda Balasubramanian
- Journal: Methods in molecular biology (Clifton, N.J.)
- DOI: 10.1007/978-1-0716-5412-5\_2
- External ID: 42763848
- Keywords: rna, single cell, scrna, cell type, cell atlas
- Source URL: <https://doi.org/10.1007/978-1-0716-5412-5_2>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1007%2F978-1-0716-5412-5_2>

Abstract: Single-cell RNA sequencing (scRNA-seq) has revolutionized the ability to resolve cellular heterogeneity within complex tissues, enabling the identification of discrete cell states. Here, we present an in silico analytical pipeline designed to characterize intestinal stem cells (ISC), transit-amplifying (TA) progenitors, and BEST4⁺ enterocyte precursors from human scRNA-seq datasets, with a focus on inflammatory contexts such as inflammatory bowel disease (IBD). The pipeline integrates dataset acquisition, quality control, normalization, dimensionality reduction, unsupervised clustering, and cell type annotation using a reference cell atlas. We implemented iterative subsetting and re-clustering of ISC and TA compartments to identify inflammation-associated subpopulations and epithelial biomarkers. While demonstrated in the context of IBD, this computational framework is broadly applicable to other tissues and pathological conditions where stem/progenitor dynamics underpin disease progression and tissue repair.

## Inferring Gene Regulatory Networks in Stem Cells: Methods and Applications.
- Source: Methods in molecular biology (Clifton, N.J.) (journals)
- Date: 2027-01-01
- Categories: Genomics & sequence analysis, Single-cell & spatial, Systems & networks
- Authors: Daniela Solano-Galarza, Simone Roeh, Thomas Walzthoeni
- Journal: Methods in molecular biology (Clifton, N.J.)
- DOI: 10.1007/978-1-0716-5539-9\_1
- External ID: 42734741
- Keywords: chromatin, dna, rna, single cell, cell type, gene regulatory
- Source URL: <https://doi.org/10.1007/978-1-0716-5539-9_1>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1007%2F978-1-0716-5539-9_1>

Abstract: Gene regulatory networks (GRNs) represent the complex interplay of transcription factors, regulatory elements, and target genes that orchestrate cellular identity and function, playing a crucial role in the differentiation and maintenance of stem cells. This chapter provides an overview of experimental and computational methodologies for inferring GRNs, with particular emphasis on single-cell approaches. We first review key experimental techniques for detecting transcription factor binding sites, chromatin accessibility, and DNA motifs, alongside essential databases that support GRN reconstruction. We then introduce computational inference methods that can be categorized into four principal frameworks: correlation-based approaches, regression and machine learning models, probabilistic and deep learning methods, and integrative or message-passing frameworks. To illustrate practical application, we present a case study applying the pySCENIC workflow to a peripheral blood mononuclear cell single-cell RNA sequencing dataset from mouse, demonstrating how regulon-based analysis can reveal cell-type-specific regulatory programs. This chapter aims to serve as a practical guide for researchers seeking to understand and implement GRN inference methodologies in stem cell biology and related fields.

## Teratoma Formation and Genomic Profiling Using Multi-Omics Approaches.
- Source: Methods in molecular biology (Clifton, N.J.) (journals)
- Date: 2027-01-01
- Categories: Genomics & sequence analysis, Single-cell & spatial
- Authors: Benjamin L Kidder
- Journal: Methods in molecular biology (Clifton, N.J.)
- DOI: 10.1007/978-1-0716-5539-9\_25
- External ID: 42734765
- Keywords: genomic, chromatin, rna, rna seq, gene expression, epigenetic, multi omics, single cell, scrna
- Source URL: <https://doi.org/10.1007/978-1-0716-5539-9_25>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1007%2F978-1-0716-5539-9_25>

Abstract: Teratoma formation is the gold standard assay for evaluating the developmental pluripotency of human and mouse embryonic stem cells (ESCs) and induced pluripotent stem cells (iPSCs). Following subcutaneous injection into immunodeficient mice, pluripotent stem cells spontaneously differentiate into derivatives representing all three embryonic germ layers-ectoderm, mesoderm, and endoderm. Beyond serving as a functional assay for pluripotency, teratomas provide a unique three-dimensional model system for studying early human development and lineage specification in vivo. This chapter describes comprehensive protocols for teratoma formation in immunodeficient mice, tissue processing for multiple downstream genomic applications, and multi-omics profiling approaches. We detail methods for embryonic stem cell culture, teratoma generation via subcutaneous injection, tissue dissection and processing for chromatin immunoprecipitation followed by sequencing (ChIP-Seq), RNA sequencing (RNA-Seq), single-cell multiome profiling combining chromatin accessibility (ATAC-Seq) and gene expression (scRNA-Seq), and histological analysis using hematoxylin and eosin (H&E) staining. Additionally, we provide bioinformatics workflows for analyzing the resulting genomic datasets to characterize the epigenetic and transcriptional landscapes of teratoma-derived tissues. These methods enable comprehensive molecular characterization of developmental processes and provide valuable resources for stem cell biologists studying pluripotency, differentiation, and early embryonic development.

## Topic-Driven Bibliometrics and Trend Intelligence for Stem Cell and Cancer Research.
- Source: Methods in molecular biology (Clifton, N.J.) (journals)
- Date: 2027-01-01
- Categories: Tools & resources
- Authors: Benjamin L Kidder
- Journal: Methods in molecular biology (Clifton, N.J.)
- DOI: 10.1007/978-1-0716-5539-9\_4
- External ID: 42734744
- Source URL: <https://doi.org/10.1007/978-1-0716-5539-9_4>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1007%2F978-1-0716-5539-9_4>

Abstract: The rapid growth of biomedical literature has created an urgent need for computational tools that enable researchers to systematically analyze publication trends, identify emerging research themes, and map the evolution of scientific fields. PubMed Atlas is a command-line and web-enabled workflow for topic-driven bibliometrics and trend intelligence using PubMed E-utilities. The pipeline executes PubMed queries, retrieves matching PMIDs, downloads full metadata records in batches, parses structured information (title, abstract, authors/affiliations, MeSH terms, publication types, grants, keywords, DOI), and stores normalized data in a local SQLite database for rapid querying and visualization. A Streamlit dashboard provides interactive exploration of publication trends, journal distributions, MeSH term summaries, geographic distributions, and recent article browsing with direct PubMed links. This protocol describes the installation, configuration, and operation of PubMed Atlas for cancer stem cell and stem cell transcriptional network research, and other fields, enabling investigators to conduct reproducible bibliometric analyses and identify knowledge gaps in rapidly evolving fields.

## TORC: Target-Oriented Reference Construction for Supervised Cell-Type Identification in scRNA-seq.
- Source: Methods in molecular biology (Clifton, N.J.) (journals)
- Date: 2027-01-01
- Categories: Genomics & sequence analysis, Single-cell & spatial, Tools & resources
- Authors: Xin Wei, Wenjing Ma, Zhijin Wu, Hao Wu
- Journal: Methods in molecular biology (Clifton, N.J.)
- DOI: 10.1007/978-1-0716-5539-9\_5
- External ID: 42734745
- Source URL: <https://doi.org/10.1007/978-1-0716-5539-9_5>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1007%2F978-1-0716-5539-9_5>
- Code: <https://github.com/weix21/TORC>

Abstract: Cell-type identification is a crucial step in single-cell RNA-seq (scRNA-seq) data analysis, for which supervised methods are preferred due to their accuracy and efficiency. The quality of the reference data plays an important role in cell-type identification performance, but systematic strategies for selecting and reconstructing reference data remain limited. We present Target-Oriented Reference Construction (TORC), a widely applicable strategy for constructing reference data from available labeled cells given a target dataset. TORC alleviates the differences in data distribution and cell-type composition between the reference and the target. TORC combines initial supervised prediction, optional reference expansion using target cells with high-confidence predicted labels, and reference reconstruction guided by estimated cell-type compositions. Here, we provide detailed, step-by-step instructions describing the input requirements, configurable parameters, and practical considerations for applying TORC in real scRNA-seq analyses. TORC is available at https://github.com/weix21/TORC , where an example implementation using an MLP-based classifier is provided.

## A curated structural dataset of peptide-protein complexes reveals biases in existing datasets and principles of peptide binding.
- Source: Protein science : a publication of the Protein Society (journals)
- Date: 2026-10-01
- Categories: Proteins & structural biology, Tools & resources
- Authors: Rahma Hamdani, Javier Delgado, Luis Serrano
- Journal: Protein science : a publication of the Protein Society
- DOI: 10.1002/pro.70779
- External ID: 42741994
- Source URL: <https://doi.org/10.1002/pro.70779>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1002%2Fpro.70779>

Abstract: Peptide-protein interactions are fundamental to many biological processes, and peptide design is gaining interest due to its therapeutic potential. This has led to the emergence of various structural databases, such as PepBDB and Peptipedia, that classify peptides as polypeptides with fewer than 50 amino acids. These databases provide valuable starting points for studying peptide recognition by protein partners and are widely used for machine learning applications, docking, and scoring functions. However, under such a length definition for peptides, we have very different cases that are likely to confound the analysis of peptide binding, including miniproteins, peptide-peptide complexes, intramolecular peptide disulfide bonds, intermolecular disulfide bridges, and proteins undergoing internal cleavage, such as serpins, as well as non-natural amino acids and covalently bound cofactors. Here, we present a rigorous classification of peptide-protein complexes to generate datasets suitable for comparative energetic analysis with a focus on peptides that are unstructured in the absence of their target protein. The analysis of this dataset shows that peptide binding is typically driven by a small number of hotspot residues mainly enriched in aromatic and bulky hydrophobic side chains. Their number of hotspots and their spatial organization depend on peptide length, secondary structure, and covalent constraints. Short peptides rely on central anchor regions, whereas longer peptides distribute hotspots more broadly, with helices showing periodic spacing and β-strands relying more on backbone-mediated stabilization. Disulfide bonds further decrease the number of hotspots per peptide length by either pre-organizing the peptide or acting as covalent anchors. This work provides a curated resource and general principles for peptide recognition. It highlights the importance of structurally classifying peptide-protein complexes to avoid bias in downstream computational and machine-learning applications.

## A Quantitative Systems Pharmacology Model of Human Leucine Metabolism.
- Source: CPT: pharmacometrics & systems pharmacology (journals)
- Date: 2026-10-01
- Categories: Systems & networks
- Authors: J Cody Herron, Anna Sher, Yingbo Ma, Elisabeth Roesch, Sebastian Miculța-Câmpeanu, Christopher Rackauckas, Kevin J Filipski, Rachel Roth Flach, Theodore R Rieger, Cynthia J Musante, Richard Allen
- Journal: CPT: pharmacometrics & systems pharmacology
- DOI: 10.1002/psp4.70328
- External ID: 42754827
- Source URL: <https://doi.org/10.1002/psp4.70328>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1002%2Fpsp4.70328>

Abstract: Branched-chain amino acids (BCAAs) are essential dietary components that humans cannot synthesize. Altered BCAA levels have been associated with biomarkers or potential risk factors in several metabolic disorders, including insulin resistance, type 2 diabetes, obesity, and cardiovascular disease. However, the underlying mechanisms regulating BCAA metabolism and how cellular or signaling modifications may alter BCAA levels are yet to be fully elucidated. To investigate the fate of plasma and intracellular leucine, we developed a mathematical model of human leucine metabolism. Through a virtual population approach, the model was constructed based on known biology and data and calibrated to fit available acute leucine and α-ketoisocaproic acid (KIC) clinical challenge data in healthy subjects. The rate-limiting step of BCAA catabolism is oxidative decarboxylation by branched-chain α-ketoacid dehydrogenase (BCKDH), a process that is negatively regulated by phosphorylation by branched-chain α-ketoacid dehydrogenase kinase (BDK). Recent preclinical studies have reported that inhibition of BDK leads to significant lowering of plasma BCAA levels. Modeling predicts that the magnitude of reduction observed in plasma BCAA and BCKA levels upon BDK inhibition may require incorporation of an additional regulatory mechanism, such as feedback on leucine and KIC uptake into tissues. This leucine systems model has implications for drug discovery and development, enables a mechanistic understanding of clinical data, and could be used as a tool for the design and analysis of therapeutic modifications of leucine and KIC metabolism.

## DisPhaseDB 2.0: Improved interpretation of disease-associated variants in liquid-liquid phase separation proteins with agent-accessible querying.
- Source: Protein science : a publication of the Protein Society (journals)
- Date: 2026-10-01
- Categories: Tools & resources
- Authors: Justo Garcia-Messina, Alvaro M Navarro, Cristina Marino-Buslje
- Journal: Protein science : a publication of the Protein Society
- DOI: 10.1002/pro.70786
- External ID: 42742059
- Source URL: <https://doi.org/10.1002/pro.70786>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1002%2Fpro.70786>

Abstract: Membraneless organelles formed through liquid-liquid phase separation (LLPS) are fundamental to cellular organization and are involved in multiple processes, including responses to stimuli and stress. The study of disease-associated variants in LLPS proteins remains vital for understanding protein dysfunction in human diseases. However, maintaining specialized, integrative resources is often hindered when database updates rely on human intervention, typically resulting in long intervals between updates. Here, we introduce a major update to DisPhaseDB (https://disphasedb.leloir.org.ar/), a comprehensive resource for disease-associated variants in LLPS proteins, integrated with an open-source Snakemake workflow organizing systematic data acquisition and parsing into traceable steps. Crucially, the automated system continuously fetches data from source databases, keeping DisPhaseDB up-to-date without the delays of manual maintenance. The updated release expands the database with additional proteins, increases disease annotation coverage by 174%, and adds clinical significance and allele frequency annotations to enhance variant interpretation. To improve accessibility, we also introduce a Model Context Protocol (MCP) server that establishes a standardized interoperability layer, enabling AI agents and large language models to directly query database records through structured operations. This architecture grounds generative workflows in a trusted source, replacing unconstrained web retrieval and reducing unsupported content. In a comparative benchmark, data retrieval through the MCP server achieved a mean F1 score of 0.99, compared to 0.30 for unguided generative retrieval. Together, these developments position DisPhaseDB2.0 as a maintainable resource for LLPS-related variants, optimizing reproducible data access for both human researchers and emerging agentic AI workflows.

## ESpma: A method for assessing biological/non-biological interfaces using point-clouds-based structural features and protein language models.
- Source: Protein science : a publication of the Protein Society (journals)
- Date: 2026-10-01
- Categories: Proteins & structural biology
- Authors: Sarah Nozawa, Kentaro Tomii, Yoshinori Fukasawa
- Journal: Protein science : a publication of the Protein Society
- DOI: 10.1002/pro.70790
- External ID: 42752891
- Source URL: <https://doi.org/10.1002/pro.70790>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1002%2Fpro.70790>
- Code: <https://github.com/fukasawa-group/espma>

Abstract: Distinguishing biological protein-protein interfaces from non-biological contacts remains an important task in structural biology, particularly as protein complex structures continue to accumulate. Recent advances in protein language models (pLMs) have expanded their use across diverse protein prediction tasks, including sequence-based protein-protein interaction prediction. Such approaches often pool representations over entire sequences, whereas structure-based methods commonly rely on more complex graph- or geometry-based integration. We therefore asked how much interface-relevant information could be extracted by simply pooling pLM embeddings over structurally defined local regions. Here, we revisited biological-versus-crystal interface classification as a structurally well-defined testbed for this question. We developed a framework that compares full-sequence pooling, interface-localized pooling of pLM embeddings, and multimodal integration of sequence-derived embeddings with point-cloud representations of protein surfaces. On two benchmark datasets, interface-localized pooling achieved stronger performance than full-sequence or non-interacting surface pooling across three pLM backbones. Despite its simplicity, the resulting representation performed within the range of established methods that rely on explicit evolutionary analysis or geometric modeling. Incorporating explicit geometric surface information changed performance slightly, without reaching statistical significance. Because the model is linear, the interface-level score decomposes exactly into per-residue contributions, which varied within amino-acid types. Together, our results indicate that the principal gain arises from localizing pLM embeddings to physically interacting residues, enabling a lightweight and accessible implementation for biological-versus-crystal interface classification. Code and scripts are available on GitHub: https://github.com/fukasawa-group/espma.

## EvoMut: A computational framework for engineering oxidative stability in proteins.
- Source: Protein science : a publication of the Protein Society (journals)
- Date: 2026-10-01
- Categories: Proteins & structural biology, Tools & resources
- Authors: Seyed Shahriar Arab, Chenlin Hsieh, Nathan E Lewis
- Journal: Protein science : a publication of the Protein Society
- DOI: 10.1002/pro.70774
- External ID: 42741978
- Source URL: <https://doi.org/10.1002/pro.70774>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1002%2Fpro.70774>

Abstract: Amino acid oxidation is a major cause of protein instability and loss of function in therapeutic and industrial settings. Although methionine, cysteine, tryptophan, tyrosine, histidine, lysine, and arginine residues are widely recognized as oxidation-prone, only a subset of such residues is dominant functional hotspots, and not all are suitable targets for mutation. Identifying these vulnerable, yet engineerable, sites remains a major challenge. Here, we present EvoMut, a residue-level analytical framework for evaluating both oxidative vulnerability and mutation feasibility. EvoMut estimates oxidation risk by integrating structural features, local functional context, intrinsic chemical susceptibility, and evolutionary conservation. A central feature of the framework is the explicit separation of oxidation risk from mutation feasibility. Specifically, candidate substitutions are evaluated only after high-risk residues are identified and ranked by evolutionary substitution patterns. Application of EvoMut to multiple proteins, and evaluation with experimental data, showed that oxidation-prone residues differ markedly in their engineering potential. EvoMut distinguishes residues that are both oxidation-sensitive and evolutionarily permissive from those that are chemically vulnerable but functionally constrained. By providing residue-level mechanistic insight, EvoMut offers a practical framework for the rational design of oxidation-resistant proteins. EvoMut is freely available as a web server at https://proteus.cmm.uga.edu/evomut.

## Expanded Stoichiometric Model of Chondrocyte Metabolism: Response to Cyclical Shear and Compressive Loading.
- Source: Journal of biomechanical engineering (journals)
- Date: 2026-10-01
- Categories: Systems & networks
- Authors: Aubrey H Kimmel, Adrienne D Arnold, Ayten E Erdogan, Ronak Kommineni, Erik P Myers, Breschine Cummins, Ross P Carlson, Ronald K June 2nd
- Journal: Journal of biomechanical engineering
- DOI: 10.1115/1.4072447
- External ID: 42552809
- Source URL: <https://doi.org/10.1115/1.4072447>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1115%2F1.4072447>

Abstract: Cartilage deterioration is a hallmark of osteoarthritis, and there is substantial interest in developing strategies for cartilage repair. Cyclical mechanical stimulation has been known for decades to drive synthesis of cartilage matrix proteins. Matrix synthesis requires activation of central metabolism for producing precursors to nonessential amino acids required for protein translation. However, there are gaps in knowledge regarding how mechanical stimuli affect chondrocyte central metabolism. Here, we find that cyclical shear and compression drive differences in chondrocyte central metabolism in a sex-dependent manner. Based on established biochemistry, we developed and tested a stoichiometric model containing 139 metabolites and 172 reactions from central metabolism that includes production of key cartilage matrix proteins. We then used experimental metabolomics data from shear and compressive stimulation of osteoarthritic chondrocytes to constrain this model and ran multiple simulations examining the potential for producing matrix proteins and ATP. Our results show that both shear and compression can stimulate osteoarthritic chondrocyte metabolism in a manner consistent with production of cartilage matrix proteins, with notable differences between male and female chondrocytes. Additionally, and importantly, our simulation results suggest that nitrogen availability is a key limitation to chondrocyte synthesis of matrix proteins. These results are a starting point for using central metabolism of chondrocytes to optimize synthesis of matrix proteins for cartilage repair. For example, increasing glutamine levels in the presence of cyclical compression has potential to increase production of both types II and VI collagen. These strategies have potential for improving cartilage tissue engineering and repair.

## Feature-Space Planes Searcher: A Universal Domain Adaptation Framework for Interpretability and Computational Efficiency.
- Source: IEEE transactions on pattern analysis and machine intelligence (journals)
- Date: 2026-10-01
- Categories: Proteins & structural biology
- Authors: Zhitong Cheng, Yiran Jiang, Yulong Ge, Yufeng Li, Zhongheng Qin, Rongzhi Lin, Jianwei Ma
- Journal: IEEE transactions on pattern analysis and machine intelligence
- DOI: 10.1109/tpami.2026.3703974
- External ID: 42301825
- Keywords: structure prediction, framework
- Source URL: <https://doi.org/10.1109/tpami.2026.3703974>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1109%2Ftpami.2026.3703974>

Abstract: Domain shift, characterized by degraded model performance during the transfer from labeled source domains to unlabeled target domains, poses a persistent challenge for deploying deep learning systems. Current unsupervised domain adaptation (UDA) methods predominantly rely on fine-tuning feature extractors-an approach limited by high computational cost, reduced interpretability, and poor scalability to modern architectures. Our analysis reveals that models pre-trained on large-scale data exhibit domain-invariant geometric patterns in their feature space, characterized by intra-class clustering and inter-class separation, thereby preserving transferable discriminative structures. These findings suggest that cross-domain performance degradation is often associated with decision-boundary misalignment, and that correcting such misalignment can serve as an effective alternative to feature adaptation, particularly when pretrained representations are sufficiently strong. Unlike fine-tuning entire pre-trained models, which risks introducing unpredictable feature distortions, we propose the Feature-space Planes Searcher (FPS): a novel domain adaptation framework that optimizes decision boundaries by leveraging these geometric patterns while keeping the feature encoder frozen. This streamlined approach enables interpretable analysis of adaptation while substantially reducing memory and computational costs through offline feature extraction, permitting full-dataset optimization in a single training cycle. Moreover, we introduce an Intra-Class Distance Metric (ICDM) that enables fully unsupervised hyperparameter selection without requiring target-domain labels. Evaluations on public benchmarks show that FPS achieves competitive performance across standard benchmarks, with notable gains in several settings and tasks. FPS scales efficiently with large multimodal models and shows versatility across diverse domains including protein structure prediction, remote sensing classification, and earthquake detection. We anticipate FPS will provide a simple, effective, and generalizable framework for domain adaptation tasks.

## Testing for Genetic Interactions in Complex Disease With Distance Correlation.
- Source: Biometrical journal. Biometrische Zeitschrift (journals)
- Date: 2026-10-01
- Categories: Genomics & sequence analysis, Mathematical biology & statistics
- Authors: Fernando Castro-Prado, Javier Costas, Dominic Edelmann, Wenceslao González-Manteiga, David R Penas
- Journal: Biometrical journal. Biometrische Zeitschrift
- DOI: 10.1002/bimj.70150
- External ID: 42703867
- Source URL: <https://doi.org/10.1002/bimj.70150>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1002%2Fbimj.70150>

Abstract: Understanding epistasis (genetic interaction) may shed some light on the genomic basis of common diseases, including disorders of maximum interest due to their high socioeconomic burden, like schizophrenia. Distance correlation is an association measure that characterizes general statistical independence between random variables, not only the linear one. Here, we propose distance correlation as a novel tool for the detection of epistasis from case-control data of single-nucleotide polymorphisms. On the methodological side, we highlight the derivation of the explicit asymptotic null distribution of the test statistic. We show that this is the only way to obtain enough computational speed for the method to be used in practice, in a scenario where the resampling techniques found in the literature are impractical. Our simulations show satisfactory calibration of significance, as well as comparable or better power than existing methodology. We conclude with the application of our technique to a schizophrenia genetics dataset, obtaining biologically sound insights.

## Detection and sequencing of Ap2N-capped RNAs in human cells
- Source: RNA-Seq Blog (feeds)
- Date: 2026-09-21T11:13:16+00:00
- Categories: Blog
- Source URL: <https://www.rna-seqblog.com/detection-and-sequencing-of-ap2n-capped-rnas-in-human-cells/>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fwww.rna-seqblog.com%2Fdetection-and-sequencing-of-ap2n-capped-rnas-in-human-cells%2F>
- Abstract: not stored for this record.

## Single-cell RNA sequencing provides a closer look at the Aedes aegypti midgut
- Source: RNA-Seq Blog (feeds)
- Date: 2026-09-21T11:13:08+00:00
- Categories: Blog
- Source URL: <https://www.rna-seqblog.com/single-cell-rna-sequencing-provides-a-closer-look-at-the-aedes-aegypti-midgut/>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fwww.rna-seqblog.com%2Fsingle-cell-rna-sequencing-provides-a-closer-look-at-the-aedes-aegypti-midgut%2F>
- Abstract: not stored for this record.

## A reduced glycosaminoglycan-linked residual-strain model captures regional opening angle changes after depletion in the porcine thoracic aorta
- Source: bioRxiv (preprints)
- Date: 2026-09-21
- Authors: Labrosse, M. R., Ghadie, N., St-Pierre, J.-P., Boodhwani, M.
- DOI: 10.64898/2026.07.15.738269
- Source URL: <https://doi.org/10.64898/2026.07.15.738269>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.07.15.738269>

Abstract: Residual stresses in arteries are commonly revealed by the spring opening of a ring after a radial cut. Glycosaminoglycans (GAGs) contribute to this response, but fixed-charge-density (FCD)-driven Donnan swelling alone does not fully explain the opening-angle reduction measured after enzymatic GAG depletion. We therefore tested whether regional mean depletion responses are consistent with a removable preferred stretch field shaped by the transmural FCD profile and superposed on a structural field retained after depletion. A reduced analytical-computational axisymmetric closure framework was applied to regional-average measurements from the ascending aorta, arch, and descending porcine thoracic aorta. Each control state was fitted using its measured circumferential opening angle, geometry, material properties, and through-wall FCD profile. GAG depletion was represented by removing the FCD-linked preferred-stretch component. One coefficient governing this removable component was selected jointly from the three measured regional post-depletion angles. A one-layer wall was the primary parsimonious model; a two-layer wall tested anatomical robustness. The one-layer model fitted a shared coefficient of -1.4226 x 10-3 (mEq/L)-1 and predicted depleted angles of 82.639\{degrees\}, 44.032\{degrees\}, and 19.979\{degrees\}, compared with measured values of 85\{degrees\}, 41\{degrees\}, and 18\{degrees\} (three-region RMSE 2.50\{degrees\}). The two-layer model fitted -1.51746 x 10-3 (mEq/L)-1 and predicted 82.972\{degrees\}, 44.273\{degrees\}, and 18.534\{degrees\} (RMSE 2.24\{degrees\}). These results show that the three regional mean depletion responses can be represented by one common FCD-linked removable preferred-stretch contribution. Because the model uses regional averages and fitted region-specific control structural fields, it does not establish specimen-level predictive validity or uniquely identify the underlying GAG-mediated mechanism. Donnan swelling remains mechanically relevant, but it is insufficient alone to explain the measured regional response.

## A statistical review of polygenic risk scores: from heuristics to model-based inference.
- Source: Statistical applications in genetics and molecular biology (journals)
- Date: 2026-09-21
- Categories: Genomics & sequence analysis
- Authors: Xuan Huang, Wei Jiang
- Journal: Statistical applications in genetics and molecular biology
- DOI: 10.1515/sagmb-2026-0007
- External ID: 42760896
- Keywords: genome, genomic, inference
- Source URL: <https://doi.org/10.1515/sagmb-2026-0007>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1515%2Fsagmb-2026-0007>

Abstract: Polygenic risk scores (PRS) were initially developed as pragmatic tools to aggregate genome-wide association study (GWAS) signals for individual-level prediction, relying on heuristic strategies such as clumping and thresholding to approximate independence among variants. Although computationally efficient and widely accessible, these early approaches were sensitive to tuning parameters and limited in their ability to capture the diffuse signal characteristic of highly polygenic traits. As GWAS sample sizes expanded and biobank-scale resources emerged, methodological priorities shifted toward statistically principled models that explicitly represent linkage disequilibrium, effect-size heterogeneity, and population structure. In this review, we examine the methodological evolution of PRS construction from threshold-based aggregation to fully model-based inference frameworks, including linear mixed models, LD-aware Bayesian shrinkage approaches, machine learning, and recent multi-ancestry extensions, and summarize practical considerations for method selection under different data-access, LD-reference, tuning, and ancestry settings. Collectively, these developments mark a transition from heuristic scoring algorithms to a mature, statistically grounded paradigm for genomic risk prediction.

## AS-OCT dataset with anatomical structure segmentation and scleral spur localization in cataract and glaucoma
- Source: Scientific Data (journals)
- Date: 2026-09-21T00:00:00+00:00
- Categories: Tools & resources
- Authors: Jiongning Zhao, Huihui Fang, Yuetong Yang, Xinyu Fu, Jingru Deng, Jinghe Yu, Xingying Yan, Xinya Hu, Xiaoqing Wang, Yuting Hu, Di Gong, Zhe Zhang, Wei Chi, Weihua Yang, Yanwu Xu
- Journal: Scientific Data
- DOI: 10.1038/s41597-026-08315-8
- Source URL: <https://doi.org/10.1038/s41597-026-08315-8>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1038%2Fs41597-026-08315-8>

Abstract: Anterior Segment Optical Coherence Tomography (AS-OCT) provides high-resolution, non-invasive visualization of the anterior eye and is widely used for clinical diagnosis and surgical planning. However, automated analysis of AS-OCT images remains limited by the lack of publicly available datasets with comprehensive anatomical annotations and disease labels. Here we present the AS-OCT Multi-Structure and Disease Classification Dataset (ASOCT-MSDC), a curated dataset designed for anatomical structure segmentation and disease classification. The dataset contains 1106 AS-OCT images from 1106 eyes of 627 individuals across four clinical categories: normal, glaucoma, cataract, and glaucoma-cataract comorbidity. Each image underwent stringent quality assessment, and three standardized image quality scores—eyelid obscuration, black-line anomaly, and central light artifact—are provided as metadata. Expert-validated annotations include segmentation masks for the anterior chamber, iris, lens, and nucleus, together with scleral spur localization points. By combining anatomical annotations, disease labels, and image quality metadata, ASOCT-MSDC supports research on anatomical segmentation, disease classification, and image quality assessment in anterior segment OCT images.

## Bayesian Efficient Coding
- Source: bioRxiv (preprints)
- Date: 2026-09-21
- Categories: Computational neuroscience
- Authors: Park, I. M., Pillow, J. W.
- DOI: 10.1101/178418
- Source URL: <https://doi.org/10.1101/178418>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1101%2F178418>

Abstract: The efficient coding hypothesis, which proposes that neurons are optimized to maximize information about the environment, has provided a guiding theoretical framework for sensory and systems neuroscience. More recently, a theory known as the Bayesian Brain hypothesis has focused on the brain's ability to integrate sensory and prior sources of information in order to perform Bayesian inference. Although pieces of a connection between these two hypotheses have appeared in prior work, a general formulation that treats the optimality criterion as an arbitrary functional of the posterior distribution -- and thereby admits both information-based and other objectives within a single framework -- has remained largely implicit. Here we make this formulation explicit, developing a Bayesian theory of efficient coding that defines Bayesian efficient codes in terms of four basic ingredients: (1) a stimulus prior distribution; (2) an encoding model; (3) a capacity constraint, specifying a neural resource limit; and (4) a loss functional, quantifying the desirability or undesirability of various posterior distributions. Classic efficient codes arise as the special case in which the loss functional is the posterior entropy, leading to a code that maximizes mutual information, but alternate loss functionals give solutions that differ dramatically from information-maximizing codes. Within this framework we introduce \{\\it covtropy\}, a novel family of losses defined on posterior distributions and parameterized by a single exponent, and use it to show that decorrelation of sensory inputs -- optimal under classic efficient codes in low-noise settings -- can be disadvantageous for objectives that penalize large errors. We then reanalyze Laughlin's seminal data on contrast coding in the blowfly large monopolar cell, and find that the measured response nonlinearity is better explained by minimizing $L\_p$ reconstruction error with $p = 1/2$ than by information maximization, overturning a forty-year-old interpretation. Bayesian efficient coding thus enlarges the family of codes that are optimal under different objectives and provides a more general framework for understanding the design principles of sensory systems.

## Benchmarking confidence estimation and rescoring for cyclic peptide-protein complex predictions
- Source: bioRxiv (preprints)
- Date: 2026-09-21
- Categories: Proteins & structural biology
- Authors: Li, Z., Yuan, Y., Hu, K., Pan, P., He, F.
- DOI: 10.64898/2026.08.20.746104
- Source URL: <https://doi.org/10.64898/2026.08.20.746104>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.08.20.746104>

Abstract: Cyclic peptides are a rapidly expanding class of therapeutics, but the reliability of deep-learning structure prediction for cyclic peptide-protein complexes has not been systematically evaluated. We assembled a curated benchmark of 111 nonredundant complexes spanning five cyclization chemistries and assessed two co-folding models, Boltz and Protenix, each generating 100 poses per target (22,200 total). Stratifying all poses by complex attributes, we found that disulfidecyclized peptides and small protein targets (200 or fewer target residues) were predicted significantly worse by both tools, with target size the largest and most consistent effect; overall accuracy nevertheless remained high (median top-pose DockQ of about 0.89, 96-98% of targets Acceptable or better), indicating that pose generation is rarely the bottleneck. Conversely, native model ranking scores correlated only moderately with pose quality (Spearman rank correlations of 0.53-0.66): approximately 12% of poses showed high model ranking score/confidence despite poor pose DockQ quality, and the highest-quality pose was not ranked first for nearly every target. We therefore augmented the native score with externally computed interface descriptors normalized by chain length, principally the per-residue density of inter-chain hydrogen bonds, in a gradient-boosted rescoring model evaluated under target-grouped cross-validation that prevents leakage, improving out-of-fold ROC-AUC for both tools, significantly so for Protenix. Together, these findings identify pose ranking, rather than pose generation, as the major limitation of current cyclic peptide-protein complex prediction and demonstrate that complementary structural features can improve confidence-based pose selection.

## Benchmarking generative models for COI DNA barcoding
- Source: Scientific Reports (journals)
- Date: 2026-09-21T00:00:00+00:00
- Categories: Genomics & sequence analysis
- Authors: Cho-I Moon, Dae Kwon Song, Jie Eun Park, Jun Yang Jeong, Chan Eui Hong, Hyeon Jun Shin, Hyeok Lee, Kyoung Won Lee, Hee-ju Hwang, Yong Seok Lee
- Journal: Scientific Reports
- DOI: 10.1038/s41598-026-63888-z
- Source URL: <https://doi.org/10.1038/s41598-026-63888-z>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1038%2Fs41598-026-63888-z>

Abstract: Cytochrome c oxidase subunit I (COI) DNA barcoding is widely used for species identification and biodiversity studies. However, COI datasets exhibit high intra-species similarity and significant inter-species imbalance, which limits sequence analyses. To address data scarcity, deep learning based generative models have been explored for sequence generation. We implemented six generative models incorporating gated recurrent unit (GRU) layers, Transformer blocks, and convolutional layers to generate species-specific COI sequences across four taxonomic groups: Cypraeidae, Drosophila, Bats, and Birds. The generated sequences were evaluated in terms of plausibility, phylogenetic consistency, and diversity. Finally, GRU-based autoregressive language model achieved the best performance. It preserved codon-level structures to real data, with GC₃ content differences (Δ) ≤ 0.004, codon bias JSD ≤ 0.013, and ORF mean length differences (Δ) < 0.05. It also reproduced genetic structures with intra-species K2P mean differences (Δ) ≤ 0.13, real–synthetic K2P mean ≤ 0.09, and barcode gap rate differences (Δ) ≤ − 0.6. Additionally, it generated sequences with minimal redundancy, indicated by JSD-kmer ≤ 0.03, Self-BLEU differences (Δ) ≤ 0.001, and AA values between 0.54 and 0.75. These results suggest that GRU-based COI sequence generation can serve as a robust simulation strategy for addressing data scarcity and imbalance in bioinformatics applications.

## CD55-CD319-CX3CR1 flow cytometry gating strategy recapitulates scRNA-seq-defined memory CD8 T cell subpopulations
- Source: bioRxiv (preprints)
- Date: 2026-09-21
- Categories: Genomics & sequence analysis, Single-cell & spatial
- Authors: Bohacova, P., Terekova, M., Shpynov, O., Francis, T., Husarcikova, K., Tsurinov, P., Kossl, J., Kleverov, M., Keppel, M., Harridge, S. D. R., Singh, N., Artyomov, M. N.
- DOI: 10.64898/2026.09.15.751807
- Keywords: transcriptomic, epigenetic, epigenetically, scrna, single cell
- Source URL: <https://doi.org/10.64898/2026.09.15.751807>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.15.751807>

Abstract: Human CD8 T cells have traditionally been classified into naive, central memory, effector memory, and terminal effector subsets using CCR7 and CD45RA expression, a framework that has guided immunological research and clinical immune monitoring for nearly three decades. However, recent single-cell studies have revealed transcriptionally distinct CD8 T cell populations, including GZMK-, GZMB-, and central memory-like states, raising important questions regarding their relationship to canonical flow cytometric subsets. Here, we systematically integrated transcriptomic, epigenetic, and phenotypic analyses to evaluate the correspondence between these classification schemes. We demonstrate that conventional CCR7-CD45RA gating generates heterogeneous populations containing extensive mixtures of transcriptionally and epigenetically distinct CD8 T cell states, resulting in poor resolution of biologically meaningful subsets. To address this limitation, we developed a surface-marker framework based on CD55, CD319, and CX3CR1 that accurately identifies transcriptionally defined human CD8 T cell populations using standard flow cytometry. This strategy enables direct isolation of viable cells, including GZMK-expressing cells increasingly implicated in aging, chronic inflammation, autoimmunity, and cancer, which previously could only be identified using intracellular staining or single-cell sequencing. Functional characterization of purified subsets revealed marked differences in proliferative capacity, cytokine production, and cytotoxic activity, demonstrating that transcriptionally defined states possess distinct immune functions. Together, these findings establish a biologically grounded framework for CD8 T cell classification and provide a practical platform for mechanistic studies, biomarker discovery, and cellular immunotherapy applications.

## celltypeEnrich: a consensus-based scRNA-seq cluster annotation tool
- Source: bioRxiv (preprints)
- Date: 2026-09-21
- Categories: Genomics & sequence analysis, Single-cell & spatial, Tools & resources
- Authors: Rutledge, S., Tuteja, G.
- DOI: 10.64898/2026.09.15.751735
- Source URL: <https://doi.org/10.64898/2026.09.15.751735>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.15.751735>

Abstract: Motivation Single-cell RNA sequencing (scRNA-seq) cluster annotation is a critical step in data analysis. Current methods are time-consuming, difficult to reproduce, or limited in tissue or species coverage. Results We developed celltypeEnrich, a cluster-level annotation tool that uses a hypergeometric test to identify enrichment of cell-type-specific genes from input gene lists. Enrichment results from up to 26 reference datasets are used to determine a consensus annotation. Benchmarking using scRNA-seq datasets from three tissues spanning two species showed 62-72% annotation accuracy for celltypeEnrich, generally outperforming other tools, which had either lower accuracy, incomplete tissue coverage, or the need for parameter optimization. The performance of celltypeEnrich remained stable when input gene lists were down-sampled to 25% of their original size. Availability and Implementation celltypeEnrich is freely available at (https://celltypeenrich.gdcb.iastate.edu) as an R Shiny web application under the MIT license for non-profit academic use.

## Code-multiplexed multi-frequency impedance cytometry with a unified deep-unfolding network.
- Source: Microsystems & nanoengineering (journals)
- Date: 2026-09-21
- Categories: Single-cell & spatial
- Authors: Wonjun Lee, Sindy K Y Tang
- Journal: Microsystems & nanoengineering
- DOI: 10.1038/s41378-026-01420-z
- External ID: 42764344
- Keywords: single cell
- Source URL: <https://doi.org/10.1038/s41378-026-01420-z>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1038%2Fs41378-026-01420-z>

Abstract: Impedance flow cytometry (IFC) is a label-free, single-cell measurement technique that captures biophysical properties beyond traditional biochemical markers. Code-multiplexing allows parallelization of IFC with simple hardware but requires advanced signal processing algorithms to resolve overlaps in signals originating from different channels. Existing methods, however, rely on multiple task-specific networks with template-based linear fitting, which loses accuracy under nonlinear or unstable conditions common in microfluidic experiments. Prior studies have also been restricted to demultiplexing single-frequency impedance measurements. To this end, we develop a unified deep-unfolding network for analyzing code-multiplexed, multi-frequency IFC data. We unfold the successive-interference cancellation (SIC) algorithm into a deep-learning network, where repeated stages of a single multitask network implement iterative signal estimation and interference cancellation that reflect the structural prior of SIC. To mitigate nonlinear signal stretching and amplification, our network recognizes events by predicting bit-level intensity and duration. For multi-frequency impedance profiling, we apply least-squares fitting to the predicted single-frequency real-impedance trace to map the real and imaginary impedance traces at other frequencies. On the cell-bead mixture evaluation dataset, our pipeline resolves overlaps from singlets to triplets reliably, reconstructs impedance-intensity distributions accurately, and enables multi-frequency impedance profiling. As a demonstration of principle, we use our pipeline to perform label-free quantification of basophil activation from code-multiplexed, multi-frequency IFC measurements. Consistent with our previous study, impedance opacity correlates well with activation levels measured by fluorescence flow cytometry. In summary, our study demonstrates the feasibility and utility of a deep-unfolding network that extends code-multiplexed, multi-frequency IFC to label-free single-cell functional assays.

## ContiTE: continuous manifold MoE for few-shot cross-tissue mRNA translation efficiency prediction
- Source: Briefings in Bioinformatics (journals)
- Date: 2026-09-21T00:00:00+00:00
- Categories: Genomics & sequence analysis, Proteins & structural biology
- Authors: Yuxiao Wei, Qi Zhang, Sainan Huo, Xuezhong Zhou
- Journal: Briefings in Bioinformatics
- DOI: 10.1093/bib/bbag518
- Source URL: <https://doi.org/10.1093/bib/bbag518>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1093%2Fbib%2Fbbag518>

Abstract: Predicting mRNA translation efficiency across diverse tissues is a critical yet challenging task due to the phenomenon of concept shift, where identical genomic sequences exhibit distinct functional profiles across varying cellular environments. Existing static models are often limited by fixed parameterization, struggling to capture these continuous regulatory variations without requiring computationally expensive fine-tuning. To address these limitations, we propose ContiTE, a continuous manifold mixture-of-experts (MoE) framework explicitly designed for efficient few-shot cross-tissue adaptation. Unlike traditional MoE architectures that rely on discrete output-space mixing, ContiTE operates on a continuous parameter manifold by dynamically synthesizing domain-specific weights through a context-aware hyper-router that linearly combines shared atomic basis kernels. This paradigm enables smooth interpolation of translational rules and provides a flexible mechanism to model complex, context-dependent biological regulation. Furthermore, we introduce a gradient-based test-time adaptation strategy that allows the model to rapidly align to new, unseen tissues by solely optimizing low-dimensional context embeddings while keeping the backbone parameters frozen. Experimental results on a comprehensive human and mouse atlas demonstrate that ContiTE significantly outperforms state-of-the-art methods in few-shot scenarios, improving average Pearson correlation coefficient by 12.09% and $R^\{2\}$ by 18.22%. By mitigating negative transfer and requiring only limited target-domain data for calibration, ContiTE provides a computational framework for tissue-specific mRNA translation-efficiency prediction and demonstrates a parameter-reconstruction strategy that may be extensible to other cross-domain sequence-learning tasks, subject to task-specific validation.

## Controlled evaluation of architectural, classifier, and training refinements in MolTrans-based drug-target interaction prediction
- Source: BMC Bioinformatics (journals)
- Date: 2026-09-21T00:00:00+00:00
- Categories: Proteins & structural biology
- Authors: Hao Pang, Fiseha Berhanu Tesema, Tianxiang Cui, Yuan Cheng, Yanwen Mao
- Journal: BMC Bioinformatics
- DOI: 10.1186/s12859-026-06634-6
- Source URL: <https://doi.org/10.1186/s12859-026-06634-6>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1186%2Fs12859-026-06634-6>

Abstract: Drug–target interaction (DTI) prediction is a central task in computational drug discovery, but performance improvements can be difficult to attribute when architectural modifications and training changes are introduced simultaneously. This study evaluates a Bidirectional Cross-Attention and Global Aggregation DTI model (BCAG-DTI) through eight controlled configurations that separate bidirectional cross-attention, global average–max pooling, classifier design, and optimisation strategy. All principal experiments use fixed training, validation, and test partitions and five matched random seeds on BindingDB, BIOSNAP, and DAVIS. Relative to the MolTrans baseline, the complete BCAG-DTI configuration improves mean area under the receiver operating characteristic curve (AUROC) from 0.8815 to 0.9063 on BindingDB, from 0.8631 to 0.8891 on BIOSNAP, and from 0.8808 to 0.8955 on DAVIS. The corresponding gains in area under the precision–recall curve (AUPRC) are 0.0831, 0.0346, and 0.0619, respectively. Controlled ablation shows that the enhanced classifier achieves the highest mean AUROC on BindingDB and BIOSNAP and the highest mean AUPRC and F1-score on all three datasets, whereas the cross-attention-plus-pooling configuration with enhanced training achieves the highest mean AUROC on DAVIS. The enhanced training strategy also provides substantial improvements, while cross-attention alone reduces performance under the baseline training configuration. In BIOSNAP robustness experiments, BCAG-DTI improves MolTrans for unseen drugs, unseen proteins, and 70–90% missing-interaction settings, whereas its AUROC and F1-score are slightly lower at 95% missing data. As an external same-split reference, CPI-GGS evaluated on the fixed BIOSNAP partitions achieves 0.8619 ± 0.0023 AUROC, 0.8645 ± 0.0042 AUPRC, and 0.7938 ± 0.0028 F1-score, compared with 0.8891 ± 0.0078, 0.8992 ± 0.0067, and 0.8178 ± 0.0094 for BCAG-DTI. This comparison is interpreted in the context of different input preprocessing pipelines and substantial differences in model capacity. Attention case analysis further indicates that cross-attention weights should be treated as model-internal allocation patterns rather than validated binding contacts. Overall, the results show that classifier design and optimisation account for a substantial portion of the improvement over MolTrans, while the contribution of cross-modal architectural components is optimisation-sensitive and dataset-dependent.

## Copy Number Variant Detection by Exome/Genome Sequencing Versus Chromosomal Microarray: A Comparative Study of Over 9,000 Clinical Cases.
- Source: Genetics in medicine : official journal of the American College of Medical Genetics (journals)
- Date: 2026-09-21
- Categories: Genomics & sequence analysis
- Authors: Sarah R Poll, Flavia M Facio, Kirsty McWalter, Patricia C Lopes, Amanda Lindy, Bethany Friedman, Kirsten Kelly, Olivia Trimmier, Jane Juusola, Paul Kruszka, Wei Wang, Lisa Dyer, Lisong Shi, Britt Johnson, Ganka Douglas
- Journal: Genetics in medicine : official journal of the American College of Medical Genetics
- DOI: 10.1016/j.gim.2026.102727
- External ID: 42765364
- Keywords: genome, variant detection
- Source URL: <https://doi.org/10.1016/j.gim.2026.102727>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1016%2Fj.gim.2026.102727>

Abstract: PURPOSE: Copy number variants (CNVs) are implicated in many health conditions. Chromosomal microarray (CMA) has traditionally been the first-tier test for CNV detection. However, exome and genome sequencing (ES/GS) can identify CNVs alongside other variant types. This study compared CNV detection using CMA versus ES/GS in a large clinical cohort. METHODS: CNV calls from CMA and ES/GS were analyzed in a diverse cohort of over 9,000 individuals tested in a high-throughput clinical laboratory. Concordance between platforms was evaluated, with discordant findings reviewed to determine their nature and causes. RESULTS: ES/GS showed >99% concordance with CMA. CMA results not detected on ES/GS were typically CNV 41%. CONCLUSION: ES/GS matched or exceeded CMA performance for CNV detection and identified additional variant types. These findings support the adoption of ES/GS as first-tier tests for CNV detection, streamlining diagnostic workflows, and improving diagnostic rate by capturing both small and large structural variants with high accuracy.

## Correlation-aware discovery of co-occurring mutational signatures in cancer
- Source: bioRxiv (preprints)
- Date: 2026-09-21
- Categories: Genomics & sequence analysis, Tools & resources
- Authors: Jin, H., Geiger, B., Glodzik, D., Gulhan, D. C., Park, P. J.
- DOI: 10.64898/2026.09.14.751548
- Source URL: <https://doi.org/10.64898/2026.09.14.751548>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.14.751548>

Abstract: Somatic mutations in cancer genomes record the activities of diverse mutational processes. Mutational signature analysis has advanced mechanistic understanding of mutagenesis and informed clinical decision-making, yet existing methods assume independence among signatures--an unrealistic assumption that can produce composite or contaminated signatures, reduce detection power, and yield inconsistent results. Here we present Cornet (CORrelated NMF ExTraction), a framework for mutational signature discovery that explicitly models co-occurring processes and jointly infers signatures and their correlation structure. Benchmarking on simulated data shows Cornet more accurately recovers distinct signatures under strong correlations. Applied to cancer genomes, Cornet enables unsupervised discovery of the colibactin-associated signature SBS88 in oral cancers and identifies the tobacco smoking signature SBS4 in bladder cancer, where it was previously thought absent. Cornet also uncovers a novel mutational process implicated in early-onset colorectal cancer and a signature arising from the interplay between tobacco smoking and ERCC2-mutation-driven nucleotide-excision repair deficiency. Together, these results demonstrate that modeling correlations among mutational processes is essential for high-resolution signature discovery and dissecting the mutational etiology of human cancer.

## Differential analysis of microbial interaction networks
- Source: Briefings in Bioinformatics (journals)
- Date: 2026-09-21T00:00:00+00:00
- Categories: Evolution & metagenomics
- Authors: Marianna Milano, Pietro Hiram Guzzi
- Journal: Briefings in Bioinformatics
- DOI: 10.1093/bib/bbag522
- Source URL: <https://doi.org/10.1093/bib/bbag522>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1093%2Fbib%2Fbbag522>
- Code: <https://github.com/mmilano87/NetMicrobiome>

Abstract: Microbiome studies increasingly indicate that disease-associated shifts cannot be understood from compositional changes alone. The functional architecture of microbial communities—encoded in patterns of association among microbial gene families—may reveal how these systems reorganize across biological conditions. Here, we present a network-based framework for characterizing microbiome rewiring across conditions. The approach combines condition-specific network inference, differential network analysis, and pathway-level network analysis to identify associations that are gained, lost, or altered between groups, with a specific focus on sex-dependent differences. We apply the framework to inflammatory bowel disease, type 2 diabetes, and atherosclerotic cardiovascular disease (ACVD), comparing male and female-specific microbial gene family networks within each disease context. Across these settings, differential networks flag large numbers of candidate rewired associations; however, permutation testing (sex labels shuffled, group sizes preserved, 500 permutations for gene-family networks, and 1000 for pathway networks) shows that the global amount of apparent rewiring is not greater than expected under the null at the global or edge level in any cohort, and that most edges exclusive to one group are induced by group-specific feature filtering rather than by a genuine change in association ($\\sim $80%–83% in ACVD). We therefore present the method as a rigorously validated framework and a cautionary case study: the differential-network machinery is sound, but the headline biological signal in a naive analysis is largely a property of correlation thresholding and, for the longitudinal inflammatory bowel disease (IBD) cohort, of pseudoreplication. The only non-null result across all validations is a SOHPIE-DNA per-taxon test in the IBD disease arm (15 taxa at FDR $< 0.05$), which we report as a single nominal finding requiring independent replication. Code, data, and supplementary information are available at https://github.com/mmilano87/NetMicrobiome.

## Estimation of stratified seroprevalence directly from raw serological assay measurements with multi-level Bayesian mixture modelling
- Source: medRxiv (preprints)
- Date: 2026-09-21
- Categories: Mathematical biology & statistics, Tools & resources
- Authors: Wymant, C., Kendall, M., Hay, J. A., Fraser, C.
- DOI: 10.64898/2026.09.17.26363301
- Source URL: <https://doi.org/10.64898/2026.09.17.26363301>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.17.26363301>
- Code: <https://github.com/BDI-pathogens/dvsb>

Abstract: Estimating seroprevalence for an infection -- the proportion of individuals with antibodies -- and its variability between subpopulations is a key first step for research and the allocation of resources for treatment and prevention. The relevant raw data is often serological assay measurements that serve as a proxy for antibody level, such as optical density values from enzyme-linked immunosorbent assays (ELISA). Analysis pipelines typically proceed through sequential steps of fitting calibration data independently in each separate batch, transforming assay measurements to antibody levels via each fitted calibration relationship, classifying antibody levels into discrete serostatus by comparison to a threshold, and finally comparing seropositive proportions between subpopulations. Such approaches have numerous limitations including loss of information, overconfidence (discarding uncertainty), and inappropriate choice of threshold. We developed a Bayesian statistical model that replaces the sequential steps of the typical pipeline by multiple levels within a single hierarchical model, treating each sample's antibody level as a model parameter rather than directly observed data. This integrates the connections from the raw assay measurements all the way through to subpopulation variability in seroprevalence, allowing partial pooling of information between related variables and the propagation of uncertainty from each part of the model throughout the whole of the rest of the model. We calculate observation probabilities for all assay measurements from both samples and calibration data, allowing easy identification of outlying calibration or sample measurements. We replace a single hard classification threshold for disease status by continuous probabilities that are adapted to each subpopulation. We allow a flexible multivariate random-effects logistic regression to capture variability in seroprevalence between subpopulations. Using simulated data we found that the typical stepwise approach gave prevalence estimates far from the true values with narrow confidence intervals. Our method, dvsb, had markedly greater accuracy. We report the run time and convergence of dvsb when applied to a real dataset for Lassa fever IgG antibodies measured with ELISA, comprising 72,863 measurements for 21,391 unique samples (results reported elsewhere). dvsb is available at https://github.com/BDI-pathogens/dvsb.

## EvSpark: Lossless Speculative Decoding for Hybrid DNA Foundation Models
- Source: bioRxiv (preprints)
- Date: 2026-09-21
- Categories: Genomics & sequence analysis, Tools & resources
- Authors: Ding, H., Nie, W., Wu, N., Qiu, T.
- DOI: 10.64898/2026.09.02.749017
- Source URL: <https://doi.org/10.64898/2026.09.02.749017>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.02.749017>

Abstract: Hybrid DNA foundation models combine convolutional, recurrent, and attention layers, making speculative decoding more difficult than truncating a KV cache. We present EvSpark, a speculative decoding system for Evo2 that verifies draft blocks in parallel and restores all three classes of inference state by selecting retained intermediate states, without replay. A compact, hidden-state-conditioned drafter proposes each block in one parallel forward pass. On Evo2 7B, a 48-prompt benchmark with three training seeds yields 2.96x on 43 real-sequence prompts and 3.27x including the five synthetic controls. Acceleration persists at 262k-token context (1.84x - 2.43x on two bacterial genomes) and over 32k generated tokens. Retraining the same drafter architecture for Evo2 20B and 40B yields 2.18x - 2.46x on real sequences and 2.51x - 2.78x on the full suite. Autoregressive drafter comparisons and batch measurements show why low draft latency, rather than acceptance alone, determines the gain. The method preserves the target distribution in exact arithmetic. In bf16, greedy tests find no non-tie divergences across 48 prompts and six checkpoints; sampling tests expose residual numerical sensitivity, especially in repetitive sequences. A cost-efficient 7B drafter requires 1.06 incremental GPU-hours of training, excluding teacher-data collection, and achieves 2.82x on real sequences. In regulatory-DNA design, EvSpark achieves a median complete-workflow speedup of 1.57x over a calibrated batched native baseline.

## Generating protein hydrogels with customizable stress relaxation behavior via deep learning-driven entanglement design
- Source: Nature Communications (journals)
- Date: 2026-09-21T00:00:00+00:00
- Categories: Proteins & structural biology
- Authors: Puqing Deng, Yutong Wu, Hong Kiu Francis Fok, Linyan Li, Wen-Bin Zhang, Fei Sun, Hanyu Gao
- Journal: Nature Communications
- DOI: 10.1038/s41467-026-77607-9
- Source URL: <https://doi.org/10.1038/s41467-026-77607-9>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1038%2Fs41467-026-77607-9>

Abstract: Protein hydrogels are promising artificial extracellular matrices (ECMs) for 3D stem cell and organoid culture due to their favorable stress relaxation behavior (a decrease in stress in response to strain). Inter-chain entangled motifs, in which different protein chains are interlaced, represent a powerful strategy to synthesize such hydrogels. However, designing these motifs with tailored properties such as binding energy remains a major challenge due to the difficulty of simultaneously controlling these properties while ensuring entanglement. Here, we introduce TangleDiff, a deep learning framework for the de novo design of homodimeric entangled proteins with programmable features. TangleDiff generates diverse foldable entangled sequences with an in-silico success rate exceeding 70%, markedly outperforming current models (~1%). By conditioning TangleDiff on inter-chain binding energy, we generate novel protein dimers whose binding energies closely match the specified ranges, with approximately 70% of successful designs conforming to expected values. We experimentally validate TangleDiff by designing nine homodimers targeting various binding energies; seven successfully form hydrogels, with stress relaxation dynamics correlated with specified binding energies. This work establishes a general strategy for entangled protein design, opening avenues for entanglement-based biomaterial innovation.

## GlycoViz: A glycoproteomics tool for validating and visualizing glycopeptide identifications
- Source: Bioinformatics (journals)
- Date: 2026-09-21T00:00:00+00:00
- Categories: Proteins & structural biology, Tools & resources
- Authors: Wenzhou Li, Niclas Olsson, Fiona E McAllister
- Journal: Bioinformatics
- DOI: 10.1093/bioinformatics/btag691
- Source URL: <https://doi.org/10.1093/bioinformatics/btag691>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1093%2Fbioinformatics%2Fbtag691>
- Code: <https://github.com/calico/glycoviz>

Abstract: Summary GlycoViz is an open-source algorithm and software tool designed to address critical bottlenecks in the post-identification workflow of mass spectrometry-based glycoproteomics. Current glyco search engines excel at glycopeptide identification but there is a lack of tools for robust post-identification validation, global glycan composition statistics, and reliable site-specific analysis, particularly concerning false positives and quantification distortion due to shallow detection depth. GlycoViz offers an alternative solution with a novel glycopeptide validation algorithm that is compatible with data from multiple search engines. Analysis of N-linked glycans is fully supported whereas O-linked glycan data analysis is limited primarily to visualization and non-optimized validation. Availability The open-source software is available on GitHub at https://github.com/calico/glycoviz. It can be launched as local software through Docker Compose or hosted as a web application in a server. A snapshot of the code is available on Zenodo: https://doi.org/10.5281/zenodo.22680123 Supplementary information Supplementary data are available at Bioinformatics online.

## Gravlax: an annotation-independent molecular evidence archive for single-cell RNA-seq
- Source: bioRxiv (preprints)
- Date: 2026-09-21
- Categories: Genomics & sequence analysis, Tools & resources
- Authors: Patro, R.
- DOI: 10.64898/2026.09.18.752708
- Source URL: <https://doi.org/10.64898/2026.09.18.752708>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.18.752708>
- Code: <https://github.com/COMBINE-lab/gravlax>

Abstract: A cell-by-gene count matrix is the artifact of a single-cell RNA-seq experiment that is most often stored, shared, and reanalyzed. It is the output of a computation whose inputs are the sequenced molecules and a gene annotation, and while the molecules never change, the annotation is revised continually. Once the matrix has been produced, the evidence behind it can no longer be reinterpreted. Recovering that evidence means returning to raw reads or alignments that are large, costly to process, and frequently unavailable. We ask whether a concise representation of the molecules themselves can be extracted once and reused indefinitely, to quantify under any future annotation, to query and discover features that no annotation yet describes, and to pool evidence across cells and samples. Our starting observation is that the procedures that turn alignments into counts, gene assignment and UMI collapse, never read most of what an alignment file contains. They consume relations among molecules such as shared genomic geometry, shared placements, barcode identity, and the equality or near-equality of UMIs. We show that these relations form a statistic that is sufficient for such consumers, and we design a compact, seekable, content-authenticated archive that stores them while deferring every annotation-dependent decision to analysis time. Archives compose into content-addressed collections that route cohort queries to the molecules that can answer them without copying molecules. We implement these ideas in a tool called gravlax. Across four human 10x 3' datasets, gravlax archives require 11--18 bits per read and are 9.0--12.7x smaller than tag-preserving CRAM. Count matrices replayed from an archive deviate from direct STARsolo quantification by 0.24--0.75% of normalized UMI mass, whereas changing GENCODE v32 to v49 moves 2.12--4.64%, and quantification replay is 34--82x faster than STARsolo at matched thread budgets. A federated index over eight archives occupies 2.96% of their size, answers a 96-query junction panel 2.59x faster than the archives alone, and screens the cohort genome-wide for unannotated splice events that recur across donors in just 9 seconds. Because the molecules are retained, the archives also answer questions the matrix has discarded. An analysis of four peripheral-blood archives recovers a validated FYB1 immune-cell splicing switch, a cross-fitted fragment model appropriate for 3' chemistry reveals an eight-donor shift in NTRK2 terminal-isoform usage from astrocyte and neural-stem-cell populations to mature neurons, and pooling evidence across cells within the context of an expectation-maximization algorithm recovers 75--98% of withheld multi-gene molecule labels. Gravlax is open source, implemented in Rust, licensed under the BSD 3-clause license, and available at https://github.com/COMBINE-lab/gravlax.

## HD-AIP: A Heterogeneous Dual-Stream Alignment-Free Framework for Anti-Inflammatory Peptide Prediction Based on Language Models and CT-Net
- Source: Bioinformatics (journals)
- Date: 2026-09-21T00:00:00+00:00
- Categories: Proteins & structural biology, Tools & resources
- Authors: Jiangli Li, Quan Zou, Yansu Wang, Yifeng Bai, Hao Zhou, Mengting Niu
- Journal: Bioinformatics
- DOI: 10.1093/bioinformatics/btag697
- Source URL: <https://doi.org/10.1093/bioinformatics/btag697>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1093%2Fbioinformatics%2Fbtag697>
- Code: <https://github.com/Zerofly0/HD-AIP>

Abstract: Motivation Anti-inflammatory peptides (AIPs) show therapeutic potential for treating chronic and autoimmune diseases. Computational screening of these peptides remains challenging because their typically short sequences limit traditional feature extraction effectiveness, while homology-based methods incur high computational costs. Results This study proposes HD-AIP, an alignment-free heterogeneous dual-stream prediction architecture. This framework extracts peptide features in parallel from both macroscopic and microscopic perspectives. Macroscopically, HD-AIP integrates global semantic features extracted by two large protein language models, ProtT5 and ESM-2 3B, with sequence-level physicochemical properties, followed by feature selection and LightGBM classification. Microscopically, an asymmetric parallel network named CT-Net utilizes BioVec embeddings and residue-level physicochemical features, using a CNN branch to capture local motifs and a Transformer branch to model long-range dependencies. The two streams are adaptively fused via a dynamic soft ensemble strategy. On an independent test set, HD-AIP outperforms baseline models across multiple metrics including accuracy, area under the receiver operating characteristic curve, and the Matthews correlation coefficient. These results indicate that HD-AIP improves AIP prediction performance without sequence alignment. This architecture serves as an effective computational tool for the high-throughput virtual screening and candidate discovery of AIPs. Availability You can find the source code and dataset needed on GitHub(https://github.com/Zerofly0/HD-AIP). The required package files have been listed in the readme. Supplementary information Supplementary data are available at Bioinformatics online.

## High-order enhancer hubs buffer allelic regulatory variation through kinetic compensation
- Source: bioRxiv (preprints)
- Date: 2026-09-21
- Authors: Tan, J., Sentmanat, M., Wu, Y., Peng, C., Fronick, C., Markovic, C., Cui, X., Fulton, R., Head, R., Wang, T., Sun, Y.
- DOI: 10.64898/2026.09.14.750771
- Source URL: <https://doi.org/10.64898/2026.09.14.750771>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.14.750771>

Abstract: Diploid genomes carry millions of heterozygous variants in cis-regulatory DNA, yet most genes produce similar RNA output from two parental alleles. How this balance is maintained is unclear. We developed Nanopore-HiChIP, a long-read method that maps high-order enhancer hubs on each haplotype. Over half of these enhancer hubs differ in chromatin architecture and transcription-factor occupancy between homologous chromosomes, but their target genes show substantially lower rates of allele-specific expression than genes lacking hub regulation. Single-cell kinetic modeling shows that burst frequency and burst size change in opposite directions, thereby preserving balanced transcriptional output. This hub-mediated kinetic buffering is enriched at haploinsufficient genes and coincides with smaller effects of expression quantitative trait loci. Enhancer hubs therefore absorb allelic regulatory variation through kinetic compensation, protecting dosage-sensitive transcription.

## Improving Data Quality, Model Transparency and Performance in Lung Histopathology with Explainable AI
- Source: bioRxiv (preprints)
- Date: 2026-09-21
- Categories: Biological imaging
- Authors: Rashed, S. K., Nilsson, M., Aits, S.
- DOI: 10.64898/2026.09.14.751577
- Keywords: histopathology
- Source URL: <https://doi.org/10.64898/2026.09.14.751577>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.14.751577>
- Code: <https://github.com/Aitslab/Histology_XAI>

Abstract: Convolutional neural networks (CNNs) have shown strong capabilities for image analysis. However, deploying these models in medical settings is complicated by their limited transparency. Over recent years, many approaches have been developed to overcome the so-called "black box" problem of deep neural networks. Here, we show how such explainable AI (XAI) approaches can be applied to not only improve transparency but also training data quality and model performance, with classification of lung damage in histopathology images as use case. First, we conducted a thorough exploratory data analysis and visually compared the compressed multi-dimensional representations of the histology images from the last CNN layers with labels given by pathologists to reveal flaws in the training data. Second, we used Gradient-based Class Activation Mapping (Grad-CAM) as well as SHapley Additive exPlanations (SHAP) values to identify image regions that significantly contributed to the model decisions. To overcome identified shortcomings, we then finetuned additional top layers of the CNNs and introduced model architectures with attention which improved model performance. In summary, we developed a practical workflow that uses interpretability analyses to examine model perception, assess label consistency, and guide model refinement on the example of lung histopathology scoring, demonstrating how XAI approaches can increase both transparency and model performance. The code for this paper is shared at https://github.com/Aitslab/Histology\_XAI.git.

## Individual-level expression deconvolution and assessment of cross-sample variation
- Source: bioRxiv (preprints)
- Date: 2026-09-21
- Categories: Genomics & sequence analysis, Mathematical biology & statistics
- Authors: Kang, K., Xie, K.
- DOI: 10.64898/2026.09.14.751599
- Source URL: <https://doi.org/10.64898/2026.09.14.751599>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.14.751599>

Abstract: Recovering cell-type-specific gene expression from bulk RNA sequencing would facilitate the study of transcriptional variation among individuals. However, accuracy can differ substantially among genes and cell types. We describe a reference-informed Bayesian deconvolution framework and a score that identifies gene--cell-type pairs likely to have more accurate estimates of cross-sample variation. The score uses bulk counts, reference expression profiles, and estimated RNA proportions. Known component expression is used to train and evaluate the score, but is not needed to calculate predictions from a trained model. We evaluated the approach in a ROSMAP-derived simulation with 40 target donors, 2,000 genes, and seven cell types. Median gene-wise correlation was 0.801 for raw allocated counts and 0.296 after normalization within each donor and cell type. To evaluate the score, we divided genes into five sets, kept linked genes together, and scored each set using a model trained on the other four. Retaining approximately 20\\% of pairs within each cell type increased the median normalized correlation to 0.622. Ranking pairs only by the estimated share of a gene's bulk RNA contributed by the cell type yielded 0.570 at the same retained count. These results show that observable information can help prioritize pairs with more accurately recovered cross-sample variation.

## Integrating structural homology with deep learning to achieve highly accurate protein-protein interface prediction for the human interactome
- Source: bioRxiv (preprints)
- Date: 2026-09-21
- Categories: Proteins & structural biology, Tools & resources
- Authors: Xiong, D., Torres, M., Murray, D., Zhang, Z., Li, L., Naravane, A. C., Fragoza, R., Honig, B., Yu, H.
- DOI: 10.1101/2025.06.09.658393
- Source URL: <https://doi.org/10.1101/2025.06.09.658393>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1101%2F2025.06.09.658393>

Abstract: A significant portion of disease-causing mutations occur at protein-protein interfaces however, the number of structurally resolved multi-protein complexes is extremely small. Here we present a computational pipeline, PIONEER2, that integrates 3D structural similarity with geometric deep learning to accurately predict protein binding partner-specific interfacial residues. We compare the performance of PIONEER2 to that of AlphaFold3 and found, using a test set of PDB structures, that their performance is quite similar. However, about 20% of AlphaFold3 predictions for protein-protein complexes in the PDB have AlphaFold3 ranking scores below 0.5, which indicates an uncertain model. For these structures, PIONEER2 outperforms AlphaFold3 at discriminating interfacial from non-interfacial residues. Further, about half of the AlphaFold3 ranking scores on high confidence protein-protein interactions (PPIs) not associated with a PDB structure are below 0.5 indicating that PIONEER2 offers superior interface prediction for a large number of PPIs for which structures are not available. We created a comprehensive 3D structurally informed interactome encompassing all 352,124 experimentally detected binary human PPIs in the current literature and made PIONEER2 interface predictions for each. We experimentally validated these predictions by generating 1,866 mutations and testing their disruptive impact on 5,010 mutation-interaction pairs. PIONEER2-predicted interfaces are found to be comparable to PDB structures in their ability to predict disruptive mutations while AlphaFold3 performance is reduced. Similarly, PIONEER2-predicted interfaces outperform AlphaFold3 in accounting for the depletion of non-deleterious common population variants and the enrichment of disease-related mutations on protein surfaces. Overall, our results suggest that PIONEER2-predicted interfaces provide a valuable tool for studying disease etiology, advancing personalized medicine and for fundamental research. We further implemented PIONEER2 as a user-friendly web server (https://pioneer2.yulab.org) platform for users to explore our 3D interactome models and conduct genome-wide functional genomics studies.

## Introns encode a vast new class of Kink-loop RNAs that autoregulate pre-mRNA splicing
- Source: bioRxiv (preprints)
- Date: 2026-09-21
- Categories: Genomics & sequence analysis, Proteins & structural biology
- Authors: Li, B., Lin, Q., Liu, A., Gan, H., Liu, S., Zheng, W., Liang, Y., Wan, G., Qu, L., Yang, J.
- DOI: 10.64898/2026.08.03.742641
- Keywords: splicing, genome, rna
- Source URL: <https://doi.org/10.64898/2026.08.03.742641>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.08.03.742641>

Abstract: Introns occupy nearly one-third of the human genome, yet whether they encode widespread regulatory functions remains unclear. Here, we identify a vast new class of intron-derived Kink-loop RNAs (klRNAs) that autoregulate host pre-mRNA splicing. We developed orthogonal sequencing methods to systematically uncover approximately 15,200 and 3,500 previously unannotated klRNAs in humans and mice, respectively. Bound by the conserved RNA-binding protein 15.5K, klRNAs are compact orphan RNAs characterized by stereotypically positioned terminal C/D motifs that form K-loop structures. Depletion of 15.5K broadly disrupts klRNA biogenesis. Functional and genetic perturbations establish that klRNAs suppress host intron excision through a K-loop-dependent mechanism. Together, our findings establish klRNA-mediated autoregulation as a widespread principle governing intron fate and reveal a previously unrecognized regulatory layer encoded within mammalian introns.

## Is level-1 blob reconstruction under the network multispecies coalescent easy?
- Source: bioRxiv (preprints)
- Date: 2026-09-21
- Categories: Evolution & metagenomics, Mathematical biology & statistics
- Authors: Dai, J., Molloy, E.
- DOI: 10.64898/2026.06.06.730607
- Source URL: <https://doi.org/10.64898/2026.06.06.730607>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.06.06.730607>

Abstract: Hybridization is an important evolutionary process, commonly modeled by the network multispecies coalescent. Reconstructing evolutionary histories under this model is notoriously costly, even for level-1 networks where hybridization events are isolated from each other. The widely used methods that combine speed with statistical guarantees rely on quartet concordance factors computed for all subsets of four species, resulting in an O(n^4k) bottleneck that severely limits scalability to large numbers of species (n) and genes (k). Among quartet-based methods, NANUQ+ is notable because it decomposes the problem into two steps: first reconstructing a tree of blobs, which compresses each non-treelike part of the network, called a blob, into a single vertex, and second reconstructing the internal structure of each level-1 blob, specifically its circular order and hybrid vertex. Here, we investigate whether level-1 blob reconstruction is difficult once the tree of blobs is known. We present a fast and statistically consistent algorithm, called NetCS, based on two simple primitives: majority voting and merge sort, circumventing the bottleneck of computing all quartet concordance factors. In simulations, NetCS achieved comparable accuracy to NANUQ+ and was dramatically faster, enabling analyses of 200 taxa and 1000 genes in only a few minutes. Both methods attained near-perfect accuracy when given the true tree of blobs; however, their performance degraded in end-to-end pipelines due to errors in tree of blobs reconstruction. Strikingly, even methods that reconstruct level-1 networks directly struggled to accurately predict hybrid ancestry. Our results suggest that reconstructing level-1 blobs is unexpectedly easy once the tree of blobs is known, and that a major challenge for phylogenetic network inference lies in accurate tree of blobs reconstruction.

## Large scale prospective evaluation of co-folding across hundreds of Mac1-ligand complexes and three virtual screens
- Source: bioRxiv (preprints)
- Date: 2026-09-21
- Categories: Proteins & structural biology
- Authors: Kim, J., Correy, G. J., Hall, B. W., Rachman, M. M., Mailhot, O., Togo, T., Gonciarz, R. L., Jaishankar, P., Neitz, R. J., Hantz, E. R., Doruk, Y. U., Stevens, M. G. V., Diolaiti, M. E., Reid, R., Gopalkrishnan, S., Krogan, N. J., Renslo, A. R., Ashworth, A., Shoichet, B. K., Fraser, J. S.
- DOI: 10.64898/2025.12.25.696505
- Source URL: <https://doi.org/10.64898/2025.12.25.696505>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2025.12.25.696505>

Abstract: Accurate prediction of ligand-bound protein complexes and ranking them by affinity are central problems in drug discovery. While deep learning co-folding methods can help address these challenges, their evaluation has been hampered by the difficulties in assessing independence from training data and insufficiently large test sets. Here we test the ability of co-folding methods to predict the structures of 551 ligands, 489 of which are of sufficient quality for evaluation, bound to the SARS-CoV-2 NSP3 macrodomain (Mac1) that were determined after the training cut-off dates. AlphaFold3 (AF3), Boltz-2, and Chai-1 each reproduced >50% of the Mac1 ligand poses to better than 2 \[A\] RMSD of experiment. Despite the potential for co-folding to describe protein conformational changes that stabilize ligand binding, we did not find that common conformational rearrangements, including peptide flip and a large loop opening, were recapitulated by the co-folding prediction. For AF3 and Chai-1, ligand pose prediction confidence weakly, but significantly, tracked experimental potency, while DOCK3.7 energies were only weakly correlated. Boltz-2 affinity predictions showed the strongest correlation with measured potency and, after calibration, achieved lower mean absolute error than a baseline predictor. We next assessed whether co-folding scores could rescore docking hit-lists to distinguish true ligands from non-binders among hundreds of molecules prospectively experimentally tested against AmpC \{beta\}-lactamase, the dopamine D4 and the \{sigma\}2 receptors. AF3 ligand pose confidence values did not separate true ligands from high-scoring false-positives as effectively as docking scores or Boltz-2 affinity predictions did. Taken together, the modest, but independent correlations of docking score and co-folding confidence or affinity suggests that integrating physics-based and deep-learning approaches may help with hit prioritization and subsequent optimization in structure-based ligand discovery.

## Macro-level causal discovery from single-time-point observations of ecological dynamical systems
- Source: bioRxiv (preprints)
- Date: 2026-09-21
- Categories: Evolution & metagenomics
- Authors: Bystrova, D., Assaad, C., Si-moussi, S., Thuiller, W.
- DOI: 10.1101/2024.10.10.608447
- Source URL: <https://doi.org/10.1101/2024.10.10.608447>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1101%2F2024.10.10.608447>

Abstract: Ecologists often aim to uncover causal relations within ecological systems using observational data. However, many studies still rely mainly on correlation-based approaches, which do not support causal interpretation. Causal discovery methods have recently gained interest, particularly for ecological time-series data. Yet such data are often scarce, as collecting repeated and consistent measurements over long periods is costly, require strict continuation of protocols and is time-consuming. Consequently, ecologists have often access only to single-time-point observational data generated by dynamical systems. In this setting, unobserved past states can induce substantial unmeasured confounding, limiting the ability of standard algorithms such as PC and FCI to recover micro-level causal relations between observed variables. We show that PC and FCI can nevertheless recover meaningful causal information. In particular, they can identify specific cluster-level structures, which we call clustered super-unshielded colliders, which provide information about the partial causal ordering of macro-level variables. We further show that both algorithms can be reduced to a simple procedure, which we call RestPC, that yields the same identifiable information. We illustrate our results using simulated data and two real-world datasets: one on bird abundance, climate, and land cover, and another on soil microbial communities, environmental, terrain, and geochemical variables.

## Noninvasive imaging-based vascular score is associated with immunotherapy outcome in non-small cell lung cancer
- Source: Scientific Reports (journals)
- Date: 2026-09-21T00:00:00+00:00
- Categories: Biological imaging
- Authors: Jun Hyeong Park, Woo Kyung Ryu, Chul-Ho Kim, Jeong-Seok Choi, Jae Won Chang, In Young Jo, Byung-Joo Lee, Jun Hyeok Lim, Jaesung Heo
- Journal: Scientific Reports
- DOI: 10.1038/s41598-026-71525-y
- Source URL: <https://doi.org/10.1038/s41598-026-71525-y>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1038%2Fs41598-026-71525-y>

Abstract: Morphological abnormality in tumor vasculature, a recently recognized mechanism of resistance to immune checkpoint inhibitors (ICIs), remains difficult to quantify due to its heterogeneity. We present the vascular risk score (VRS), a deep learning–based imaging metric that quantifies abnormalities in tumor vasculature on CT scans. We trained a deep learning model to learn representations of vascular morphology. Abnormality was then quantified using Gaussian mixture modeling as the degree to which each patient’s tumor vasculature deviated from the learned distribution of normal morphology. We validated VRS in a cohort of 321 NSCLC patients treated with ICIs. Patients with low VRS showed significantly longer progression-free and overall survival, and lower VRS was observed in patients with disease control compared with progressive disease. Combining VRS with PD-L1 expression provided modest additional discrimination over either marker alone. VRS enables noninvasive, objective quantification of tumor vascular abnormality and may serve as a prognostic imaging marker in NSCLC.

## OpenCRS: an open-source regulated human cardiorespiratory model with large-scale calibration and global sensitivity analysis at rest and during exercise
- Source: bioRxiv (preprints)
- Date: 2026-09-21
- Categories: Tools & resources
- Authors: Wang, S.-Y., Saxton, H., Balmus, M., Niederer, S. A.
- DOI: 10.64898/2026.09.14.751454
- Source URL: <https://doi.org/10.64898/2026.09.14.751454>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.14.751454>

Abstract: Cardiopulmonary exercise testing reveals cardiovascular and respiratory limitations not apparent at rest, but similar measurements can arise from different interacting regulatory mechanisms, preventing physiological causal inference from data alone. Mechanistic computational models can separate these mechanisms in silico. However, whole-body cardiorespiratory models contain hundreds of parameters, making global sensitivity analysis and calibration challenging. We present OpenCRS, an open-source Python cardiorespiratory model coupling closed-loop 0D lumped-parameter circulation with gas exchange and integrated autonomic and respiratory control. The framework incorporates baroreflex, chemoreflex, pulmonary stretch receptors, central command, and neuromuscular drive, together with novel representations of atrial dynamics and exercise baroreflex set-point resetting. We introduce a scalable calibration pipeline that (i) applies a derivative-based global sensitivity measure (DGSM) directly to the simulator, reducing 272 parameters to 72 influential ones; (ii) trains Gaussian process emulator surrogates within iterative History Matching to exclude implausible parameter regions; and (iii) performs Bayesian calibration (MCMC), inferring a joint posterior with Hamiltonian Monte Carlo (No-U-Turn sampler) under a Gaussian copula prior. A single maximum a posteriori parameter set simultaneously reproduced 50 literature-derived rest and exercise targets (45/50 within 1 SD, all within 1.83 SD), with the rest-to-exercise transition emerging from the models embedded feedback rather than independent fitting. Simulator-based DGSM agreed with constrained Sobol indices (mean Spearman rank correlation 0.79). Sensitivity analysis identified influential physiological mechanisms. Only 2-16 parameters contributed >2% of the Sobol total-effect sensitivity per target. Resting cardiovascular targets were driven by unstressed volumes and cardiac mechanics, while exercise shifted influence towards autonomic efferent regulation. Respiratory outputs remained most sensitive to chemoreflex and gas exchange parameters. These shifts capture coordinated cardiorespiratory adaptation to metabolic demand. More broadly, the framework provides a population-level prior with quantified uncertainty for cardiovascular digital twins and a reusable route for calibrating high-dimensional, regulated physiological models.

## Optimized Multiple Circular Sequence Alignment for Cyclic Peptide Motif Discovery
- Source: bioRxiv (preprints)
- Date: 2026-09-21
- Categories: Genomics & sequence analysis, Proteins & structural biology, Tools & resources
- Authors: Yuan, Y., Li, Z., Hu, K., Pan, P., He, F.
- DOI: 10.64898/2026.07.28.741376
- Source URL: <https://doi.org/10.64898/2026.07.28.741376>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.07.28.741376>
- Code: <https://github.com/IVB-Generative-Biology/mars-turbo>

Abstract: Head-to-tail (H2T) cyclized peptides are an increasingly important modality in drug discovery, combining high target affinity and selectivity with metabolic sta-bility. Because their underlying chemistry is still that of a linear amino-acid chain, their linear sequence representation is the native input format of main- stream sequence generative models now driving de novo peptide design (Slough et al., 2018; Rettie et al., 2025a;b). Discovering the conserved motifs responsible for a family's function across a library of such candidates requires a multiple sequence alignment (MSA). Because a cyclic peptide can be linearised at any residue, the alignment must additionally solve for the unknown rotation of each sequence, which is the multiple circular sequence alignment (MCSA) problem. However, leading MCSA heuristics (e.g. Ayad & Pissis, 2017) were tuned for the genomic regime (a few tens of long sequences) and become prohibitively slow on the cyclic peptide library regime (hundreds to thousands of shorter sequences). We close this gap by identifying quality-preserving optimisation opportunities, notably the library-scale preset tailored to short-sequence inputs (algorithmic details in Appendix A), and by adding an orthogonal multi-core and SIMD backend for further performance tuning, which gives near-linear thread scaling on the pairwise-comparison stage. We validate the pipeline on a library of 1,000 H2T cyclized peptides of length 18 targeting the oncoprotein Mouse double minute 2 human homolog (MDM2) produced by an internal peptide-design engine. In this practical setup, the optimised MCSA recovers the underlying positional motif of MDM2 binders at the same fidelity as the original MCSA implementation while running over 650x faster. Our optimised MCSA tool thus enables library-scale cyclic peptide sequence alignment and is publicly available at https://github.com/IVB-Generative-Biology/mars-turbo.

## Physiological variability in key Alzheimer's biomarkers in amyloid-positive clinical trial cohorts and the mechanistic basis of biomarker ratios
- Source: medRxiv (preprints)
- Date: 2026-09-21
- Authors: Tasiudi, E., Hawellek, D. J., Aponte, E., Tonietto, M., Soares, H., Diack, C., Soubret, A., Kam-Thong, T., Ribba, B., Boareto, M.
- DOI: 10.64898/2026.09.18.26363412
- Source URL: <https://doi.org/10.64898/2026.09.18.26363412>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.18.26363412>

Abstract: Background: Protein biomarkers in cerebrospinal fluid (CSF) and plasma have established themselves as essential tools for the diagnosis of neurological disorders and for disease monitoring, thanks to their accuracy and clinical validity. Outside their original intended context of use, protein biomarkers are increasingly used in clinical trials to assess the effects of novel therapies on brain biology. Physiological differences across individuals - such as CSF or plasma volume and elimination kinetics - also contribute to biomarker variability. A deeper understanding of these sources of variability is essential to further improve biomarker interpretation, particularly in clinical trials where populations are by design more homogeneous than in real-world diagnostic settings. This study aimed to quantify the contributions of physiological factors to inter-individual biomarker variability, and to identify the mechanistic basis for why ratio-based normalization can lead to enhanced biomarker performance. Methods: Using a mechanistic kinetic framework and paired CSF and plasma baseline data from four Phase III clinical trials (GRADUATE I and II, CREAD and CREAD2), we quantified physiological and neurobiological contributions (hereafter, physiological and neurobiological variability) to inter-individual variability of commonly used AD biomarkers (A\{beta\}40, A\{beta\}42, p-tau181, t-tau, NfL, GFAP, sTREM2, YKL-40), and evaluated the conditions under which ratio-based normalization can reduce physiological variability. Results: In these amyloid-positive clinical-trial populations, a large fraction of inter-individual variability could be attributed to physiological variability. While normalizing biomarker values by A\{beta\}40 or A\{beta\}42 reduced physiological variability for some biomarkers, the effects of the normalization were compartment-, and cohort-dependent. Using our framework, we identified two conditions under which ratio-based normalization is most likely to improve biomarker performance: (1) the target and reference biomarkers strongly share physiological variability, and (2) they maintain independent neurobiological variability. These conditions provide a mechanistic explanation for why A\{beta\}-based ratios are useful for some biomarkers but are not optimal for others. Discussion: These findings advance our understanding of the sources of biomarker variability in clinical trials and provide a framework for better understanding the mechanistic basis of ratio-based normalization. This work aims to strengthen the utility of using biomarkers by clarifying when and why ratio-based approaches are most informative.

## POMS enhances open spectral library search for identification of modified peptides
- Source: BMC Bioinformatics (journals)
- Date: 2026-09-21T00:00:00+00:00
- Categories: Proteins & structural biology, Tools & resources
- Authors: Younghee Seo, Eunok Paek, Seungjin Na
- Journal: BMC Bioinformatics
- DOI: 10.1186/s12859-026-06672-0
- Source URL: <https://doi.org/10.1186/s12859-026-06672-0>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1186%2Fs12859-026-06672-0>
- Code: <https://github.com/icp-kbsi/POMS>

Abstract: Background Peptide identification from tandem mass spectrometry (MS/MS) data is a central task in proteomics. Spectral library searching combined with open modification search (OMS) has emerged as an effective strategy for identifying peptides carrying unexpected post-translational modifications (PTMs). However, existing methods do not adequately account for modification-induced fragment-ion shifts during candidate selection, leading to missed identifications when mass shifts distort spectral similarity. Results We present POMS, a modification-aware framework for open spectral library search that improves candidate retrieval by integrating complementary spectral representations with sequence-derived theoretical fragment features. These representations compensate for modification-induced fragment-ion shifts and improve alignment between query and library spectra, increasing robustness to modification site variability. Benchmarking on large-scale human MS/MS datasets showed that POMS yielded up to ~ 8% more peptide identifications during the open search phase than conventional methods. Cross-validation with independent database search engines further demonstrated a > 5% increase in consistent peptide identifications. Notably, POMS substantially improved the identification of peptides carrying C-terminal modifications, for which conventional candidate retrieval methods are particularly susceptible to fragment-ion shifts. Conclusions POMS improves the sensitivity and reliability of modification-tolerant spectral library searching by incorporating modification-aware candidate retrieval while preserving computational efficiency. The framework provides a practical solution for large-scale PTM discovery and is readily applicable to existing spectral library search workflows. POMS is publicly available under the CC BY-NC-SA 4.0 license at https://github.com/icp-kbsi/POMS .

## Questioning the G2 phase in the budding yeast cell cycle with a qualitative and possibilistic model
- Source: bioRxiv (preprints)
- Date: 2026-09-21
- Categories: Systems & networks, Mathematical biology & statistics
- Authors: Faure, A., Liakopoulos, D., Gaucherel, C.
- DOI: 10.64898/2026.02.06.704310
- Source URL: <https://doi.org/10.64898/2026.02.06.704310>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.02.06.704310>

Abstract: The budding yeast S. cerevisiae, a foundational model for cell cycle studies, exhibits a complex phase organisation (G1, S, G2/M) governed by checkpoints ensuring faithful cellular inheritance. However, the existence of a distinct G2 phase in yeast remains debated, with some advocating for a prometaphase instead. To address this issue, we developed a discrete-event, qualitative, and possibilistic model, the first one to our knowledge, to integrate organelle-level components (replication forks, sister chromatids, mitotic spindle, bud) while remaining parsimonious. Unlike molecular-centred or overly complex whole-cell models, this approach bridges broad systemic and finer mechanistic scales. Our results demonstrate that the model faithfully recapitulates cell cycle progression and supports the dispensable G2 phase. This possibilistic model inspired from recent applications in ecology advocates in favor of the necessity of prometaphase. This study thus provides a unifying and flexible framework to resolve long-standing ambiguities in yeast cell dynamics, while avoiding the pitfalls of excessive complexity or reductionism.

## Rapid volumetric reconstruction and tracking for Fourier light-field microscopy enables real-time calcium imaging in freely behaving Hydra.
- Source: bioRxiv (preprints)
- Date: 2026-09-21
- Categories: Biological imaging, Tools & resources
- Authors: Adkins, R., Hausen, R., Noss, J., Lemson, G., Howard, J.
- DOI: 10.64898/2026.09.17.752193
- Source URL: <https://doi.org/10.64898/2026.09.17.752193>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.17.752193>

Abstract: Fourier light field microscopy (FLFM) enables high-speed volumetric imaging by encoding multiple angular perspectives of a three-dimensional sample onto a single image. For this reason, FLFM is well-suited to sparse and rapidly evolving biological systems. To aid in the adoption of FLFM, we present OpenFLR, an open-source software framework for real-time volumetric reconstruction, three-dimensional particle tracking, and calcium image processing using FLFM. OpenFLR reconstruction is distributed as four interchangeable interfaces: a Python library, a command-line script, an interactive web application, and an ImageJ/micromanager plugin, so that the pipeline is accessible to both developers and bench biologists. Building on established Richardson-Lucy deconvolution, we use a hybrid experimental-computational PSF calibration strategy and a triangulation approach to tracking to extract particle positions in 3D directly from raw light field frames, bypassing reconstruction. We validate the complete pipeline on GCaMP6s recordings of freely behaving Hydra vulgaris, tracking sparse populations of neurons as they undergo large three-dimensional displacements.

## Robust High-Throughput Flickering Spectroscopy for Measurements of Red Blood Cell Membrane Mechanics
- Source: bioRxiv (preprints)
- Date: 2026-09-21
- Categories: Biological imaging
- Authors: Ayazi, F., Kotar, J., Rayner, J. C., Cicuta, P.
- DOI: 10.64898/2026.09.16.751779
- Source URL: <https://doi.org/10.64898/2026.09.16.751779>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.16.751779>

Abstract: The mechanical properties of red blood cell (RBC) membranes are critical to their function in oxygen delivery, and changes to these properties as RBCs age affect their journey around the circulatory system, including clearance by the spleen. Such changes can also have significant health effects,including in cardiovascular disease and on the interactions between RBCs and malaria parasites. The function of RBCs requires them to be very soft, which together with the small size of the cells brings the energy required for measurable deformation of the membrane into the range of typical thermal energies. Red Blood Cells can be observed to flicker under optical microscopy, and these shape fluctuations can be quantified and used to obtain key biophysical parameters such as the tension and bending modulus of the membrane. Typically the shape of the cell's equator is extracted, and the mean power spectrum is obtained by time-averaging the power present in the normal modes of the thermal fluctuations. Flickering spectroscopy has been used extensively on RBCs, but has so far been very low throughput and with some technical limitations. Here we address issues related to active versus passive fluctuations, focusing and optics, camera exposure and sampling, and fast contour detection. In combination with an automated imaging system it is possible to measure thousands of cells in one day, with no user input. We validate this new pipeline by chemical modification of RBC mechanics, and by comparison with simulated fluctuations including each confounding effect. The methods and codes for this robust flickering analysis will allow consistent measurements across labs.

## Survey of transcription initiation in the streamlined genomes of Paramecium
- Source: bioRxiv (preprints)
- Date: 2026-09-21
- Categories: Genomics & sequence analysis
- Authors: Jimenez-Marin, B., Stickling, D., Swenty, T., Gout, J.-F., Miller, S., Lynch, M.
- DOI: 10.64898/2026.09.19.752804
- Keywords: genomes, genome, gene expression, survey
- Source URL: <https://doi.org/10.64898/2026.09.19.752804>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.19.752804>

Abstract: In the genus Paramecium, the macronuclear genome is remarkably compact and optimized for gene expression. As a means to explore eukaryotic transcription in the context of a streamlined genome and shed light on the role of sequence architecture on gene expression and loss, we analyzed the distribution and diversity of candidate transcription initiation sites (TISs) in Paramecium sexaurelia, Paramecium tetraurelia and their outgroup, Paramecium caudatum. Our analysis suggests that for Paramecium, most genes have very short 5 prime UTRs (40 bp or less) and their transcription initiation regions (TIRs) have a median dispersion (akin to width) of 8-10 bp. The TIRs for the three species have high AT content. TIR dispersion is not to gene expression. However, mean TIS position relative to the translation start site per gene does in gene expression for the three species, and is often conserved between them. While mean TIS position and gene expression are linked, gene expression itself is the main driver of paralog retention in the aurelias. As compared to other eukaryotes, Paramecium has a uniquely well-defined and short main TIS region, and sequence motifs that likely diverge from the consensus in multicellular eukaryotes.

## SVPG: a pangenome-based structural variant detection approach and rapid augmentation of pangenome graphs with new samples
- Source: Nature Methods (journals)
- Date: 2026-09-21T00:00:00+00:00
- Categories: Genomics & sequence analysis, Tools & resources
- Authors: Tao Jiang, Heng Hu, Runtian Gao, Shuqi Cao, Zhongjun Jiang, Murong Zhou, Wentao Gao, Shengming Zhou, Guohua Wang
- Journal: Nature Methods
- DOI: 10.1038/s41592-026-03219-2
- Source URL: <https://doi.org/10.1038/s41592-026-03219-2>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1038%2Fs41592-026-03219-2>

Abstract: Breakthrough advances in long-read sequencing have opened unprecedented opportunities to study genetic variations through pangenome analysis, yet tools that effectively leverage such frameworks for structural variant (SV) detection remain limited. In addition, efficient construction of pangenome graphs becomes increasingly challenging with the acquisition of larger numbers of samples. Here we present SVPG, an approach that leverages haplotype-resolved pangenome reference for accurate SV detection and rapid pangenome graph augmentation from long-read sequencing data. Compared with state-of-the-art SV callers, SVPG maintained superior overall performance across different sequencing technologies and coverages. SVPG also achieved notable improvements in calling individual-specific SVs, including rare and somatic SVs. Furthermore, in a benchmark involving 20 samples, SVPG accelerated pangenome graph augmentation by nearly tenfold compared with traditional augmentation strategies. These results indicate that SVPG has the potential to improve SV detection and serve as an effective tool, offering new possibilities for advancing pangenomic research.

## The Dresden Dataset for 4D Reconstruction of Non-Rigid Abdominal Surgical Scenes
- Source: Scientific Data (journals)
- Date: 2026-09-21T00:00:00+00:00
- Categories: Tools & resources
- Authors: Reuben Docea, Rayan Younis, Yonghao Long, Maxime Fleury, Jinjing Xu, Chenyang Li, André Schulze, Ann Wierick, Johannes Bender, Micha Pfeiffer, Qi Dou, Martin Wagner, Stefanie Speidel
- Journal: Scientific Data
- DOI: 10.1038/s41597-026-08289-7
- Keywords: dataset
- Source URL: <https://doi.org/10.1038/s41597-026-08289-7>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1038%2Fs41597-026-08289-7>

Abstract: The D4D Dataset provides paired endoscopic video and high-quality structured-light geometry for evaluating 3D reconstruction of deforming abdominal soft tissue in realistic surgical conditions. Data were acquired from six porcine cadavers using a da Vinci Xi stereo endoscope and a Zivid structured-light camera, registered via optical tracking and manually curated iterative alignment methods. Three session types (whole deformations, incremental deformations, and moved-camera sessions) probe algorithm robustness to non-rigid motion, deformation magnitude, and out-of-view updates. Each clip provides rectified stereo images, per-frame instrument masks, stereo depth, start/end structured-light point clouds, curated camera poses and camera intrinsics. In postprocessing, ICP and semi-automatic registration techniques are used to register data, and instrument masks are created. The dataset enables quantitative geometric evaluation in both visible and occluded regions, alongside photometric view-synthesis baselines. Comprising over 300,000 frames and 369 point clouds across 98 curated sessions, this resource can serve as a comprehensive benchmark for developing and evaluating non-rigid SLAM, 4D reconstruction, and depth estimation methods.

## Transformer-based multi-modal representation learning and hybrid interaction modeling for miRNA–disease association prediction
- Source: BMC Bioinformatics (journals)
- Date: 2026-09-21T00:00:00+00:00
- Categories: Systems & networks
- Authors: Jihwan Ha
- Journal: BMC Bioinformatics
- DOI: 10.1186/s12859-026-06668-w
- Keywords: mirna, representation learning
- Source URL: <https://doi.org/10.1186/s12859-026-06668-w>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1186%2Fs12859-026-06668-w>
- Abstract: not stored for this record.

## Two Comparators May Be All We Need
- Source: bioRxiv (preprints)
- Date: 2026-09-21
- Categories: Proteins & structural biology, Tools & resources
- Authors: Muskal, S. M.
- DOI: 10.64898/2026.09.15.751854
- Source URL: <https://doi.org/10.64898/2026.09.15.751854>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.15.751854>

Abstract: A compound in a cell meets a spectrum of proteins drawn from many families at once, while screening most often interrogates one target at a time. Two questions asked many times in a rank ordering workflow ultimately guide decisions on what gets made and what gets counter-screened, and both are comparative: which of two targets does a compound prefer, and which of two compounds does a target prefer. We built one model for each, over a roster of 1,879 human proteins covering 34 protein families. Each model is given two chemical structures and a sequence, or two sequences and a chemical structure, and returns which member of the pair is preferred together with how firmly it holds that view. No conformational analysis, protein structure, binding site or docked pose is used. Across families, asked which of two targets a compound prefers, the model is correct 0.75 of the time over 8,689 held-out comparisons, and 0.93 of the time on the third of them it holds most confidently. Asked which of two compounds a single target prefers, it is correct 0.71 of the time over 65,725 held-out comparisons, rising to 0.96 on the most confidently held. Neither compound in any of those comparisons appeared anywhere in training. Accuracy in both tracks the size of the real difference between the two measurements, from near chance where they fall within half a log unit to about 0.90 where they differ by more than two logs, and it holds across 31 protein families, not only the best-measured one. Within the chemistry and the targets they were built on, these models rank compounds and rank targets well. Both models can be explored and downloaded at familyfoundationmodel.com. Keywords: target preference; compound preference; pairwise comparison; polypharmacology; off-target triage; ESM2; random forest; structure-free prediction; ChEMBL

## Uniformly processed transcriptome-wide alternative splicing profiles for pediatric cancer research
- Source: bioRxiv (preprints)
- Date: 2026-09-21
- Categories: Tools & resources
- Authors: Liang, C. E., Shapiro, J. A., Beale, H. C., Taroni, J. N., Vaske, O. M.
- DOI: 10.64898/2026.09.14.750241
- Source URL: <https://doi.org/10.64898/2026.09.14.750241>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.14.750241>

Abstract: Alterations in regulatory processes like alternative splicing contribute to pediatric cancer development. Although splicing aberrations have been observed in pediatric leukemias, alternative splicing has yet to be studied in pediatric cancers at scale, due to a lack of uniformly processed, sample-level pediatric cancer splicing profiles with non-diseased tissue comparators. We address this need by quantifying splice event usage for a curated set of bulk RNA-seq datasets from the NCI's Therapeutically Applicable Research to Generate Effective Treatments (TARGET, n = 1152) and Genotype-Tissue Expression (GTEx, n = 1098) as a comparator. This Treehouse Splice Compendium is accompanied by a reproducible workflow that was used to generate the data in the compendium and reflects the largest known RNA-seq dataset processed by the splice quantification tool Shiba. The compendium is part of a suite of large, uniformly processed datasets aggregated by the UCSC Treehouse Childhood Cancer Initiative and Alex's Lemonade Stand Foundation's Childhood Cancer Data Lab, which include the Treehouse Expression Compendia, refine.bio, and the Single-cell Pediatric Cancer Atlas.

## A lifespan single-cell atlas of the human developing hippocampus benchmarks familial Alzheimer's disease brain organoids.
- Source: bioRxiv (preprints)
- Date: 2026-09-20
- Categories: Single-cell & spatial, Tools & resources
- Authors: Ivleva, E., Kruikov, E., Arboleda-Velasquez, J. F., Baranov, P.
- DOI: 10.64898/2026.09.18.752796
- Source URL: <https://doi.org/10.64898/2026.09.18.752796>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.18.752796>

Abstract: Familial Alzheimer's disease (fAD) is an early-onset form of AD caused by autosomal-dominant variants in APP, PSEN1, or PSEN2, with PSEN1 accounting for most genetically defined cases \[1\]. The hippocampus is among the earliest and most severely affected brain regions in AD \[2,3\]. Human induced pluripotent stem cell (iPSC)-derived brain organoids recapitulate key features of early human brain development and provide a tractable model for studying how fAD mutations perturb neurodevelopmental processes \[4\]. However, their interpretation is complicated by heterogeneous regional identity, variable maturation state, and cell-type composition across protocols \[5,6\]. Existing single-cell studies of human hippocampus cover prenatal \[7\] and postnatal \[8-10\] stages but do not provide a continuous developmental reference. By elevating the atlas approach in utilizing single-cell RNA-sequencing data, we obtain standardized information on the organoid cell class and type composition and maturation states. Here, we constructed the Human Developing Hippocampus Atlas (HuDeHA), an integrated single-cell reference comprising 658,059 cells spanning post-conceptional week 3 to 15.3 years, and used it to benchmark iPSC-derived brain organoids carrying PSEN1 E280A which is associated with fAD in a large Colombian population. Reference-based mapping revealed altered cellular composition in PSEN1 E280A organoids, including reduced radial glia and increased neural crest-derived neurons. These changes were accompanied by cross-lineage transcriptional alterations, including broad upregulation of the ventral patterning factor MEIS2 and reduced expression of the \{beta\}-binding protein transthyretin (TTR) in choroid-plexus and ependymal-associated populations. Reconstructed neuronal-lineage trajectories showed a shift toward mature states in PSEN1 E280A organoids. Together, these findings establish HuDeHA as a resource for developmental benchmarking of hippocampus-relevant organoid systems and describe cell-lineage-specific developmental changes in PSEN1 E280A organoids that may inform interpretation of early cellular alterations in fAD.

## AlphaGenome Atlas: in silico mutagenesis of the entire human genome improves prioritization and interpretation of non-coding variants
- Source: medRxiv (preprints)
- Date: 2026-09-20
- Categories: Genomics & sequence analysis, Tools & resources
- Authors: Cheng, J., Taylor, K. R., Nicolaisen, L., Pan, J., Bycroft, C., Perino, M., Ward, T., Hawkes, G., Covill, L. E., Weilert, M., Thomas, R. W., Latysheva, N., Hirschmann, M. J., Chen, X. D., Beaumont, R. N., Chundru, V. K., Weedon, M. N., Bourdareau, S., Chu, H., Hariharan, D., Kagohara, T., Tenorio, L., Ushigome, Y., Shearer, C. A., Ikica, B., Fang, A., Naciri, M., Johnston, V., Green, R., Wong, L. H., Dutordoir, V., Mottram, A., Gayoso, A., Arvaniti, E., Novati, G., Rehm, H. L., Chen, F., Lareau, C. A., Wright, C. F., O'Donnell-Luria, A., Zeitlinger, J., Kohli, P., Avsec, Z.
- DOI: 10.64898/2026.09.16.26363192
- Source URL: <https://doi.org/10.64898/2026.09.16.26363192>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.16.26363192>

Abstract: A major challenge in genomics is deciphering the functional consequences of non-coding genetic variation. Here we present AlphaGenome Atlas, a comprehensive resource that enables the joint interpretation and prioritization of variant effects across the entire human genome. Using AlphaGenome, we predicted the regulatory effects across thousands of molecular phenotypes for every possible human single nucleotide variant and many observed indels. These predictions were then used to derive a unified and interpretable AlphaGenome Variant Impact (AVI) score and to map cis-regulatory motifs across the genome. AVI achieved state-of-the-art performance across diverse benchmarks with improved prioritization of deleterious non-coding variants. Application of the combined Atlas resource helped solve an epileptic encephalopathy rare disease case, increased the statistical power to detect rare non-coding variants driving population-level phenotypes, and enhanced the mechanistic interpretation of these variants. Thus, AlphaGenome Atlas improves the prioritization and molecular interpretation of non-coding variants with genetic and clinical significance.

## ASTRAL-X: Scaling Coalescent-Based Species Tree Inference to 300,000 Taxa
- Source: bioRxiv (preprints)
- Date: 2026-09-20
- Categories: Genomics & sequence analysis, Evolution & metagenomics, Tools & resources
- Authors: Saha, A., Bayzid, M. S.
- DOI: 10.64898/2026.07.31.742122
- Source URL: <https://doi.org/10.64898/2026.07.31.742122>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.07.31.742122>
- Code: <https://github.com/aaniksahaa/ASTRAL-X>

Abstract: Advances in genome sequencing have enabled phylogenomic studies involving tens or even hundreds of thousands of species. However, scalability remains a major computational challenge for statistically consistent species tree inference at this scale. ASTRAL, the most widely used coalescent-based species tree estimator, remains limited by computational and memory bottlenecks that make ultra-large analyses impractical. Here we present ASTRAL-X, a complete algorithmic redesign of the ASTRAL framework that overcomes these computational limitations. By fundamentally redesigning the underlying data representations, algorithms, and computational framework, ASTRAL-X dramatically reduces running time while lowering memory requirements to nearly the size of the input--the asymptotically optimal bound--thereby enabling statistically consistent species tree inference directly from unrooted gene trees at an unprecedented scale. ASTRAL-X preserves ASTRAL's statistical guarantees and achieves accuracy comparable to state-of-the-art methods across simulated and empirical datasets while reconstructing species trees containing 200,000 and 300,000 taxa in only 5 hours and 12 hours, respectively, using modest computational resources. Notably, ASTRAL-X reconstructed the evolutionary history of 9\{,\}524 angiosperm species in only 16 minutes. These results enable statistically consistent coalescent-based species tree inference at the scale demanded by emerging Tree of Life initiatives. ASTRAL-X is publicly available at \\url\{https://github.com/aaniksahaa/ASTRAL-X\}.

## BELL: Biomodel Evidence and LLM-based Logic
- Source: bioRxiv (preprints)
- Date: 2026-09-20
- Categories: Tools & resources
- Authors: Arazkhani, N., Cochran, B. H., Miskov-Zivanov, N.
- DOI: 10.64898/2026.09.18.721350
- Source URL: <https://doi.org/10.64898/2026.09.18.721350>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.18.721350>

Abstract: Building accurate and predictive mechanistic models requires careful biological interaction curation and verification against existing knowledge. When done manually, these tasks become impractical, especially with massive extraction of interactions facilitated by advanced natural language processing methods and large language models (LLMs). We present BELL (Biomodel Evidence and LLM-based Logic), a biocuration support framework that automates evidence retrieval, scoring, and explanation for interaction-level verification. BELL processes each interaction through a five-step pipeline: entity grounding, database ranking, evidence retrieval from seven biological databases, a heuristic four-dimension programmatic scoring, and chain-of-thought explanation with a recommended curator action generated by LLMs. We applied BELL on 210 protein-protein interactions from a curated Glioblastoma Multiforme (GBM) model. Results show that 49.5% of interactions achieved HIGH confidence and 72.4% received a positive curator recommendation, while qualitative flags precisely directed curator attention to evidence gaps. BELL is integrated into the KALIMBA curation platform and available at www.boheme.pitt.edu/Kalimba.

## BRIDGE-AD reveals Alzheimer's disease effectors through interpretable large-scale omics integration
- Source: bioRxiv (preprints)
- Date: 2026-09-20
- Categories: Genomics & sequence analysis, Single-cell & spatial, Tools & resources
- Authors: Cerneckis, J., Baltusyte, G., Convey, H., Sun, G., Abela, Z. C. E., Ramirez, M., Wang, D., Sun, G., Zhou, T., Spring, D., Saeb-Parsy, K., Han, N., Shi, Y.
- DOI: 10.64898/2026.09.14.750802
- Source URL: <https://doi.org/10.64898/2026.09.14.750802>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.14.750802>

Abstract: The growing landscape of Alzheimer's disease (AD) datasets creates opportunities to integrate heterogeneous evidence and systematically discover disease effectors. We present BRIDGE-AD, an interpretable network medicine framework that transforms multimodal data into a unified, disease-specific gene representation for AD effector prioritisation. We integrated more than 30 datasets and curated resources spanning omics, functional, genetic and prior disease knowledge layers. BRIDGE-AD outperformed recently published pretrained and modality-specific gene embeddings in recovering AD-associated genes and produced a genome-wide resource of candidate AD effectors. Established and newly prioritised effectors formed 19 functional clusters, revealing a global molecular landscape of AD biology. BRIDGE-AD supported an SPP1-centred cross-compartment hypothesis and nominated SCARB2, a poorly characterised candidate, for functional validation. SCARB2 rewired lysosomal, lipid-handling and autophagic programmes in microglia, whereas disrupted SCARB2 glycosylation in AD implicated altered SCARB2 processing and function. The accompanying website, explore-bridgead.com, enables users to trace the curated evidence and generate mechanistic hypotheses.

## CUBE: Multimodal Representation Learning Reveals Biological Structure Across Histomorphology, Spatial Protein Phenotypes, and Transcriptome-Associated Signals
- Source: bioRxiv (preprints)
- Date: 2026-09-20
- Categories: Single-cell & spatial, Biological imaging
- Authors: Ge, Z., Cai, H.
- DOI: 10.64898/2026.09.14.751379
- Source URL: <https://doi.org/10.64898/2026.09.14.751379>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.14.751379>

Abstract: Integrating histological, spatial protein, and transcriptomic information into a biologically grounded representation remains challenging because these modalities are rarely available as fully paired measurements, while existing computational approaches are commonly developed around individual modality pairs. To address this, we present CUBE (Colorectal Universal Representation & Bridge Encoder), a multimodal representation-learning framework that uses hematoxylin and eosin (H&E) histology as a bridge to integrate spatial protein phenotypes with transcriptome-associated information from incompletely paired data. CUBE independently learns representations from H&E-multiplex immunohistochemistry (mIHC) and H&E-pseudo-ST relationships and integrates them through attention-based fusion with biological grounding from mIHC-derived concepts. The H&E-mIHC representation supported competitive spatial protein reconstruction, while the pseudo-ST-supervised representation transferred to experimentally measured Visium HD spatial transcriptomics data and improved further after decoder-only calibration with the encoder frozen. Importantly, the fused representation retained biological information beyond its direct training targets, capturing immune-epithelial spatial organization and an independently measured ECM-receptor interaction transcriptomic program. Together, these findings demonstrate that separately paired spatial modalities can be organized through histology into a biologically structured and testable multimodal representation, providing a proof-of-concept strategy for multimodal tissue learning without requiring fully paired molecular measurements.

## Development and evaluation of a core genome multi-locus sequence typing scheme for the Enterobacter cloacae complex
- Source: bioRxiv (preprints)
- Date: 2026-09-20
- Categories: Genomics & sequence analysis
- Authors: Miller, H. C., Bakker, S., Dyet, K., Winter, D.
- DOI: 10.64898/2026.09.14.751580
- Source URL: <https://doi.org/10.64898/2026.09.14.751580>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.14.751580>

Abstract: The Enterobacter cloacae species complex (ECC) comprises a group of closely related, opportunistic Gram negative bacteria of major public health concern due to their frequent involvement in healthcare-associated infections and their capacity to acquire and disseminate multidrug resistance, including carbapenemases. Because members of this complex are often difficult to distinguish phenotypically and their taxonomic status is subject to debate, there is a need for a standardized, species-complex wide typing scheme based on whole genome sequencing (WGS) that can be applied to any species within the complex. In this study we have developed and evaluated a core genome multi-locus sequence typing (cgMLST) scheme suitable for species within the ECC. Using 3442 publicly available genomes from 27 ECC species or subspecies we developed a scheme with 1812 loci, comprising loci present in 99% of all genomes. Among the 3442 isolates in our study, 99.9% had 95% or more of the cgMLST targets, indicating that the schema is well-defined and representative for the breadth of ECC species in our study. On two independent evaluation datasets, the scheme reliably resolved epidemiologically linked isolates with 0-3 allelic differences and returned the same outbreak clusters defined previously by higher resolution core genome SNP (cgSNP) analysis. Hierarchical clustering analysis at different levels of resolution showed that the cgMLST profiles could potentially be used to differentiate between species and sub-lineages in the complex. The cgMLST schema will improve the ability of public health laboratories to perform WGS-based surveillance of ECC species.

## Dynamic Pocketome of Trace Amine-Associated Receptors
- Source: bioRxiv (preprints)
- Date: 2026-09-20
- Categories: Proteins & structural biology
- Authors: Rienaecker, C., Nicoli, A., Selent, J., Di Pizio, A.
- DOI: 10.64898/2026.09.14.751399
- Source URL: <https://doi.org/10.64898/2026.09.14.751399>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.14.751399>

Abstract: Trace amine-associated receptors (TAARs) are class A GPCRs that span two distinct physiological roles: TAAR1 is a CNS drug target, whereas TAAR2 to TAAR9 detect volatile amines in the olfactory epithelium. Recent experimental structures resolve their architecture and ligand-binding mode, but capture only static snapshots, which cannot address how the binding site and the overall pocketome respond to ligand binding. Here, we present a simulation library comprising 26 experimental structures of four human and murine genes in the apo and holo states, each in triplicate (total aggregate time of 156 microseconds). Cavities were detected and analysed across the entire receptor surface throughout each trajectory. Orthosteric changes did not follow a single direction when comparing the states: apo sites were neither uniformly smaller nor uniformly more flexible than their holo counterparts, instead pointing to a receptor- specific ligand-receptor interplay that propagates beyond the orthosteric pocket. This plasticity is further highlighted by the size composition and stability of allosteric pockets, which proves that some regions are larger in the apo state while others are larger when a ligand is present. Resolving such trends required pockets to be comparable across trajectories, which a novel global identifier (GID) provides. Although agnostic to functional annotation, the GID recovered the orthosteric site as a single region in both states and matched pockets across replicates of the same receptor-state pair. The result is a dynamic pocketome map of the TAAR family obtained by an approach transferable to other membrane proteins.

## Identifying Putative Pathogenic Non-Coding Variants in Unresolved Rare Disease Patients Using Topologically Associated Domains
- Source: bioRxiv (preprints)
- Date: 2026-09-20
- Categories: Genomics & sequence analysis, Tools & resources
- Authors: Gacita, A. M., Pahl, M., Torres, M. D., Ganesan, S., Blair, J. J., Patel, K., Ramakrishnan, R., Conlin, L., Helbig, I., Grant, S. F. A.
- DOI: 10.64898/2026.09.17.752339
- Source URL: <https://doi.org/10.64898/2026.09.17.752339>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.17.752339>

Abstract: Unresolved rare disease is a major public health challenge affecting ~300 million people worldwide. At least 50% of these individuals remain genetically unresolved after applying exome sequencing and/or whole genome sequencing. One source of these missing diagnoses is the presence of rare variants within the non-coding genome that are detected but not interpreted by whole genome sequencing. In order to systematically evaluate candidate pathogenic non-coding variants, we created the Genomic Analysis of Variants in Unresolved Rare Disease (GAVURD) system. GAVURD leverages trio whole genome sequencing alignment data to produce a short list of putative pathogenic non-coding variants for a given proband. GAVURD uses best practices for de novo and rare inherited variant identification, links variants to human disease genes harnessing topologically associated domain (TAD) data, and rank prioritizes variants based on phenotypic overlap. As a proof-of-concept, we applied GAVURD to ten probands with unresolved rare disease and implicated six potentially causal non-coding variants based on a confluence of evidence supportive of pathogenicity. The GAVURD system serves an important role in prioritizing candidate non-coding causal variants for unresolved rare disease that can serve as the high value and informed focus of additional functional follow-up studies.

## Impact of mass oral cholera vaccination in an endemic area of the Democratic Republic of the Congo: a surveillance-based counterfactual modeling analysis
- Source: medRxiv (preprints)
- Date: 2026-09-20
- Categories: Proteins & structural biology
- Authors: Perez-Saez, J., Malembaka, E. B., Bouman, J. A., Bugeme, P. M., Hutchins, C., Jackson, J., Tshiwedi-Tsilabia, E., Hulse, J. D., Saidi, J. M., Rumedeka, B. B., Itongwa, M., Cumming, O., German, E., Kulondwa, J.-C., Dighe, A. B., Clutter, C., Lessler, J. T., Leung, D. T., Gallandat, K., Lee, E. C., Okitayemba, P. W., Mukadi-Bamuleka, D., Knee, J., Azman, A. S.
- DOI: 10.64898/2026.09.17.26363298
- Keywords: antibody
- Source URL: <https://doi.org/10.64898/2026.09.17.26363298>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.17.26363298>

Abstract: Background Oral cholera vaccines (OCVs) are a key component of cholera control recommended in cholera-endemic areas. Yet evidence of their population-level impact is limited, especially in Africa where most cholera deaths occur. Here, we estimate the impact of mass administration of OCVs in the cholera-endemic city of Uvira, Democratic Republic of the Congo. Methods We conducted enhanced cholera surveillance at the two official cholera treatment facilities in Uvira between Jan 2017 and Dec 2023, centered around a mass vaccination campaign that achieved 66% coverage with at least one dose of Euvichol Plus vaccine in 2020. We combined systematic rapid diagnostic case testing with repeated, representative population surveys capturing healthcare-seeking behavior, vaccination, population mobility, and antibody profiles to estimate seroincidence. We developed a Bayesian framework that integrates these data into an ensemble of mechanistic cholera transmission models that account for time-varying transmissibility, realistic immunity dynamics and loss of vaccination coverage to population turnover. OCV impact was assessed through ensemble counterfactual modeling of alternative vaccination scenarios. Findings We estimate that the 2020 mass vaccination averted 56% (95% Credible Interval: 34-81) of infections and deaths over the subsequent three years. This corresponds to 2,350 (median, 95% CrI: 890-8,960) averted facility-attended cases, and 44 (median, 95% CrI: 16-150) averted facility and community deaths. Vaccination of the entire eligible population of Uvira would have averted 64% (median, 95% CrI: 43-87) of cases and deaths, with a negative but limited influence of vaccine coverage loss because of population turnover (71% median, 95% CrI: 52-90 at half the population turnover). Interpretation Although mass vaccination averted a significant fraction of cholera cases that would have otherwise occurred in this endemic setting, imperfect coverage, population turnover, and high transmission rates contributed in offsetting the larger potential benefits of OCV. Successful cholera control in Uvira hinges on multisectorial approaches including provision of safe water and sanitation. Setting and communicating realistic expectations for mass vaccination programs in highly endemic areas is critical for maintaining confidence in the current generation of OCVs. Funding Gavi (M&E 9166 09 20 A16) and the Wellcome Trust (221688/Z/20/Z).

## Language-Model-Based Detection of Genetic Editing in Bacteria
- Source: bioRxiv (preprints)
- Date: 2026-09-20
- Categories: Genomics & sequence analysis
- Authors: Gabay, E., Burstein, D.
- DOI: 10.64898/2026.09.17.751846
- Source URL: <https://doi.org/10.64898/2026.09.17.751846>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.17.751846>

Abstract: Recent advances in genome editing allow easy genetic manipulation of bacteria, providing them with new traits, some of which could be hazardous, e.g. enhanced virulence or extended resistance to antibiotics. The ability to detect artificially modified bacteria is crucial for identifying potential bio-threats. However, malicious genome editing could be challenging to trace due to the natural exchange of genes among bacteria through horizontal transfer. After curating extensive datasets including natural genomes and simulated edited genomes, we utilized a natural language processing approach to detect edited genomes. We developed a transformer-encoder-based machine-learning classifier that, instead of analyzing words in sentences, models gene families in genomes. After training the model on our datasets, it is able to accurately detect genes artificially added to bacterial genomes due to their unnatural context. Our approach provides a scalable method for identifying engineered sequences without relying on specific marker genes, with potential applications in biosecurity, agriculture, GMO regulation and more.

## Learning interpretable kinetic models for biomolecular interaction networks
- Source: bioRxiv (preprints)
- Date: 2026-09-20
- Categories: Proteins & structural biology, Systems & networks
- Authors: Eliasian, R., Hazan, Y., Tzori, T., Raveh, B.
- DOI: 10.64898/2026.09.14.751382
- Source URL: <https://doi.org/10.64898/2026.09.14.751382>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.14.751382>

Abstract: Many cellular machines operate through weak, transient, and multivalent interactions whose functional states are governed by recurring interaction patterns rather than persistent molecular geometries. Here, we introduce interaction-based Markov state models (iMSMs), which construct interpretable kinetic models using unsupervised clustering of time-averaged, identity-resolved interaction distributions around a focal entity. Applied to nucleocytoplasmic transport at two molecular resolutions, iMSMs resolve graded interaction states spanning strong, partial, and weak engagement. During pore transport, partially engaged states provide faster routes to disengagement than strongly bound states, while the networks reveal transport pathways, interaction hubs, bottlenecks, and kinetic commitment. At a finer resolution, FG motifs exchange contacts while remaining associated, recovering established slide-and-exchange dynamics. iMSMs reproduce free energies, permeabilities, and multi-step kinetics, with nearly fourfold faster permeability convergence from truncated trajectories than direct transport-event counting. Thus, iMSMs connect rapidly exchanging contacts to graded interaction states and their kinetics across molecular resolutions.

## Leveraging Large Language Models for Colorectal Cancer Symptom Extraction from MIMIC-IV Clinical Notes
- Source: medRxiv (preprints)
- Date: 2026-09-20
- Authors: Lee, Y., Dinov, I., Hu, X., Jiang, Y.
- DOI: 10.64898/2026.09.15.26362961
- Source URL: <https://doi.org/10.64898/2026.09.15.26362961>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.15.26362961>

Abstract: Background: Much of the symptom burden in colorectal cancer (CRC) patients is documented in unstructured discharge-note narrative, and manual extraction is not scalable. Whether large language models (LLMs) outperform rule-based and named entity recognition (NER) methods has not been rigorously benchmarked. Objective: To benchmark rule-based, NER, and zero-shot LLM methods for extracting 46 cancer-related symptoms from CRC discharge notes against an adjudicated ground truth. Methods: We analyzed 2,704 discharge notes from CRC patients in MIMIC-IV. A 46-symptom target list was built from the Memorial Symptom Assessment Scale and the EORTC QLQ-CR29. Four approaches -- dictionary-based rule matching, pretrained clinical NER, and zero-shot Claude Haiku and Gemini 3.5 Flash -- plus two hybrid variants (LLM output with post-hoc rule-based negation filtering) were evaluated against a 200-note gold standard adjudicated by two raters (pooled kappa=0.71, macro kappa=0.49), using Macro/Micro F1, precision, and recall. Results: Gemini 3.5 Flash performed best (Macro F1=0.70, Micro F1=0.86, Macro Precision=0.74), followed by Claude Haiku (Macro F1=0.63, Macro Recall=0.71); both substantially outperformed rule-based (Macro F1=0.44) and NER (Macro F1=0.38) methods. Post-hoc negation filtering paradoxically degraded LLM performance (Gemini+Hybrid Macro F1=0.58; Claude+Hybrid Macro F1=0.54) by overriding correct predictions through rigid, fixed-window matching. Conclusions: Zero-shot LLMs substantially outperform rule-based and NER approaches for CRC symptom extraction; post-hoc negation correction should not be applied to LLM outputs without syntactic scope validation. Implications for Practice: Zero-shot LLM extraction offers a scalable, accurate alternative to manual chart review and traditional NLP pipelines for oncology symptom surveillance, without institution-specific rule development or model training.

## MAT-classifier: A memory-efficient pipeline for accurate genus level profiling from ancient metagenomic data
- Source: bioRxiv (preprints)
- Date: 2026-09-20
- Categories: Genomics & sequence analysis, Evolution & metagenomics, Tools & resources
- Authors: Dhibar, A., Matz, M. V.
- DOI: 10.64898/2026.01.28.702372
- Source URL: <https://doi.org/10.64898/2026.01.28.702372>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.01.28.702372>

Abstract: Background: Advances in sequencing technology have expanded opportunities to recover microbial DNA from ancient samples and reconstruct past environments and host-microbe interactions. However, the field remains constrained by computational challenges and accuracy problems, as rare ancient microbial DNA must be distinguished from abundant modern contaminants. Moreover, existing pipelines demand substantial computational resources, particularly memory, limiting their accessibility. Results: Here, we present MAT-classifier, a genus-level profiling workflow for detecting ancient microbial taxa from metagenomic projects, designed to reduce computational requirements while increasing accuracy. Unlike its counterparts, MAT-classifier first consolidates candidate references at the genus level and then performs independent alignments using conventional short-read aligners instead of metagenomic aligners. Using simulated datasets, we showed that this approach achieves more accurate classification of ancient taxa while requiring substantially less memory and shorter runtime than a modern counterpart, the aMeta pipeline. Benchmarking on multiple empirical ancient datasets further confirmed its low memory footprint and practical utility. Conclusions: MAT-classifier provides a reliable, computationally efficient, and accessible framework for ancient microbiome profiling. It lowers computational barriers while maintaining robust classification performance, facilitating broader application of ancient microbial DNA analysis.

## Measurement reliability bounds functional benchmarks and relocates where variant effect prediction fails
- Source: bioRxiv (preprints)
- Date: 2026-09-20
- Categories: Genomics & sequence analysis
- Authors: Zhang, N.
- DOI: 10.64898/2026.09.14.751496
- Source URL: <https://doi.org/10.64898/2026.09.14.751496>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.14.751496>

Abstract: Background. Variant effect predictors are increasingly benchmarked against multiplexed assays of variant effect (MAVEs) rather than clinical labels, which removes label circularity but introduces a new problem: a correlation against a measurement cannot exceed the measurement's own reproducibility, and precision varies sharply across the territories compared. Results. We scored nineteen predictors across sixteen strata of a frozen atlas of 64,178 saturation genome editing variants in seven cancer-susceptibility genes. From published replicate scores and standard errors we estimated each territory's reliability ceiling, and showed by simulation that the correction reduces error above a ceiling of about 0.45 and amplifies it below. Ceilings vary more across territory than predictors do, and correcting for them redraws the map at the splice extremes. The collapse at canonical splice sites is largely a property of the assay: the median shortfall relative to coding narrows from 1.7- to 1.4-fold; this convergence survives dropping BARD1 or PALB2 but inverts when BRCA1 is dropped, so we report all three leave-one-gene-out folds rather than claim gene independence, and the frontier parity rests on one deposit. Genuine failure lies 11-50 bp into the intron, which the uncorrected map presents as modest. Across MaveDB, 2,452 of 2,803 score sets carry, at the upper bound, what a reliability estimate needs, though a conventional column-name search finds only a tenth; among 674 human deposits with a computable ceiling, 29.9-51.8% fall below 0.90. Scored as classification against the assays' own functional calls in three genes, the same predictors separate damaging from tolerated better than their correlations suggest, though none reaches the strongest evidence band at the 95%-specificity operating point. Conclusions. Territory-resolved benchmarks should report a per-stratum reliability estimate, or state that the assay permits none. It asks nothing of depositors and applies today, at the upper bound, to most (87%) of MaveDB.

## Mechanical cues stabilize a conserved morphometric state associated with radial glia-like competence across systems
- Source: bioRxiv (preprints)
- Date: 2026-09-20
- Authors: Soriano-Esque, J. P., Borau, C., Sortino, R., Garma, L. D., Ortega, A., Asin, J., Alcantara, S.
- DOI: 10.1101/2025.01.09.632137
- Source URL: <https://doi.org/10.1101/2025.01.09.632137>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1101%2F2025.01.09.632137>

Abstract: Mechanical cues influence neural development, yet how tissue architecture is integrated into progenitor cell states remains poorly understood. Here, we show that defined microtopographies induce an early nuclear remodeling program associated with radial glia (RG)-like competence. Aligned microgrooves promote nuclear elongation, reduced Lamin A/C to B1 ratio, sustained \{beta\}-catenin activity, and distinct patterns of nuclear calcium dynamics, all of which merge before peak Pax6 expression. Pharmacological inhibition of mechanosensitive calcium signaling abolishes RG-associated marker induction while preserving nuclear remodeling, indicating that calcium-dependent pathways are required for Pax6 induction but are dispensable for the establishment of the underlying morphometric state. To quantitatively describe these transitions, we developed an interpretable model based on nuclear geometry and local cell density that identifies morphometric states associated with RG-like competence across experimental conditions. Application of this framework to embryonic mouse and human cortex revealed analogous signatures in native RG populations. Together, these findings indicate that tissue architecture influences RG-like competence through conserved nuclear morphometric states across developmental contexts.

## MIMoSA: A Tool for Model-Independent Comparison of Transcription Factor Binding Motifs
- Source: bioRxiv (preprints)
- Date: 2026-09-20
- Categories: Genomics & sequence analysis, Tools & resources
- Authors: Tsukanov, A. V., Levitsky, V. G.
- DOI: 10.64898/2026.05.13.725009
- Source URL: <https://doi.org/10.64898/2026.05.13.725009>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.05.13.725009>
- Code: <https://github.com/ubercomrade/mimosa>

Abstract: Transcription factors (TFs) regulate gene expression by binding specific DNA sequences, called transcription factor binding sites (TFBSs), and motifs summarize the sequence specificity of these interactions. Although the position weight matrix (PWM) remains the most widely used motif model, alternative models can capture dependencies between nucleotide positions. Available tools for motif comparison are designed only for PWM motifs, and converting a motif from an alternative model into a PWM often leads to a loss of information. We propose MIMoSA (Model-Independent Motif Similarity Assessment), a tool that compares motif models independently of their representation. MIMoSA compares recognition profiles produced by different motifs on the same DNA sequence set rather than their internal parameters. Comparison of MIMoSA with PWM-based tools TomTom and MACRO-APE with the HOCOMOCO motif collection ensured comparable performance of all tools. A case study of a ChIPseq dataset for ATF3 TF further supported the reliability of MIMoSA application. The tool is available at \\url\{https://github.com/ubercomrade/mimosa\}.

## miRstring: An RNA language model enables mature miRNA decoding and artificial small RNA design across species
- Source: bioRxiv (preprints)
- Date: 2026-09-20
- Categories: Genomics & sequence analysis, Proteins & structural biology, Tools & resources
- Authors: Peng, R., Li, X., fang, t., Yu, X.
- DOI: 10.64898/2026.09.17.752258
- Source URL: <https://doi.org/10.64898/2026.09.17.752258>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.17.752258>

Abstract: MicroRNAs (miRNAs) are processed from structured precursors and subsequently loaded into Argonaute proteins to repress target mRNAs. Yet generalized computational frameworks capable of decoding mature miRNAs from precursor context remain limited, constraining both cross-species annotation and rational artificial miRNA design. Here, we present miRstring, a biogenesis-aware RNA language framework that decodes the four boundaries defining the miRNA/miRNA duplex. Trained on 77,708 miRNA precursors spanning 414 species, miRstring outperforms existing methods under family- and species-held-out evaluations and accurately identifies the first nucleotide of mature miRNAs. Importantly, its attention mechanism highlights the miRNA/ miRNA\* boundary sites cleaved by endonucleases, indicating that the model captures biologically meaningful features. Furthermore, we employed miRstring to design optimal pre-miRNA scaffolds for artificial miRNAs and validated its efficacy in repressing target mRNAs. Taken together, miRstring establishes a scalable route from cross-species mature-miRNA annotation to predictive design, enabling artificial intelligence-driven miRNA engineering and extending computational miRNA analysis toward broadly applicable small-RNA biotechnology.

## Myocardial Stiffness Tracking & Assessment Toolkit Driven by Artificial Intelligence For Reproducible Benchmarking of Temporal Segmentation and Shear-Wave Velocity Stabilization on Synthetic Data
- Source: medRxiv (preprints)
- Date: 2026-09-20
- Categories: Biological imaging
- Authors: Tsebro, T., Safi, A., Wang, B., Hutchison, L., Chaudhry, K., Alzoubi, L., Noge, M., Hua, A., Ray, N., Chianale, A., Baghban-Bashi, P., Karayunusoglu, M., Xu, Y., Saed, A., Ghobriel, W., Malik, A., Fernandes, D. D.
- DOI: 10.64898/2026.09.17.26363353
- Source URL: <https://doi.org/10.64898/2026.09.17.26363353>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.17.26363353>

Abstract: Propose: Routine echocardiographic assessment can be affected by operator variability and time-consuming post-processing, which may limit rapid quantitative analysis. Real-time cine ultrasound demands low-latency automated processing for clinical usage. Existing approaches lack standardized evaluation frameworks, verified edge deployment, and structured reproducibility guarantees. To address these limitations, we built Myocardial Stiffness Tracking & Assessment Toolkit driven by Artificial Intelligence (MyoSTAT.AI). We report a deterministic, fully reproducible pipeline that performs real-time segmentation and SWE velocity estimation, evaluated entirely on synthetic data. Approach: Four U-Net-based architectures spanning single-frame and temporal designs were evaluated via systematic ablation: a 2D baseline, a 2.5D stacked-frame model, a 3D volumetric model, and a ConvLSTM variant. Experiments used synthetic reference corpora with split accounting the segmentation corpus contained 1,200 synthetic frames, and the SWE corpus contained 900 synthetic velocity-field cases. The SWE branch used Radon-transform-based propagation-direction estimation, time-domain shear-wave speed estimation, and configurable temporal stabilization on synthetic velocity fields. TensorRT FP16 deployment benchmarks were performed separately for the exported UNet2.5D segmentation model on RTX 3060 and Jetson Orin Nano hardware. Results: On the synthetic evaluation quantities, the ConvLSTM variant achieved the highest segmentation accuracy (Dice 0.994, IoU 0.987), an 18.6% improvement over the 2D baseline (Dice 0.808). A 2.5D model with a 3-frame temporal window achieved Dice 0.983 at substantially lower latency (433 ms vs. 1193 ms). TensorRT FP16 deployment yielded 341 FPS on the RTX 3060 and 90.4 FPS on the Jetson Orin Nano. Penalized least-squares temporal stabilization (smoothn, \{lambda\}=10) reduced shear-wave speed CoV by 63.8% at 2.7 ms latency overhead per frame, with no loss of edge preservation. Conclusions: We demonstrate a deterministic, reproducible computational benchmarking framework for cardiac segmentation and SWE velocity-field stabilization, evaluated entirely on synthetic data, together with preliminary inference feasibility on workstation and embedded hardware. These results represent a synthetic proof-of-concept rather than a validated clinical or operator-independent tool; acquisition of real echocardiographic data and clinical validation are required before any claim of diagnostic or deployment readiness can be made.

## Prioritizing Peptides for Targeted Mass Spectrometry Experiments Using Deep Learning
- Source: Journal of Proteome Research (journals)
- Date: 2026-09-20T00:00:00+00:00
- Categories: Proteins & structural biology, Tools & resources
- Authors: Shreyash Sonthalia, Bo Wen, Priank Dasgupta, Chris Hsu, Michael J. MacCoss, William Stafford Noble
- Journal: Journal of Proteome Research
- DOI: 10.1021/acs.jproteome.6c00457
- Source URL: <https://doi.org/10.1021/acs.jproteome.6c00457>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1021%2Facs.jproteome.6c00457>

Abstract: One critical step in any targeted mass spectrometry experiment is selecting, from each protein of interest, a small number of peptides that respond well in the mass spectrometer and can serve as reliable proxies for protein quantification. Existing methods select target peptides either by relying on prior empirical measurements, limiting their applicability to previously observed peptides, or using machine learning to predict peptide behavior from sequence alone. However, current machine learning tools suffer from various limitations, including using detectability as an indirect proxy for intensity, relying on small training sets, or ignoring the precursor charge state. In this study, we introduce Bromo, a transformer-based deep learning model that ranks peptide precursors from a given protein by their relative response, taking charge state into account. Trained on millions of annotated peptide pairs derived from large-scale, publicly available data-independent acquisition mass spectrometry data, Bromo consistently outperforms existing sequence-based methods across diverse, independent data sets. Furthermore, we show that fine-tuning Bromo on experiment-specific data can account for differences in sample preparation, sample matrix, and instrument platform, all of which influence which peptides serve as optimal targets. This adaptability makes Bromo a practical tool for selecting target peptides for selected reaction monitoring and parallel reaction monitoring assay development across a wide range of experimental conditions.

## Prototype-based explainable deep learning for sex classification of Mountain Gazelles in the wild
- Source: Scientific Reports (journals)
- Date: 2026-09-20T00:00:00+00:00
- Authors: Tali Shitrit, Amir Kedem, Efrat Yagur, Ilan Shimshoni, Anna Zamansky
- Journal: Scientific Reports
- DOI: 10.1038/s41598-026-71679-9
- Source URL: <https://doi.org/10.1038/s41598-026-71679-9>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1038%2Fs41598-026-71679-9>

Abstract: Wildlife preserves are critical for countering the decline of biodiversity, yet monitoring species distribution and demographics via direct observation is time-consuming and prone to observer disturbance. Deep learning has emerged as a scalable alternative for analyzing field data, yet the ’black-box’ nature of standard models hinders their adoption. This is a major obstacle in ecological research, where transparent reasoning is essential for trust and meaningful interpretation. We address this challenge through the lens of sex classification in the endangered Mountain Gazelle, utilizing a dataset of 1833 cropped bounding boxes extracted from unconstrained camera trap imagery. Reliable automated classification from field imagery is essential for assessing population structure and welfare metrics for effective conservation management. While Mountain Gazelles exhibit distinct sexual dimorphism in horn morphology and body size, inferring sex from unconstrained camera trap data remains challenging due to significant variability in pose, illumination, and occlusion. To enhance explainability in Mountain Gazelle sex classification, we employed a prototype-based deep learning framework. Unlike standard “black-box” models, this approach represents classes through learned visual exemplars (prototypes) that correspond directly to crucial image features, ensuring transparent reasoning. Building upon the PIP-Net architecture, we introduced a novel enhancement: the ability to learn prototypes of variable sizes and aspect ratios, moving beyond rigid, fixed-size patches. These prototypes are learned autonomously without manual trait annotation, allowing the model to discover distinct regional features. By optimizing across a list of candidate scales, our non-fixed prototype approach captures larger, more semantically coherent regions, achieving a global accuracy and F1-score of 75% alongside detailed explanations. Our analysis indicates that prominent prototypes frequently capture regions consistent with known morphological traits: both sexes are often identified via central body, leg, and head features, with the model distinguishing between them based on sex-specific morphological proportions. For males, horn-related prototypes serve as a distinct cue, reflecting the main dimorphic trait cited in literature. These representations suggest that the model’s decision-making relies on relevant morphological features rather than background artifacts, offering a transparent and accurate tool for wildlife classification. The dataset will be made publicly available upon publication.

## Recursive feedback between Piezo1 conformation and membrane mechanics drives self-organization into finite clusters
- Source: bioRxiv (preprints)
- Date: 2026-09-20
- Authors: Guo, Z., Bagchi, A., Dhankhar, M., dehghany dahaj, M., Shenoy, V.
- DOI: 10.64898/2026.09.14.751624
- Source URL: <https://doi.org/10.64898/2026.09.14.751624>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.14.751624>

Abstract: Piezo1 is a major mechanosensitive ion channel through which cells convert physical force into calcium-dependent signaling programs. In living membranes, this conversion depends not only on channel activation, but also on whether Piezo1 channels remain dispersed, assemble into finite clusters, or concentrate at sites where receptor signaling and mechanical forces reorganize the membrane. How single-channel force sensing is amplified into these collective spatial states remains unknown. Here we identify a membrane-feedback mechanism that converts single-channel mechanosensing into self-organized Piezo1 clusters. Coupling channel shape to membrane-cortex elasticity reveals that neighboring channels relax shared deformation fields, generating an effective interaction with short-range attraction opposed by longer-range repulsion. As channel density or membrane tension increases, this balanced interaction shifts Piezo1 from dispersed channels into mesoscale finite clusters. Brownian-dynamics simulations reproduce experimentally observed Piezo1 cluster geometries and swelling-induced cluster growth, while comparisons across distinct cellular systems place Piezo1 organization within a common density-tension framework. Applying the same mechanism to LPS-activated macrophages shows how receptor-induced membrane reorganization locally concentrates Piezo1 above the clustering threshold. Overall, these results recast Piezo1 mechanotransduction from isolated-channel force sensing to a membrane-driven self-organization process that spatially biases force-dependent calcium signaling within cells.

## Scaling coalescent-based species tree inference to 100,000 taxa with STELAR-X
- Source: bioRxiv (preprints)
- Date: 2026-09-20
- Categories: Genomics & sequence analysis, Evolution & metagenomics, Tools & resources
- Authors: Saha, A., Bayzid, M. S.
- DOI: 10.1101/2025.11.22.689894
- Source URL: <https://doi.org/10.1101/2025.11.22.689894>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1101%2F2025.11.22.689894>

Abstract: Summary methods reconstruct species trees from collections of gene trees while accounting for gene tree discordance and provide a statistically consistent framework for phylogenomic inference under the multispecies coalescent model. While existing triplet- and quartet-based approaches such as ASTRAL and STELAR have provable statistical consistency, their running time and memory usage restrict their applicability to ultra-large datasets. We introduce STELAR-X, a statistically consistent and highly scalable triplet-based phylogenetic inference algorithm that achieves an asymptotically optimal memory complexity of $O(nk)$ for \\textit\{n\} species and \\textit\{k\} gene trees, essentially matching the input size and allowing analyses to remain feasible as long as the input trees fit in memory, while also substantially reducing running time. STELAR-X achieves this through a compact integer tuple-based encoding of tree bipartitions, efficient precomputation of bipartition weights, and GPU parallelism. These innovations substantially reduce computational overhead in the underlying dynamic programming framework. Experiments demonstrate that STELAR-X achieves unprecedented scalability. On simulated datasets with 10,000 taxa and 1,000 gene trees, STELAR-X runs 3,576$\\times$ faster than ASTRAL-MP (the most scalable variant of ASTRAL) while using 13.9$\\times$ less CPU memory. STELAR-X analyzed a dataset of 100,000 taxa and 1,000 genes in 34.37 minutes using 58.40 GB RAM, and a 100,000-gene dataset with 1000 taxa in just 3.32 minutes using 74.10 GB RAM, scales that were previously intractable for statistically consistent summary methods. Moreover, applying STELAR-X to two large-scale avian datasets produced trees highly consistent with established bird phylogenies, demonstrating its robustness on biological data.

## Scaling Network Medicine with LLMs for Combinatorial Drug Repurposing in ER+ Breast Cancer
- Source: bioRxiv (preprints)
- Date: 2026-09-20
- Categories: Systems & networks
- Authors: Hamed, A. A., Fandy, T. E., Rocha, L. M.
- DOI: 10.64898/2026.09.14.728753
- Keywords: pathway
- Source URL: <https://doi.org/10.64898/2026.09.14.728753>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.14.728753>

Abstract: Drug repurposing can accelerate therapy discovery for ER+ breast cancer, but combination selection remains difficult. We developed an LLM-driven network medicine framework that extracts drug--target relationships from 595,122 PubMed abstracts, builds a cross-model consensus network, overlays it onto the KEGG estrogen signaling pathway, and enumerates complementary drug pairs. Candidate combinations are ranked by ComboRank, which aggregates pathway coverage, LLM consensus, RAG validation, cross-method agreement, and ClinicalTrials.gov precedent. The framework identified 166 significant pairs at FDR <= 0.05, including 31 with clinical-trial precedent, with 393 shared drug--target pairs across extraction strategies and 11 exact pairs additionally supported by the pathway overlay.

## SEAHORSE: A Serendipity Engine Assaying Heterogeneous Omics-Related Sampling Experiments
- Source: bioRxiv (preprints)
- Date: 2026-09-20
- Categories: Tools & resources
- Authors: Quackenbush, A., Kolluri, J., Biju, R., Nhong, S., DeConti, D., Shutta, K. H., Wu, H., Quackenbush, J., Saha, E., Eicher, T. D.
- DOI: 10.1101/2025.08.15.670514
- Source URL: <https://doi.org/10.1101/2025.08.15.670514>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1101%2F2025.08.15.670514>

Abstract: Large public molecular atlases such as the Genotype-Tissue Expression (GTEx) project and The Cancer Genome Atlas (TCGA) invite systematic discovery, yet most analyses remain hypothesis-driven and interrogate a tiny fraction of possible relationships among phenotypic, clinical, and molecular variables. We developed SEAHORSE (Serendipity Engine Assaying Heterogeneous Omics-Related Sampling Experiments), a discovery engine and accompanying R package that exhaustively precomputes all pairwise associations across heterogeneous data types and presents them as a searchable association landscape. Using GTEx (948 donors, 43 tissues, 154 phenotypes), SEAHORSE generated 341,008 phenotype-phenotype associations, 125,246,938 phenotype-gene associations, and 10,269,910,410 gene-gene correlations. In parallel analyses spanning 33 tumor types in TCGA, SEAHORSE generated 625,042 phenotype-phenotype associations, 183,080,369 phenotype-gene associations, and 12,096,948,950 gene-gene correlations. Across GTEx, height was repeatedly associated with enrichment of transcriptional programs, most strikingly the Kyoto Encyclopedia of Genes and Genomes (KEGG) term "Pathways in Cancer," significant in 17 tissues, offering a molecular entry point into long-reported links between stature and cancer risk. Height was also associated with immune, cardiovascular, and neurologic programs. In TCGA, age was consistently associated with WNT signaling, translation, cell differentiation, and cell cycle programs across tumors. These findings illustrate a new paradigm: large cohorts should be treated not merely as repositories for testing preconceived hypotheses but as association landscapes that can generate unexpected biological hypotheses.

## SMORE: joint dimension reduction and cell population discovery on single-cell methylome data
- Source: bioRxiv (preprints)
- Date: 2026-09-20
- Categories: Genomics & sequence analysis, Single-cell & spatial, Tools & resources
- Authors: Deng, J., Wang, Z., Tang, W., Hu, G., Feng, H.
- DOI: 10.64898/2026.09.14.751514
- Source URL: <https://doi.org/10.64898/2026.09.14.751514>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.14.751514>

Abstract: Single-cell DNA methylation profiling technology captures novel epigenetic data modality but are challenging to analyze because of their heterogeneity, high dimensionality, and ultra-sparsity. Here we present SMORE (Single-cell MethylOme Reduction and Embedding), a computational method for joint dimensionality reduction and cell population discovery dedicated to single-cell DNA methylation data. SMORE operates on a Bayesian framework that converts methylation proportions into ordered methylation states and jointly infers a low-dimensional representation, cell populations and their number. By using low-rank latent Gaussian factorization and adopting a mixture-of-finite-mixtures prior on latent cell scores, SMORE infers cell assignments without requiring a prespecified cluster number, and propagates uncertainty from methylation measurements to cell assignments. Across simulations spanning varying sample sizes, population imbalance, signal strengths and model misspecification, SMORE accurately recovered latent population structure and outperformed existing methods. Applied to human single-cell methylation datasets from lung, peripheral blood and primary motor cortex, SMORE recovered biologically supported cell population structures. SMORE provides an uncertainty-aware framework for dimension reduction and population discovery for single-cell methylomes.

## StressNET: an adaptable deep-learning model for mechanical stress inference in tissues
- Source: bioRxiv (preprints)
- Date: 2026-09-20
- Categories: Biological imaging, Tools & resources
- Authors: Chara, O., Aldecoa Rodrigo, N., Borges, A., Miranda-Rodriguez, J. R., Ventura, G., Sedzinski, J., Lopez-Schier, H.
- DOI: 10.64898/2026.09.14.751456
- Source URL: <https://doi.org/10.64898/2026.09.14.751456>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.14.751456>

Abstract: Mechanical interactions between cells are fundamental to tissue morphogenesis during development and regeneration. Computational methods that infer intercellular stresses from microscopy images of cell shapes offer a non-invasive alternative to experimental perturbation techniques, yet all existing approaches rely on explicit physical models. Here we present StressNET, a Graph Neural Network (GNN) that infers intercellular mechanical stresses directly from tissue geometry, without assuming any underlying physical model. We generated synthetic datasets to train and benchmark StressNET, and demonstrate that its predictions achieve state-of-the-art correlation with experimental stress proxies in zebrafish neuromasts and Xenopus embryos. Analysis of the network's latent space reveals that StressNET learns global organizational principles of mechanical stress distribution, beyond local cell-cell interactions. StressNET is open-source and provides pre-trained models that can be fine-tuned on in vivo data, making it a broadly adaptable tool for studying tissue mechanics across biological systems.

## Structural stability of mutualistic networks over large geographic and temporal scales
- Source: bioRxiv (preprints)
- Date: 2026-09-20
- Categories: Evolution & metagenomics
- Authors: Perez-Lamarque, B., Andreoletti, J., Morillon, B., Pion-Piola, O., Lambert, A., Morlon, H.
- DOI: 10.1101/2025.10.08.681159
- Source URL: <https://doi.org/10.1101/2025.10.08.681159>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1101%2F2025.10.08.681159>

Abstract: Mutualistic interactions form species-rich, complex networks that play essential roles for ecosystem function. Over macroevolutionary time scales, global- and continental-level networks change as species emerge and go extinct, yet the stability of their structural organization remains poorly understood. Here, we develop ELEFANT, a novel method for reconstructing ancestral interaction networks based on the hypothesis that species interactions are shaped by unobserved traits that evolve over evolutionary time. We show that ancestral interaction networks can be reliably reconstructed from present-day phylogenetic and interaction data. We infer the ancestral networks of plant mutualisms involving arbuscular mycorrhizal fungi, bat pollinators, and bird seed dispersers at large biogeographic scales. We find that these mutualistic networks exhibit a modular structure that seems to have persisted for millions of years, in part maintained by the evolutionary conservatism of species interactions. As species diversify, they tend to show limited shifts in mutualistic partners, which results in a remarkable long-term stability of mutualistic network structure at large spatial scales.

## Structure-based Antibody Renumbering
- Source: bioRxiv (preprints)
- Date: 2026-09-20
- Categories: Proteins & structural biology, Tools & resources
- Authors: del Alamo, D.
- DOI: 10.64898/2026.09.16.751788
- Source URL: <https://doi.org/10.64898/2026.09.16.751788>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.16.751788>

Abstract: Antibody numbering schemes like IMGT and Chothia assign each residue in the variable domain a consistent index based on substructural position. These annotations standardize sequences with different lengths, facilitating tasks ranging from engineering of individual molecules during drug development to large-scale curation of training data for de novo antibody design. Yet existing algorithms for performing this annotation process, termed renumbering, rely exclusively on amino acid sequence for inference. Consequently, these can fail when presented with unnatural or unusual features such as long CDRs or engineered insertions. To address this gap, this work introduces Structure-based Antibody Renumbering, abbreviated SAbR, a method that assigns these annotations from structure alone. SAbR shows comparable performance to sequence-based renumbering methods on held-out expert-annotated structures, as well as high agreement with sequence-based methods in a larger benchmark of diverse structures. It also outperforms peer methods on de novo-designed molecules, and shows higher success rates than renumbering by structural alignment. However, limited generalization performance is observed in more distantly related systems. Overall, these results establish structure-based renumbering as a robust alternative for natural and engineered antibodies when such data is available. Code and model weights are available on GitHub.

## The activation function of CA3 pyramidal neurons is optimal for the stable recall of memory patterns
- Source: bioRxiv (preprints)
- Date: 2026-09-20
- Categories: Computational neuroscience
- Authors: Cohen, U., Picher, M. M., Navas-Olive, A., Lengyel, M.
- DOI: 10.64898/2026.09.17.752025
- Source URL: <https://doi.org/10.64898/2026.09.17.752025>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.17.752025>

Abstract: Hippocampal area CA3 is widely believed to serve a core memory function: retrieving distributed patterns of neural activity that were previously stored in the recurrent connections between neurons. However, it remains unknown how the physiological properties of individual neurons contribute to this function. Here, we present a mathematical analysis of a canonical recurrent neural network model of CA3. Our analysis provides three main, experimentally testable predictions for recurrent circuits performing stable memory recall. Two of these predictions, that pyramidal cells should be characterized by elevated intrinsic excitability and by inhibition-dominated recurrent connections, are consistent with previous experimental findings. We thus focused on testing the third prediction: that neuronal activation functions have an exponent slightly above $1$. For this, we performed \\textit\{in vivo\} intracellular recordings of 49 pyramidal neurons from the CA3 area of behaving mice. Fitting a range of parametric models to these voltage traces and performing statistical model comparison revealed an average neural activation exponent slightly above $1$, confirming our theoretical predictions, with substantial heterogeneity across cells, which remains to be accounted for by the theory. An analysis of 133 further cells from two other previous studies provided neural activation exponent estimates highly consistent with those we found in our data. Our results suggest that the properties of hippocampal CA3 area are tuned to support reliable memory recall even at the level of single neurons.

## The interplay between detection and localization in human vision
- Source: bioRxiv (preprints)
- Date: 2026-09-20
- Authors: Coupette, F., Brainard, D. H., Smithson, H. E., Read, D. J.
- DOI: 10.64898/2026.07.06.736811
- Source URL: <https://doi.org/10.64898/2026.07.06.736811>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.07.06.736811>

Abstract: Detection and localization are fundamental but distinct tasks of sensory systems: detecting a signal requires accumulating evidence for its presence, whereas localizing it requires extracting spatial information. How active sampling should be organized when both tasks must be performed simultaneously remains unclear. Here, using an analytically tractable model of early visual processing and Bayesian ideal-observer inference, we establish a direct relation between the two tasks: localizing a stimulus is equivalent to detecting its spatial gradient. Consequently, detection and localization favor fundamentally different sampling dynamics when the stimulus size exceeds the effective blur scale of the visual system. Applied to fixational eye movements (FEMs), which continually translate stationary visual stimuli across the adapting retina, this distinction yields two competing optimal movement scales. Localization is optimized when the eye moves approximately one retinal blur length during the adaptation time, whereas detection is optimized by motion on the scale of the stimulus size. The scale of physiological FEMs lies near the predicted localization optimum, while stimulus detection remains comparatively robust to variations in eye motion. Our model recovers established laws of temporal and spatial summation while predicting additional scaling regimes arising from FEMs. We propose an experimental paradigm that tests these predictions by measuring localization at matched detectability, eliminating the unknown internal-noise amplitude. Our results reveal a fundamental trade-off between acquiring evidence for the presence and position of spatially extended signals and suggest that human fixational eye movements preferentially support spatial localization.

## Z-Hunt-DP: accelerating thermodynamic Z-DNA prediction with dynamic programming
- Source: bioRxiv (preprints)
- Date: 2026-09-20
- Categories: Genomics & sequence analysis, Tools & resources
- Authors: Al Jumaily, M., Qureshi, H., Yan, H., Li, Y.
- DOI: 10.64898/2026.09.14.751448
- Source URL: <https://doi.org/10.64898/2026.09.14.751448>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.14.751448>
- Code: <https://github.com/Aljumaily/Z-Hunt-DP>

Abstract: Z-DNA is a left-handed DNA conformation implicated in gene regulation and chromatin dynamics. Because it is usually less thermodynamically favorable than canonical B-DNA under physiological conditions, computational tools are needed to identify sequences likely to adopt the Z conformation. Legacy Z-Hunt uses a dinucleotide thermodynamic model, but searches every anti/syn assignment in a window, causing its conformation search to grow exponentially with window size. We present Z-Hunt-DP, an exact dynamic programming reformulation that preserves the original thermodynamic objective while reducing this search from O(2^d) to O(d) for a window of d dinucleotide positions. On benchmark windows, Z-Hunt-DP matched the brute-force minimum energy within numerical tolerance and achieved a 3.30x10^5 speedup at 20 dinucleotides. In an interval-localization benchmark on public human loci with experimentally mapped Z-DNA, it recovered the clipped reference interval in all 24 cases and was the most stable localizer under midpoint-core and expanded-panel analyses. Since the comparison set mixes thermodynamic, heuristic, and learned models, these cross-tool results are interpreted as localization comparisons rather than direct thermodynamic score tests. The source code and benchmark materials are available at https://github.com/Aljumaily/Z-Hunt-DP.

## A Bottom-Up Platform for Quantitative Single-ParticleTracking Through Bacterial Biofilm Mimics
- Source: bioRxiv (preprints)
- Date: 2026-09-19
- Categories: Genomics & sequence analysis, Proteins & structural biology, Biological imaging
- Authors: Shepherd, J. W., Howard, J. A. L.
- DOI: 10.64898/2026.07.02.736016
- Keywords: dna, microscopy
- Source URL: <https://doi.org/10.64898/2026.07.02.736016>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.07.02.736016>

Abstract: Chronic infections persist in large part thanks to protection that biofilms afford their bacterial creators. The extracellular polymeric substance of biofilms is a hydrated matrix of DNA, polysaccharides, and structural proteins, amongst other components, through which nutrients, signalling molecules, and antimicrobial agents must diffuse to reach the bacteria within. Quantitative measurement of transport on the nanoscale within in vivo biofilms remains challenging due to optical heterogeneity, autofluorescence, active remodelling of biofilms and the ambiguity in trajectory reconstruction during single-particle tracking (SPT). Here, we present a methodological framework for measuring molecular transport in defined minimal extracellular matrix models using quantum dots as fluorescent nanoscale probes imaged with high-speed SlimVar microscopy. To establish conditions in which high-diffusivity particle trajectories can be reliably reconstructed, upper limits to quantum dot concentrations were estimated from Brownian motion. The 99th-percentile inter-frame jump distance was estimated from the three-dimensional Brownian jump distance distribution and used to define a target average nearest neighbour distance, and therefore a per-particle volume, used for calculating a concentration which minimises the probability of trajectory collision during data acquisition. Quantum dot movement was imaged at sub-millisecond frame rates and diffusion coefficients were calculated in a 20% glycerol control and in DNA nanostar hydrogels modelling minimal extracellular matrix scaffolds assembled at 250 M and 500 M. Median diffusion coefficients decreased from 94.9 m2\*s-1 in glycerol to 15.9 m2\*s-1 and 8.3 m2\*s-1 in the 250 M and 500 M hydrogels, respectively. More broadly, this work establishes a workflow for quantitative SPT in minimal biofilm models. Rather than attempting to reproduce the full biological complexity of native biofilms, this approach provides the basis of a modular experimental framework in which individual extracellular matrix components can be incorporated sequentially and their effects on molecular transport quantified.

## A DNA methylation-based machine learning model for early and accurate diagnosis of cervical HSIL+ lesions
- Source: Scientific Reports (journals)
- Date: 2026-09-19T00:00:00+00:00
- Categories: Genomics & sequence analysis
- Authors: Yanfang Zhi, Ya Li, Jingjing Ren, Luqi Zhou, Yawen Yang, Yanmei Li, Canyu Li, Yannan Chen, Xin Zhao
- Journal: Scientific Reports
- DOI: 10.1038/s41598-026-68562-y
- Keywords: dna, methylation
- Source URL: <https://doi.org/10.1038/s41598-026-68562-y>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1038%2Fs41598-026-68562-y>

Abstract: Current cervical cancer screening methods lack accuracy in early diagnosis and risk prediction. We developed a DNA methylation-based diagnostic model for cervical high-grade squamous intraepithelial lesions or more severe lesions (HSIL+). This study systematically collected 172 liquid-based cytology samples from patients with positive human papillomavirus (HPV) test results. Bisulfite conversion-based next-generation sequencing (NGS) methylation sequencing technology was employed to quantitatively assess the methylation levels of consecutive CpG sites within specific segments in MIR9-3HG, TERT, GATA3, and CDKN2A genes across various grades of cervical lesions. Machine learning algorithms (LASSO regression, random forest, and support vector machine \[SVM\]) identified methylated characteristic CpG sites.The methylation levels of the four genes detected in the HSIL/ cervical squamous cell carcinoma(SCC) group were significantly higher than those in the Negative for Intraepithelial Lesion or Malignancy(NILM)/ low - grade squamous intraepithelial lesions (LSIL) group. Receiver Operating Characteristic (ROC) curve analysis showed that the areas under the curve (AUCs) for CDKN2A, MIR9-3HG, GATA3 and TERT in diagnosing HSIL+ were 0.880 (95% CI: 0.824–0.937), 0.779 (95% CI: 0.704–0.854), 0.769 (95% CI: 0.684–0.855) and 0.713 (95% CI: 0.627–0.800), respectively. The CpG methylation sites selected by the support vector machine (SVM), random forest algorithm and LASSO regression analysis were further cross-validated by Venn diagram. Finally, four CpG sites (all located in CDKN2A) were selected to successfully construct an efficient diagnostic model for HSIL+, with a sensitivity of 0.704 and a specificity as high as 0.929. The diagnostic model constructed in this study can accurately diagnose HSIL + of the cervix at an early stage, and it has significant clinical application value.

## A Gaussian process approach facilitates the identification of robust biomarkers for exposure to complex pesticide mixtures.
- Source: Environmental toxicology and chemistry (journals)
- Date: 2026-09-19
- Categories: Genomics & sequence analysis
- Authors: Ruben Bakker, Yuliya Shapovalova, Tjeerd M H Dijkstra, Tom Heskes, Cornelis A M van Gestel, Katja Hoedjes
- Journal: Environmental toxicology and chemistry
- DOI: 10.1093/etojnl/vgag252
- External ID: 42765347
- Source URL: <https://doi.org/10.1093/etojnl/vgag252>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1093%2Fetojnl%2Fvgag252>

Abstract: Biomarkers can provide a high-throughput and accurate assessment of the impact of complex chemical mixtures in the environment on organisms, but their identification through gene expression analysis is hindered by noise, synergistic interactions, and non-linear expression patterns. We generated finely resolved transcriptomic data from the ecotoxicological model species Folsomia candida exposed to two binary pesticide mixtures: One combining two neonicotinoid insecticides (imidacloprid and clothianidin) and the other a neonicotinoid (imidacloprid) with an azole fungicide (cyproconazole). Using these datasets, we developed a Gaussian Process (GP) framework to identify robust gene expression biomarkers, accounting for non-linear and synergistic interaction effects across experiments. Joint analysis of two binary mixtures increased the overlap of differentially expressed genes (DEGs) compared to separate analyses, improving robustness. In simulations, GP models outperformed linear models, accurately fitting complex, non-linear concentration-response relationships. Four biomarkers, three for neonicotinoids (ARRD, SMCT and nAchR) and one for azole fungicides (CYP), identified through this framework, were empirically validated and confirmed to be specifically responsive to their target pesticide, even under co-exposure. These findings highlight the effectiveness of GP models for mixture exposure transcriptomics and their broader applicability to other omics data and research fields.

## A reference genome without a virus: cDNA reconstruction reveals the provenance and function of the MS2 phage sequence
- Source: bioRxiv (preprints)
- Date: 2026-09-19
- Categories: Genomics & sequence analysis
- Authors: Small, E., Lasley, G., Layton, E., Wiwi, A., Del Curto, D., Weinstock, L. D., Thongchol, J., Zhang, J., CAHILL, J.
- DOI: 10.64898/2026.09.18.752701
- Keywords: genome, genomes, rna
- Source URL: <https://doi.org/10.64898/2026.09.18.752701>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.18.752701>

Abstract: Reference genomes are often treated as faithful representations of experimentally validated viral genomes, yet the relationship between historically curated reference sequences and infectivity is rarely tested experimentally. Here, we developed a cDNA-based reconstruction platform for the canonical RNA phage MS2 and used it to compare the current NCBI reference genome (RefSeq) with closely related published isolate sequences. We found that isolate-derived sequences reproducibly yielded infectious phage, whereas the current MS2 RefSeq-derived construct did not, showing that the present reference does not represent a single experimentally validated infectious genome but instead reflects sequence curation across multiple studies. We then compared conventional and AI-enabled approaches to identify minimal changes that restore infectivity to MS2 RefSeq; a human experimentalist correctly prioritized corrective changes, whereas the genome language model Evo2 did not. We also observed that closely related corrected reference-derived constructs showed a ~4-log difference in phage output, and subsequent analysis indicated that this difference was associated with an apparent replicase frameshift in the lower-output background. This suggests that the low output construct class represents rare mutations from genomes that are one mutational step away from true function, rather than uniform function of the dominant construct population. A complementary cell-free assay provided a lower-background orthogonal readout of construct-level function, yielding ~1 x106 PFU/mL from the high-output background within 2 hours while showing no detectable recovery from the low-output background. Together, these results establish a robust platform for RNA phage reconstruction and raise the possibility that historical reference genomes, especially for RNA viruses, may not always remain faithful to experimentally validated biological function. More broadly, these findings underscore the need to verify the infectivity of reference genomes, particularly when they were assembled non-contiguously or shaped by cumulative human curation. They also highlight the importance of clearly distinguishing historically curated reference sequences from experimentally validated infectious genomes when such data are used to train or evaluate AI/ML models.

## A systemic neuroendocrine immune axis in breast cancer revealed by MMD regularized cross tissue latent alignment across four independent cohorts.
- Source: Computers in biology and medicine (journals)
- Date: 2026-09-19T00:00:00Z
- Categories: Genomics & sequence analysis
- Authors: Hezil Nabil, A. Bouridane, Sumaya Al-Máadeed, Iman M. Talaat, R. Hamoudi
- Journal: Computers in biology and medicine
- DOI: 10.1016/j.compbiomed.2026.111936
- External ID: 52d6dea83cc12ed5581a5384f3867ea735bdcbf3
- Source URL: <https://doi.org/10.1016/j.compbiomed.2026.111936>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1016%2Fj.compbiomed.2026.111936>

Abstract: Understanding systemic determinants of breast tumor immunity requires bridging transcriptomically distinct tissue compartments that cannot be sampled simultaneously in a single patient. We developed an MMD-regularized Domain Adaptation Autoencoder (DAA) to align unpaired RNA-seq profiles from GTEx neuroendocrine tissues (n=189) and TCGA-BRCA tumors (n=1391) within a shared 128-dimensional latent space, enabling the first cross-tissue transcriptomic interrogation of the neuroendocrine-breast tumor immune interface. The dominant cross-tissue axis was identified by Pearson correlation and rigorously validated by permutation testing (n=1000 iterations), then independently assessed in METABRIC microarray (n=1980) and SCAN-B RNA-seq (n=3273) cohorts via a strict gene-intersection protocol that eliminated zero-padding artefacts. The DAA achieved stable cross-domain alignment (mixing score =26.58%), and Latent Dimension 31 emerged as a significant systemic immune-inflammatory axis (p=0.001; aggregate correlation 13.9× above the permutation null), driven by T-cell receptor variable chains, immunoglobulin genes, and the tolerogenic phospholipase PLA2G2D. METABRIC validation recovered a mechanistically concordant acute-phase secretory signature (LBP, SAA1, PLA2G2A), while SCAN-B confirmed PLA2G2D and CCL18 on a unified cross-platform latent axis. The latent score significantly stratified overall survival (p=0.0036) and relapse-free survival (p=0.0084), and precisely reproduced the established breast cancer immune topology across all six molecular subtypes (Kruskal-Wallis H=139.4, p<0.0001). External validation in the independent neoadjuvant GEO cohort GSE25066 (n=508; Affymetrix GPL96) via a Strict Intersection Protocol (604-gene intersection, zero-padding eliminated) confirmed axis recovery (Latent Dimension 115; PLA2G2D |r|=0.266), significant distant relapse-free survival stratification (log-rank p=0.022), and non-significant pathological complete response to chemotherapy (p=0.241), establishing the axis as a prognostic but not predictive biomarker. Functional annotation in GSE25066 revealed significant correlation with all 12 curated immune cell signatures (Spearman ρ=0.10-0.42; all padj<0.05), and GSEA pre-ranked analysis across 13,236 genes identified 34 significantly enriched Hallmark pathways (FDR < 0.25), led by Interferon Gamma Response (NES =2.90) and opposed by Estrogen Response Early (NES =-2.71). These findings, validated across 7152 patients in four independent cohorts, provide a computational transcriptomic framework linking systemic neuroendocrine regulation to breast tumor immunobiology and nominate PLA2G2D, SAA1, and LBP as candidate circulating biomarkers warranting prospective proteomic validation.

## AASIA: A comprehensive protein structural interactome database for agricultural animals.
- Source: Journal of advanced research (journals)
- Date: 2026-09-19
- Categories: Tools & resources
- Authors: Jiajun Li, Mengdi Yuan, Linyang Jiang, Dianke Li, Zhongtao Yin, Feng Zhu, Wenyu Shi, Zhuocheng Hou, Ziding Zhang
- Journal: Journal of advanced research
- DOI: 10.1016/j.jare.2026.09.015
- External ID: 42762846
- Source URL: <https://doi.org/10.1016/j.jare.2026.09.015>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1016%2Fj.jare.2026.09.015>

Abstract: INTRODUCTION: Agricultural animals are essential for global food production and sustainable agriculture. Improving disease resistance and economically important traits is critical for ensuring food security. Proteome-wide protein-protein interactions (PPIs), termed as interactomes, underlie key biological processes in cells, making them critical for deciphering the genome-phenome relationship. However, species-specific PPI resources remain limited for most agricultural animals. OBJECTIVES: Motivated by the rapid development of deep learning and breakthroughs in AI-driven protein structure prediction, we attempted to develop an integrated structural interactome resource for agricultural animals. METHODS: We predicted species-specific interactomes through an integrative prediction pipeline that combines interolog mapping, domain-domain interaction inference, and deep learning. The 3D complex structures of all predicted PPIs were further generated using ESMFold. RESULTS: We established AASIA (Agricultural Animals Structural Interactome Atlas; https://aasia.zzdlab.com), a user-friendly database comprising 410,421 high-confidence PPIs and corresponding complex structures across five key species, including Anas platyrhynchos, Bos taurus, Gallus gallus, Ovis aries, and Sus scrofa. CONCLUSIONS: AASIA provides the first proteome-scale structural interactome resource for five key agricultural animals. By integrating interaction prediction with structural modeling, it enables mechanistic insight into agriculturally relevant traits and supports applications in variant interpretation, multi-omics integration, and molecular breeding.

## Accurate RNA-Ligand Binding Site Prediction Based on a Multi-Channel Graph Neural Network.
- Source: Interdisciplinary sciences, computational life sciences (journals)
- Date: 2026-09-19
- Categories: Genomics & sequence analysis, Proteins & structural biology
- Authors: Na Li, Jingran Niu, Zhendong Liu, Jiamin Jiang, Bingbing Guo, Yujie Li, Jiafeng Yu, Dongqing Wei, Rongjun Man
- Journal: Interdisciplinary sciences, computational life sciences
- DOI: 10.1007/s12539-026-00883-y
- External ID: 42762429
- Source URL: <https://doi.org/10.1007/s12539-026-00883-y>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1007%2Fs12539-026-00883-y>

Abstract: RNA-ligand binding-site prediction is a challenging task in RNA molecular analysis. Binding regions are often sparse, structurally heterogeneous, and difficult to delineate accurately at the nucleotide level. Existing sequence-based methods lack explicit structural modeling, while conventional graph neural networks tend to mix signals around binding/non-binding transition regions. In this paper, BC-GNN, a multi-channel graph neural network for nucleotide-level RNA-ligand binding-site prediction, is proposed. BC-GNN integrates sequence-informed auxiliary transition estimation, boundary-aware propagation (BAP), microenvironment-aware channel recalibration (MACR), and hierarchical multi-scale integration (HMSI) to improve structural representation learning. When evaluated on a benchmark derived from RNAmigos2 using the official leakage-controlled 0.75 split, BC-GNN achieves an AUC of 0.8280, an F1-score of 0.6086, and an MCC of 0.4511, outperforming multiple re-evaluated baselines under the same rigorous protocol. These results demonstrate that BC-GNN is effective for RNA-ligand binding-site prediction.

## An automated deep learning pipeline for assessing aortic remodeling after frozen elephant trunk repair: a single-center pilot study
- Source: Scientific Reports (journals)
- Date: 2026-09-19T00:00:00+00:00
- Categories: Biological imaging
- Authors: Yu Nakano, Ikki Kojima, Masaki Kano, Toshiki Fujiyoshi, Yusuke Shimahara
- Journal: Scientific Reports
- DOI: 10.1038/s41598-026-72475-1
- Source URL: <https://doi.org/10.1038/s41598-026-72475-1>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1038%2Fs41598-026-72475-1>

Abstract: Computed tomography (CT) is the gold standard for assessing aortic remodeling following aortic dissection treatment. Clinical studies utilizing CT data are needed to evaluate the effectiveness of interventions such as the frozen elephant trunk (FET) procedure. However, large-scale studies are currently limited by the time-consuming and labor-intensive nature of manual image analysis. To address this, we developed an AI-based automated pipeline using open datasets to assess morphological changes in aortic dissection. We evaluated the pipeline’s efficacy using preoperative and postoperative CT scans from 14 patients who underwent FET repair. Two surgeons independently measured the aorta and true lumen manually. To separate segmentation error from plane selection error, one surgeon repeated the delineation on the plane selected by the pipeline. Because measurements were clustered within patients, agreement was assessed using linear mixed-effects models and repeated-measures Bland–Altman analysis. Agreement at a single time point was good to excellent (intraclass correlation coefficients: 0.90 to 0.95). On the identical plane, AI segmentation agreed closely with manual delineation (Dice: 0.951 for the aorta, 0.913 for the true lumen), with systematic bias arising largely from plane selection. This pipeline is feasible for future cohort-level research.

## Ancestral Sequences Cannot be Accurately Reconstructed via Interpolation in a Variational Autoencoder’s Latent Space
- Source: Bulletin of Mathematical Biology (journals)
- Date: 2026-09-19T00:00:00+00:00
- Authors: Evan Gorstein, Mengze Tang, Hailey Bruzzone, Claudia Solís-Lemus
- Journal: Bulletin of Mathematical Biology
- DOI: 10.1007/s11538-026-01749-6
- Source URL: <https://doi.org/10.1007/s11538-026-01749-6>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1007%2Fs11538-026-01749-6>

Abstract: Standard methods for ancestral sequence reconstruction (ASR) rely on substitution models for the residues in a biological sequence and assume independent evolution across these sites, ignoring the epistatic interactions that shape molecular evolution. In contrast, deep learning models like variational autoencoders (VAEs) can learn low-dimensional representations (“embeddings") of sequences in a protein family that may implicitly handle these dependencies, raising the possibility of performing more accurate ASR by interpolating between extant sequence embeddings within the VAE’s latent space. In this study, we test this hypothesis by developing and evaluating a VAE-based ASR pipeline. Benchmarking this approach against established likelihood-based and parsimony methods using various simulations of protein evolution, including scenarios with and without epistasis, we find that the VAE-based approach is consistently and significantly outperformed by standard methods, even in epistatic regimes where it was hypothesized to have an advantage. We further show that this failure is not due to a lack of phylogenetic structure in the latent space, which does contain evolutionary signal. Rather, the primary limitation is the information loss inherent to the autoencoding process: the VAE’s decoder cannot generate sequences with sufficient fidelity for the precise demands of ASR.

## Bayesian bilevel operator learning with low-rank adaptation for efficient uncertainty quantification of PDE inverse problems
- Source: Nature Communications (journals)
- Date: 2026-09-19T00:00:00+00:00
- Authors: Ray Zirui Zhang, Christopher E. Miles, Xiaohui Xie, John S. Lowengrub
- Journal: Nature Communications
- DOI: 10.1038/s41467-026-77768-7
- Source URL: <https://doi.org/10.1038/s41467-026-77768-7>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1038%2Fs41467-026-77768-7>

Abstract: Uncertainty quantification in PDE inverse problems is essential in many applications. Scientific machine learning and AI enable data-driven learning of model components while preserving physical structure, and provide the scalability and adaptability needed for emerging imaging technologies and clinical insights. We develop a Bilevel Local Operator Learning framework for Bayesian inference in PDEs (B-BiLO). At the upper level, we sample parameters from the posterior via Hamiltonian Monte Carlo, while at the lower level we fine-tune a neural network via low-rank adaptation (LoRA) to approximate the solution operator locally. B-BiLO enables efficient gradient-based sampling without synthetic data or adjoint equations and avoids sampling in high-dimensional weight space, as in Bayesian neural networks, by optimizing weights deterministically. We analyze errors from approximate lower-level optimization and establish their impact on posterior accuracy. Numerical experiments across PDE models, including tumor growth, demonstrate that B-BiLO achieves accurate and efficient uncertainty quantification.

## Benchmarking foundation models for tumor segmentation across multiple cancer types
- Source: Scientific Reports (journals)
- Date: 2026-09-19T00:00:00+00:00
- Categories: Biological imaging, Tools & resources
- Authors: Matteo Tortora, Elena Mulero Ayllón, Filippo Ruffini, Valerio Guarrasi, Paolo Soda
- Journal: Scientific Reports
- DOI: 10.1038/s41598-026-71014-2
- Source URL: <https://doi.org/10.1038/s41598-026-71014-2>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1038%2Fs41598-026-71014-2>
- Code: <https://github.com/arco-group/tumor_benchmarking>

Abstract: Tumor segmentation is a core task in medical image analysis, with direct implications for diagnosis, treatment planning, and disease monitoring. Whether promptable foundation models are mature enough for heterogeneous oncological scenarios remains an open question. We present a multi-cancer benchmark spanning lung, liver, kidney, brain, and breast tumor settings. Conventional supervised models (U-Net, DeepLabV3, Swin UNETR, nnU-Net) are compared against SAM-based foundation models (MedSAM and Medical SAM 2) under a common evaluation protocol. Prompt robustness is assessed by perturbing input bounding boxes through isotropic scaling and spatial shifting at inference time. Fine-tuned Medical SAM 2 with bounding-box prompting achieves the strongest benchmark-level profile, with the best results on Lung1, HCC, and KiTS23, while its zero-shot bounding-box configuration performs best on ATLAS. Swin UNETR ranks first on all three BraTS targets, and nnU-Net 3D full resolution performs best on ISPY1. Bounding-box prompting outperforms point-based guidance throughout the SAM-based family, and fine-tuning has a strong effect on performance. The robustness analysis reveals a trade-off: MedSAM tolerates prompt perturbations, while Medical SAM 2 achieves higher accuracy but degrades under box tightening and spatial shifts. These results support the use of SAM-based models for multi-cancer segmentation, while showing that reliability depends on prompt quality and anatomical context. Code is available at https://github.com/arco-group/tumor\_benchmarking .

## Challenging selective vulnerability in Parkinson's disease: a systematic review and meta-analysis
- Source: bioRxiv (preprints)
- Date: 2026-09-19
- Categories: Computational neuroscience
- Authors: Lunt, W., Moore, J. A., Cottard, E., Murphy, A. E., Shah, M., Sang, J., Choi, J., Dash, H., Dawson, S., Green, N., Nagaeva, E., Burke, S., Higgins, J. P. T., Skene, N. G.
- DOI: 10.64898/2026.05.13.724902
- Keywords: neuronal, systematic review
- Source URL: <https://doi.org/10.64898/2026.05.13.724902>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.05.13.724902>

Abstract: Selective vulnerability is widely assumed in Parkinson's disease (PD), but whether dopaminergic neurons of the substantia nigra are uniquely vulnerable has not been established. Explanations of neuronal degeneration in Parkinson's disease often centre on the distinctive properties of substantia nigra dopaminergic neurons. Whether these properties are necessary for substantial neuronal loss remains unclear. We preregistered a systematic review of 166 post-mortem case-control studies published between 1963 and 2025 and synthesised neuronal counts and densities from 152 studies using a multilevel meta-analysis. After six decades, only 18% of countable brain atlas labels had been examined, and most populations were represented by a single study. Substantial loss beyond dopaminergic and classically pigmented populations shows that neither dopaminergic identity nor neuromelanin is necessary for marked degeneration, challenging key tenets of selective vulnerability in PD. Comparisons across other anatomical and physiological features remain limited by sparse sampling and uncertainty. We identify less-studied populations with substantial estimated loss and estimate the additional sampling needed to improve precision under specified assumptions. These findings direct replication and comparative counting towards uncertainties that limit explanations of neuronal loss across affected and potentially spared populations.

## DNA extraction from sweetpotato (Ipomoea batatas) root tissues supports routine genotyping
- Source: Molecular Breeding (journals)
- Date: 2026-09-19T00:00:00Z
- Categories: Genomics & sequence analysis, Evolution & metagenomics
- Authors: Simon Fraher, Alexander M. Sandercock, Dong-Yan Zhao, Tyler Slonecki, Katarzyna Heller-Uszynska, Andrzej Kilian, Yasmin Cummins, Vidushi Patel, C. Beil, Moira J. Sheehan, G. Yencho
- Journal: Molecular Breeding
- DOI: 10.1007/s11032-026-01718-w
- External ID: 1b642ecbc59e536e1f37ebbdff9d20775a08b00f
- Keywords: dna, genomic, genotyping
- Source URL: <https://doi.org/10.1007/s11032-026-01718-w>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1007%2Fs11032-026-01718-w>

Abstract: Sweetpotato (Ipomoea batatas) breeders increasingly rely on genomic tools to enhance selection decisions. However, regrowing plants with sufficient leaf material for sampling requires several months, delaying genotyping and downstream decisions. Sampling storage root tissue would provide an earlier genotyping option, but root versus leaf tissue DNA extractions have not been compared for genotyping applications. Here, we compared genomic DNA from two storage root tissues, the cambium (“flesh”) and periderm (“skin”), and evaluated their performance against fresh leaf tissue. Root samples were taken at two storage times postharvest: 4 and 16 months. The approach produced adequate DNA, as assessed by sequencing depth and missing data rates, across all tissue types and storage times. Genotyping with a targeted 3,120 DArTag SNP panel revealed highly similar allele frequencies between root and leaf tissues (R2 > 0.96). Within-line dosage calls showed mean concordance of 81.3–85.5% for exact matches, increasing to 97.3–98.5% when allowing a ± 1 dose difference. While leaf tissue had higher read depth and lower missing rates, all root tissue types exceeded the minimum 90 mean read depth recommended for hexaploid dosage calling and fell within 5% of leaf tissue missing rates. Within-root tissue comparisons did not differ significantly across tissue type or storage time. Genetic relationships in principal component analysis were consistent across tissue types, supporting repeatability. Root tissues are therefore a suitable replacement for leaf tissue in routine genotyping workflows. This methodology enables faster selection decisions, resource savings, and genotyping outside the busy growing season.

## Genetic diversity of Legionella species in culture-negative clinical and environmental specimens by sequencing the 23S-5S ribosomal intergenic spacer region
- Source: bioRxiv (preprints)
- Date: 2026-09-19
- Categories: Genomics & sequence analysis, Tools & resources
- Authors: Jacqueline, C., Peticca, A., Lannes, J., Curtil-dit-Galin, M., Ibranosyan, M., Beraud, L., Descours, G., Jarraud, S., Ginevra, C.
- DOI: 10.64898/2026.09.18.752540
- Source URL: <https://doi.org/10.64898/2026.09.18.752540>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.18.752540>

Abstract: The diagnosis of Legionnaires' disease (LD) caused by Legionella non-pneumophila species is likely to increase with broader use of PCR targeting Legionella spp. In this context, accurate species identification in PCR-positive but culture-negative samples is essential to improve understanding of disease epidemiology and to support source attribution. Here, we presented a validated and user-friendly bioinformatic pipeline compatible with next-generation sequencing (NGS) for analyzing the hypervariable 23S-5S region, paired with a curated database encompassing all described Legionella species as of January 2026. Parameters were optimized for sensitivity and specificity using both strains and culture-positive clinical and environmental samples. We then applied the pipeline retrospectively to 92 culture-negative PCR-positive samples collected from 2023 to 2025. Legionella species were successfully assigned in 60% (55/92) of tested samples and revealed a high diversity. Co-infections were detected in clinical samples, including combinations of L. pneumophila with L. longbeachae or L. bozemanii, while environmental samples contained up to six different species. These results demonstrate that 23S-5S amplicon NGS enables species-level identification in the absence of cultured isolates, improving surveillance of non-pneumophila Legionella cases. The proposed pipeline, implemented in QIIME2 and accompanied by a publicly available database, provides a practical framework for routine molecular monitoring and outbreak investigation.

## HemaViT: transformer-based deep learning for automated non-invasive anemia detection using conjunctival imaging
- Source: Scientific Reports (journals)
- Date: 2026-09-19T00:00:00+00:00
- Categories: Biological imaging
- Authors: Gourishetty Sindhusha, Rupesh Kumar Mishra, R. Jegadeesan
- Journal: Scientific Reports
- DOI: 10.1038/s41598-026-70965-w
- Source URL: <https://doi.org/10.1038/s41598-026-70965-w>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1038%2Fs41598-026-70965-w>

Abstract: Anemia is a common hematological disorder that requires timely diagnosis to reduce the risk of severe health complications, particularly in resource-limited healthcare settings. This study proposes HemaViT, a transformer-based deep learning framework for automated, non-invasive detection of anemia from conjunctival images. The proposed framework integrates Dual-Attention PSPNet for accurate conjunctiva segmentation, a Feature Pyramid Network (FPN) with Multiscale Feature Exposure (MFE) for hierarchical feature extraction, a Vision Transformer (ViT) for global contextual representation, and the Improved Waterwheel Plant Algorithm (IWPA) for automated hyperparameter optimization. Experiments were conducted on the publicly available Eyes Defy Anemia dataset, which contains 1,320 conjunctival images, using stratified five-fold cross-validation. HemaViT achieved an average accuracy of 95.8%, precision of 97.6%, recall of 96.9%, F1-score of 97.2%, specificity of 98.7%, and an AUC-ROC of 0.98. Comparative experiments demonstrated that the proposed framework consistently outperformed widely used deep learning models, including ResNet50, DenseNet121, EfficientNet-B0, MobileNetV3, and a CNN + RNN hybrid model, while maintaining a favorable balance between classification performance and computational complexity. Ablation studies further confirmed the contributions of conjunctiva segmentation, multi-scale feature extraction, transformer-based contextual learning, and IWPA-driven hyperparameter optimization to the overall performance. Although additional validation on larger, more diverse clinical datasets is required, the proposed framework demonstrates strong potential for automated, non-invasive anemia screening in mobile health and resource-constrained clinical environments.

## Identifying cohorts at elevated risk of cancers using generative modeling of patient health states
- Source: medRxiv (preprints)
- Date: 2026-09-19
- Authors: Khan, A., Forster, D. T., Harsh, M., Zheng, C., Warner, E. T., Ritter, D., Chang, A., Wei, Q., Sorensen, T. K., Marks, D. S., Sequist, L. V., Hadlock, J. J., Fillmore, N. R., Sander, C.
- DOI: 10.64898/2026.09.09.26362676
- Source URL: <https://doi.org/10.64898/2026.09.09.26362676>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.09.26362676>

Abstract: While large language models are powerful generators of new text, forecasting disease progression from longitudinal health histories remains a challenging problem. We introduce GenEHR, an autoregressive generative model trained on electronic health records (EHRs) from millions of patients that explicitly represents the irregular time intervals between visits when forecasting future clinical events. We combine the general-purpose patient representation learned during foundational training with parameter-efficient supervised adaptation for the task of pan-cancer risk stratification. In five large EHR cohorts supervised adaptation substantially improved prediction performance of a first cancer diagnosis within a five year horizon window. Our retrospective results support the evaluation of GenEHR-CancerRisk as a prospective clinical decision-support tool for prioritizing patients for risk-based screening for aggressive cancer types, such as pancreatic and ovarian cancer.

## Mechanistic Interpretability of Fine-Tuned Protein Language Models for Nanobody Thermostability Prediction
- Source: Bioinformatics (journals)
- Date: 2026-09-19T00:00:00+00:00
- Categories: Proteins & structural biology
- Authors: Taihei Murakami, Yuki Hashidate, Yasuhiro Matsunaga
- Journal: Bioinformatics
- DOI: 10.1093/bioinformatics/btag685
- Source URL: <https://doi.org/10.1093/bioinformatics/btag685>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1093%2Fbioinformatics%2Fbtag685>
- Code: <https://github.com/matsunagalab/paper_nanobody-thermostability-sae>

Abstract: Motivation While Protein Language Models (PLMs) fine-tuned on biophysical data achieve high predictive accuracy, the physical principles underlying their predictions remain obscure. Deciphering these representations offers a unique opportunity to not only interpret model decisions but also to discover novel biophysical insights governing protein properties. Here, we present a framework using Sparse Autoencoders (SAEs) to extract mechanistic knowledge from PLMs fine-tuned for nanobody thermostability. Results We fine-tuned the ESM-2 model on the nanobody thermostability dataset, achieving superior performance compared to significantly larger state-of-the-art models. SAE analysis successfully decomposed the model's dense embeddings into sparse, interpretable features without loss of predictive accuracy. We characterized these features through both global and local analyses. Global analysis provided an aggregate map of position-dependent feature contributions, whereas local analysis identified specific residue-level patterns, including known determinants such as the VHH-tetrad and critical disulfide bonds, as well as candidate stabilizing residues. Free Energy Perturbation calculations supported the structural plausibility of selected residue-level hypotheses. These results show that SAE-based interpretation can generate testable, structurally grounded hypotheses for rational protein engineering. Availability The data and source code of the proposed method are available at GitHub (https://github.com/matsunagalab/paper\_nanobody-thermostability-sae) and Zenodo (DOI: 10.5281/zenodo.18012027). Supplementary information Supplementary data are available at Bioinformatics online.

## Multi-model biological and sequence information fusion for gene regulatory network inference from single-cell transcriptomics
- Source: bioRxiv (preprints)
- Date: 2026-09-19
- Categories: Genomics & sequence analysis, Single-cell & spatial, Systems & networks
- Authors: zhong, l., Yan, B., Wang, J., xie, m.
- DOI: 10.64898/2026.09.13.751326
- Source URL: <https://doi.org/10.64898/2026.09.13.751326>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.13.751326>

Abstract: Identification of transcription factor target gene interactions and construction of the gene regulatory networks (GRNs) are essential for understanding the molecular mechanisms underlying transcriptional gene regulation. Large scale single cell transcriptomics across different tissues offers unprecedented resolution of cellular diversity and regulatory dynamics by capturing gene expression heterogeneity. However, existing methods often lack effective multimodal integration and fail to fully exploit the hierarchical structure in Gene Ontology (GO) and gene sequence level representations, which limits their ability for predictive performance and biological interpretability. We present scMGFGRN, a multi-model deep learning framework that integrates single-cell transcriptomic profiles with GO hierarchical relationships, gene sequences by leveraging denoising auto encoders, graph attention feature extraction and pertained DNA language model to capture multi-source dependencies within multi-model biological knowledge, while its gated multi head attention module effectively identifies informative regulatory signatures and integrate complementary features from different sources to predict accurate gene regulatory networks. Benchmarking on the seven datasets of human and mouse demonstrates that scMGFGRN outperforms state of the art methods in identifying GRNs. Further analyses reveal that scMGFGRN effectively identifies novel TF gene interactions (TGIs) and reconstructs cell type specific GRNs. Interpretability analysis reveals the contribution patterns of heterogeneous biological sources, demonstrating the ability of scMGFGRN to integrate transcriptomic profiles with multi model structure information.

## Network-based gene prioritization using hybrid scoring for complex disease module discovery
- Source: Scientific Reports (journals)
- Date: 2026-09-19T00:00:00+00:00
- Categories: Systems & networks, Tools & resources
- Authors: Sveva Bonomi, Loris Bottelli, Elisa Oltra, Tiziana Alberio
- Journal: Scientific Reports
- DOI: 10.1038/s41598-026-71419-z
- Source URL: <https://doi.org/10.1038/s41598-026-71419-z>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1038%2Fs41598-026-71419-z>

Abstract: Complex diseases arise from perturbations in interconnected biological networks rather than isolated genetic defects. Network-based approaches provide a systematic framework for understanding disease mechanisms through protein-protein interaction data, yet most existing methods are developed and validated on a single disease or pathogenic mechanism, leaving their generalizability largely untested. We present a disease-agnostic computational framework that integrates five biological databases to construct robust ground truth gene sets, applies sensitivity analysis to identify high-confidence disease genes, and combines network topology with diffusion algorithms for systematic gene prioritization, requiring only a disease name as input. Our approach employs a noisy-OR fusion strategy to integrate evidence from DISEASES, GeneCards, OpenTargets, Gene2Phenotype, and NCBI Gene, followed by genetic algorithm optimization and random walk with restart for gene prioritization. Disease-relevant subnetworks were extracted and analyzed using complementary clustering algorithms (Leiden and MCL) to identify modules validated through pathway enrichment analysis. Without disease-specific tuning, the same workflow was applied unchanged to two neurodegenerative disorders and one autoimmune disease, chosen to span markedly different pathogenic mechanisms: Alzheimer disease (AD), Parkinson disease (PD), and rheumatoid arthritis (RA). In each case the framework recovered known disease genes and identified biologically coherent, disease-specific modules: in AD, 17 stable high-confidence genes and modules enriched in lipid and cholesterol metabolism and amyloid precursor protein processing; in PD, 17 stable genes and modules associated with mitophagy, ubiquitin-proteasome signaling, and mitochondrial dysfunction; in RA, 20 stable genes and modules enriched in antigen presentation and JAK-STAT cytokine signaling. This consistent recovery of mechanistically distinct, biologically appropriate signatures from a single unmodified pipeline demonstrates that the framework generalizes across disease categories rather than being tuned to any one of them. By prioritizing biological validation through pathway enrichment over predictive accuracy, and by demonstrating consistent performance across neurodegenerative and autoimmune contexts without disease-specific adaptation, the framework offers a versatile, readily extensible tool for exploratory analyses of disease mechanisms, including diseases for which prior mechanistic knowledge is limited. All code, data, and results are publicly available to ensure reproducibility.

## Optimizing the connectivity of protein conformations to untangle ensemble refinement
- Source: bioRxiv (preprints)
- Date: 2026-09-19
- Categories: Proteins & structural biology
- Authors: Passmore, S. K., Holton, J. M., Zatsepin, N. A., Martin, A. V.
- DOI: 10.64898/2026.09.17.752399
- Source URL: <https://doi.org/10.64898/2026.09.17.752399>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.17.752399>

Abstract: Proteins naturally adopt multiple conformations in mediating cellular processes, and ensemble models are used to fit X-ray crystallography data that captures this heterogeneity. In practice, ensemble refinement produces only minor improvements in agreement with experimental data (R-free) over single-conformation models. It has recently been shown that ensemble models are universally trapped, or "tangled"; refinement algorithms strain each individual conformation in the model to fit the electron density in its immediate vicinity, missing more harmonious ways to arrange the collection of protein conformations to fit the electron density. Here, we demonstrate that this type of trap may be escaped by formulating the construction of low-energy conformations from individual conformer coordinates as an integer linear programming problem. The method successfully recovered the two original protein conformations from a previously published synthetic dataset that traps current refinement methods. Inclusion of the method in an automated refinement procedure with real data is shown to improve R-free and reduce geometric strain in a four-conformation model by comparison with controls. Applying this method in combination with human input and fitting low-occupancy waters to density features in the bulk solvent, we produce models for deposited datasets of three separate 14-19 kDa proteins with greatly improved geometry and R-factors. This includes a 0.77 \[A\] six-conformation model of the SARS-CoV-2 macrodomain Mac1 (PDB ID: 44PS) with an R-work of 4.7% and an R-free of 6.4%.

## Reconstruction of FACS-partitioned Adaptive Immune Receptor Repertoires from FACS-partitioned B and T Cell Subsets
- Source: bioRxiv (preprints)
- Date: 2026-09-19
- Categories: Genomics & sequence analysis
- Authors: Zhao, H., Morgan, A., Yasuda, M., Mirebrahim, H., Schlecht, U., McNamara, S., Adachi, R., Rubelt, F., Kumar, D., Utiramerur, S., Arnaout, R., Asgharian, H.
- DOI: 10.64898/2026.09.13.746584
- Source URL: <https://doi.org/10.64898/2026.09.13.746584>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.13.746584>

Abstract: Adaptive immune-receptor repertoire sequencing (AIRRseq) is crucial for understanding immune system diversity and its relationship to disease dynamics. Partitioning of total B and T cells into their major subsets with distinct immunological functions - IgM+ vs. class-switched B cells (IgG+ > IgA+) and CD4+ vs. CD8+ T cells, respectively - allows for AIRRseq-based analysis of the unique contributions of each compartment to the overall immune response, a major advantage over traditional bulk sequencing workflows. However, data from these subsets is not directly comparable with the vast majority of publicly available AIRRseq data, which comes from unfractionated B and T cells, an important incompatibility. Here we investigate computational methods for reconstructing complete AIRRseq repertoires from partitioned B and T cell subsets in diverse individuals. Peripheral blood mononuclear cells (PBMCs) were partitioned via positive selection of IgM+ B-cell subsets and CD4+ T cells using immunomagnetic beads; genomic DNA was then extracted and B- and T-cell receptors were sequenced. Four reconstruction methods are introduced and evaluated for concordance with matching unpartitioned repertoires to assess preservation of key repertoire characteristics. Results show that these methods enable accurate estimates of overall immune-repertoire diversity from B- and T-cell subsets in a way that simply pooling the sequence data from sub-repertoires cannot.

## Redefining Non Invasive Post Transplant Surveillance: A Bayesian Meta Analysis and Decision Curve Framework for Donor Derived Cell Free DNA in Heart Transplantation
- Source: medRxiv (preprints)
- Date: 2026-09-19
- Categories: Genomics & sequence analysis
- Authors: John, J. D., Henna, F., Waseem, F., Hassan, M. A., Bacha, Z., Mukhlis, M., Mohammed, B. K., Cheema, S., Shah, K.
- DOI: 10.64898/2026.05.15.26353184
- Keywords: dna, framework
- Source URL: <https://doi.org/10.64898/2026.05.15.26353184>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.05.15.26353184>

Abstract: Donor-derived cell-free DNA (dd-cfDNA) is increasingly used for post transplantation non- invasive surveillance; however, its clinical interpretation remains inconsistent, with widely ranging thresholds and is typically applied as a single binary cutoff in literature. The optimal decision framework for rule-out and rule-in decisions, and whether a single threshold remains clinically meaningful, are currently uncertain. We performed a Bayesian hierarchical summary receiver operating characteristic (HSROC) meta-analysis of 14 studies (1,763 patients) evaluating dd-cfDNA against endomyocardial biopsy. To account for serial testing within individuals, we applied a cluster-corrected design effect, reducing 6,103 observations to 2,518 effective tests. Threshold-dependent sensitivity and specificity were modelled continuously. We compared a conventional single-threshold approach with a data-driven adaptive framework defining rule-out and rule-in thresholds and evaluated clinical utility by decision-curve analysis across rejection prevalences from 1% to 50%, incorporating repeat-testing strategies. The pooled area under the HSROC curve was 0.78 (95% CrI, 0.67-0.84). The Youden-optimal threshold (0.20%) yielded balanced sensitivity (0.77) and specificity (0.77) but failed to support clinical objectives of diagnosis. An adaptive framework identified a rule-out threshold of 0.16% (sensitivity 0.80) and a rule-in threshold of 0.48% (specificity 0.90), defining a indeterminate / grey zone. The residual one-in-five false-negative rate at the rule-out anchor reflects low-grade, non-cytolytic rejection, the imperfect histological reference standard and fractional suppression of the donor signal, rather than the statistical model; a result above the rule-in anchor carries a positive predictive value of approximately 38% at 10% prevalence and denotes an indication for tissue diagnosis and multimodal investigation, not for empiric treatment. Across low-to-intermediate prevalence, dd-cfDNA-guided strategies exceeded both the biopsy-all and monitor-all reference strategies; among testing strategies, repeat-if-borderline achieved the highest net benefit across the majority of the prevalence-threshold space and sustained positive net benefit over the widest operating range of any strategy, reducing false-positive biopsies without materially compromising detection. At high prevalence, where a first elevated result is usually true, biopsy-all became competitive. A single threshold is therefore clinically inadequate for post-transplant surveillance. Our tri-state, prevalence-aware framework integrating rule-out, indeterminate, and rule-in zones with selective repeat testing, more accurately reflects biomarker behavior and yields greater net benefit than any single cutoff because these anchors are pooled, population-level estimates rather than universal constants, programs should adopt this architecture and calibrate their own high-sensitivity rule-out and high-specificity rule-in thresholds to their local assay and population.

## Reproducible Agent-Based Simulations of Protein Aggregation: A FAIR Implementation.
- Source: Bulletin of mathematical biology (journals)
- Date: 2026-09-19
- Categories: Proteins & structural biology, Tools & resources
- Authors: Isabella V Gimón, Conner Sandefur, Santiago Schnell
- Journal: Bulletin of mathematical biology
- DOI: 10.1007/s11538-026-01739-8
- External ID: 42762374
- Source URL: <https://doi.org/10.1007/s11538-026-01739-8>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1007%2Fs11538-026-01739-8>

Abstract: Aberrant protein aggregation is implicated in many neurodegenerative diseases and is strongly modulated by intracellular spatial constraints such as macromolecular crowding and clearance. Computational studies of aggregation, however, frequently lack the documentation and provenance required for independent reproduction. We present a spatial, lattice-based agent‑based model of intracellular protein aggregation implemented in Julia and packaged as a research software object aligned with the FAIR (Findable, Accessible, Interoperable, and Reusable) principles. Individual monomers undergo reversible activation, oligomer nucleation, aggregate growth, and optional oligomer clearance; stochastic movement and local encounters on a 3D face-centered cubic lattice capture spatial heterogeneity and crowding. The accompanying repository includes centralized parameters, machine-readable metadata, version-pinned dependencies, example runs, and automated post-simulation analysis. Ensemble simulations reproduce canonical aggregation phases (lag, nucleation, growth, saturation) and illustrate that oligomer removal reduces the final aggregate burden. The model also supports configurable macromolecular crowding via spherical obstacles, enabling systematic exploration of crowding effects on aggregation kinetics. Runtime benchmarking shows that 300 independent simulations (1,000 monomers; 5,000 timesteps), executed as 15 concurrent single-threaded jobs on institutional high-performance computing resources, completed in approximately 25 h of wall-clock time, enabling parameter sweeps and ensemble averaging. Together, the model and its FAIR packaging provide a reproducible template for transparent, extensible agent-based computational biology.

## Scaling Functional Annotation Across Proteomes, Pangenomes and Metagenomes with Sma3s v3
- Source: bioRxiv (preprints)
- Date: 2026-09-19
- Categories: Genomics & sequence analysis, Proteins & structural biology, Tools & resources
- Authors: Rubio, A., Garcia-Junco, J. L., Luque-Jimenez, E., Martin Dominguez, A., Dopazo, J., Perez-Pulido, A. J., Casimiro-Soriguer, C. S.
- DOI: 10.64898/2026.09.14.748778
- Source URL: <https://doi.org/10.64898/2026.09.14.748778>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.14.748778>

Abstract: High-throughput sequencing has generated protein datasets whose scale increasingly exceeds the practical limits of conventional functional annotation workflows. We present Sma3s v3, a scalable reimplementation of the Sma3s three-step annotation strategy, which combines transfer from highly similar homologs, orthology-based inference, and functional enrichment among homologous proteins. Sma3s v3 replaces BLAST-based searches with MMseqs2 and introduces parallel processing, reusable SQLite caches, taxonomic filtering, and traceable outputs that retain the evidence underlying each assignment. We evaluated the method on a Vibrio cholerae pangenome comprising 50,415 gene clusters from 11,295 quality-filtered genomes and on a metagenomic catalogue containing 843,935 proteins. After excluding non-informative assignments, Sma3s v3 annotated 30,662 pangenome clusters (60.8%), comparable to InterProScan (60.2%) and exceeding eggNOG-mapper (41.4%), while providing 5,747 annotations not recovered by either comparator. Gene Ontology comparisons showed broad semantic agreement between methods, with Sma3s v3 frequently contributing more non-redundant information in Molecular Function and Biological Process. Within the pangenome, annotation coverage reached 97.1% for core clusters and approximately 59% for accessory and unique clusters. Exact protein matches to non-Vibrio genera identified 1,838 candidate horizontally transferred clusters enriched in genetic mobility, antimicrobial resistance, and metal tolerance functions. In the metagenomic catalogue, Sma3s v3 annotated 728,014 proteins (86.3%), compared with 616,895 (73.1%) using InterProScan 2026, and recovered approximately 20,000 unique functional terms. These results establish Sma3s v3 as a scalable and interpretable tool for functional annotation and re-annotation of proteomes, pangenomes, and metagenomic protein catalogues.

## Somatic haplotype reconstruction and variant recalibration from tumor-only long-read sequencing
- Source: bioRxiv (preprints)
- Date: 2026-09-19
- Categories: Genomics & sequence analysis, Tools & resources
- Authors: Chen, Z.-Y., Zheng, Z., Luo, R., Fu, H.-F., Yang, Y.-J., Huang, Y.-T.
- DOI: 10.64898/2026.09.14.751225
- Source URL: <https://doi.org/10.64898/2026.09.14.751225>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.14.751225>

Abstract: Separating somatic from germline variants and reconstructing somatic haplotypes are the two central problems of tumor-only cancer genome analysis. Long reads carry the linkage needed to solve both, but chromosome-scale loss of heterozygosity (LOH) and an unknown degree of normal-cell admixture blur the distinction between somatic and germline haplotypes. Here we present LongPhase-TO, the first method to reconstruct somatic haplotypes from a tumor sample alone. Rather than mapping somatic variants onto germline haplotypes, LongPhase-TO co-phases germline and somatic alleles in a unified graph, in which LOH and tumor DNA fraction are resolved internally from heterozygosity depletion and haplotype imbalance rather than a copy-number and ploidy model. Across eight datasets from six cancer cell lines, LongPhase-TO increased haplotype block N50 by a median of 2.9-fold relative to germline phasers. It also consistently improved somatic single-nucleotide variant (SNV) and indel calls from ClairS-TO and DeepSomatic-TO, raising mean F1 from 0.55 to 0.62 and 0.65 for SNVs and from 0.19 to 0.23 for indels, with the largest gains at low tumor DNA fraction. Across breast, melanoma and lung cancer cell lines, LongPhase-TO improves the accuracy of existing somatic callers and reconstructs megabase-scale somatic haplotypes.

## The building blocks of social structure: simulating constraints of social network analysis for inference about group-level properties
- Source: bioRxiv (preprints)
- Date: 2026-09-19
- Authors: Brooks, J., Badihi, G., Samuni, L.
- DOI: 10.64898/2026.09.17.752303
- Source URL: <https://doi.org/10.64898/2026.09.17.752303>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.17.752303>

Abstract: The interdisciplinary usage of Social Network Analysis (SNA) means that, when researchers calculate global network metrics, the research interest can range from the ultimate ecological and evolutionary pressures that select for specific group-level structures or the proximate mechanisms by which those structures emerge. Despite major methodological advances, social structures measured via SNA metrics are rarely connected to the underlying forces shaped by latent individual dispositions, ecological constraints, and their interplay with group-level affordances. We build a model to simulate fission-fusion spatial association data from a specified set of parameters in order to address the "black box" of animal SNA. We broke down animal social systems into three building blocks: individual dispositions, external factors, and the social structure that emerges. Each of these building blocks was represented by one or more model parameters, which together describe a set of underlying processes that lead to different measurable features of animal social structure, which we quantified with commonly used global network metrics within an SNA framework (e.g., clustering coefficient, density, modularity). In doing so, our model directly links global network metrics to the underlying behavioural processes that generate them and demonstrates how global network metrics can be influenced by changes to their underlying behavioural processes. Holding all other parameters constant, we find that subtle changes to 1) demography (i.e., group size) and ecology (i.e., size and variability of foraging parties), 2) variation of individual behaviour and related observation bias, 3) imposed group sub-structures, and 4) observation effort, affect all network metrics calculated non-linearly and in some cases non-monotonically. Our findings indicate that failure to account for the underlying behavioural processes can bias inference drawn from SNA. This model and conceptual framework provide a novel tool with which to begin opening the black box of SNA and to systematically address the evolution of diverse group-level social structures.

## Translational efficiency guides microbial community remodeling.
- Source: Gut microbes (journals)
- Date: 2026-09-19T00:00:00Z
- Authors: Qi Xiang, Ya-Nan Li, Jing-Peng Yang
- Journal: Gut microbes
- DOI: 10.1080/19490976.2026.2736907
- External ID: 279d462eac9dd8dac992c9f618ffd44402efd451
- Source URL: <https://doi.org/10.1080/19490976.2026.2736907>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1080%2F19490976.2026.2736907>

Abstract: While metagenomics provides compositional insights, its correlative nature limits causal community remodeling. Overcoming this, in a recent Cell study, Moyne et al. introduced the Microbial Interaction and Niche Determination (MIND) framework. By leveraging translational efficiency to map resource competition and niche partitioning, MIND establishes a mechanistic blueprint for rational engineering.

## Updated Transposable Element Libraries for Drosophila melanogaster in Dfam 4.0
- Source: bioRxiv (preprints)
- Date: 2026-09-19
- Categories: Tools & resources
- Authors: Goubert, C., Gray, A., Hubley, R., Wheeler, T. J., Smit, A. F. A.
- DOI: 10.64898/2026.09.13.751287
- Source URL: <https://doi.org/10.64898/2026.09.13.751287>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.13.751287>

Abstract: Drosophila melanogaster repeatome, comprising roughly 20% of the genome, is characterized by a large fraction of active TE families counterbalanced by efficient purifying selection. Consequently, many TE families persist at low copy numbers and are frequently population specific. Hybridization and horizontal transfer provide new families, which can spread through natural populations within decades. These dynamics, together with years of independent curation efforts, left the D. melanogaster mobilome distributed across several, partly redundant, libraries. Prompted by submissions of population-specific data, we undertook a complete overhaul of the D. melanogaster TE libraries, begun in Dfam 3.9 and finalized in Dfam 4.0. We cross-referenced the new submissions against Repbase, FlyBase, the Berkeley Drosophila Genome Project, and our own Dfam 3.8 to resolve redundancy and reconcile names, then rebuilt or newly constructed the seed alignment for most families. Seeds came from four sources: the dm6 reference itself, which supported the majority of models; insertions >100 bp from 13 samples of a recently published D. melanogaster pangenome; the genomes of other members of the D. melanogaster subgroup, which supplied copies for older families too degraded in dm6 alone; and diverged matches recovered during iterative curation, which resolved into subfamilies and previously undescribed relatives. Rebuilding the seeds corrected consensus sequences that were truncated, chimeric, or skewed by co-duplicated fragments, and lowered the mean Kimura divergence of annotated copies from their consensus. The revision also added families with no prior Dfam representation, including the DNA P-element (absent from dm6) and a collection of novel families that have recently invaded natural populations. Following manual curation, the new library contains 398 models, up from 226 in Dfam 3.8, and annotates an additional 1.5% of the dm6 reference.

## Why Life is Hot
- Source: bioRxiv (preprints)
- Date: 2026-09-19
- Authors: Schilling, T., Warren, P., Poon, W. C. K.
- DOI: 10.64898/2026.02.13.705721
- Source URL: <https://doi.org/10.64898/2026.02.13.705721>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.02.13.705721>

Abstract: The process of evolution by natural selection leads to phenotypes of increasing fitness. For cellular chemical reaction networks, this means optimising a variety of fitness functions such as robustness, precision, or sensitivity to external stimuli. We argue that these diverse goals can be achieved by a versatile, generic mechanism: coupling chemical reaction networks to reservoirs that are strongly out of equilibrium. Using theory and numerics we show that this mechanism of optimisation comes at the price of significant heat dissipation. We compute the heat flux caused by kinetic proofreading in Escherichia coli and show that it constitutes a significant fraction of the total heat flux experimentally measured in this model organism. We then demonstrate that the degree of optimality achievable saturates, and that Nature appears to operate near saturation despite high energetic costs. We argue that \`life is hot' largely because of the need for a versatile mechanism to optimise a variety of fitness functions.

## BrainWideBench: Benchmarking large-scale pretraining and across-animal transfer in multi-region neural recordings
- Source: arXiv (preprints)
- Date: 2026-09-18T17:54:32Z
- Categories: Computational neuroscience, Tools & resources
- Authors: Alexandre Andre, Shivashriganesh P. Mahato, Vinam Arora, Keshav Balaji, Divyansha Lachi, Nanda H. Krishna, Jingyun Xiao, Yizi Zhang, Ximeng Mao, Wenrui Ma, Han Yu, International Brain Laboratory, Daniel Birman, Niccolò Bonacchi, Gaelle A. Chapuis, Joana A. Catarino, Felicia Davatolhagh, Mayo Faulkner, Laura Freitas-Silva, Fei Hu, Julia M. Huntenburg, Anup Khanal, Inês Laranjeira, Petrina Lau, Guido T. Meijer, Nathaniel J. Miska, Jean-Paul Noel, Alejandro Pan-Vazquez, Georg Raiser, Cyrille Rossant, Karolina Z. Socha, Anne E. Urai, Miles J. Wells, Steven J. West, Olivier Winter, Blake Richards, Guillaume Lajoie, Cole Hurwitz, Mehdi Azabou, Matthew R. Whiteway, Liam Paninski, Eva L. Dyer
- External ID: 2609.22064v1
- Source URL: <https://arxiv.org/abs/2609.22064v1>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Farxiv.org%2Fabs%2F2609.22064v1>
- PDF: <https://arxiv.org/pdf/2609.22064v1>

Abstract: Advances in large-scale neural recording have made it possible to collect data across many animals and distributed brain regions, raising the question of whether this scale can be exploited to learn general-purpose neural representations transferable across diverse downstream tasks. Yet, progress toward this goal has been limited by fragmented evaluation protocols and a narrow focus on individual task domains. Here, we present BrainWideBench, a benchmark for evaluating across-animal transfer on multi-region neural recordings, built on the International Brain Laboratory Brainwide Map dataset of neural and behavioral recordings spanning 276 brain regions from 139 mice performing a sensory-guided decision-making task. The benchmark is organized around three complementary task suites that evaluate whether learned representations support downstream decoding of behavior, can predict masked or future neural activity, and can recover biologically meaningful anatomical organization. With this benchmark, we systematically evaluate pretraining methods across transfer settings, including finetuning on downstream objectives and zero-shot generalization to unseen animals. Our results confirm pretraining improves performance over matched single-session baselines, but we show current methods exhibit heterogeneity in transfer capabilities: gains depend strongly on the alignment between pretraining objectives and downstream tasks. No single approach performs uniformly well across all three suites, and most methods are designed to only address a subset of them. Together, these findings suggest that learning representations that jointly generalize across behavior, dynamics, and anatomy remains an open challenge. By providing a unified and reproducible evaluation suite, BrainWideBench establishes a framework for measuring progress toward general-purpose models of the mouse brain.

## Learning Cardiac Features: ECG Biometrics Across Time and~Exercise
- Source: arXiv (preprints)
- Date: 2026-09-18T16:22:31Z
- Authors: Luca Thiebaud, Paul Chauchat, Mustapha Ouladsine, Stéphane Delliaux
- External ID: 2609.21962v1
- Source URL: <https://arxiv.org/abs/2609.21962v1>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Farxiv.org%2Fabs%2F2609.21962v1>
- PDF: <https://arxiv.org/pdf/2609.21962v1>

Abstract: Electrocardiograms (ECGs) carry subject-specific patterns enabling reliable individual discrimination, forming the basis of ECG biometrics. Beyond authentication, this paradigm holds significant potential to secure sensitive cardiac data and to serve as a pretext task in self-supervised learning. Yet, most studies remain confined to singlesession, resting data, leaving robustness to temporal and physiological variations largely untested. We address this gap by evaluating ECG biometrics under realistic conditions involving exercise-induced stress and cross-session variability. A Siamese ResNet with late multi-lead fusion strategy is trained on a large ECG dataset extracted from cardiopulmonary exercise tests and evaluated with a exercise-and time-aware protocol, as well as on public benchmarks. This first extensive assessment of ECG biometrics under combined physiological and temporal variability achieves an intra-session rest-to-peak EER of 1.7% and stateof-the-art 3.9% on the CYBHi dataset. Findings support the presence of an intrinsic cardiac signature resilient to physiological and temporal drift.

## Service Notice: 18 Sep 2026 – Ensembl archives failing to load or loading slowly
- Source: Ensembl (feeds)
- Date: 2026-09-18T15:33:22+00:00
- Categories: Blog
- Source URL: <https://www.ensembl.info/2026/09/18/service-notice-18-sep-2026-ensembl-archives-failing-to-load-or-loading-slowly/?utm_source=rss&utm_medium=rss&utm_campaign=service-notice-18-sep-2026-ensembl-archives-failing-to-load-or-loading-slowly>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fwww.ensembl.info%2F2026%2F09%2F18%2Fservice-notice-18-sep-2026-ensembl-archives-failing-to-load-or-loading-slowly%2F%3Futm_source%3Drss%26utm_medium%3Drss%26utm_campaign%3Dservice-notice-18-sep-2026-ensembl-archives-failing-to-load-or-loading-slowly>
- Abstract: not stored for this record.

## Updates to Phyloxml files and schema source
- Source: Ensembl (feeds)
- Date: 2026-09-18T15:19:41+00:00
- Categories: Blog
- Source URL: <https://www.ensembl.info/2026/09/18/updates-to-phyloxml-files-and-schema-source/?utm_source=rss&utm_medium=rss&utm_campaign=updates-to-phyloxml-files-and-schema-source>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fwww.ensembl.info%2F2026%2F09%2F18%2Fupdates-to-phyloxml-files-and-schema-source%2F%3Futm_source%3Drss%26utm_medium%3Drss%26utm_campaign%3Dupdates-to-phyloxml-files-and-schema-source>
- Abstract: not stored for this record.

## Catena: A Comprehensive Software Suite for Large-Scale Connectomics
- Source: arXiv (preprints)
- Date: 2026-09-18T15:16:37Z
- Categories: Biological imaging, Computational neuroscience, Tools & resources
- Authors: Samia Mohinta, Pedro Gómez-Gálvez, Shi Yan Lee, Daniel Franco-Barranco, Michael Clayton, Stephan Preibisch, Jan Funke, Albert Cardona
- External ID: 2609.21887v1
- Source URL: <https://arxiv.org/abs/2609.21887v1>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Farxiv.org%2Fabs%2F2609.21887v1>
- PDF: <https://arxiv.org/pdf/2609.21887v1>
- Code: <https://github.com/Mohinta2892/catena>

Abstract: The gold standard datasets for mapping connectomes are electron microscopy volumes of densely labeled neural tissue at nanometer resolution. Yet reconstructing and proofreading neuronal arbors and annotating all synapses requires pipelining multiple software tools that are often fragmented, inconsistently maintained, or proprietary, hindering reproducibility and automation. Here, we introduce Catena, an open-source, comprehensive, developer-centric software suite for connectomics that integrates modules for 3D neuron and organelle segmentation, synapse detection, microtubule tracking, and neurotransmitter inference. Catena organizes its modules in composable, chunk-wise processing pipelines in a completely documented, extensible, and adaptable design. We further reduce compute and ground-truth data requirements with pretrained machine learning models, facilitating fine-tuning. Catena ships fully containerized modules that encapsulate evolving dependencies for consistent execution across workstations and clusters. By consolidating open components, shareable models, and containerized runtimes, Catena delivers a reproducible and scalable approach to mapping cellular connectomes from electron microscopy volumes. Code and documentation: https://github.com/Mohinta2892/catena.git

## MIST: Multimodal Survival Prediction with Genomic-Guided Histology Attention
- Source: arXiv (preprints)
- Date: 2026-09-18T14:21:37Z
- Categories: Biological imaging
- Authors: Muhammet Sami Yavuz, Sabri Mustafa Kahya, Richard R. Chen, Jana Lipkova, Benedikt Wiestler
- External ID: 2609.21811v1
- Source URL: <https://arxiv.org/abs/2609.21811v1>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Farxiv.org%2Fabs%2F2609.21811v1>
- PDF: <https://arxiv.org/pdf/2609.21811v1>
- Code: <https://github.com/samiyavuuz/MIST>

Abstract: Multimodal survival models can combine complementary prognostic information from whole-slide images and genomic profiles, but effective fusion remains challenging amid external cohort shift and computational complexity. To address these challenges, we propose MIST, multimodal survival prediction with genomic-guided histology attention. MIST represents genomic features as tokens and allows them to query compact foundation-model-derived histology context tokens before survival prediction. This design enriches molecular information with histology context rather than merging separately encoded modalities only at the final stage. Training combines discrete-time survival prediction with genomic feature masking, WSI dropout, and paired WSI-genomics contrastive alignment. Across four external evaluations in colon, renal, lung, and glioblastoma cohorts, MIST improves external C-index over standard fusion baselines in the primary comparisons. These results support genomic-guided histology attention as a compact and effective strategy for multimodal oncology outcome prediction. Our code is available at https://github.com/samiyavuuz/MIST .

## Exact Counts of Binary Phylogenetic Networks with Four Reticulations
- Source: arXiv (preprints)
- Date: 2026-09-18T13:41:17Z
- Categories: Evolution & metagenomics, Mathematical biology & statistics
- Authors: Hao Yu, Louxin Zhang
- External ID: 2609.21772v1
- Source URL: <https://arxiv.org/abs/2609.21772v1>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Farxiv.org%2Fabs%2F2609.21772v1>
- PDF: <https://arxiv.org/pdf/2609.21772v1>

Abstract: Phylogenetic networks provide a flexible framework for representing reticulate evolutionary processes, such as hybridization, introgression, recombination, and horizontal gene transfer. However, their combinatorial complexity makes even basic enumeration problems difficult. Building on our previous work for networks with up to three reticulations, we derive an explicit closed-form formula for the number of unrestricted rooted binary phylogenetic networks with four reticulations on \\(n\\) labeled taxa. Our approach is based on tree-component graphs. We classify the 79 possible component graphs corresponding to networks with four reticulations into ten groups. We then enumerate the networks associated with each group by combining known counts of one-component networks, forests, and networks with fewer reticulations. Summing these contributions yields the desired formula. This result extends the exact enumeration of unrestricted binary phylogenetic networks to four reticulations and further demonstrates the effectiveness of component graphs for systematically organizing and counting increasingly complex network classes.

## Signature of mechanically induced cell extrusions in cell size distribution
- Source: arXiv (preprints)
- Date: 2026-09-18T12:37:04Z
- Authors: Marko Popović, Ali Tahaei
- External ID: 2609.21702v1
- Source URL: <https://arxiv.org/abs/2609.21702v1>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Farxiv.org%2Fabs%2F2609.21702v1>
- PDF: <https://arxiv.org/pdf/2609.21702v1>

Abstract: How a growing tissue organizes its own homeostatic state is a central question in the physics of living matter. We show that when a growing epithelial sheet counteracts increasing cell density by mechanically squeezing cells out of its plane, a homeostatic in-plane pressure emerges as a generalization of a yield stress. We find that in the quasistatic growth limit the homeostatic state is marginally stable, with a pseudogap in the distribution of local distances to the extrusion threshold pressure. Because such mechanically induced extrusions arise from an instability of individual cells, the pseudogap is imprinted in the distribution of cell areas. This provides an image-based way to test for presence of mechanically induced extrusions and we identify this signature in the developing wing epithelium of \\textit\{D.~melanogaster\}. We expect the same principles to apply to confined three-dimensional tissues.

## Best Matches in Phylogenetic Networks
- Source: arXiv (preprints)
- Date: 2026-09-18T12:35:20Z
- Categories: Evolution & metagenomics, Mathematical biology & statistics
- Authors: Patricia A. Ebert, Peter F. Stadler, Marc Hellmuth
- External ID: 2609.21700v1
- Source URL: <https://arxiv.org/abs/2609.21700v1>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Farxiv.org%2Fabs%2F2609.21700v1>
- PDF: <https://arxiv.org/pdf/2609.21700v1>

Abstract: Best match graphs (BMGs) were introduced in mathematical phylogenetics to describe the concept of closest relatives for related genes (leaves of rooted tree) in different organisms (defining leaf colors). We generalize this concept here to leaf-colored rooted networks, where least common ancestors are in general neither unique nor comparable. We characterize BMGs of rooted networks as those vertex-colored digraphs that are properly colored and satisfy an easy-to-check condition that we call the sicor-in-hub property. BMGs can be recognized in linear time and an explaining network can be constructed in quadratic time. Analogous results are obtained for reciprocal best match graphs (RBMGs), where an edge $\\\{x,y\\\}$ corresponds to pairs of vertices with different color that are mutually closest relatives.

## Extending Decoupled Attention to Dense Prediction and Masked Training for Multi-Channel Images
- Source: arXiv (preprints)
- Date: 2026-09-18T11:11:57Z
- Categories: Biological imaging
- Authors: Umar Marikkar, Sameed Husain, Muhammad Awais, Sara Atito
- External ID: 2609.21629v1
- Source URL: <https://arxiv.org/abs/2609.21629v1>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Farxiv.org%2Fabs%2F2609.21629v1>
- PDF: <https://arxiv.org/pdf/2609.21629v1>

Abstract: Multi-Channel imaging (MCI) data differs fundamentally from natural images, as each channel records a semantically distinct signal rather than a colour band. To adapt vision encoders to MCI data, Multi-Channel Vision Transformers (MC-ViTs) tokenize each channel independently and concatenate the resulting tokens into one sequence, and the channel count is no longer fixed by the architecture. Self-attention is then computed across all channel-patch tokens with no restriction on which channels attend to which, which dilutes the features of individual channels. The Decoupled Vision Transformer (DC-ViT) regulates this by separating updates computed within a channel from updates computed across channels, and by forming a representation per channel before the channels are combined. Its formulation, however, pairs tokens by spatial position, and thus requires the same visible tokens in every channel. Correspondence under independent per-channel masking is recovered by solving a linear assignment between the retained patches of each channel, which allows decoupled attention to be combined with current masked multi-channel training in its standard configuration rather than a restricted one. Across three classification and three segmentation benchmarks spanning fluorescence microscopy, imaging mass cytometry and satellite imaging, including dense prediction at high channel counts, the resulting formulation outperforms the strongest MC-ViT baseline.

## High Reconstruction Quality and Restart Repeatability Do Not Guarantee Recovery of Ground-Truth Muscle Synergies
- Source: arXiv (preprints)
- Date: 2026-09-18T07:06:29Z
- Authors: Ye Ma, Dongwei Liu, Meijin Hou, Chenyi Guo
- External ID: 2609.21394v1
- Source URL: <https://arxiv.org/abs/2609.21394v1>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Farxiv.org%2Fabs%2F2609.21394v1>
- PDF: <https://arxiv.org/pdf/2609.21394v1>

Abstract: High reconstruction quality and agreement across repeated fits do not necessarily establish recovery of muscle synergies. We tested whether a variance-accounted-for (VAF)/elbow rule recovers the generating synergy count and spatial vectors, whether high restart repeatability indicates recovery, and how five design factors affect recovery. Non-negative matrix factorisation was applied to 4,320 synthetic 16-muscle datasets varying generating rank, noise, trial count, spatial similarity and activation overlap. Combined recovery required the correct rank and cosine similarity of at least 0.80 for every matched spatial vector. Factor effects and two-factor interactions were assessed using exploratory heteroscedastic Wald tests with Benjamini-Hochberg adjustment. Rank selection was exact in 17.6% of datasets, too low in 54.9% and too high in 27.5%; combined recovery was 13.9%. Among fits with VAF at least 0.90, only 11.3% achieved combined recovery. Among 3,762 datasets with spatial repeatability at least 0.95, 19.6% had the correct rank and 15.7% achieved combined recovery. All five factors were associated with recovery (adjusted p < 0.001). Recovery declined from 26.2% to 1.7% with increasing spatial similarity and from 26.2% to 2.2% with increasing activation overlap. It was lower at ranks 7-9 than at 3-5, increased from 11.0% with 3 trials to 15.8% with 80 trials, and varied non-monotonically with noise. Five noiseless signals synthesised from measured-sEMG reference factors also showed under-selection despite VAF above 0.918. Under this selector, high reconstruction quality and restart agreement were insufficient indicators of correct rank and spatial recovery. Muscle-synergy interpretation should account for rank sensitivity and the separability of spatial and activation patterns.

## Robust Dual-Regularized Variable Selection under Outlier Contamination
- Source: arXiv (preprints)
- Date: 2026-09-18T05:56:33Z
- Categories: Genomics & sequence analysis
- Authors: Abdul-Nasah Soale, Adewale F. Lukman, Essoham Ali
- External ID: 2609.21342v1
- Keywords: genomic
- Source URL: <https://arxiv.org/abs/2609.21342v1>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Farxiv.org%2Fabs%2F2609.21342v1>
- PDF: <https://arxiv.org/pdf/2609.21342v1>

Abstract: Real data often contain unusual observations that can exert disproportionate effects on variable selection, especially in complex predictor settings. We propose a two-stage \{\\it sparse median outer product of gradients (smOPG)\} method for variable selection in single index models with outlier contamination. We first estimate sparse local gradients via \\(\\ell\_1\\)-penalized local median regression and then recover the active predictor set from a rank-one sparse approximation of the resulting gradient matrix using regularized singular value decomposition. The combination of median regression and local weighting provides robustness to both response outliers and leverage points. Extensive simulations across varying dimensions and contamination mechanisms demonstrate the favorable variable selection performance of smOPG relative to existing methods. Applications to air pollution and genomic data demonstrate practical utility, while theory establishes active-set recovery without requiring selection consistency of individual local regressions.

## AI Can Help Find New Uses for Old Drugs; Advancing Them to Patients Is the Hard Part
- Source: Bio-IT World (feeds)
- Date: 2026-09-18T05:01:13+00:00
- Categories: Blog
- Source URL: <https://www.bio-itworld.com/news/2026/09/18/ai-can-help-find-new-uses-for-old-drugs--advancing-them-to-patients-is-the-hard-part>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fwww.bio-itworld.com%2Fnews%2F2026%2F09%2F18%2Fai-can-help-find-new-uses-for-old-drugs--advancing-them-to-patients-is-the-hard-part>
- Abstract: not stored for this record.

## MIRCID: Inferred Hub-miRNAs Drive Cross-Task Improvements in Drug Mechanistic Modeling
- Source: arXiv (preprints)
- Date: 2026-09-18T03:45:42Z
- Categories: Genomics & sequence analysis, Tools & resources
- Authors: Xin Cao, Yigang Chen, Jiatong Xu, Ziyue Zhang, Xiang Cheng, Shenyu Wang, Yangyi Zhang, Xiaoxuan Cai, Shidong Cui, Zihao Zhu, Xiang Ji, Hsi-Yuan Huang, Yang-Chi-Dung Lin, Hsien-Da Huang
- External ID: 2609.21280v1
- Source URL: <https://arxiv.org/abs/2609.21280v1>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Farxiv.org%2Fabs%2F2609.21280v1>
- PDF: <https://arxiv.org/pdf/2609.21280v1>

Abstract: Drug mechanism-of-action (MoA) modeling commonly relies on perturbational transcriptomes, but matched microRNA (miRNA) measurements are often unavailable. Inferred regulatory features offer a scalable way to reuse these data. Here, we present MIRCID, a framework comparing gene expression with inferred transcription factor (TF) activity and miRNA expression across pathway classification and similarity-based MoA retrieval. HubmiRNet infers 414 pan-cancer hub miRNAs (HubmiRs) from 977 L1000 landmark genes, achieving a Pearson correlation coefficient of 87.72\\%; its 1,298-output variant also outperformed SiCmiR on the full-miRNA task (71.21\\% versus 67.30\\%). In the evaluated comparisons, miRNA augmentation provided more consistent gains than TF activity. Generic embedding controls showed model-dependent utility, while complementarity analyses identified a distinct, partially linearly recoverable representation that retained gene-derived structure. Illustrative rescue cases linked improved classification to biologically plausible miRNA patterns in samples with weak transcriptional signatures. These findings support inferred HubmiRs as a biologically informed recoding of transcriptomic data for perturbational drug modeling, while leaving recovery of measured perturbational miRNA responses to further validation.

## Identifying Neural State Changes due to Gain versus Off-Manifold Displacement
- Source: arXiv (preprints)
- Date: 2026-09-18T03:36:30Z
- Categories: Computational neuroscience
- Authors: Sam McKenzie
- External ID: 2609.21272v1
- Source URL: <https://arxiv.org/abs/2609.21272v1>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Farxiv.org%2Fabs%2F2609.21272v1>
- PDF: <https://arxiv.org/pdf/2609.21272v1>

Abstract: Memory segmentation is thought to arise from rapid decorrelation in neural activity, often quantified by Euclidean distance or cosine angle. Although these metrics detect a transition, they do not reveal how the new state relates to the repertoire represented by the neural manifold. This matters because neuromodulators that drive state transitions also alter excitability, and learning may repurpose existing representations or create new ones. Here, I introduce a geometric decomposition that separates changes attributable to gain modulation of a nearby manifold state from movement within the manifold and genuine off-manifold displacement. The approach uses the radial axis of neural population activity to partition the normal space of a local manifold region. A central challenge is identifiability: given only a static reference manifold and a single test state, neither the state from which a perturbation began nor its gain magnitude and mechanistic decomposition can generally be recovered uniquely. I therefore formulate identifiability as a cascade of geometric gates specifying when each component can be interpreted. The gates distinguish structural failures, including the absence of a local chart or incorrect intrinsic dimensionality, from estimation error and systematic bias caused by reference sampling, tangent-frame error, gain-axis misalignment, anchor displacement, and poor ratio conditioning. Simulations show that neighborhood size, curvature, sampling density, ambient dimension, and noise act through a small set of geometric quantities. The framework specifies when assignments to gain or novelty are identifiable, how they become biased, and which diagnostics reveal the relevant failure regime. By quantifying the nature rather than only the magnitude of neural state change, it provides a clear, readily interpretable framework for evaluating mechanisms of neural state transitions.

## Reliability-Centered Evaluation of Sparse Longitudinal CT Lesion-Size Forecasting with Conformal Interval Calibration and Gompertz-Inspired Regularization
- Source: arXiv (preprints)
- Date: 2026-09-18T01:25:48Z
- Authors: Lingfei Kong
- External ID: 2609.21197v1
- Source URL: <https://arxiv.org/abs/2609.21197v1>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Farxiv.org%2Fabs%2F2609.21197v1>
- PDF: <https://arxiv.org/pdf/2609.21197v1>

Abstract: Sparse longitudinal CT follow-up limits lesion-size forecasting when only a few prior observations are available. We constructed a five-visit DLT-derived same-lesion trajectory benchmark from DeepLesion and Deep Lesion Tracker (DLT), yielding 205 trajectories from 129 patients. We compared an exploratory conventional sparse-to-final analysis with a primary fixed visit-index horizon design predicting the common log change from T3 to T4 while progressively adding earlier observations, evaluating predictive accuracy, uncertainty reliability, post-hoc conformal interval calibration, subgroup performance, and Gompertz-inspired trajectory regularization. The evaluated methods showed partially overlapping point-prediction accuracy but distinct uncertainty behavior. Mean held-out RMSE across ten training seeds was 0.4726, 0.4305, 0.4499, and 0.4513 for m = 1, 2, 3, 4, indicating the lowest mean RMSE at m = 2; additional history did not improve RMSE. At m = 4, raw Cohort-Level Feature GP coverage was near the 95% nominal level, whereas MC Dropout, Deep Ensemble, and residual-scale intervals were conservative. Patient-level conformal calibration generally produced near-nominal or conservative coverage at the cost of wider intervals. Patient-grouped development cross-validation selected lambda\* = 0 for the Gompertz-inspired term. A global population reference frequently opposed lesion-level change directions, and prediction difficulty varied across anatomical subgroups. Overall, additional historical observations provided limited predictive benefit once the prediction horizon was controlled, while predictive accuracy, uncertainty reliability, and trajectory consistency did not necessarily improve together, and should be evaluated jointly in sparse longitudinal imaging.

## SpecOpt: Contact-Diff Reasoning for Agentic Molecule Optimization Toward Binding Specificity
- Source: arXiv (preprints)
- Date: 2026-09-18T00:16:35Z
- Categories: Proteins & structural biology
- Authors: Thao Nguyen, Heng Ji
- External ID: 2609.21165v1
- Source URL: <https://arxiv.org/abs/2609.21165v1>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Farxiv.org%2Fabs%2F2609.21165v1>
- PDF: <https://arxiv.org/pdf/2609.21165v1>

Abstract: Off-target protein binding is a major source of adverse effects for small-molecule drugs, yet most structure-based molecular design methods focus on generating selective compounds de novo rather than improving the selectivity of existing, well- characterized drugs. We introduce specificity optimization (SpecOpt), a molecular design task that seeks constrained structural modifications to an existing compound that increase its binding preference for an intended target over known off-targets while preserving its structural identity and drug-like properties. To enable systematic evaluation, we construct a ChEMBL-derived benchmark from compound-target interaction data, identifying intended targets through curated drug-mechanism annotations and off- targets through measured activities. We then develop an agentic framework that docks each compound against its intended target and off-targets, compares the resulting poses through residue-aware atom-protein contacts, and provides these differential interactions to a large language model to propose targeted structural modifications. Candidates are retained only if they satisfy molecular similarity, ADMET, and target-off-target docking selectivity criteria. On 915 compounds, the agent improves the target- off-target binding gap for 84.8% of compounds, shifting the mean gap from -0.72 to +0.47 kcal/mol while maintaining a mean Tanimoto similarity of 0.72 to the starting compounds. Ablation studies identify residue-specific contact information as the critical optimization signal: replacing residue identities with binary contact indicators eliminates improvement on all 29 ablation compounds. These results establish SpecOpt as a distinct molecular design problem and demonstrate residue-aware differential interactions as an effective signal for improving the specificity of existing compounds.

## A Dataset of Temporally Consistent Instance Annotations for 4D Plant Phenotyping
- Source: Scientific Data (journals)
- Date: 2026-09-18T00:00:00+00:00
- Authors: Jonas Bömer, Elias Marks, Facundo Ramón Ispizua Yamati, Cyrill Stachniss, Stefan Paulus, Anne-Katrin Mahlein
- Journal: Scientific Data
- DOI: 10.1038/s41597-026-08326-5
- Source URL: <https://doi.org/10.1038/s41597-026-08326-5>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1038%2Fs41597-026-08326-5>

Abstract: Spatio-temporal 4D plant phenotyping requires high-quality datasets with precise and temporally consistent annotations. However, such datasets are currently limited due to the substantial effort required for data acquisition and manual annotation. To address this limitation, we present Sugar4D, a publicly available 4D plant phenotyping dataset of sugar beet acquired using a terrestrial LiDAR scanner. The dataset comprises 768 point clouds from 48 individual plants representing twelve genotypes, recorded semiweekly across 16 consecutive time points during the growing season. All point clouds are provided with temporally consistent, pointwise instance annotations at the organ level, enabling the tracking of individual leaves over time. Sugar4D includes 6778 annotated leaves corresponding to 675 unique leaf instances. In addition to the annotated point cloud data, we provide 58 plant- and five leaf-related morphological parameters for each plant and leaf at each time point, validated using a 3D-printed plant reference model and invasive manual reference measurements. Sugar4D supports the development and evaluation of methods for plant instance segmentation, temporal registration, organ tracking, and spatio-temporal morphological analysis.

## A diffusion model of viral evolution predicts mutation fitness and evolutionary trajectories
- Source: bioRxiv (preprints)
- Date: 2026-09-18
- Categories: Genomics & sequence analysis, Proteins & structural biology
- Authors: Wu, J., Ding, X., Wu, A.
- DOI: 10.64898/2026.09.16.752245
- Source URL: <https://doi.org/10.64898/2026.09.16.752245>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.16.752245>

Abstract: Viral evolution arises from random mutations and natural selection, yet computational approaches rarely model these two forces in a unified way. We present Viral Evolution Simulator (VES), a diffusion model-based framework that mirrors this duality by design: forward noise injection simulates stochastic mutation, and reverse denoising recapitulates selective filtering. Trained solely on viral protein sequences, VES predicts mutational fitness without functional data, measuring fitness as the reconstruction difficulty of a mutated sequence relative to its wild-type counterpart. Across immune escape, receptor binding, and deep mutational scanning datasets, VES outperforms state-of-the-art generative models, achieving a 31.78% error reduction over the best baseline in immune escape mutation fitting evaluation. When trained on sequences collected before June 2024 and evaluated against H1N1 strains that later emerged, VES assigned high scores to 16 of 20 mutations that subsequently showed the sharpest frequency shifts. Extending to avian influenza H5, the framework reveals a dynamic interplay between antigenic escape and human-type receptor binding. Both functions dropped sharply in 2021, followed by a sustained rise in receptor affinity that could connect to recent epidemiological trends. VES offers a generalizable, sequence-only foundation for tracing evolutionary trajectories and prioritizing mutations for surveillance and experimental validation, pointing toward where functional efforts might matter most.

## A higher-order equivalence of Lotka-Volterra and replicator dynamics reveals tight connections between ecology and evolution
- Source: bioRxiv (preprints)
- Date: 2026-09-18
- Categories: Evolution & metagenomics
- Authors: Gokhale, C. S., Traulsen, A.
- DOI: 10.1101/2025.03.28.645916
- Source URL: <https://doi.org/10.1101/2025.03.28.645916>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1101%2F2025.03.28.645916>

Abstract: The Lotka-Volterra equations are foundational in ecology, modelling logistic growth in isolated populations and complex dynamics in interacting species. Hofbauer and Sigmund established that these equations are equivalent to the replicator dynamics of evolutionary games, so dynamic patterns in ecology are mirrored in evolutionary games and vice versa. Both fields have since turned to non-linearities and higher-order interactions, where it is unclear whether this equivalence still holds. Here, we demonstrate the general equivalence and illustrate it in classical non-linear models from theoretical ecology. Non-linearities in either field leave this foundational connection intact. Yet directly modelling ecological dynamics with evolutionary games, or the reverse, risks misinterpretation, since species in Lotka-Volterra dynamics cannot be equated with strategies in evolutionary games. Our study reveals the tight interplay between ecology and evolutionary game theory, and the robustness of their mathematical connection even in complex scenarios, alongside its caveats.

## A Hybrid Residual-Swin Transformer Design with Attention for Prostate Cancer Segmentation
- Source: bioRxiv (preprints)
- Date: 2026-09-18
- Categories: Biological imaging
- Authors: Singh, R., Gupta, S., Juneja, S., Gupta, D., Maggu, S., Wang, M., Mallik, S.
- DOI: 10.64898/2026.09.14.751449
- Source URL: <https://doi.org/10.64898/2026.09.14.751449>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.14.751449>

Abstract: Prostate cancer, a leading cause of cancer-related deaths among men globally, necessitates the development of precise diagnostic and treatment strategies. Accurate segmentation of prostate cancer in medical imaging, particularly in MRI scans, is crucial for early diagnosis and clinical decisions. Traditional manual segmentation techniques, while efficient, are labor-intensive and necessitate significant expertise, resulting in an increasing demand for automated alternatives. The RSAUNet is a novel deep learning architecture designed to improve prostate cancer segmentation. It incorporates essential components, including Residual Blocks, Swin Transformer Blocks, and Attention Mechanisms within a U-Net architecture. This markedly enhances the model's capacity to discern complex anatomical features and accurately segment malignant tissues. Using sophisticated deep Learning methods, RSAUNet addresses the complexities of prostate imaging, delivering reliable, consistent segmentation results. The model was evaluated against various cutting-edge techniques on extensive multi-parametric MRI datasets, attaining an impressive Dice coefficient (DC) of 0.998 and a Jaccard Index (IoU) of 0.965. These findings highlight the innovative characteristics of RSAUNet and its capacity to transform prostate cancer diagnosis and treatment strategies. The proposed model surpasses current methods and shows potential for practical clinical applications, providing an efficient, precise, and scalable solution for automated prostate cancer segmentation and fostering optimism for the future of medical imaging and diagnosis.

## A model of the locust visual system under diverse stimuli highlights functionality of various mechanisms and suggests a parsimonious structure
- Source: bioRxiv (preprints)
- Date: 2026-09-18
- Categories: Computational neuroscience
- Authors: Olson, E. G. N., Wiens, T. K., Gray, J. R.
- DOI: 10.64898/2026.09.14.751446
- Source URL: <https://doi.org/10.64898/2026.09.14.751446>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.14.751446>

Abstract: Locusts possess a collision-sensitive pathway in their visual system culminating in a neuron, the lobula giant movement detector (LGMD), which preferentially responds to looming stimuli. The LGMD is also notable for its reduced or absent response to other stimuli, such as translating objects, wide-field motion, and incoherent images of looming objects. While experimental and modelling work have both shed light on the various mechanisms underpinning these properties, broad examinations of how these mechanisms interact with each other and with diverse stimulus types have been limited. To address this, we developed a model incorporating mechanisms from recent literature, and subjected to an array of looming stimuli with varying size-to-speed ratio, polarity (OFF and ON), and coherence; wide-field background motion, translating stimuli, and trajectory changes were also examined. The model showed quantitative and qualitative fidelity to biological data in its replication of a wide variety of stimulus responses, with its preference for incoherent visual stimuli being a notable exception. Based on investigations with various inhibition types removed, it was shown that lateral inhibition predominantly suppressed wide-field motion responses. Global inhibition both normalized growing excitation over the course of looming, and improved recognition of looming stimuli against moving backgrounds. Feedforward inhibition was found to have diverse roles, including shaping peak response time and its variability, and improving coherence selectivity. The success of the model in replicating multiple stimuli also showed the plausibility of several underlying assumptions, including sharing of input between the LGMD and multiple other neuron types.

## A novel framework to modelling regulation of prey populations by a generalist predator
- Source: BMC Biology (journals)
- Date: 2026-09-18T00:00:00+00:00
- Categories: Evolution & metagenomics, Mathematical biology & statistics
- Authors: Andrew Y. Morozov, Boris W. Berkhout, Donald DeAngelis
- Journal: BMC Biology
- DOI: 10.1186/s12915-026-02732-2
- Source URL: <https://doi.org/10.1186/s12915-026-02732-2>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1186%2Fs12915-026-02732-2>

Abstract: Background In ecological communities, a predator often feeds on several food sources. Such a predator is known as a generalist, and its feeding is traditionally modelled by a functional response. However, the conventional concept of the functional response is not applicable to predators for which feeding niches of individual predators are much narrower than that of the entire predator population. Therefore, the predator population effectively consists of cohorts of specialists, each of which has its specific diet. In this case, modelling predator-prey interactions requires an alternative approach. The objectives of this study are twofold. First, we provide an empirical example of a generalist predator consisting of cohorts of specialists to motivate our theoretical study. Second, we develop a mathematical framework to modelling food consumption of a generalist predator consisting of specialised feeders. Results Firstly, we experimentally show that the freshwater predatory snail Anentome helena , feeding on non-predatory snails, has individual feeding niches that are much narrower than that of the entire predator population. Then, we present a new generic framework to model food webs with a generalist predator consisting of individuals with very narrow feeding niches. The proposed modelling approach allows for switching among specialist cohorts, governed by variation in profitability of each foraging strategy. Using a model of trophic interactions, we show that structuring within the predator population promotes coexistence of competing prey species; however, the outcome depends on the initial configuration of specialist cohorts within the predator population. The system exhibits oscillations of prey densities while maintaining an approximately constant predator density; this pattern was not reported in previous predator-prey models. Conclusions We critically revisit the long-standing concept of the functional response of a generalist predator. We argue that if individual foragers within the predator population develop a stable preference for a particular food resource, feeding cannot be described using the traditional functional response modelling approach based on the total predator density. Using our new modelling framework, accounting for individual preferences, we demonstrate the existence of novel patterns of predator-prey dynamics. These patterns predict high biodiversity in ecosystems, long-term ecological transients, and the possibility of a new type of predator-prey cycles.

## A novel virus lineage is abundant in metaviromes from Dehalococcoides-containing mixed cultures
- Source: bioRxiv (preprints)
- Date: 2026-09-18
- Categories: Genomics & sequence analysis, Proteins & structural biology
- Authors: Nesbo, C. L., Morson, N., Molenda, O., Lomheim, L., Lossouarn, J., Maxwell, K. L., Edwards, E. A.
- DOI: 10.64898/2026.09.15.751708
- Keywords: genomes, peptides
- Source URL: <https://doi.org/10.64898/2026.09.15.751708>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.15.751708>

Abstract: Dehalococcoides mccartyi are obligately anaerobic organohalide-respiring bacteria that play important roles in the detoxification of chlorinated pollutants in groundwater and sediments. Despite having small genomes, they host a diverse set of mobile elements. Here we characterize a family of mobile elements, termed integrative and mobilizable element 1 or IME1, comprising 20 from Dehalococcoides and one from Dehalogenimonas alkenigignens. IME1s are 20,930 - 28,058 bp and are found both integrated in the genomes and as circular episomes. Bioinformatic characterization of IME1 encoded proteins revealed a highly conserved structure with 14 hierarchical orthologous groups (HOGs) found in all 21 IME1s. IME1s lack recognizable hallmark proteins of tailed bacterial viruses (or tailed phages) but encode proteins with similarities to those of filamentous bacterial viruses. In particular, one conserved HOG shows sequence similarity to the pI-like ATPase, the only conserved marker protein identified across filamentous bacterial viruses. Additionally, IME1s encode several small proteins with predicted transmembrane domains and signal peptides, another feature used to identify filamentous bacterial viruses. Both features are also found in budding archaeal viruses of various morphotypes. IME1s dominated metaviromes obtained from the Dehalococcoides-containing KB-1 mixed culture and electron micrographs of the corresponding viral fractions revealed abundant filamentous virus-like particles. We therefore propose that the IME1s represent a novel lineage of double stranded budding, likely filamentous, viruses. Database searches suggest IME1s are found in Dehalococcodia and other Chlorofexota but are so far restricted to this phylum.

## A Robust DIA-Based Platform for Large-Scale Plasma Glycoproteomics and Biomarker Discovery
- Source: Journal of Proteome Research (journals)
- Date: 2026-09-18T00:00:00+00:00
- Categories: Proteins & structural biology
- Authors: Chi-Hung Lin, Sayee Sawale, Mark Marispini, Adam Poltorak, Joon-Yong Lee, Natalie Smith, Wan-Fang Chou, Hao Qian, Philip Ma, Bruce Wilcox
- Journal: Journal of Proteome Research
- DOI: 10.1021/acs.jproteome.6c00522
- Keywords: glycoproteomics, glycopeptide
- Source URL: <https://doi.org/10.1021/acs.jproteome.6c00522>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1021%2Facs.jproteome.6c00522>

Abstract: Structural changes in protein glycosylation are recognized as phenotypes in numerous diseases, including cancer. Despite the promise of glycoproteomics for diagnostic biomarker discovery, large-scale studies remain limited by challenges in reproducibility, throughput, and quantitative precision. Here, we present a robust data-independent acquisition (DIA)-based plasma glycoproteomics platform that enables high-throughput and reproducible glycopeptide quantification suitable for large-cohort studies. The workflow integrates automated in-solution digestion and enrichment, optimized DIA-LC-MS acquisition, and a customized bioinformatics pipeline to achieve reproducible glycopeptide identification and quantitation. Across 560 replicates of a pooled human plasma sample processed in seven batches, the platform achieved high reproducibility, with intra-batch coefficients of variation (CVs) below 15% (n = 80 per batch) and an inter-batch CV of 22.9%. The method accurately captured expected fold changes from spiked-in glycoprotein standards. In the setting of a proof-of-concept pilot study to demonstrate the platform’s ability to detect disease-associated glycosylation differences, the platform successfully identified glycoform-specific changes and distinguished between noncancer individuals (n = 39) and those with stage III/IV lung cancer (n = 39). Together, these results establish a high-throughput DIA-based glycoproteomics workflow suitable for large-cohort studies to discover glycopeptide biomarker candidates.

## A tree-based kernel for densities and its applications in clustering DNase-seq profiles
- Source: Biometrics (journals)
- Date: 2026-09-18T00:00:00+00:00
- Categories: Genomics & sequence analysis, Mathematical biology & statistics
- Authors: Yuliang Xu, Kaixuan Luo, Li Ma
- Journal: Biometrics
- DOI: 10.1093/biomtc/ujag154
- Source URL: <https://doi.org/10.1093/biomtc/ujag154>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1093%2Fbiomtc%2Fujag154>

Abstract: Modeling multiple sampling densities within a hierarchical framework enables borrowing of information across samples. These “density random effects” can act as kernels in latent variable models to represent exchangeable subgroups or clusters. A key feature of these kernels is the (functional) covariance they induce, which determines how densities are grouped in mixture models. Our motivating problem is clustering chromatin accessibility profiles from high-throughput DNase-seq experiments to detect transcription factor (TF) binding. TF binding typically produces footprint profiles with spatial patterns, creating long-range dependency across genomic locations. Existing nonparametric hierarchical models impose restrictive covariance assumptions and cannot accommodate such dependencies, often leading to biologically uninformative clusters. We propose a nonparametric density kernel that is flexible enough to capture diverse covariance structures and adapts to various spatial patterns of TF footprints. The kernel specifies dyadic tree splitting probabilities via a multivariate logit-normal model with a sparse precision matrix. Bayesian inference for latent variable models using this kernel is implemented through Gibbs sampling with Pólya–Gamma augmentation. Extensive simulations show that our kernel substantially improves clustering accuracy. We apply the proposed mixture model to DNase-seq data from the Encyclopedia of DNA Elements project, which results in biologically meaningful clusters corresponding to binding events of two common TFs.

## A Unified 3D Generative Model for Synthesizable Structure-Based Drug Design
- Source: bioRxiv (preprints)
- Date: 2026-09-18
- Categories: Proteins & structural biology
- Authors: Igashov, I., Schneuing, A., Dobbelstein, A. W., Morozova, I., Neeser, R. M., Zielinski, K., Abriata, L. A., Petruzzella, A. S., Pavel Iosub, D. R., Gampp, O., Lyubimov, A. Y., Elizarova, E., Ferrara, I., Sousa, P. M. F., Lemos, A. R., Testori, F., Miranda Herrera, P. A., Kanis, L., Schmidt, J., Braza, M. K. E., Amaro, R. E., Thoma, N., Ferraris, D. M., Riek, R., Fraser, J. S., Schwaller, P., Bronstein, M., Correia, B.
- DOI: 10.64898/2026.09.15.751537
- Keywords: peptides
- Source URL: <https://doi.org/10.64898/2026.09.15.751537>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.15.751537>

Abstract: Traditional screening-based drug discovery is inherently limited by the astronomical scale of the chemical space. Generative modelling offers a compelling alternative to the classical search paradigm and enables rational, bottom-up design of novel and target-specific small molecules. However, its impact has been hampered by challenges in synthetic accessibility of the designed compounds and lack of large-scale experimental validation. Here, we introduce LDDM (Large Drug Discovery Model), a generative framework that supports a range of drug discovery tasks, including constrained and unconstrained docking, fragment linking and growing, and de novo design. We further introduce a programmable design algorithm that enables accurate design of synthetically accessible compounds satisfying various fine-grained objectives. We experimentally validated the designed or optimised ligands for five therapeutically relevant protein targets. In all cases, LDDM achieved high success rates, allowing us to identify molecules with confirmed binding affinity while synthesizing only a small number of generated compounds. The best designs were structurally characterised through NMR spectroscopy and X-ray crystallography, demonstrating high prediction accuracy. Overall, LDDM provides a scalable and flexible platform for the rapid and tailored design of small molecules and non-natural peptides for therapeutic applications.

## A universal power law optimizes energy and representation fidelity in visual adaptation
- Source: bioRxiv (preprints)
- Date: 2026-09-18
- Categories: Computational neuroscience
- Authors: Mariani, M., Moosavi, A. S., Ringach, D., Dipoppa, M.
- DOI: 10.1101/2025.03.20.643406
- Source URL: <https://doi.org/10.1101/2025.03.20.643406>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1101%2F2025.03.20.643406>

Abstract: Sensory systems continuously adapt their responses based on the probability of encountering a given stimulus. In the mouse primary visual cortex (V1), the population response magnitude is a power law of the stimulus probability in the environment. For a given stimulus type (e.g., oriented gratings), the power law's exponent is invariant to changes in statistical environments, enabling predictions of population responses to new environments. Here, we aim to provide a normative explanation for the power law behavior. We develop an efficient coding model where neurons adjust their firing rates through optimization of a weighted objective, hypothesizing that the neural population adapts to enhance stimulus detection and discrimination while reducing overall neural activity. We show that a model balancing representational fidelity and energy efficiency matches the power law observed experimentally for a wide range of parameters, while models of adaptation with alternative coding objectives and resource constraints are unable to reproduce this empirical observation. Furthermore, we account for the invariance of the power law's exponent across environmental changes by linking it to the dependence of tuning curve modulation on stimulus probability. Finally, we explain how variations in the exponent with different stimulus types (e.g., natural stimuli) result from changes in the minimal distances between neural representations, in agreement with experimental findings. We conclude that a universal power law of adaptation can be explained as a trade-off between representation fidelity and energy cost.

## A Variational Modeling Framework for Population Genetic Dynamics
- Source: bioRxiv (preprints)
- Date: 2026-09-18
- Categories: Evolution & metagenomics, Mathematical biology & statistics
- Authors: Li, C.
- DOI: 10.64898/2026.09.17.750900
- Source URL: <https://doi.org/10.64898/2026.09.17.750900>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.17.750900>

Abstract: Population genetic dynamics are shaped by multiple evolutionary processes, including mutation, recombination, selection, and genetic drift. With the rapid growth of genomic data, modeling multilocus evolutionary dynamics and the resulting patterns of genetic variation has become increasingly important. Existing approaches often face challenges in jointly describing multiple evolutionary processes and modeling multilocus systems, motivating the development of more flexible and extensible frameworks. Here, we develop a variational framework for population genetic dynamics based on generalized gradient-flow theory, drawing on nonequilibrium thermodynamics. In our model, mutation, recombination, and selection are represented as modular variational components, with different life-cycle stages connected through a gamete-individual two-state system. Mutation and recombination act on the gamete state, selection acts on the individual state, and finite-population genetic drift is represented by a stochastic extension. Our model represents multilocus genetic variation directly in the full haplotype-frequency space, recovers classical mutation, recombination, and selection dynamics in the corresponding limits and provides a natural stochastic extension for finite populations. Numerical experiments and SLiM forward simulations show that our model captures the dynamics of allele frequencies, haplotype frequencies, and linkage disequilibrium in multilocus systems. The variational formulation opens avenues for future extensions to additional evolutionary processes and more complex multilocus systems, as well as for developing differentiable computational methods for gradient-based parameter inference and scalable genomic modeling.

## Accelerated discovery of thermostable vaccines using data-efficient AI
- Source: bioRxiv (preprints)
- Date: 2026-09-18
- Categories: Genomics & sequence analysis, Proteins & structural biology
- Authors: Tian, J., Tran, K. T. M., Pogostin, B. H., Sheridan, O., Mursalova, S., Lee, A. H., Liu, S., Hamkins, J., Antov, D., Power, A. L., Dash, Z. S., Yun, D., Konakovic Lukovic, M., Langer, R., Jaklenec, A.
- DOI: 10.64898/2026.09.17.752370
- Keywords: rna, antibody
- Source URL: <https://doi.org/10.64898/2026.09.17.752370>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.17.752370>

Abstract: The inherent instability of mRNA- lipid nanoparticles (LNPs) necessitates ultra-cold storage, creating significant barriers for global distribution and limiting their broader application in advanced delivery systems. Solid-state, water-free formulations offer a promising solution by enhancing thermostability and enabling integration into emerging delivery modalities such as microneedle (MN) patches. Prior efforts to stabilize mRNA-LNPs have been constrained by narrow formulation scope and low-throughput screening methods. Here, we introduce AGENT (Algorithm-Guided Experimental design for lipid Nanoparticle Thermostabilization), an AI-driven framework that couples high-throughput experimentation with Bayesian optimization to rapidly identify thermostable mRNA-LNP formulations. Manual exploration of the formulation space required months of screening and yielded suboptimal candidates. In contrast, AGENT extracted maximal information from sparse experimental datasets, enabling efficient formulation optimization in only six iterations completed within one month. Using AGENT, we stabilized mRNA vaccines with diverse LNPs, including those in clinical use, into solid state formulations that retained 100% bioactivity after storage at 37 degree C for over two months. The thermostable vaccines induced antigen-specific IgG and germinal center B cell responses that were non-inferior to those elicited by freshly prepared soluble vaccines. The solid-state formulations were further incorporated into dissolvable MN patches and administered to rodents and nonhuman primates, yielding comparable neutralizing antibody titers compared to conventional intramuscular delivery of fresh vaccines. To our knowledge, this study presents the first demonstration of AI-driven design of thermostable RNA vaccines, offering a scalable, cold-chain-free solution for global immunization. By addressing both stability and delivery challenges, AGENT provides a potentially transformative platform for developing accessible next-generation therapeutics.

## Accurate reconstruction of spatial cell-type maps and characterization of domain-specific functions based on a gene-aware heterogeneous network.
- Source: Genome research (journals)
- Date: 2026-09-18
- Categories: Genomics & sequence analysis, Single-cell & spatial
- Authors: Zilin Li, Zhaoyang Huang, Yan Li, Chenguang Zhao, Liang Yu
- Journal: Genome research
- DOI: 10.1101/gr.282246.126
- External ID: 42642330
- Source URL: <https://doi.org/10.1101/gr.282246.126>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1101%2Fgr.282246.126>

Abstract: Spatial transcriptomic (ST) profiles gene expression with spatial context, but most platforms capture multicellular spots containing mixed cell types, making accurate deconvolution essential. Existing reference-based methods using scRNA-seq often ignore spatial dependency and gene-level contribution, yielding fragmented maps and limited insight into domain-specific programs. Here, we propose a gene-aware heterogeneous graph attention network called STGnet for ST deconvolution and functional annotation. Leveraging a hybrid pseudospot generation strategy that captures realistic spatially enriched cell-type patterns, STGnet accurately integrates spatial adjacency, transcriptional similarity, and gene-spot associations within a unified heterogeneous network. Attention weights highlight domain-specific genes for interpretable domain annotation. Importantly, STGnet can characterize spatially ordered functional programs across domains that may be associated with disease progression. These insights may facilitate the discovery of spatial disease mechanisms and improve understanding of pathological tissue organization. Experiments on simulated and real data sets show that STGnet achieves the best overall performance compared with state-of-the-art methods.

## An Artificial Intelligence Model for Longitudinal Assessment of TCR Repertoires in SARS-CoV-2 Vaccine Recipients
- Source: bioRxiv (preprints)
- Date: 2026-09-18
- Categories: Proteins & structural biology
- Authors: Wang, Z., Zhao, Y., Xiao, X., He, B., Sun, Y., Xiong, S., Qin, C., Zhou, Z., Chang, L., Bai, J., Zhao, W., Liang, W., Yao, J.
- DOI: 10.64898/2026.09.16.752071
- Keywords: antibody, epitopes, antibodies
- Source URL: <https://doi.org/10.64898/2026.09.16.752071>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.16.752071>

Abstract: T cells play a crucial role in reducing disease severity during SARS-CoV-2 infection and in shaping long-term immune memory. However, the precise molecular immune responses, particularly involving T-cell receptor (TCR) repertoire changes after full vaccination, and the use of TCR analysis to evaluate vaccine efficacy, remain incompletely understood. In this study, we developed the DeepAir-Cov19 model (AUC=0.94), a large language model tailored to identify SARS-CoV-2-specific TCRs, and observed significant differences in the TCR profiles between antibody-negative and antibody-positive populations before and after vaccination, indicating that the immune status of pre-vaccine recipients can directly assess the efficacy of vaccination. Notably, SARS-CoV-2-specific TCRs expanded to peak levels after the second dose and remained detectable in most subjects up to 10 months post-vaccination. Meanwhile, we also identified three specific V genes, 20 V-J combinations, and 3 epitopes associated with these responses. Finally, by leveraging vaccine-specific TCRs as novel biomarkers, we developed a vaccine efficacy model that predicts antibody levels with a mean AUC of 0.96. These findings underscore the accuracy of the large model in predicting SARS-CoV-2-specific TCRs and reveal a strong correlation between SARS-CoV-2-specific TCR responses and antibody levels. This highlights the complementary and synergistic roles of T cells and antibodies in providing vaccine-mediated protection. Our results offer valuable insights into the longitudinal dynamics of SARS-CoV-2-specific TCRs and illustrate the potential for developing potency evaluation models based on these TCR insights.

## An atlas of transcription factor cooperation reveals how motif readers shape regulatory output
- Source: bioRxiv (preprints)
- Date: 2026-09-18
- Categories: Genomics & sequence analysis
- Authors: Xiong, H., Liu, J., Wang, W.
- DOI: 10.64898/2026.09.14.751590
- Source URL: <https://doi.org/10.64898/2026.09.14.751590>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.14.751590>

Abstract: Regulatory motifs are conventionally associated with named transcription factors (TFs), yet a motif label need not identify the protein that reads the sequence or the regulatory consequence that follows in a given cell. We analyzed 1,552 TF binding datasets in 10 cell types using ARES, a multi-agent system that tests competing mechanisms of TF-motif dependencies in a specific cellular context against multi-omic data. We found that the inferred mechanisms converged on three operating routes: direct sequence recognition, protein-mediated recruitment or exclusion, and regulatory context. Importantly, the predictive motifs of the target TF binding were read by their conventionally "canonical" TFs in only one third of resolved dependencies, and these "canonical" TFs were expressed much less often than the inferred readers. Furthermore, we observed that motif similarity was associated with shared regulatory region type but not shared transcriptional outcome, whereas reader identity was associated with both and the only feature among the examined associated with outcome. In validation case studies where an inferred reader was perturbed, target TF occupancy fell in proportion to reader binding before perturbation, and a natural variant disrupting the predictive motif altered target TF binding at every intermediate step of the inferred mechanism. These observations were further supported by single-cell perturbation, in vitro cooperativity and evolutionary constraint. Together, these results separate motif identity from reader identity and regulatory output, suggesting that a motif acts as an address whose regulatory consequence is shaped in trans by the protein that interprets it.

## An automated high-resolution screening platform identifies regulators of anchor cell invasion in C. elegans
- Source: Science Advances (journals)
- Date: 2026-09-18T00:00:00+00:00
- Authors: Simon Berger, Silvan Spiri, Evelyn Lattmann, Stefanie Engleitner, Mitchell P. Levesque, Andrew deMello, Alex Hajnal
- Journal: Science Advances
- DOI: 10.1126/sciadv.aef6546
- Source URL: <https://doi.org/10.1126/sciadv.aef6546>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1126%2Fsciadv.aef6546>

Abstract: Microfluidic devices are valuable tools for live imaging. However, widespread adoption of microfluidic-based screening methods has been limited by the complexity of the existing techniques. Here, we introduce a user-friendly, high-throughput, and high-resolution automated imaging system for C. elegans . We demonstrate the system’s capabilities in an RNA interference (RNAi) screen, combined with neural network–based phenotypic scoring. We evaluated the effects of RNAi targeting 193 candidate genes on anchor cell (AC) invasion, a model for basement membrane (BM) breaching that shares similarities with tumor cell invasion during cancer metastasis. Over 40,000 animals were imaged at subcellular resolution and scored using a custom neural network classifier with an accuracy of over 92%. The screen identified 41 of 52 genes previously known to control AC invasion, along with 51 additional regulators of invasion. This automated imaging and classification system enables researchers to perform forward mutagenesis, RNAi, and drug screens in C. elegans with much greater speed and higher resolution than previously possible.

## An extended Kalman filter for large-volume path positioning of aquatic animals within acoustic telemetry arrays
- Source: bioRxiv (preprints)
- Date: 2026-09-18
- Authors: Campbell, J. A., Elings, J., Lundberg, P., Mawer, R., Pauwels, I., Hölker, F.
- DOI: 10.64898/2026.09.17.752304
- Source URL: <https://doi.org/10.64898/2026.09.17.752304>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.17.752304>

Abstract: Acoustic telemetry is a core methodology for collecting fine-scale movement data for aquatic animals. When telemetry receivers are set up in closely spaced arrays with overlapping detection ranges, the detection times of a tagged animal can be used to estimate its position and movement paths. In practice, estimating these paths can be challenging. Traditional time-difference-of-arrival methods generally provide positioning accuracy too poor for inferring fine-scale behaviours, while more robust state-space positioning models can be computationally intensive and practically infeasible to run on large datasets. Here, a novel telemetry positioning method is presented where time-of-arrival positioning is implemented as a state-space model within an extended Kalman filter. The resulting model, termed EK-TOA, provides closed-form solutions to track estimation. Simulated datasets of fish movement within a 2D telemetry array are used to verify the models performance and a real case study is provided where EK-TOA is utilized for the long-term tracking of a tagged fish. In comparison to currently available positioning models, EK-TOA provides a fast and accurate solution for tracking fine-scale movement behaviours of aquatic animals over long, continuous periods of time.

## An lncRNA-aware single-cell framework with donor-level validation identifies reproducible MEG3 enrichment in human liver sinusoidal endothelial cells.
- Source: Functional & integrative genomics (journals)
- Date: 2026-09-18
- Categories: Genomics & sequence analysis, Single-cell & spatial, Tools & resources
- Authors: Hidenori Tani
- Journal: Functional & integrative genomics
- DOI: 10.1007/s10142-026-02051-3
- External ID: 42758357
- Source URL: <https://doi.org/10.1007/s10142-026-02051-3>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1007%2Fs10142-026-02051-3>

Abstract: Long non-coding RNAs (lncRNAs) associated with metabolic liver disease are usually identified from bulk tissue, which cannot resolve the hepatic cell types that express them, and default single-cell pipelines discard most lncRNAs at feature selection. We present an lncRNA-aware single-cell analysis framework - retaining all detectable GENCODE v45 lncRNAs during highly variable gene selection - combined with donor-level validation that guards against pseudoreplication. The framework applies to already published data and needs no lncRNA-specific protocol. Applying it to the human Liver Cell Atlas (Gene Expression Omnibus accession GSE192742; 152,559 annotated cells, 16 donors), and using the source publication's own cell-type annotation, we found MEG3 enriched in liver sinusoidal endothelial cells (LSECs) in abundance as well as in detection frequency: donor-level pseudobulk expression was 6.27 counts per 10,000 versus 1.05 in the next-ranked cell type, and the LSEC pseudobulk value exceeded the pooled non-LSEC value in all nine informative donors (paired Wilcoxon signed-rank: two-sided P = 0.0039, one-sided P = 0.0020). The LSEC-enriched detection pattern reproduced in an independent five-donor atlas (all five donors concordant). KCNQ1OT1 was not LSEC-specific. Benchmarking our clustering against the published annotation showed that two marker-scored lineages contained none of their nominal cell type; separately, the candidates returned by a within-lineage pseudotime screen were driven entirely by non-endothelial cells contaminating the LSEC lineage: correlations of |ρ| > 0.31 fell below 0.10 within correctly annotated endothelial cells and reversed sign in one quarter of alternative roots. We therefore report the pseudotime screen as a negative result, and provide the framework, the annotation audit and the two-cohort analysis as a reproducible resource.

## An open field phenomics resource for multimodal maize yield prediction across divergent environments
- Source: bioRxiv (preprints)
- Date: 2026-09-18
- Categories: Tools & resources
- Authors: DeSalvio, A. J., Mohseni, P., Adak, A., Murray, S. C., Arik, M. A., Wong, R. K. W., Jung, J., Lima, D. C., Aviles, A. C., Buckler, E. S., Duffield, N., Edwards, J., Ertl, D., Flint-Garcia, S., Gore, M. A., Hirsch, C. N., Holland, J. B., Kaeppler, S. M., Miller, J., Romay, C., Schnable, J. C., Singh, M. P., Sparks, E. E., Thompson, A., Washburn, J. D., Weldekidan, T., Winans, N. D., de Leon, N.
- DOI: 10.64898/2026.09.17.752474
- Source URL: <https://doi.org/10.64898/2026.09.17.752474>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.17.752474>

Abstract: Temporal drone phenotyping captures crop development, but irregular flight schedules complicate comparisons across environments. We release curated imagery from 356 flights across 19 Genomes to Fields environments containing 1,180 maize (Zea mays L.) hybrids. To evaluate its utility, we integrated functional principal components of vegetation index and weather trajectories with genomic information. Combinined genomic and phenomic kernels improved yield prediction, reaching correlations up to r = 0.501 for held-out hybrids in environments represented in training and 0.408 when environments were also withheld. Accumulated growing degree days offered no consistent predictive advantage over days after planting, and weather contributed modest, task-dependent gains. A transformer neural process learned directly from irregular observations, serving as a novel application of neural process models in agriculture. Mapping vegetation index functional principal components identified recurrent quantitative trait loci on chromosomes 3 and 7. This resource and its reproducible analyses guide the use of temporal spectral data for crop prediction and genetic discovery.

## Application of deep learning to estimate blue and fin whale call density in the southern California Current Ecosystem
- Source: Scientific Reports (journals)
- Date: 2026-09-18T00:00:00+00:00
- Authors: Michaela N. Alksne, Marie A. Roch, Kaitlin E. Frasier, John A. Hildebrand, Shane Andres, Dolapo Adesanya, Lauren M. Baggett, Joshua M. Jones, Ana Širović, Joshua Zingale, Simone Baumann-Pickering
- Journal: Scientific Reports
- DOI: 10.1038/s41598-026-71037-9
- Source URL: <https://doi.org/10.1038/s41598-026-71037-9>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1038%2Fs41598-026-71037-9>

Abstract: Blue ( Balaenoptera musculus ) and fin whales ( Balaenoptera physalus ) are dominant contributors to low-frequency ocean soundscapes, yet reliably extracting their calls from long-term passive acoustic recordings is methodologically challenging. Here, we train a multi-class deep-learning detector to identify five principal blue and fin whale call types (A, B, D, 20 Hz, and 40 Hz) from low-frequency spectrograms using a Faster R-CNN architecture combined with three rounds of iterative human review and hard-negative mining, progressively expanding and rebalancing the training set using California Cooperative Oceanic Fisheries (henceforth, CalCOFI) sonobuoy and moored hydrophone recordings from the southern California Current Ecosystem. The detector was evaluated on four independent test datasets spanning multiple years, seasons, and recording platforms and then deployed on CalCOFI sonobuoy recordings collected quarterly over two decades (2004–2024). The final model achieved consistently high mean precision, recall and F1 scores for most call types (e.g., A: 0.71/0.71/0.71; B: 0.83/0.59/0.63; D: 0.79/0.84/0.80; 20 Hz: 0.87/0.74/0.78), while 40 Hz calls remained challenging (0.42/0.69/0.51), primarily due to confusion with spectrally overlapping humpback whale downsweeps. Detections were post-processed using call-specific characteristics and received-level thresholds and normalized by recording effort and detection area to derive standardized indices of call density $$\\left( \\frac\{\\text \{calls\}\}\{\\text \{h\} \\cdot 1000 \\text \{ km\}^2\}\\right) $$ with uncertainty estimates. Densities were aggregated annually and show call-specific differences between inshore and offshore habitats and interannual variability associated with periods of anomalous oceanographic conditions. Inter-call interval analyses suggested seasonal stability in blue whale song, high variability in blue and fin whale social calls, and seasonal and interannual variability in fin whale song repetition rates. This study is among the first to use deep-learning to estimate baleen whale call density from decades of passive acoustic recordings in a complex soundscape.

## Autoantibody reactome profiling reveals distinct humoral immune signatures induced by simulated microgravity, radiation and combined exposure in mice.
- Source: Molecular & cellular proteomics : MCP (journals)
- Date: 2026-09-18
- Categories: Genomics & sequence analysis, Proteins & structural biology
- Authors: Lining Wu, Kaipeng Zheng, Bomiao Yu, Bingxin Gao, Pancheng Xiao, Saiya Wang, Yanjun Li, Ruimin Liu, Xiaomei Zhang, Mansheng Li, Chun-Ping Cui, Weiming Tian, Xiaobo Yu
- Journal: Molecular & cellular proteomics : MCP
- DOI: 10.1016/j.mcpro.2026.101665
- External ID: 42759701
- Keywords: dna
- Source URL: <https://doi.org/10.1016/j.mcpro.2026.101665>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1016%2Fj.mcpro.2026.101665>

Abstract: Microgravity (μG) and space radiation are major environmental stressors leading to immune dysfunction in astronauts, but the mechanisms governing humoral immunity and effective interventions remain unclear. In this study, we used a high-throughput autoantigen microarray derived from the AAgAtlas 1.0 database to profile humoral immune responses in mice exposed to μG and proton irradiation. We generated 70,490 autoantibody reactivity measurements and observed combined effects of radiation and μG on humoral immunity. Specifically, 38 IgM and 57 IgG autoantibody candidates showed dose-associated changes across 0, 1 and 5 Gy proton irradiation. We further established a global autoantibody landscape for mice under μG, 1 Gy, 5 Gy and μG & 1 Gy. Autoantibody targets altered under μG were enriched for cardiovascular-associated proteins, whereas proton irradiation-associated targets were enriched for DNA repair-related proteins and co-exposure-associated targets were enriched for inflammation-related proteins, revealing distinct humoral immune recognition signatures across exposure conditions. Additionally, a glutamine-free diet (GFD) was associated with changes in autoantibody profiles and shifts of selected bone and hematological parameters toward control levels under μG and μG & 1 Gy conditions. This work provides a resource for investigating humoral immunity under space-related stress and identifies candidate autoantibody signatures and a potential role of dietary modulation that warrant further validation.

## Automated pancreatic segmentation and regional fat quantification suggest tail fat association with type 2 diabetes
- Source: Scientific Reports (journals)
- Date: 2026-09-18T00:00:00+00:00
- Categories: Biological imaging
- Authors: Li YingHao, Wang LiHui, Zhu ZhongQi, Wang SuCheng, Huang ChangDong, Li RenFeng, Cao KaiMing, Hu HaiYang, Jia YiMing, Liang SongTao, Yang Guang, Lu Qing, Wang Hongzhi
- Journal: Scientific Reports
- DOI: 10.1038/s41598-026-57789-4
- Source URL: <https://doi.org/10.1038/s41598-026-57789-4>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1038%2Fs41598-026-57789-4>

Abstract: Accurate pancreatic segmentation and quantitative assessment are crucial for investigating diabetes pathogenesis. Despite advancements in deep learning, regional segmentation remains challenging due to the organ’s anatomical complexity. This study established a two-stage analytical framework: (1) nnU-Net-based whole-pancreas segmentation followed by (2) a novel algorithm for semi-automated head-body-tail partitioning with expert verification, enabling regional quantification of volume and fat content. The segmentation network achieved a Dice coefficient of 0.92. Kruskal-Wallis H tests indicated a significant between-group difference in pancreatic tail fat content across glycemic status groups (p<0.05). Using region-specific fat metrics, we constructed five random forest classifiers showing optimal performance with total fat content (Area Under the Curve, AUC = 0.72) and composite fat indices (AUC = 0.73) for distinguishing healthy controls, prediabetic, and diabetic cohorts.

## Bacpipe: A Python package to make bioacoustic deep learning models accessible
- Source: Methods in Ecology and Evolution (journals)
- Date: 2026-09-18T00:00:00+00:00
- Categories: Tools & resources
- Authors: Vincent S. Kather, Sylvain Haupert, Burooj Ghani, Dan Stowell
- Journal: Methods in Ecology and Evolution
- DOI: 10.1111/2041-210x.70406
- Source URL: <https://doi.org/10.1111/2041-210x.70406>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1111%2F2041-210x.70406>

Abstract: Natural sounds have been recorded for millions of hours over the previous decades using passive acoustic monitoring. Improvements in deep learning models have vastly accelerated the analysis of large portions of this data. While new models advance the state‐of‐the‐art, accessing them using tools to harness their full potential is not always straightforward. Here we present bacpipe , a collection of bioacoustic deep learning models and evaluation pipelines accessible through a graphical and programming interface, designed for both ecologists and computer scientists. Bacpipe streamlines the usage of state‐of‐the‐art models on custom audio datasets, generating acoustic feature vectors (embeddings) and classifier predictions. A modular design allows evaluation and benchmarking of models through interactive visualizations, clustering and probing. We believe that access to new deep learning models is important. By designing bacpipe to target a wide audience, researchers will be enabled to answer new ecological and evolutionary questions in bioacoustics.

## Balanced contractility and adhesion drive polarization in a minimal elastic actomyosin network
- Source: PLOS Computational Biology (journals)
- Date: 2026-09-18T00:00:00+00:00
- Categories: Mathematical biology & statistics
- Authors: Zeno Messi, Franck Raynaud, Nathan W. Goehring, Alexander B. Verkhovsky
- Journal: PLOS Computational Biology
- DOI: 10.1371/journal.pcbi.1014750
- Source URL: <https://doi.org/10.1371/journal.pcbi.1014750>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1371%2Fjournal.pcbi.1014750>

Abstract: Polarization of migrating cells involves chemical and mechanical interactions of signaling networks, cytoskeleton, plasma membrane, and substrate adhesions. Still, it is not fully understood which mechanisms and components are sufficient for symmetry breaking, and if they work independently or together. Here, we use a discrete active network model to investigate if and how an elastic cytoskeletal network is capable of breaking symmetry solely through mechanical interactions. Our minimal model consists of elastic bonds, attractive force dipoles, and force-sensitive anchor points, initially distributed uniformly and subject to simple turnover rules. We find that these features are sufficient to produce different cell behaviors, and, remarkably, to drive symmetry breaking and directed (polarized) motion. Network behavior was primarily determined by the turnover rate of anchor points, which, itself, is a function of the ratio between dipole force and the threshold force required for anchor removal. Directional motion emerged at intermediate turnover rates, at which tension in the network accumulated through several turnover cycles before eventually exceeding the adhesion removal threshold locally at the edge, mirroring our recent experimental findings on the correlation of the traction force with protrusion-retraction transitions in the cell. At high turnover rates, forces were unable to build up to sufficiently high levels, while at low turnover rates, anchors hindered motion. These results demonstrate how directed motion can emerge as an intrinsic property of a simple mechanical network, independently of external cues or complex signaling networks. Given the concordance between this model and recent experimental findings, we suggest that polarization by contraction-adhesion dynamics could be a fundamental emergent behavior of actin-myosin networks.

## Bayesian Additive Regression Trees for Modeling Multiple Exposures With Measurement Error
- Source: Statistics in Medicine (journals)
- Date: 2026-09-18T00:00:00+00:00
- Authors: Madeleine E. St. Ville, Yaeji Lim, Ruijin Lu, Katherine L. Grantz, Zhen Chen
- Journal: Statistics in Medicine
- DOI: 10.1002/sim.70739
- Source URL: <https://doi.org/10.1002/sim.70739>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1002%2Fsim.70739>

Abstract: We develop a Bayesian Additive Regression Trees with Measurement Error (BART‐ME) model for flexibly estimating exposure‐response functions when multiple covariates are measured with classical error. Unlike existing approaches, BART‐ME accommodates nonlinearities, interactions, correlated exposures, and correlated measurement errors, while exploiting replicate measurements to estimate the error variance‐covariance structure. Posterior inference is obtained via a Metropolis‐within‐Gibbs algorithm, yielding estimates of exposure‐response functions, variable importance, and measurement reliability. Simulation studies demonstrate that BART‐ME reduces bias and improves coverage compared with Naïve BART analyses that ignore measurement error, even with sparse replication. In an application to first‐trimester ultrasound data from the NICHD Fetal Growth Studies, BART‐ME identified crown‐rump length as the most reliable predictor of gestational age at delivery and revealed nonlinear patterns attenuated by Naïve methods. These results illustrate the potential of BART‐ME as a general framework for correcting measurement error in complex exposure settings.

## BEAR-GRN: Systematic assessment of single-cell multi-omics-based gene regulatory network inference methods
- Source: Nature Communications (journals)
- Date: 2026-09-18T00:00:00+00:00
- Categories: Single-cell & spatial, Systems & networks
- Authors: Karamveer Karamveer, Eric Moeller, Hannah Valensi, Ewura-Esi Manful, Yasin Uzun
- Journal: Nature Communications
- DOI: 10.1038/s41467-026-77838-w
- Keywords: single cell, multi omics, gene regulatory, inference
- Source URL: <https://doi.org/10.1038/s41467-026-77838-w>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1038%2Fs41467-026-77838-w>
- Abstract: not stored for this record.

## Beyond Steady-State Adaptation: Evaluating the Dynamic Performance of Biological Feedback Control
- Source: bioRxiv (preprints)
- Date: 2026-09-18
- Categories: Systems & networks, Mathematical biology & statistics
- Authors: Tamayo-Luisce, A., Gomez-Schiavon, M.
- DOI: 10.64898/2026.09.17.752162
- Source URL: <https://doi.org/10.64898/2026.09.17.752162>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.17.752162>

Abstract: Organisms rely on feedback control mechanisms to maintain key biological variables within functional ranges despite persistent perturbations. Understanding how effectively these mechanisms maintain homeostasis is fundamental, yet quantitative evaluation of feedback performance remains challenging in nonlinear biological systems. Our previously developed framework, Control Ratio (CoRa), addresses this challenge by isolating the contribution of a feedback interaction through controlled comparison with an otherwise identical system in which that interaction has been removed. However, because CoRa evaluates adaptation solely through steady-state responses, it cannot distinguish controllers that ultimately recover to the same state but follow markedly different transient trajectories, despite the potentially profound physiological consequences of those dynamics. Here, we introduce CoRaDyn, a framework for evaluating adaptation as a dynamic process rather than solely as a steady-state outcome. Building on the comparative strategy introduced in CoRa, CoRaDyn quantifies the cumulative effect of feedback throughout the post-perturbation response, generating a time-dependent characterization of feedback performance. This approach reveals how the contribution of feedback depends not only on system parameters but also on the time horizon over which adaptation is evaluated, allowing distinct physiological objectives -- such as rapid recovery or the generation of transient pulses -- to be systematically compared. We demonstrate CoRaDyn using a gene regulatory circuit implementing proportional-integral-derivative (PID) control. Whereas CoRa predicts identical perfect adaptation for all controllers containing integral feedback, CoRaDyn discriminates their transient performance, quantifies the dynamic contributions of proportional and derivative control, and reveals trade-offs that are invisible to steady-state analyses. By extending feedback evaluation from endpoints to the full adaptation process, CoRaDyn broadens the scope of biological questions that can be addressed using the CoRa framework and provides a general approach for comparing feedback architectures when transient dynamics are central to biological function.

## BFVD v3-UniProt-complete, improved viral protein structure predictions
- Source: bioRxiv (preprints)
- Date: 2026-09-18
- Categories: Tools & resources
- Authors: Kim, R. S., Pimenova, O., Levy Karin, E., Mirdita, M., Steinegger, M.
- DOI: 10.64898/2026.09.16.752260
- Source URL: <https://doi.org/10.64898/2026.09.16.752260>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.16.752260>

Abstract: The Big Fantastic Virus Database (BFVD) v3 is the most comprehensive resource for ColabFold-AlphaFold2-predicted viral protein structures. It holds 5,776,417 structures of nearly all viral sequences in UniProt 2025\_03, a 16.4-fold increase compared to the representative-only catalogs of BFVD v1 and v2. Structure prediction quality has also improved, with high-confidence predictions accounting for 75.3% of BFVD v3 entries. This is due to Logan's enormous sequence assembly, now mined for the entire BFVD v3 instead of only for shallow alignments, alongside continued prediction quality improvements. BFVD v3 covers 72.7% of ICTV's virus species, 1.54-fold more than BFVD v2. With this version, the fraction of fully covered viral reference proteomes increases from 1.5% to 72.6%. BFVD v3 thus brings us closer to a structural catalog of the known virosphere and is freely available at bfvd.steineggerlab.workers.dev and bfvd.foldseek.com.

## BGC Atlas v2: biosynthetic gene clusters with taxonomic and environmental context at scale
- Source: bioRxiv (preprints)
- Date: 2026-09-18
- Categories: Tools & resources
- Authors: Bagci, C., Talamas-Tanner, A., Ziemert, N.
- DOI: 10.64898/2026.09.14.751540
- Source URL: <https://doi.org/10.64898/2026.09.14.751540>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.14.751540>

Abstract: Genomic and metagenomic studies have revealed that microbes have the capacity to produce an enormous diversity of secondary metabolites. These compounds play important roles in microbial interactions and are also a major source of medicines and other useful natural products, yet only a small fraction of this biosynthetic potential has been experimentally characterized. BGC Atlas was developed to explore this largely uncharacterized diversity by placing biosynthetic gene clusters (BGCs) into genomic and environmental context. Here, we present BGC Atlas v2, which expands the collection nearly ninefold to more than 16 million predicted BGCs and extends it from metagenomic assemblies to MAGs, single-amplified genomes and isolate genomes. The new release provides taxonomic assignments for nearly all BGCs, harmonized environmental metadata, nested searches across biosynthetic, taxonomic, environmental and geographic properties, and protein-sequence searches against BGC-encoded genes. Despite its scale, the collection reveals how much microbial biosynthetic diversity remains unexplored. By connecting pathways to related families, organisms and environments, BGC Atlas v2 supports natural-product discovery and investigations of microbial biosynthetic diversity across taxa and ecosystems. BGC Atlas v2 is freely available at https://bgc-atlas.cs.uni-tuebingen.de.

## BoneGraph: A Domain-Specialised, Self-Correcting Reasoning System for Bone Science Retrieval, Grounded Inference, and Image Mechanics
- Source: bioRxiv (preprints)
- Date: 2026-09-18
- Categories: Biological imaging, Tools & resources
- Authors: Valijonov, J., Soar, P., Le Houx, J., Tozzi, G.
- DOI: 10.64898/2026.09.17.752482
- Source URL: <https://doi.org/10.64898/2026.09.17.752482>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.17.752482>

Abstract: Bone science literature spans biology, mechanics, materials science, and clinical medicine, and its volume makes reliable knowledge synthesis increasingly difficult. General-purpose large language models (LLMs) answer fluently but under-represent this niche domain, cannot cite specific evidence, and offer no mechanism to be corrected durably. Here we present BoneGraph, a domain-specialised system for bone science delivered as a five-tab web application over a shared substrate: a curated full-text corpus of 7,449 documents embedded into 248,629 passage vectors using SPECTER2, a scientific-paper embedding model, and a bone knowledge graph of 1,597 concepts with 1,699 causal relations. The five tabs are: (I) Chat, retrieval-augmented question answering with server-rebuilt inline citations; (II) Search, raw semantic retrieval with no LLM in the loop; (III) Reasoning, a self-correcting loop in which a deterministic physics check and a literature/knowledge-graph critic constrain the answer, and a user's feedback becomes a durable, per-user rule; (IV) Vision, a bone-region classifier trained on frozen BiomedCLIP features that grounds a vision-language model, guarded against out-of-distribution inputs and augmented with image-embedding correction memory; and (V) Mechanics, integrating our previous data-driven image mechanics (D2IM) model that predicts displacement and strain fields from a single undeformed micro-CT image. All inference is performed locally, without third-party API calls, and the public beta is served at bonegraph.org. Retrieval attains a mean reciprocal rank (MRR) of 0.928 on a 30-question, seven-domain benchmark, and the Vision classifier attains 92.6% accuracy on the held-out MURA (MUsculoskeletal RAdiographs) dataset. A grounded-reasoning benchmark shows that, with the correct passage, BoneGraph raises answer accuracy from 42% to 78%. BoneGraph makes a major contribution to bone-science informatics: to our knowledge it is the first domain-specialised system to unify curated retrieval, deterministic physics-grounded self-correction, and durable per-user learning for bone science.

## Characterization of recombinase-based genetic parts and circuits using nanopore sequencing
- Source: Nature Communications (journals)
- Date: 2026-09-18T00:00:00+00:00
- Categories: Genomics & sequence analysis
- Authors: F. Veronica Greco, Sarah K. Cameron, Shivang Hina-Nilesh Joshi, Sarah Guiziou, Jennifer A. N. Brophy, Claire S. Grierson, Thomas E. Gorochowski
- Journal: Nature Communications
- DOI: 10.1038/s41467-026-77605-x
- Keywords: dna
- Source URL: <https://doi.org/10.1038/s41467-026-77605-x>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1038%2Fs41467-026-77605-x>

Abstract: Recombinases are versatile enzymes able to perform the precise insertion, deletion, and rearrangement of DNA and can act as a foundation for programmable genetic logic and memory. Crucial for their use are accurate measurements of function. However, these are often laborious, time-consuming, and costly to collect. To address this, we develop a semi-automated workflow that combines low-cost liquid handling robotics, multiplexed long-read nanopore sequencing, and a supporting computational analysis tool to enable the high-throughput and detailed characterization of recombinase parts and circuits when used in a variety of contexts and organisms. Our approach overcomes the limitations of typically used fluorescence-based assays and is able to monitor temporal dynamics, observe structural changes at nucleotide resolution, and unravel the internal workings of complex multi-state circuits. The ability to scale up and automate genetic circuit characterization is an essential step towards more rigorous biological metrology that can support the construction of predictive models for efficiently engineering biology.

## Comparative Mitogenomics and Molecular Phylogeny of Agriculturally Significant Tephritid Fruit Fly Pests in Bangladesh
- Source: bioRxiv (preprints)
- Date: 2026-09-18
- Categories: Genomics & sequence analysis, Proteins & structural biology, Evolution & metagenomics
- Authors: Rahman, S., Shormi, F. A.
- DOI: 10.64898/2026.09.12.751165
- Keywords: genome, genomic, phylogeny, phylogenetic, 16s
- Source URL: <https://doi.org/10.64898/2026.09.12.751165>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.12.751165>

Abstract: Tephritid fruit flies of the tribe Dacini rank among the most economically destructive agricultural pests globally, with several Bactrocera Macquart and Zeugodacus Hendel species causing severe losses to fruit and vegetable production across South and Southeast Asia. Bangladesh harbors five dacine species of primary agricultural significance: Bactrocera dorsalis (Hendel), B. carambolae Drew & Hancock, B. zonata (Saunders), B. correcta (Bezzi), and Zeugodacus cucurbitae (Coquillett). Here we present a comparative mitogenomic and molecular phylogenetic framework based on 21 unique mitogenome records retrieved from NCBI GenBank, representing 19 dacine ingroup taxa and two outgroups. Direct parsing of the GenBank sequences confirmed the canonical complement of 13 protein-coding genes (PCGs) and two rRNA genes in all 21 records. Among the 14 Bactrocera ingroup taxa, genome size ranged from 15,273 to 15,977 bp. Whole-genome AT content ranged from 66.6% in Bactrocera tsuneonis to 82.2% in Drosophila melanogaster; the five Bangladesh-relevant dacine pest species showed tightly clustered AT content of 72.9-73.6%. A concatenated alignment of 13,600 nucleotide sites (13 PCGs + 12S + 16S rRNA) was analyzed by maximum-likelihood inference under the GTR+FO model (IQ-TREE v1.6.11; 1,000 ultrafast bootstrap replicates). The ML tree recovers the B. dorsalis complex (UFBoot = 79-100) and places B. correcta and B. zonata as a maximally supported sister pair (UFBoot = 100). This study provides a sequence-verified mitogenomic reference framework for molecular identification and pest surveillance in Bangladesh, and should be interpreted as a curated comparative baseline rather than a population-genomic analysis, as no newly collected Bangladeshi specimens were sequenced.

## Compartment-Aware Benchmarking of Respiratory-Virus Transcriptomes for Nasal and Blood Host-Response Modules: A Computational Study.
- Source: Journal of visualized experiments : JoVE (journals)
- Date: 2026-09-18
- Categories: Genomics & sequence analysis, Tools & resources
- Authors: Shude Han, Sinan Jin
- Journal: Journal of visualized experiments : JoVE
- DOI: 10.3791/73334
- External ID: 42762140
- Source URL: <https://doi.org/10.3791/73334>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.3791%2F73334>

Abstract: Public respiratory-virus transcriptomes are valuable for studying host responses, but differences in tissue source, control definition, and study design can confound pooled analyses. We developed a compartment-aware computational workflow to determine whether reproducible host-response activity can be identified while preserving nasal and blood biological context. The paired GSE117827 paediatric cohort served as the anchor dataset, comprising nasal-swab and whole-blood transcriptomes from symptomatic picornavirus infection, symptomatic respiratory syncytial virus infection, asymptomatic picornavirus detection, and virus-negative controls. After HUGO Gene Nomenclature Committee (HGNC) filtering, 27,685 genes were analyzed. Separate 50-gene protein-coding nasal and blood modules were defined from the top positive responses and locked before external evaluation. Gene-level nasal and blood effects were nearly independent (Pearson r = 0.015), and the modules shared six exploratory rank-overlap genes (Jaccard index = 0.064). Nevertheless, the nasal module separated infection from controls in independent upper-airway cohorts, with areas under the receiver operating characteristic curve (AUROCs) of 0.749, 0.693, and 0.609, whereas the blood module achieved AUROCs of 0.832, 0.924, and 0.870 in external blood cohorts. In longitudinal natural-infection data, matched scores decreased from acute illness to discharge, with paired deltas of 0.436 for nasal samples and 0.330 for blood. Three additional benchmark datasets comprising 666 external samples, together with random-gene nulls, module-size sweeps, bootstrap stability, marker-program correlations, and variance partitioning, defined the robustness and limitations of the workflow. The Pandya 33-messenger ribonucleic acid (mRNA) set remained stronger for viral-versus-bacterial discrimination, while the blood module also increased in bacterial pneumonia. These findings support compartment-specific modules as reusable host-response activity scores for cohort comparison and recovery tracking, rather than universal pan-tissue biomarkers or stand-alone pathogen classifiers.

## Comprehensive investigation of intermittent large-amplitude excursions in the memristive Hindmarsh-Rose neuron model
- Source: Scientific Reports (journals)
- Date: 2026-09-18T00:00:00+00:00
- Authors: Dinesh Vijay Sadhasivam, Mohanasubha Ramasamy, Abirami Karunanidhi, Gaudence Nyiranzeyimana
- Journal: Scientific Reports
- DOI: 10.1038/s41598-026-71866-8
- Source URL: <https://doi.org/10.1038/s41598-026-71866-8>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1038%2Fs41598-026-71866-8>

Abstract: This study systematically investigates intermittent large-amplitude oscillations across all three state variables of the memristive Hindmarsh–Rose neuron model. The large-amplitude excursions, arising from interior crisis-induced intermittency, are first identified through local and two-parameter bifurcation analysis. All three state variables are then statistically characterized using the significant-height threshold criterion, probability distribution functions, probability of exceedance, and the density of threshold exceeding peaks over successive time-windows. The results reveal that the membrane potential exhibits a consistently lower probability and density of extreme events compared to the recovery and slow variables, and that these extreme events are confined to the chaotic regions identified via two-parameter Lyapunov exponent maps. To examine the robustness of this behavior under memory effects, the analysis is extended to the fractional-order memristive Hindmarsh-Rose neuron model, where the period-doubling route to chaos and the corresponding exceedance statistics of all state variables are investigated. The membrane potential retains its comparatively lower susceptibility to extreme events across fractional orders, confirming that this variable-dependent behavior is a robust feature of the neuron model rather than an artifact of the integer-order formulation.

## Conditional Generation And Inpainting Of Non-coding RNA Sequences With Masked Discrete Diffusion
- Source: bioRxiv (preprints)
- Date: 2026-09-18
- Categories: Genomics & sequence analysis, Proteins & structural biology, Tools & resources
- Authors: Upadhyay, U., Dai, C., Herold, J., Sato, K., Schug, A.
- DOI: 10.64898/2026.09.17.752279
- Source URL: <https://doi.org/10.64898/2026.09.17.752279>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.17.752279>

Abstract: Designing functional non-coding RNA (ncRNA) is fundamental to synthetic biology and RNA therapeutics, yet generative modelling for ncRNA has received far less attention than protein design. We present RNA-MDLM, a framework that extends Masked Discrete Language Models (MDLM) to the conditional generation and inpainting of ncRNA. We make two additions: first, conditioning on RNA-type representations from a pretrained RNA language model, and second, a modified classifier-free guidance scheme (Mod-CFG) that interpolates among conditional, unconditional, and random-sequence probabilities for better control. We also introduce REPAINT GAMES, a benchmark of seven structured masking tasks to probe a model's performance on sequence patterns, structural motifs, and base-pairing per RNA type. Our model is trained on 4.6 million ncRNA sequences spanning six evaluable classes. It produces sequences whose composition and folding statistics closely match natural RNAs. Through extensive ablation studies, we find that the embedding-conditioned model achieves the best balance of structural fidelity, biological novelty, and inpainting accuracy, and its class label steers generation far more strongly than a plain label baseline. We further show that a model trained on a smaller, class-balanced subset can appear more realistic mainly by copying abundant natural sequences rather than learning their rules. We also benchmark against a masked-diffusion model and a family-specific VAE on ribozyme families, and find that a type-conditioned model like ours and a per-family model are solving different tasks, which must be accounted for in a fair comparison. We will release the code and the trained models.

## Controlled Blood-Brain Barrier Modulation by a High-Affinity Claudin-5 Peptide Binder
- Source: bioRxiv (preprints)
- Date: 2026-09-18
- Categories: Proteins & structural biology, Systems & networks
- Authors: Berselli, A., Trevisani, M., Alberini, G., Pastore, A., Di Fonzo, A., Armirotti, A., Castagnola, V., Maragliano, L., Benfenati, F.
- DOI: 10.64898/2026.09.12.751133
- Keywords: peptide, peptides, proteomic, pathway
- Source URL: <https://doi.org/10.64898/2026.09.12.751133>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.12.751133>

Abstract: The blood-brain barrier (BBB) is a specialized interface that tightly regulates the exchange of molecules between the bloodstream and the brain. Its barrier function relies on a monolayer of brain endothelial cells sealed by tight junctions (TJs) that restrict paracellular flux through claudin-5 (CLDN5) multimeric complexes. To improve the delivery of nutrients and drugs to the brain, CLDN5-competitive peptides are promising carriers for creating size-controlled, temporary openings in the paracellular pathway. Here, we combine generative protein design and atomistic simulations to design ST9, a peptide with high nanomolar affinity for CLDN5. Compared with f1-C5C2, a CLDN5-binding peptide that we previously reported, ST9 induces a rapid, transient, size-controlled, and fully reversible increase in paracellular permeability, without altering CLDN5 expression or the proteomic profile of brain endothelial cells, indicating distinct mechanisms of TJ destabilization. This work provides a promising approach for developing next-generation BBB-opening agents to effectively treat neurological diseases.

## ConvexGating infers gating strategies from clusters in single cell cytometry data
- Source: Nature Communications (journals)
- Date: 2026-09-18T00:00:00Z
- Categories: Tools & resources
- Authors: V. Friedrich, Karola Mai, T. Hofer, E. Nössner, L. Bonaguro, Celia L. Hartmann, A. Frolov, C. Carraro, Doaa Hamada, Mehrnoush Hadaddzadeh Shakiba, Heidi Theis, Dalila Silva Ribeiro, D. Wachten, F. T. Wunderlich, M. Scholz, F. Theis, M. Becker, M. Beyer, J. Schultze, M. Büttner
- Journal: Nature Communications
- DOI: 10.1038/s41467-026-77360-z
- External ID: 1a4a12467004d00ee944bdcf40cfc6a9c499a115
- Source URL: <https://doi.org/10.1038/s41467-026-77360-z>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1038%2Fs41467-026-77360-z>

Abstract: Manual expert gating remains common practice for defining specific cell populations in flow cytometry data, but increasing numbers of measured parameters and high inter-rater variability limit consistency across studies. Cluster-based approaches use the full marker space to define cell populations more consistently, but their outputs cannot be directly implemented on a cell sorter. Here we develop ConvexGating, an artificial intelligence tool to address this gap by automatically learning interpretable gating strategies for sorting in an unbiased, data-driven manner, generating low-contamination strategies for both known and previously unknown cell populations, including plasmacytoid dendritic cells identified solely as CD57-CD13-CD45RA+ CD123+ cells. We show that ConvexGating derives sorting strategies for CD8+ subtypes and adipose progenitor cell populations, which we validate experimentally by single-cell sequencing of sorted cells. We also demonstrate that the method transfers effectively to Cytometry by Time of Flight and Cellular Indexing of Transcriptomes and Epitopes by Sequencing data and improves marker panel design for cell sorting. Here, the authors develop ConvexGating, an AI tool that learns interpretable, data-driven gating strategies for cell sorting. They show that when analyzing scRNA-seq data, it yields low-contamination populations and also works across flow cytometry, cyTOF and CITE-seq data.

## Data-driven modeling of spatiotemporal dynamics using multimodal imaging data
- Source: PLOS Computational Biology (journals)
- Date: 2026-09-18T00:00:00+00:00
- Categories: Computational neuroscience
- Authors: Chunyan Li, Yutong Mao, Xiao Liu, Wenrui Hao
- Journal: PLOS Computational Biology
- DOI: 10.1371/journal.pcbi.1014751
- Source URL: <https://doi.org/10.1371/journal.pcbi.1014751>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1371%2Fjournal.pcbi.1014751>

Abstract: Understanding how biological systems evolve across space and time remains a fundamental challenge, particularly when dynamic processes vary substantially across individuals. We present a personalized graph-based dynamical modeling framework for characterizing spatiotemporal biological dynamics from longitudinal multimodal imaging data. The framework constructs individualized brain graphs from MRI and PET measurements and learns patient-specific dynamical parameters governing regional structural and molecular changes. Applied to 1,891 participants from the Alzheimer’s Disease Neuroimaging Initiative, the model captures the coordinated evolution of amyloid- β , tau, neurodegeneration, and cognition and accurately predicts their future trajectories, outperforming established clinical and neuroimaging benchmarks. Patient-specific dynamical parameters reveal distinct patterns of biological progression and provide improved prediction of future cognitive decline compared with standard biomarkers. Sensitivity analysis further identifies regional network features associated with the propagation of pathological and structural changes, recovering known temporolimbic and frontal vulnerability patterns. These results demonstrate how data-driven dynamical modeling can integrate multimodal longitudinal measurements to uncover individualized spatiotemporal patterns and latent mechanisms of biological change. The framework provides a quantitative approach for studying complex biological dynamics across heterogeneous individuals and establishes a foundation for personalized modeling of progressive biological processes.

## DCUSV: deep clustering of ultrasonic vocalizations in rodents
- Source: Scientific Reports (journals)
- Date: 2026-09-18T00:00:00+00:00
- Authors: Sabah Shahnoor Anis, Devin M. Kellis, Kris Ford Kaigler, Marlene A. Wilson, Christian O’Reilly
- Journal: Scientific Reports
- DOI: 10.1038/s41598-026-67804-3
- Source URL: <https://doi.org/10.1038/s41598-026-67804-3>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1038%2Fs41598-026-67804-3>

Abstract: Analyzing ultrasonic vocalizations (USVs) is critical for understanding rodents’ emotional states and social behaviors. This work presents Deep Clustering of USVs (DCUSV), an automated deep clustering pipeline for analyzing preprocessed USV contours that addresses key challenges in effectively clustering USVs and revealing distinct patterns in rodent vocal behavior. DCUSV employs a dense autoencoder to compress high-dimensional spectrograms into a latent space suitable for clustering, followed by a combination of Uniform Manifold Approximation and Projection (UMAP), Hierarchical Density-Based Spatial Clustering of Applications with Noise (HDBSCAN), and Agglomerative clustering, with hyperparameter optimization, to group USVs based on their spectro-temporal features. Clustering is evaluated using the Silhouette Coefficient, Calinski-Harabasz Index, and Davies-Bouldin Index. In addition, Principal Component Analysis (PCA), t-distributed Stochastic Neighbor Embedding (t-SNE), and UMAP are used to visualize and analyze the clustering results. Performance benchmarks against six baseline methods—K-Means, Deep Embedded Clustering (DEC), Improved Deep Embedded Clustering, Spectral Clustering, Gaussian Mixture Models, and HDBSCAN—revealed that DCUSV outperforms all baselines across every metric, achieving up to 2.62× higher Silhouette Coefficient scores, 22.95× higher Calinski-Harabasz scores, and 3.62× lower Davies-Bouldin scores (lower is better for the latter). Applying DCUSV revealed four distinct call-type families, closely aligning with manually defined categories without requiring manual grouping. Furthermore, DCUSV identified statistically significant shifts in call-type distributions across experimental conditions, demonstrating its ability to capture behaviorally meaningful changes in vocal expression. Thus, DCUSV enables robust analysis of USV structure and uncovers novel patterns in rodent vocal behavior.

## Deconvolving DNA mixtures with Demixtify.
- Source: Forensic science international. Genetics (journals)
- Date: 2026-09-18T00:00:00Z
- Categories: Genomics & sequence analysis, Tools & resources
- Authors: August E. Woerner
- Journal: Forensic science international. Genetics
- DOI: 10.1016/j.fsigen.2026.103624
- External ID: 75f4fb1d9cd0f97b96e5fa76a104eb8f03393c49
- Source URL: <https://doi.org/10.1016/j.fsigen.2026.103624>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1016%2Fj.fsigen.2026.103624>

Abstract: While there are many tools to detect DNA mixtures, few can deconvolve mixed genetic profiles at genomic scales. Demixtify is one such application. Demixtify characterizes and deconvolves DNA mixtures in whole genome sequencing data. Demixtify is also fully amenable to modern imputation strategies, which in turn allows samples to be accurately characterized even when the information content is limited (<1×). In the present study Demixtify is applied to two-person in silico DNA mixtures. Genotypes are often accurately inferred, with noticeable increases in performance when genotypes are also refined. Likewise, kinship coefficients and IBD segments are generally well-recovered, especially in imbalanced mixtures. Last, a high throughput (~30×), highly imbalanced DNA mixture (~0.7%) in a public genomic resource is deconvolved and the minor contributor is traced to another individual in the study.

## Decreased Damage for proton FLASH vs Conventional Dose Rates in Mouse Jejunum Shown by Quantitative Assessment of γ-H2AX
- Source: bioRxiv (preprints)
- Date: 2026-09-18
- Categories: Genomics & sequence analysis
- Authors: Curtis, N., Kim, M. M., Verginadis, I., Zou, W., Diffenderfer, E. S., Koumenis, C., Koch, C. J., Wiersma, R. D.
- DOI: 10.64898/2026.09.11.750404
- Keywords: dna
- Source URL: <https://doi.org/10.64898/2026.09.11.750404>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.11.750404>

Abstract: Purpose: FLASH radiation with ultra-high dose rate delivery is less damaging to normal tissue than conventional radiation ( <1 Gy/s). Since radiation depletes oxygen (ROD), this damage reduction might occur via the oxygen effect. ROD experiments have shown an oxygen-independent reduction in dose effectiveness at FLASH dose rates. However, prior in vivo ROD measurements relied on extracellular oxygen probes that could not penetrate cell membranes, leaving intracellular effects unresolved. To investigate the ROD hypothesis more directly, we developed a novel three-component immunohistochemical assay with algorithmic image processing to quantitatively compare DNA damage following FLASH and conventional irradiation in mouse jejunum. Methods: Mice received intravenous EF5 2 hours before proton irradiation at FLASH (103.63 +/- 17.2 Gy/s) or conventional (0.73 +/- 0.1 Gy/s) dose rates of 2.5 Gy or 5 Gy, with unirradiated controls. Mice were euthanized 30 minutes post-irradiation, and 10 cm of jejunum was frozen as a 'Swiss Roll', sectioned, stained, and imaged. Tissue sections were stained for \{gamma\}-H2AX, DRAQ5, and EF5 to assess DNA double-strand breaks, total DNA content, and hypoxia, respectively. An in-house algorithm identified individual cell nuclei and registered each nucleus with its corresponding \{gamma\}-H2AX and EF5 signals, enabling quantitative measurement of DNA damage as a function of local tissue hypoxia. Results: Hypoxia was greatest in the villi and, to a lesser extent, the outer jejunal musculature, with substantial inter-animal variation. DNA damage decreased in hypoxic regions. FLASH enhanced the hypoxia-associated reduction in DNA damage compared with conventional dose rate and, separately, revealed an oxygen-independent reduction in DNA damage, suggesting an additional FLASH sparing mechanism. Conclusion: Current results suggest that FLASH compared to conventional dose rate radiation caused less DNA damage with increasing effect at low oxygen levels, a result consistent with ROD as a mechanism. Pronounced tissue heterogeneity in murine jejunum requires further studies to segment the effect for each tissue type.

## DeepSpaceDB 2.0: an interactive spatial transcriptomics database for large-scale Xenium data exploration
- Source: bioRxiv (preprints)
- Date: 2026-09-18
- Categories: Tools & resources
- Authors: Honcharuk, V., Takemoto, K., Masalunga, M. C., Zhao, H., Diez, D., Kawaoka, S., Vandenbon, A.
- DOI: 10.64898/2026.01.15.699623
- Source URL: <https://doi.org/10.64898/2026.01.15.699623>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.01.15.699623>

Abstract: The 10x Genomics Xenium platform enables high-resolution spatial transcriptomics at single-cell and subcellular scales, but effective reuse of public Xenium datasets is hindered by large data sizes and heterogeneous file formats. We previously developed DeepSpaceDB, a spatial transcriptomics database designed for interactive, in-depth analysis of tissues and tissue microenvironments. Here, we present a major expansion of DeepSpaceDB that integrates large-scale single-cell spatial transcriptomics data generated by the Xenium platform. In this update, we systematically collected 1,539 public Xenium datasets from multiple repositories and processed them through a robust, standardized pipeline that validates, repairs, and harmonizes heterogeneous inputs into a unified representation. To support efficient exploration of these data, we introduced a redesigned DeepSpaceDB interface and complementary Zarr-based storage formats optimized for gene-centric visualization and spatially localized queries, enabling sub-second response times for common interactive operations. The updated platform supports real-time visualization of spatial data and analysis of regions of interest directly in the web browser. Together, this expansion establishes DeepSpaceDB as a unified resource for single-cell spatial transcriptomics, substantially lowering the barrier to accessing, exploring, and reusing large-scale public Xenium datasets.

## DeepVir: A reproducible workflow for large-scale viral dark matter discovery
- Source: bioRxiv (preprints)
- Date: 2026-09-18
- Categories: Genomics & sequence analysis, Tools & resources
- Authors: Cosentino, M., FERNANDEZ NUNEZ, N., Soares, M. A., Ayouba, A., Santos, A., D'arc, M.
- DOI: 10.64898/2026.09.15.751837
- Source URL: <https://doi.org/10.64898/2026.09.15.751837>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.15.751837>

Abstract: High-throughput sequencing (HTS) has revolutionized virosphere exploration. However, characterizing highly divergent viral sequences remains a bottleneck known as Viral Dark Matter (VDM). Numerous tools were developed to unravel such diversity, but they usually require complex prior HTS data analysis processes. Consequently, a common bottleneck to VDM exploration is the manual, chained execution of complex command-line applications. To address the need for automated and scalable viral discovery, we developed DeepVir, a reproducible Snakemake pipeline that integrates classical homology-based alignments with profile Hidden Markov Model (HMM) mining of the RNA-dependent RNA polymerase (RdRp). To validate the pipelines efficacy, we analyzed 385.23 GB of publicly available transcriptomic data (203 Sequence Read Archive libraries) from 49 American bat species. DeepVir successfully identified 179 distinct viral groups. This included 903 contigs spanning nine known viral families, enabling the characterization of novel genomes within Orthomyxoviridae (Influenza A H7N9), Picornaviridae, Alphaflexiviridae, Retroviridae (Spumaretrovirinae), Papillomaviridae, Herpesviridae, and Adenoviridae. Furthermore, the pipeline uncovered 170 putative novel VDM lineages. By employing deep homology searches and Sequence Similarity Network (SSN) visualization, we contextualized these highly divergent VDM sequences, revealing significant evolutionary relationships with the orders Mononegavirales and Bunyavirales. Notably, human-driven curation of the pipelines outputs allowed for the cross-library assembly of the first putative exogenous Spumavirus in the Americas. Ultimately, by automating complex bioinformatic processing steps, scalable pipelines like DeepVir empower researchers to prioritize the biological and epidemiological interpretation of their findings, accelerating the characterization of wildlife virospheres and enhancing pathogen genomic surveillance.

## Developing SCL2205 : A Protein Sequence-based Spatial Modelling Dataset for the Protein Language Model Frontier
- Source: Bioinformatics (journals)
- Date: 2026-09-18T00:00:00+00:00
- Categories: Proteins & structural biology, Tools & resources
- Authors: Daniel Ouso, Gianluca Pollastri
- Journal: Bioinformatics
- DOI: 10.1093/bioinformatics/btag666
- Source URL: <https://doi.org/10.1093/bioinformatics/btag666>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1093%2Fbioinformatics%2Fbtag666>
- Code: <https://github.com/ousodaniel/scldata>

Abstract: Motivation Deep learning (DL) has substantially advanced protein subcellular localisation (SCL) prediction, yet its potential remains constrained by suboptimal input preparation and limited high-quality reference data. Furthermore, existing state-of-the-art (SoTA) predictors suffer from performance metric inflation due to unmitigated training-to-testing data leakage during homology augmentation. We address these challenges by introducing SCL2205, a leak-minimised benchmark dataset and pipeline curated specifically to support trustworthy, scalable, and reproducible DL-based SCL modelling. Results SCL2205 was constructed from the universal protein knowledgebase (UniProtKB) using rigorous preprocessing, manual label mapping, and stringent partitioning. When evaluated on independent test sets, SCL2205 yielded up to a 10.8 percentage point improvement in macro area under the precision–recall curve (PR-AUC) over SoTA baselines (mean Δ 95% CI=0.07−0.12 ), with maximum benefits observed when paired with modern protein language models (PLMs). Crucially, we quantify for the first time a systemic 5.2%±0.32 data leakage rate in conventional homology augmentation workflows—even when restricting sequence similarity searches to just 10% of the training set. Availability and Implementation The dataset is openly available on Dryad under a CC0 1.0 Universal licence (https://doi.org/10.5061/dryad.2ngf1vj1t). The dataset interface is available as an installable Python package, p-scldata (v2026.2.0), under the MIT licence on the Python Package Index (PyPI). Full code and data repositories are hosted on GitHub (https://github.com/ousodaniel/scldata) and archived on Zenodo (https://doi.org/10.5281/zenodo.21796423). Supplementary Information Supplementary File S1 contains code snippets, per-class PR-AUC breakdowns, class prevalence details, statistical comparison tests, and supplementary figures. Supplementary File S2 contains the exact mapping used in curation.

## Direct microhaplotype genotyping for GT-seq (Genotyping-in-Thousands by Sequencing) using a diploid abundance model
- Source: PLOS Computational Biology (journals)
- Date: 2026-09-18T00:00:00+00:00
- Categories: Genomics & sequence analysis
- Authors: Nathan R. Campbell, Amanda R. Campbell, Shannon K. Blair, Amanda J. Finger
- Journal: PLOS Computational Biology
- DOI: 10.1371/journal.pcbi.1014808
- Source URL: <https://doi.org/10.1371/journal.pcbi.1014808>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1371%2Fjournal.pcbi.1014808>

Abstract: GT-seq (Genotyping-in-Thousands by Sequencing) is widely used for high-throughput amplicon genotyping, but most analytical pipelines focus on single SNPs or rely on alignment-based variant calling. Here we present a direct microhaplotype genotyping framework that leverages the high read depth and low error rates typical of paired-end Illumina and Element sequencing. The pipeline first identifies primer-bounded reads and resolves paired-end sequences into quality-aware consensus amplicon sequences. Within each sample and locus, unique sequences are ranked by read abundance and the top one or two sequences are retained as directly observed haplotypes. These alleles are aggregated across samples to construct a catalog of observed haplotypes for each locus. In a second pass, reads are assigned to catalog haplotypes by exact sequence matching to produce diploid genotypes. Finally, catalog haplotype sequences are compared to identify phased SNP and collapsed indel variation. Optionally, catalog haplotypes may be aligned to a reference genome to project observed variants onto genomic coordinates and generate standards-compliant VCF output. This framework enables robust, microhaplotype genotyping directly from high-depth amplicon sequencing data. Comparison with an independent BWA/BCFtools alignment-based workflow demonstrated 99.67% genotype concordance across 102,520 genotype comparisons spanning 1,085 SNPs in 96 individuals. Genotype concordance remained above 99.4% even at the minimum supported sequencing depth of 10 reads per locus, demonstrating robust performance across a broad range of sequencing depths.

## Elucidating the mechanism of amikacin-induced acute kidney injury: a multilevel analysis based on the FAERS database and network toxicology.
- Source: Naunyn-Schmiedeberg's archives of pharmacology (journals)
- Date: 2026-09-18
- Categories: Tools & resources
- Authors: Chao Li, Wei Wu, Jingli Liao, Fenfen Gu, Lixia Li
- Journal: Naunyn-Schmiedeberg's archives of pharmacology
- DOI: 10.1007/s00210-026-05913-6
- External ID: 42754716
- Keywords: database
- Source URL: <https://doi.org/10.1007/s00210-026-05913-6>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1007%2Fs00210-026-05913-6>

Abstract: Amikacin, an aminoglycoside antibiotic for severe Gram-negative infections, is limited by dose-dependent nephrotoxicity. However, its acute kidney injury (AKI) risk profile and underlying molecular mechanisms remain insufficiently characterized in real-world settings. This study integrated real-world data with computational biology approaches. Pharmacovigilance analysis was performed using the FDA Adverse Event Reporting System (FAERS) to identify the risk of acute AKI associated with amikacin. Network toxicology was utilized to screen shared targets, while molecular docking and dynamics simulations were conducted to evaluate binding interactions. The expression of core genes was validated using GEO datasets. Disproportionality analysis indicated a significant amikacin-AKI association. Injectable formulation posed higher risk than inhalation (OR = 7.47). Male sex and age ≤ 65 years were independent risk factors. Network toxicology identified IL1B, CXCL8, SIRT1, and PTGS2 as hub genes. Molecular docking showed strong binding (SIRT1, - 7.789 kcal/mol; PTGS2, - 9.467 kcal/mol), with dynamics indicating stability over 100 ns. GEO analysis corroborated the predicted upregulation of IL1B and CXCL8 in AKI, and further supported the involvement of PTGS2, which was significantly upregulated in a cisplatin‑induced AKI model. This study delineates the risk profile of amikacin-associated AKI and elucidates a molecular mechanism involving multi-target interactions in renal injury induction, thereby offering a theoretical basis and identifying potential molecular targets for further investigation into clinical risk mitigation strategies.

## EM3DFold: accurate de novo protein and nucleic acid model building for cryo-EM maps using language model-powered deep learning
- Source: bioRxiv (preprints)
- Date: 2026-09-18
- Categories: Proteins & structural biology, Biological imaging, Tools & resources
- Authors: Li, T., Cao, H., Huang, S.-Y.
- DOI: 10.64898/2026.09.16.752067
- Source URL: <https://doi.org/10.64898/2026.09.16.752067>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.16.752067>
- Code: <https://github.com/huang-laboratory/EM3DFold>

Abstract: Cryo-electron microscopy (cryo-EM) has become one of the most powerful techniques for macromolecular structure determination. However, accurate model building from cryo-EM maps remains challenging, particularly for nucleic acids. Here, we present EM3DFold, a unified de novo model-building framework for accurate structure determination of proteins, nucleic acids, and protein-nucleic acid complexes from cryo-EM maps using a density-aware, large language model-powered three-track attention (TTA) network. The TTA network effectively integrates sequence, density, and structural information to enable accurate all-atom model building. EM3DFold was extensively evaluated on independent benchmarks of 298 experimental cryo-EM maps at < 4.0 \[A\] resolutions, and achieves an unprecedentedly high median accuracy of 75% completeness (88% coverage and 94% sequence accuracy) for 178 nucleic acid maps, 90% completeness (95% coverage and 96% sequence accuracy) for 124 protein-nucleic acid complexes, and 95% completeness (97% coverage and 98% sequence accuracy) for 120 protein-only targets, substantially outperforming state-of-the-art methods including ModelAngelo, EM2NA, CryoREAD, and EMProt. In addition, EM3DFold also produces models with superior model-to-map fit and stereochemical quality, providing a robust and reliable solution for automated cryo-EM model building. The EM3DFold package is freely available at https://github.com/huang-laboratory/EM3DFold.

## Evolutionary rescue model informs strategies for driving cancer cell populations to extinction
- Source: PLOS Computational Biology (journals)
- Date: 2026-09-18T00:00:00+00:00
- Categories: Evolution & metagenomics, Mathematical biology & statistics
- Authors: Amjad Dabi, Joel S. Brown, Robert A. Gatenby, Corbin D. Jones, Daniel R. Schrider
- Journal: PLOS Computational Biology
- DOI: 10.1371/journal.pcbi.1014803
- Source URL: <https://doi.org/10.1371/journal.pcbi.1014803>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1371%2Fjournal.pcbi.1014803>

Abstract: Cancers exhibit a remarkable ability to develop resistance to a range of treatments, often resulting in relapse following first-line therapies and significantly worse outcomes for subsequent treatments. While our understanding of the mechanisms and dynamics of the emergence of resistance during cancer therapy continues to advance, questions remain about how to minimize the probability that resistance will evolve, thereby improving long-term patient outcomes. Here, we present an evolutionary simulation model of a clonal population of cells that can acquire resistance mutations to one or more treatments. We leverage this model to examine the efficacy of a two-strike “extinction therapy” protocol, in which two treatments are applied sequentially to first contract the population to a vulnerable state and then push it to extinction, and compare it to a combination therapy protocol. We investigate how factors such as the timing of the switch between the two strikes, the rate of emergence of resistant mutations, the dose effects of the applied drugs, the presence of cross-resistance, and whether resistance is a discrete or a quantitative trait affect the outcome. Our results show that the timing of switching to the second strike has a marked effect on the likelihood of driving the cancer to extinction, and that extinction therapy outperforms combination therapy when cross-resistance is present. We conduct an in silico trial that reveals when and why a second strike will succeed or fail. Finally, we demonstrate that our conclusions hold whether we model resistance as a discrete trait or as a quantitative, multi-locus trait.

## Fast and Flexible Flow Decompositions in General Graphs via Dominators
- Source: Journal of Computational Biology (journals)
- Date: 2026-09-18T00:00:00+00:00
- Categories: Genomics & sequence analysis
- Authors: FRANCISCO SENA, ALEXANDRU I. TOMESCU
- Journal: Journal of Computational Biology
- DOI: 10.1177/15578666261486472
- Source URL: <https://doi.org/10.1177/15578666261486472>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1177%2F15578666261486472>

Abstract: Multi-assembly methods rely at their core on a flow decomposition problem, namely, decomposing a weighted graph into weighted paths or walks. However, most results over the past decade have focused on decompositions over directed acyclic graphs (DAGs). This limitation has led to either purely heuristic methods or, in applications, transforming a graph with cycles into a DAG via preprocessing heuristics. In this article, we show that flow decomposition problems can also be solved in practice on general graphs with cycles via a framework that yields fast and flexible mixed-integer linear programming (MILP) formulations. Our key technique relies on the graph-theoretical notion of a dominator tree , which we use to find all safe sequences of edges that are guaranteed to appear in some walk of any flow decomposition. We generalize previous results from DAGs to cyclic graphs by showing that maximal safe sequences correspond to extensions of common leaves of two dominator trees, and that all such sequences can be found in time linear in their size. Using these, we can accelerate MILPs for any flow decomposition into walks in general graphs by setting suitable variables encoding solution walks to (at least) 1 and by setting to 0 other walk variables that are nonreachable to and from safe sequences. This reduces model size and eliminates costly linearizations of MILP variable products. We experiment with three decomposition models (minimum flow decomposition, least absolute errors, and minimum path error) on four bacterial datasets. Our preprocessing enables up to 1000-fold speedups and solves many instances that would otherwise time out in under 30 seconds. We thus hope that our dominator-based MILP simplification framework, together with the accompanying software library, can serve as building blocks for multi-assembly applications.

## Fear, refuge, and adaptive harvesting stabilise a five-dimensional tri-trophic predator–prey model exhibiting a transcritical bifurcation
- Source: Scientific Reports (journals)
- Date: 2026-09-18T00:00:00+00:00
- Categories: Evolution & metagenomics, Mathematical biology & statistics
- Authors: G. Ramraj, T. Poornima
- Journal: Scientific Reports
- DOI: 10.1038/s41598-026-67595-7
- Source URL: <https://doi.org/10.1038/s41598-026-67595-7>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1038%2Fs41598-026-67595-7>

Abstract: We propose and analyse a five-dimensional continuous-time tri-trophic predator–prey model comprising a prey species, an intermediate predator, a top predator, and two adaptive harvesting efforts. The model integrates three ecologically relevant mechanisms: fear-induced reduction in prey growth, prey refuge, and dynamic harvesting effort governed by economic payoffs. Positivity and uniform boundedness of all solutions are rigorously established. All biologically feasible equilibrium points are identified, and their local asymptotic stability conditions are derived via Jacobian linearisation and the Routh–Hurwitz criterion for the interior coexistence equilibrium. A Lyapunov candidate is analysed, and extensive numerical integration from a wide range of initial conditions indicates that the interior equilibrium is globally attracting whenever it exists. The system undergoes a transcritical bifurcation at the top-predator-free planar equilibrium with respect to the predator attack rate: the top predator invades and a stable interior coexistence equilibrium emerges once the attack rate exceeds a critical threshold, established analytically through a transversal eigenvalue-crossing argument (exchange of stability) and located numerically. Extensive simulations confirm this analysis: below the threshold the top predator is excluded and the system settles to the top-predator-free state, while above it all trajectories converge to coexistence. Within the parameter ranges studied, the interior equilibrium remains asymptotically stable, with no Hopf bifurcation or sustained oscillation observed, indicating that fear, refuge, and adaptive harvesting jointly stabilise the tri-trophic chain. We further quantify how fear intensity, refuge level, and harvesting effort modulate coexistence densities and the invasion threshold.

## FreqFuseNet: Scale-Normalized Dual-Frequency Fusion for Thin-Wall Head-and-Neck OAR Segmentation
- Source: medRxiv (preprints)
- Date: 2026-09-18
- Categories: Biological imaging
- Authors: Chen, W.-Y., Lin, G.-Y., Wan, S.-Y.
- DOI: 10.64898/2026.07.09.26357642
- Source URL: <https://doi.org/10.64898/2026.07.09.26357642>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.07.09.26357642>

Abstract: Accurate segmentation of thin-wall organs-at-risk (OARs) - the cochlea, vestibular semicircular canals, internal auditory canal, tympanic cavity, and middle ear - is clinically relevant for head-and-neck radiotherapy planning, yet these small, thin-wall structures remain among the most challenging targets for automated delineation. Dual-frequency feature fusion is a promising direction for boundary-sensitive representation, but under the investigated FP16 FFT-FcaNet setting, we observe an approximately 863x activation-scale mismatch between the FFT and FcaNet branches, causing a nominal 5% residual coefficient to behave as an approximately 43x dominant term. We propose FreqFuseNet, which resolves this mismatch by normalizing the FcaNet branch to the FFT activation scale before residual injection with a fixed low-amplitude coefficient (beta = 0.05), restoring beta as an interpretable 5% residual-amplitude coefficient relative to the FFT feature scale. Under a controlled binary per-OAR ROI protocol on the SegRap2023 head-and-neck CT benchmark across 10 clinically prioritized thin-wall OARs, FreqFuseNet achieves Dice of 0.849, HD95 of 0.824 mm, and SDice@1mm of 0.959 in the primary seed, with comparable performance in an independent second seed (Dice 0.843, HD95 0.823 mm). FreqFuseNet yields statistically significant case-level aggregate improvements over 3D U-Net and MedNeXt-S (Wilcoxon p < 0.01 and p < 0.05, respectively), using only 29.7 M parameters versus 414.6 M for the full wavelet baseline.

## From Movement to Spread: Generating Livestock Contact Networks that Preserve Infection Dynamics
- Source: bioRxiv (preprints)
- Date: 2026-09-18
- Authors: Qin, T., Atamer Balkan, B., Schmid, B. V., ten Bosch, Q. A.
- DOI: 10.64898/2026.09.15.751711
- Source URL: <https://doi.org/10.64898/2026.09.15.751711>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.15.751711>

Abstract: Animal trade links livestock holdings through contacts that change from day to day, making movement networks vital for epidemic analysis and control. Official movement records reveal these transmission routes, but privacy concerns often restrict access to the original data. Furthermore, epidemiological studies frequently require synthetic networks that accurately preserve the structural and temporal dynamics driving disease spread. Here we introduce NetForge, a mechanism-informed generative framework that learns recurring sender--receiver roles from movement and farm information of Dutch national pig-movement records. We compared generators with varying structural and temporal constraints. Among them, the Operational Stochastic Block Model regime performed best by combining learned trade partner structure with constraints on how contacts persist, return, or first appear. It closely matched the accumulation of potential spreading routes in the observed network and reproduced its simulated disease transmission trajectories. Together, we show that pairing trade structure with temporal constraints is essential for capturing infection dynamics. Because NetForge models trade data as a sequence of time-framed networks, it aligns well with routinely collected movement records and provides practical guidance for building more realistic synthetic movement networks from empirical trade records.

## GABAergic interneuron pathology in schizophrenia: a systematic review and meta-analysis across cell-types, brain areas, and cortical layers
- Source: bioRxiv (preprints)
- Date: 2026-09-18
- Categories: Single-cell & spatial
- Authors: Mulvey, A. G., Gabhart, K. M., Grent-T-Jong, T., Herculano-Houzel, S., Uhlhaas, P. J., Bastos, A. M.
- DOI: 10.1101/2025.05.23.655812
- Keywords: cell type, systematic review
- Source URL: <https://doi.org/10.1101/2025.05.23.655812>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1101%2F2025.05.23.655812>

Abstract: Introduction GABAergic interneurons are implicated in the pathophysiology of schizophrenia, yet evidence regarding the nature of deficits across brain areas and interneuron subtypes remains conflicting. We adopted a meta-analytic, multi-level linear modeling approach to identify interneurons, cortical layers, and brain areas involved, and the implications for circuit functions in schizophrenia. Methods Following PRISMA guidelines, we conducted a systematic search and meta-analysis from inception to November 2025, for studies examining parvalbumin, somatostatin, calbindin, and calretinin interneuron density or mRNA expression in schizophrenia. We included data from 44 studies, comprising 736 individuals with schizophrenia and 814 healthy controls. Non-cell-specific, non-human, or indirect proxy studies were excluded. Linear mixed-effects models quantified deficits while accounting for cortical layer, cell-type, and brain area, providing a map of interneuron pathology. We further analyzed changes in GABAergic interneurons to determine whether deficits preferentially target cortical layers, cell-types, and brain areas more associated with top-down or bottom-up processing. Results Parvalbumin and somatostatin interneurons showed robust reductions, particularly in layers 3/4, whilst calbindin and calretinin interneurons were less affected. Deficits were widespread across cortical and subcortical regions. Contrasts revealed that schizophrenia is characterized by interneuron deficits preferentially affecting bottom-up signaling -- notably parvalbumin and somatostatin interneurons in layers 3/4, which are critical for gamma-band synchronization and feedforward sensory processing. Discussion These findings provide the most comprehensive meta-analysis on GABAergic interneurons in schizophrenia to-date, as well as a novel perspective on circuit dysfunctions, with implications for computational models. These results highlight the need for more widespread sampling across the brain using methodologies that can pinpoint deficits in molecularly more precise ways.

## GENESIS-SHIELD: an interpretable ensemble for anomaly detection in CRISPR genomic-workflow security
- Source: Frontiers in Artificial Intelligence (journals)
- Date: 2026-09-18T00:00:00Z
- Categories: Genomics & sequence analysis, Tools & resources
- Authors: Prabakaran C., K. R
- Journal: Frontiers in Artificial Intelligence
- DOI: 10.3389/frai.2026.1893433
- External ID: 51b43bddbff870cf0d18bc2dc578481655d06be6
- Keywords: genomic
- Source URL: <https://doi.org/10.3389/frai.2026.1893433>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.3389%2Ffrai.2026.1893433>

Abstract: The rapid advancement of CRISPR-based gene editing has introduced digital-workflow security and integrity challenges: unauthorized modifications, temporal inconsistencies, and duplicated provenance records can compromise the auditability of editing logs. These are distinct from the biological risks of editing itself; our focus is the security of the digital record . Existing single-component detectors capture only one facet of these threats. We present GENESIS-SHIELD, a multi-component ensemble whose novelty lies in the integration of established techniques for CRISPR-workflow security rather than in any single new algorithm. It combines four components—a Hierarchical Blockchain Merkle Tree with Bloom filters and a train-registry integrity check, an adaptive cross-layer entropy/KL analyzer, an interpretable decision-tree ethics rule system, and a structural temporal-consistency detector—whose weights are set by Bayesian optimization of a recall-oriented (F 2 ) validation objective. We benchmark against classic (Isolation Forest, One-Class SVM, LOF) and modern (XGBoost, autoencoder, and Deep-SVDD) baselines trained on identical features, and report a leave-one-component-out ablation, bootstrap confidence intervals, and isotonic calibration. On a synthetic benchmark of 50,000 records (5% anomalies, eight categories), the ensemble attains AUC-ROC 0.982 (95% CI 0.973–0.990) and AUC-PR 0.841, with the proposed weighted detector reaching recall 0.978 at precision 0.500 (F 1 0.661). A supervised XGBoost on the same features is competitive-to-superior (AUC-PR 0.892, F 1 0.854), and the deep Attention-LSTM remains ineffective under class imbalance (F 1 0.096). Ablation shows the ethics and structural components carry the signal while the entropy component is redundant (removing it leaves metrics unchanged). Isotonic calibration reduces expected calibration error from 0.028 to 0.002. These results are a proof-of-concept on rule-defined synthetic data; the high per-category detection reflects the benchmark's construction and does not establish real-world security. Validation on real editing logs, adversarial testing, and multi-objective coverage remain necessary before deployment.

## GenoME: a MoE-based generative model for individualized, multimodal prediction and perturbation of genomic profiles
- Source: Nucleic Acids Research (journals)
- Date: 2026-09-18T00:00:00+00:00
- Categories: Genomics & sequence analysis, Single-cell & spatial
- Authors: Jiachen Wei, Yue Xue, Hao Chai, Yi Qin Gao
- Journal: Nucleic Acids Research
- DOI: 10.1093/nar/gkag902
- Source URL: <https://doi.org/10.1093/nar/gkag902>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1093%2Fnar%2Fgkag902>

Abstract: The non-coding genome operates through a complex, multiscale regulatory system where regulated gene expressions are closely associated with cell-type-specific histone modifications, transcription factor binding, and 3D conformation. Developing computational models that can integrate these patterns to predict and interpret the regulatory system remains challenging. Here, we present GenoME, a Mixture of Experts (MoE)-based generative model that uses DNA sequence and cell-type-specific ATAC-seq signals to predict a unified genomic profiles encompassing epigenomics, transcriptomics, and chromatin architecture at base-pair to kilobase resolutions. GenoME enables multiscale predictions for held-out genomic regions and, critically, generalizes to predict the full regulatory landscape of unseen or individualized cell types from a single ATAC-seq input. We equip GenoME with an in silico perturbation framework that accurately forecasts the multimodal consequences of genetic perturbations and identifies functional enhancer–promoter connections, outperforming specialized models like Activity-by-Contact. These predictions can also be used to decipher the transcription factor grammar of cell-type-specific enhancers. GenoME thus provides a versatile, all-in-one platform for generative modeling, cross-cell-type generalization, and causal mechanistic investigation of the multiscale regulatory genome.

## High-Throughput Observational Evidence Generation Using Linked Electronic Health Record and Claims Data
- Source: medRxiv (preprints)
- Date: 2026-09-18
- Authors: Coyle, J., Shah, N., Hubbard, A., Mukerji, A., Chappelka, M., Sanghavi, N., Gombar, S.
- DOI: 10.64898/2026.04.07.26350300
- Source URL: <https://doi.org/10.64898/2026.04.07.26350300>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.04.07.26350300>

Abstract: Background: Many consequential treatment decisions involve off-label or head-to-head choices in complex, comorbid populations routinely excluded from randomized trials, or other prospective analyses. For these decisions, comparative evidence is often entirely absent. Although observational data include these patients, findings are difficult to synthesize because studies differ in cohort definitions, confounder measurement, follow-up periods, and reported outcomes. Prior systems for large scale evidence generation have largely stopped at data preparation, limiting the usefulness of their outputs to decision makers. Methods: We developed a high-throughput evidence-generation workflow using linked EHR and claims data. A prespecified causal-measurement architecture was applied consistently across scenarios, including three post-index follow-up windows through two years; 28 comorbidities; 14 healthcare resource utilization categories; 30 laboratory measures with 57 binary thresholds; 43 adverse-event categories; and evaluated 1038 clinically important outcomes. Scalable collaborative targeted learning (C-TMLE) generated confounding-adjusted, actionable treatment comparisons, with unadjusted estimates reported for transparency. Results: Across 135 clinical scenarios, the workflow generated 210,584,509 outcome evaluations. Each evaluation represented an outcome, follow-up window, treatment contrast, population stratum, and estimator, accompanied by diagnostic information. These results were synthesized into 7,500 narrative summaries and underwent structured clinical and statistical quality control. Conclusions: Standardized, high-throughput workflows can move evidence generation beyond fragmented individual studies toward comprehensive evidence packages. By making treatment-effect heterogeneity visible across clinically meaningful subgroups, this shared evidence base can support precision medicine and reduce redundant stakeholder-specific studies.

## HisTrader identifies nucleosome-free regions within ChIP-based profiling of histone post-translational modifications.
- Source: Cell reports methods (journals)
- Date: 2026-09-18
- Categories: Genomics & sequence analysis, Tools & resources
- Authors: Eftyhios Kirbizakis, Yifei Yan, Ansley Gnanapragasam, Juliana Cavalcante de Moura, Xiaoyang Zhang, Swneke D Bailey
- Journal: Cell reports methods
- DOI: 10.1016/j.crmeth.2026.101607
- External ID: 42759522
- Source URL: <https://doi.org/10.1016/j.crmeth.2026.101607>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1016%2Fj.crmeth.2026.101607>

Abstract: Enhancers and promoters regulate cell identity through the binding of transcription factors (TFs) to specific DNA motifs within accessible chromatin. These regulatory regions are often identified using chromatin immunoprecipitation (ChIP)-based assays targeting histone modifications, such as ChIP sequencing (ChIP-seq), HiChIP, and proximity-ligation-assisted ChIP-seq (PLAC-seq). However, the large size of the enriched regions, or peaks, can make it difficult to pinpoint the precise DNA sequence where TFs act or where trait- or disease-associated variants exert their effects. We present HisTrader, a computational approach that identifies nucleosome-free regions (NFRs) within ChIP-based profiling of histone modification peaks, which reduces the target sequence length for motif discovery and genetic variant prioritization. By focusing on TF accessible sites, HisTrader improves motif-based detection of the regulatory mechanisms linking cellular transitions and disease states. In addition, HisTrader enables more accurate characterization of regulatory elements affected by genetic variation contributing to disease.

## How optimal control of cellular cost shapes population-level tumor growth dynamics
- Source: bioRxiv (preprints)
- Date: 2026-09-18
- Categories: Mathematical biology & statistics
- Authors: Shrestha, P., George, J. T.
- DOI: 10.64898/2026.09.17.751892
- Source URL: <https://doi.org/10.64898/2026.09.17.751892>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.17.751892>

Abstract: Tumor progression is often modeled as a passive response to external therapy or immune pressure, but tumor populations may also exhibit population-level regulation of proliferation and apoptosis. We develop a continuous-time Markov decision framework in which a controlled birth--death process represents a tumor population modulating the balance between proliferation and susceptibility to apoptosis in the presence of extrinsic death pressure. We examine threshold and quadratic costs, an unbounded linear reward, and constrained linear and quadratic formulations to determine how objective structure shapes optimal policies and induced population drift. Threshold and quadratic penalties generate restoring dynamics, with transitions from growth to suppression and regions of near-neutral drift associated with regulated or near-dormant behavior. An unbounded linear reward instead produces sustained or near-neutral growth without a restoring regime. Under constraints, a linear reward expands the region of positive drift as capacity increases, whereas a quadratic reward generates restoring, logistic-like drift around an interior population scale. These results show that regulated tumor dynamics depend on how growth incentives, extrinsic death pressure, penalties, and constraints scale with population size.

## Img2EEG: A Scalable and Interpretable Encoding Framework for Simulating Human EEG Responses to Visual Inputs
- Source: bioRxiv (preprints)
- Date: 2026-09-18
- Categories: Computational neuroscience
- Authors: Lu, Z., Golomb, J. D.
- DOI: 10.64898/2026.09.16.751610
- Source URL: <https://doi.org/10.64898/2026.09.16.751610>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.16.751610>

Abstract: Understanding how visual information processing unfolds over time requires models that not only predict neural responses but also expose the representations that support them and generalize beyond sampled stimulus spaces. Here we introduce Img2EEG, a participant-specific image-to-EEG encoding framework that integrates hierarchical visual and semantic representations to generate temporally resolved multichannel EEG responses. Trained on THINGS EEG2, Img2EEG generalized to unseen images while preserving stimulus-specific and participant-specific response structure. Controlled perturbations of internal representations and visual inputs revealed distinct temporally structured contributions of visual and semantic information, and in silico experiments reproduced classic human neural responses, such as the face-sensitive N170, while enabling targeted representational interventions. Scaling Img2EEG to 1.28 million ImageNet images produced over 12 million synthetic EEG responses that supported cross-dataset visual reconstruction and improved the behavioral alignment of an artificial vision model. Img2EEG provides an interpretable and scalable framework for experimentally manipulable modeling of visual neural dynamics.

## Improving Lipid Identification and Quantification: Chromatogram Deconvolution for LC-MS/MS Workflows
- Source: Bioinformatics (journals)
- Date: 2026-09-18T00:00:00+00:00
- Categories: Tools & resources
- Authors: Felix Niedermaier, Denise Wolrab-Frühauf, Robert Ahrends, Dominik Kopczynski
- Journal: Bioinformatics
- DOI: 10.1093/bioinformatics/btag695
- Source URL: <https://doi.org/10.1093/bioinformatics/btag695>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1093%2Fbioinformatics%2Fbtag695>
- Code: <https://gitlab.com/computational-multiomics/mixture-model-deconvolution>

Abstract: Motivation Lipidomics relies on mass spectrometry-based workflows to identify and quantify complex lipid species. Due to the modular architecture of lipids, including headgroups, backbones, and fatty acyl chains, distinct precursor ions often produce isobaric or identical fragment ions. This problem is amplified in data-independent acquisition (DIA), where wide isolation windows (e.g., 25 Da) allow co-eluting precursors with different m/z values to generate highly chimeric MS/MS spectra. Consequently, fragments originating from multiple precursors, including isobars, isomers, and lipid-class-specific ions, are merged into a single MS/MS spectrum. Current lipid identification strategies often process such chimeric spectra in an uncontrolled manner, assigning them to one or more candidate lipids, thereby increasing false-positive identifications. Results Here, we introduce an algorithm that deconvolutes chimeric MS/MS spectra and chromatograms by exploiting their temporal correlation with associated precursor chromatographic profiles, independent of elution peak shape. Using simulated and real experimental lipidomics data, we demonstrate that this approach substantially improves lipid fragment assignment, reduces false-positive identifications, and enables more reliable fragment-level quantification, leading to more robust downstream statistical analyses and biological interpretations. Availability and Implementation Source code of the software library: GitLab (Apache 2.0 License): https://gitlab.com/computational-multiomics/mixture-model-deconvolution; Data: Zenodo (Apache 2.0 License): https://doi.org/10.5281/zenodo.21218594. Supplementary information Supplementary data are available at Bioinformatics online.

## Improving multi-trait genomic prediction using synthetic traits from hyperspectral data based on co-heritability.
- Source: TAG. Theoretical and applied genetics. Theoretische und angewandte Genetik (journals)
- Date: 2026-09-18
- Authors: Ashmita Upadhyay, Ruhana Azam, Meilu Yuan, Stefan Ivanovic, John N Ferguson, Rachel E Paul, Sanmi Koyejo, Mohammed El-Kebir, Alexander E Lipka, Andrew D B Leakey, Samuel B Fernandes
- Journal: TAG. Theoretical and applied genetics. Theoretische und angewandte Genetik
- DOI: 10.1007/s00122-026-05334-2
- External ID: 42760412
- Source URL: <https://doi.org/10.1007/s00122-026-05334-2>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1007%2Fs00122-026-05334-2>

Abstract: Synthetic traits, wavelength ratios selected by co-heritability, raised multi-trait genomic predictive ability for leaf nitrogen and specific leaf area in sorghum by up to 17% over single-trait models. Genomic prediction (GP) is an essential tool in plant breeding as it can accelerate cultivar development by predicting the performance of unphenotyped lines. When using single-trait GP models, the precision of prediction is constrained by the heritability of the target trait. Multi-trait genomic prediction can be used to improve accuracy but requires identifying secondary traits with high heritability and genetic correlation with target traits. This study assessed the efficiency of multi-trait genomic prediction models powered by secondary traits derived from high-throughput phenotyping data when predicting leaf nitrogen content (N) and specific leaf area (SLA) in diverse sorghum accessions. We hypothesized that wavelength ratios from hyperspectral data could serve as synthetic secondary traits. Therefore, we developed models for direct measures of N and SLA, plus partial least squares regression (PLSR) predictions of them (Leaf N-PLSR, SLA-PLSR), totaling four target traits. Three synthetic traits (S1, S2, S3), each a ratio of two wavelengths within the hyperspectral data, were identified based on high co-heritability with target traits. Single-trait (genomic best linear unbiased prediction (GBLUP) served as the baseline model, followed by multi-trait GBLUP models combining synthetic and target traits. Model performance was assessed using fivefold cross-validation under single-trait, CV1, and CV2 schemes. Our approach improves multi-trait genomic prediction of target traits by using synthetic traits with no intrinsic biological meaning selected through co-heritability estimation. This demonstrates that there's more useful information in the spectra than is typically utilized, and this information can be leveraged to improve multi-trait prediction of a target trait.

## Incomplete references leave bulk deconvolution targets non-identifiable, but identification has a computable precision price
- Source: bioRxiv (preprints)
- Date: 2026-09-18
- Categories: Genomics & sequence analysis, Mathematical biology & statistics
- Authors: Jiang, H., Gao, F., Liu, P., Wu, Y., Jie, Y., Li, Y., Jiang, Y.
- DOI: 10.64898/2026.09.10.750775
- Source URL: <https://doi.org/10.64898/2026.09.10.750775>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.10.750775>

Abstract: Background. Reference-based deconvolution estimates cell-type proportions from bulk profiles. Incomplete references compromise these estimates, yet many remedies return point estimates. We separate the operator provenance that reproduces an estimate from the information and precision that identify its target. Results. Two operator histories at one reduced reference give different estimators: 28.2% of 196,420 sample-deletion pairs differed by >0.1 total variation; locking every learned component made paths identical. Across nine cohorts, operator history reversed 160 of 952 association signs, 31 with a significant path, and changed significance for 126. Observationally equivalent completions can fill the open simplex and reverse retained-type rankings. Under a joint zero-exposure condition, shared structure leaves inherited bounds unchanged; one to six uncalibrated views also gave identical bounds. A profile library contracted estimator-output envelopes by 98.93% yet covered the full-reference effect for 42.77%, with 19 wrong-sign certificates; conditional sharp bounds stayed at \[-1, 1\]. Treating RNA yields as exact collapsed intervals to points covering none of seven flow-measured targets: contraction without coverage is false certainty. Calibrated cross-modal anchors contract width to 0.51 at six types but certify no valid sign. Exact RNA yields identify cell fractions from RNA contributions; a decisive sign in blood requires \{+/-\}2.6% proxy accuracy with near-total contraction of donor-heterogeneity and dynamic-range envelopes. Conclusions. Incomplete-reference deconvolution is an identification problem, not only an estimation problem. Remedies must be scored on shrinkage, coverage, certification and false certification against held-out targets. fitdrop implements this scoring and a precision frontier for planning decisive measurements.

## Independent noise realizations enable morphologically agnostic image reconstruction in single-photon-sensitive microscopy
- Source: bioRxiv (preprints)
- Date: 2026-09-18
- Categories: Biological imaging
- Authors: Cuneo, L., Zunino, A., Le, L., Agostini, S., Salzo, S., Calatroni, L., Pontil, M., Vicidomini, G.
- DOI: 10.64898/2026.09.12.751176
- Source URL: <https://doi.org/10.64898/2026.09.12.751176>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.12.751176>

Abstract: Advances in single-photon sensitive detectors are rapidly expanding the adoption of photon-counting fluorescence microscopy. Under the Poisson photon-counting statistics, iterative Richardson-Lucy (RL) -type algorithms are statistically optimal for image deconvolution but suffers from a fundamental semi-convergent behaviour: prolonged iterations inevitably amplify noise, requiring heuristic early stopping or regularisation typically based on assumptions about object morphology. Here we present a morphology-agnostic regularisation framework for RL deconvolution that exploits independent noise realisations instead of structural object priors. Preserving the Poisson statistics of the acquisition process, we formulate the Regularized by Noise (RbN): a regularized variational framework and its corresponding iterative minimization algorithm that exploit the statistical consistency of independent noise realisations to distinguish reproducible image features from stochastic noise. The regularisation strength is selected automatically using the Poisson residual whiteness principle, resulting in a fully data-driven reconstruction without heuristic parameter tuning. Photon-timing-resolved systems naturally provide the independent noise realisations exploited by our framework, whereas computational photon splitting provides a statistically equivalent implementation for photon-counting systems. We validate the approach experimentally using photon-timing-resolved confocal microscopy and image scanning microscopy across eight morphologically distinct subcellular targets, and further demonstrate its applicability to photon-counting microscopes through computational photon splitting. Across imaging modalities, detector technologies, biological structures and signal-to-noise regimes, our method eliminates RL semi-convergence, removes sensitivity to the stopping criterion, and consistently outperforms conventional RL while preserving fine structural detail. More broadly, our results establish independent noise realisations as a general source of morphology-agnostic regularisation. Although demonstrated here for SPAD-based laser-scanning microscopy and Poisson statistics, the underlying principle could be extended to other imaging modalities and, more generally, to statistical inverse problems through appropriate noise-specific formulations.

## Inferring leakage in imports of frozen seafood allowing for censoring, testing accuracy and a minimum positive prevalence
- Source: Biometrics (journals)
- Date: 2026-09-18T00:00:00+00:00
- Authors: Sumonkanti Das, Robert G Clark, Mahdi Parsa, Belinda Barnes
- Journal: Biometrics
- DOI: 10.1093/biomtc/ujag155
- Source URL: <https://doi.org/10.1093/biomtc/ujag155>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1093%2Fbiomtc%2Fujag155>

Abstract: Many countries screen import consignments to guard against the entry of exotic pests, contaminants, and pathogens. A widely used strategy is to pool individual units into groups and test each group for presence or absence of contamination. Consignments are typically rejected if there are any detections. Screening samples are commonly designed to give a high chance of detection assuming a design prevalence. What is less common, however, is to analyze the history of testing outcomes to infer how many accepted consignments contain contaminated units, and the number of contaminated units—jointly referred to as “leakage.” We build on existing censored beta-binomial models to answer these questions for the importation of frozen seafood into Australia, allowing for unknown test sensitivity and specificity. We present new theory and empirical results demonstrating that specificity is identifiable from test data in this context, but sensitivity is not. We also develop a new class of models in which consignment propensity is either zero or above a minimum positive prevalence threshold, motivated by our case study but applicable more widely. Past testing data are modeled under multiple scenarios using both hierarchical Bayes and maximum likelihood methods, revealing new insights into the risk of leakage.

## Inferring the latent network of pairwise mutualistic preferences from observed plant-pollinator interactions
- Source: bioRxiv (preprints)
- Date: 2026-09-18
- Categories: Evolution & metagenomics, Mathematical biology & statistics
- Authors: Federici, L., Matechou, E., Iacopini, I.
- DOI: 10.64898/2026.09.17.751682
- Source URL: <https://doi.org/10.64898/2026.09.17.751682>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.17.751682>

Abstract: Plant-pollinator communities are typically represented as bipartite networks, whose edges are taken directly from field records of visits. These visits, however, are only a proxy for the object of ecological interest: the latent mutualistic preference between two species. While counts are shaped by preference, they also carry confounding factors such as species abundances, sampling effort, and site- or time-specific conditions. We introduce a hierarchical Bayesian framework that treats visit counts as a realisation of a Poisson process and, on the log scale, decomposes the corresponding pairwise rate into a baseline (community-wide activity together with sampling effort), individual species effects representing abundance, and pairwise mutualistic preferences. The model extends to data replicated across sites and time points, and to the inclusion of environmental or experimental covariates. Because the whole system is fitted jointly, we obtain posterior not only for the preferences but for every latent quantity, each carrying ecological signal of its own, with uncertainty propagated through every level of the model, down to any derived network metric. On synthetic data, we show that common practices, such as reading preferences off raw counts or aggregating replicated observations into a single network, confound abundance with preference. In contrast, our framework recovers the underlying preference structure. On empirical datasets, including a seasonal multi-site pollination study where urbanisation level enters as a covariate, the inferred preference network departs markedly from the observed visits, revealing structure hidden in the raw counts: how species vary across sites and time, and which parts of the community respond most to the covariate. When communities are compared along the urbanisation gradient, standard network metrics on the preference layer revise the conclusions drawn from visits alone. The framework offers a principled way to move from networks of observed visits to networks of underlying mutualistic preferences, carrying uncertainty from the data through to the ecological conclusions and accommodating the spatial, temporal, and covariate structure of modern plant-pollinator datasets. Because it acts on the foundational step of network construction, its implications are broad, placing network-based approaches on firmer ground.

## Input-space geometry shapes adaptation dynamics underlying repetition suppression in a neural network model of relatedness priming
- Source: bioRxiv (preprints)
- Date: 2026-09-18
- Categories: Computational neuroscience
- Authors: Bartolini, D., Reber, T. P., Tchumatchenko, T., Voigt, M.
- DOI: 10.64898/2026.09.13.751213
- Source URL: <https://doi.org/10.64898/2026.09.13.751213>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.13.751213>

Abstract: A key function of the human brain is its ability to dynamically adapt to novel contexts by integrating prior experience. While neural adaptation is widely observed across cortical systems, the microcircuit-level mechanisms governing its intensity and temporal dynamics remain unclear. To bridge this gap, we develop a recurrent network of adaptive exponential integrate-and-fire neurons governed by a triplet spike-timing-dependent plasticity rule, designed to reproduce neural dynamics recorded via intracranial electrophysiology and single-unit recordings from the medial temporal lobe of neurosurgical patients performing a priming task. Using input organizations inspired by the experimental paradigm, we systematically vary input geometry to investigate its impact on adaptation dynamics. We find that neural adaptation emerges from recurrent dynamics shaped by learned connectivity, with both its magnitude and temporal profile depending on the geometry of the input space. This dependence gives rise, at the simulated single-neuron level, to a continuum of response regimes ranging from sharpening-like to fatiguing-like dynamics. Sharpening-like responses dominate when inputs exhibit high within-meta-category similarity and strong between-meta-category separation, whereas fatiguing-like responses emerge under the opposite regime.

## Integrated single-cell and bulk tissue analyses reveal distinct macrophage subtypes and a candidate prognostic signature in colorectal cancer: implications for tumor immune characterization
- Source: Frontiers in Immunology (journals)
- Date: 2026-09-18T00:00:00Z
- Categories: Genomics & sequence analysis, Single-cell & spatial, Proteins & structural biology, Systems & networks, Biological imaging
- Authors: Hui Chen, Zhi-Peng Li, Zhen Lin, Li-Jun Wan, Yu-Hang Gong, Zhi-Bin Lv, Jin-Feng Hu, Dun Pan
- Journal: Frontiers in Immunology
- DOI: 10.3389/fimmu.2026.1902264
- External ID: ad97496bb6dc0d1f2bf6914a4ce6c247a3b15ce7
- Keywords: rna, transcriptomic, single cell, pathway, leukocyte
- Source URL: <https://doi.org/10.3389/fimmu.2026.1902264>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.3389%2Ffimmu.2026.1902264>

Abstract: Colorectal cancer (CRC) represents a major global health burden, marked by high morbidity and mortality rates that place a considerable strain on healthcare systems. This study leveraged integrated bioinformatic analyses, single-cell RNA sequencing, and clinical sample validation to investigate the role of macrophage-related genes (MRGs) in CRC, with the goal of deepening our understanding of the complex interplay within the tumor microenvironment. Differential expression analysis comparing CRC tumor and normal tissues in the TCGA-COADREAD cohort identified 1,962 differentially expressed genes (DEGs). In parallel, a predefined set of 1,719 MRGs was curated from public databases and the literature to define the macrophage-related biological context. By integrating the bulk transcriptomic DEGs, single-cell macrophage subtype-specific genes, and the predefined MRG set, we identified eight hub macrophage-related DEGs (MRDEGs) implicated in CRC. Functional enrichment analysis of these eight MRDEGs revealed significant roles in immune-regulatory processes, including leukocyte chemotaxis, eosinophil chemotaxis, chemokine receptor binding, and the chemokine signaling pathway. From these eight hub MRDEGs, we selected CCL24 and MMP12 via LASSO-Cox regression to construct a prognostic risk model. Immune infiltration analysis using CIBERSORT revealed significant differences ( p < 0.05) in the abundance of nine immune cell types (including M0/M2 Macrophages, CD8+ T cells, T follicular helper cells, regulatory T cells (Tregs), memory B cells, monocytes, eosinophils, and neutrophils) between the risk groups defined by this macrophage-related signature. Drug sensitivity analysis further revealed that the high-risk group was significantly less responsive to sorafenib ( p < 0.05). The model provides candidate molecular clues for further investigation of the CRC immune microenvironment and prognostic stratification related to macrophage biology. We further performed experimental validation of the differentially expressed genes using clinical CRC samples and assessed their expression at both the transcriptional and protein levels. This study provides a candidate prognostic stratification model that requires further validation in independent cohorts and prospective studies, offering a macrophage-related prognostic clue for future investigations.

## Integrating Multi-Modal Biological Knowledge via Contrastive Dual-View Graph Learning for Phosphorylation Site-Disease Association Prediction
- Source: Bioinformatics (journals)
- Date: 2026-09-18T00:00:00+00:00
- Categories: Proteins & structural biology
- Authors: Xiangyu Chen, Lijun Quan, Yexuan Mao, Siyuan Wang, Guozheng Zhang, Siqi Li, Yelu Jiang, Liangpeng Nie, Tingfang Wu, Qiang Lyu
- Journal: Bioinformatics
- DOI: 10.1093/bioinformatics/btag690
- Source URL: <https://doi.org/10.1093/bioinformatics/btag690>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1093%2Fbioinformatics%2Fbtag690>
- Code: <https://github.com/ljquanlab/CDVGL-PDA>

Abstract: Motivation Accurately characterizing the associations between phosphorylation sites (psites) and diseases is essential for elucidating pathogenic mechanisms and guiding therapeutic discovery. However, existing computational approaches for phosphorylation site–disease association (PDA) prediction remain limited, as they often overlook the integration of biological knowledge across multiple modalities. Results We propose CDVGL-PDA, a contrastive dual-view graph learning framework for PDA prediction. Specifically, we construct a multi-modal heterogeneous graph encompassing nine node types and ten edge types, enabling comprehensive representation of phosphorylation-centric biological networks. For phosphorylation site nodes, CDVGL-PDA incorporates sequence-derived embeddings, while disease nodes are initialized with semantic features from BioBERT. The framework then performs dual-view heterogeneous graph encoding, aligns representations through contrastive learning, and adaptively integrates them via attention-based fusion to capture informative embeddings. Extensive evaluations demonstrate CDVGL-PDA’s strong predictive performance across balanced, imbalanced, and low-similarity datasets. Analysis of embeddings reveals that the model effectively captures latent biological relationships. Ablation and visualization studies validate the contributions of each module, while case studies highlight its ability to uncover potential PDAs, illustrating its promise for advancing disease mechanism research and therapeutic target discovery. Availability and Implementation The source code and datasets for CDVGL-PDA are publicly available in the GitHub repository at https://github.com/ljquanlab/CDVGL-PDA/. Supplementary information Supplementary data are available at Bioinformatics online.

## Integration of a smooth mesh-based contact pressure model into tracking and predictive simulations
- Source: bioRxiv (preprints)
- Date: 2026-09-18
- Authors: Harba, M., Serrancoli, G.
- DOI: 10.64898/2026.09.14.751442
- Source URL: <https://doi.org/10.64898/2026.09.14.751442>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.14.751442>

Abstract: Musculoskeletal simulations are widely used to estimate joint loading, yet most musculoskeletal models estimate knee contact forces as resultant forces or as normal medial and lateral contact forces, without resolving pressure distributions across the articular surfaces. This paper presents a smoothed mesh-based knee contact pressure model that computes continuously differentiable tibiofemoral contact pressures, enabling its direct integration into full-body movement simulations. Built on an elastic foundation formulation, the model introduces smooth approximations ensuring that all contact functions and their derivatives remain continuous throughout the simulation. A systematic sensitivity analysis was performed across five key parameters: mesh resolution, joint damping and three smoothing parameters. Tracking simulations across eight gait trials demonstrated that the nominal configuration achieved mean RMSE values for medial and lateral knee contact forces of 51.6 N and 75.2 N, respectively, with a mean RMSE for joint angles of 1.45\{degrees\} and (r = 0.97), converging in less than three hours on a standard computer. Mesh resolution was identified as the dominant factor that affected both accuracy and convergence, while damping variations had negligible influence. As a proof of concept, the model was also incorporated into predictive simulations, demonstrating that increasing the weight on the contact pressure term in the cost function leads to reduced tibiofemoral loading, particularly in the lateral compartment. The proposed formulation provides a computationally efficient and numerically robust framework for simulating knee contact mechanics within full-body musculoskeletal models.

## Interpretable spherical geometry of single-cell state transitions from dominant principal components
- Source: bioRxiv (preprints)
- Date: 2026-09-18
- Categories: Genomics & sequence analysis, Single-cell & spatial
- Authors: Yuan, L., Li, X., Le, M., Hicks, S. C., Deshpande, A., Taube, J. M., Szalay, A. S.
- DOI: 10.64898/2026.09.11.751061
- Source URL: <https://doi.org/10.64898/2026.09.11.751061>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.11.751061>

Abstract: Single-cell RNA-seq atlases are commonly explored with nonlinear embeddings that preserve neighborhoods but provide limited coordinate-level interpretation. We asked whether projecting the dominant principal components (PCs) of single-cell gene expression onto a unit sphere would yield an interpretable coordinate system. SPHERE-PCA L2-normalizes the first three PC coordinates, aligns a biologically defined root to the north pole, and represents each cell by three coordinates: root-aligned geodesic distance (\{theta\}), angular position (\{phi\}), and pre-projection radial magnitude (r). Across developmental and disease-associated datasets, this representation reveals structured spherical geometry, ranging from near-great-circle trajectories to multi-arc manifolds. In developmental atlases, root-aligned geodesic distance increases as CytoTRACE-inferred stemness decreases, while gene-coordinate analyses separate programs associated with angular position from those associated with radial magnitude. Fixed-loading perturbations decompose each gene's effect on cell position into progression, branch- or state-position, and radial activity components. SPHERE-PCA therefore provides a deterministic, loading-preserving coordinate framework for interpreting dominant transcriptomic variance and establishing a transparent geometric coordinate framework for perturbation analysis and virtual-cell models.

## Interrogating contrastive learning embeddings for structure-based virtual screening: a case study on DrugCLIP
- Source: bioRxiv (preprints)
- Date: 2026-09-18
- Categories: Proteins & structural biology
- Authors: Sanchez Utges, J., Jones, D. T., Orengo, C.
- DOI: 10.64898/2026.09.16.752029
- Source URL: <https://doi.org/10.64898/2026.09.16.752029>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.16.752029>

Abstract: Virtual screening has become central to early-stage drug discovery, and structure-based approaches have recently been reframed as a retrieval problem through contrastive learning methods such as DrugCLIP, which project protein pockets and ligands into a shared embedding space. However, what these abstract representations exactly encode, and how they relate to conventional notions of structural and chemical similarity, remains unclear. Here we present a systematic dissection of DrugCLIP's latent space. We show that its pocket embeddings, despite not being explicitly trained for the task, set a new state of the art in pocket similarity search while running over 100 times faster than existing structural descriptors, and that this embedding space is structurally coherent and robust to conformational variation. Ligand embeddings, by contrast, encode a pocket-aware notion of chemical similarity that only partially mirrors fingerprint-based measures. Using a rigorous de-leakage benchmark, we further show that DrugCLIP generalises to unseen proteins and chemistries rather than memorising training data, recovering the correct bound ligand within the top 1% of 50,000 candidates for 55-75% of novel targets. Performance nonetheless declines under increasingly realistic screening conditions, a drop attributable to sidechain reorientation across apo, holo and AlphaFold-derived structures, and to residue mismatch when using predicted pockets. These findings clarify the practical boundaries of DrugCLIP's applicability, identify pocket prediction accuracy as a key factor for improving performance, and offer a transferable framework for interpreting the latent spaces of related contrastive pocket-ligand encoders. Together, these results support the improvement of existing methods and the development of a new generation of contrastive screening approaches.

## iPscDB: a comprehensive platform for plant single-cell transcriptomic data integration and analysis
- Source: Nucleic Acids Research (journals)
- Date: 2026-09-18T00:00:00+00:00
- Categories: Genomics & sequence analysis, Single-cell & spatial, Tools & resources
- Authors: Peng Lu, Jingjing Jin, Jiemeng Tao, Linggai Cao, Shizhou Yu, Sujie Wang, Huan Su, Qiao Wang, Wentao Cui, Runtong Hou, Zefeng Li, Jianfeng Zhang, Yalong Xu, Yangyang Wu, Xuwu Sun, Peijian Cao
- Journal: Nucleic Acids Research
- DOI: 10.1093/nar/gkag909
- Source URL: <https://doi.org/10.1093/nar/gkag909>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1093%2Fnar%2Fgkag909>

Abstract: The rapid advancement of single-cell technologies has significantly enhanced our ability to investigate cellular heterogeneity within plant tissues. However, deciphering these intricate cellular landscapes requires processing high-dimensional gene expression matrices and integrating diverse datasets to enable accurate marker selection, cell identification, and other complex computational operations. These processes typically require broad programming expertise, posing a challenge for researchers with a limited computational background. To address this, we present integrated Plant single-cell Database (iPscDB), an integrated and multifunctional platform that facilitates the integration and analysis of plant single-cell data. iPscDB combines 4 688 428 cells and 288 139 curated cell markers derived from 946 experiments across 38 plant species. Wherever raw data were available, datasets were reprocessed through a single uniform pipeline, and both integration quality and automated cell-type annotation were benchmarked quantitatively. The platform introduces a Marker Confidence Level scheme that grades each cell-type marker by the strength and independence of its supporting evidence (from manually curated classic markers to database-derived associations), allowing users to judge marker reliability directly. The platform also provides a user-friendly online analysis pipeline and modules capable of processing raw FASTQ files or Cell Ranger-processed files. Users can configure parameters via an intuitive interface and utilize an integrated image editor to customize visualization outputs. Additionally, iPscDB supports various analyses, including cross-species gene expression, electronic Single-Cell Pictograph, and developmental trajectory. By streamlining the complex workflows of single-cell transcriptomics, iPscDB offers a practical and accessible resource for researchers with diverse technical backgrounds. iPscDB is accessible at https://www.tobaccodb.org/ipscdb/homePage.

## Joint inference of paired dynamical gene regulatory networks reveals distinct cell-state landscapes of neutrophil reprogramming
- Source: bioRxiv (preprints)
- Date: 2026-09-18
- Categories: Genomics & sequence analysis, Single-cell & spatial, Systems & networks, Mathematical biology & statistics, Tools & resources
- Authors: Ren, A., You, Y., Lu, M.
- DOI: 10.64898/2026.09.16.752194
- Source URL: <https://doi.org/10.64898/2026.09.16.752194>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.16.752194>

Abstract: Disease reprograms cells through changes in gene regulation, yet identifying these changes remains a major challenge. We introduce NetDes-Duo, a computational method that jointly infers transcription factor regulatory network models for two related conditions using scRNA-seq data. The networks are optimized to have minimal topological differences, while the associated ODE models recapitulate single-cell gene expression trajectories for both conditions. On synthetic benchmarks, NetDes-Duo outperformed methods that infer each network independently. NetDes-Duo was applied to neutrophil reprogramming in naive and tumor-bearing mice, and the network-simulated dynamics reproduced the observed cell state transitions. The naive landscape had two well-separated basins, whereas the tumor-bearing landscape was more continuous, with three shallower basins. Perturbation and driving simulations also identified Cebpb as a key driver of the tumor-bearing transition, consistent with emergency granulopoiesis literature. We expect NetDes-Duo to be a broadly applicable framework for uncovering the regulatory logic of disease-associated cell state transitions.

## Latent generative search unlocks de novo design of untapped biomolecular interactions at scale
- Source: bioRxiv (preprints)
- Date: 2026-09-18
- Categories: Proteins & structural biology
- Authors: Didi, K., Reidenbach, D., Penner, M., Ravichandran, S., Case, M., Nichols, M., Swanson, E., Reis, A., Prescott, M., Qian, Y., Qian, D., Yang, J., Li, W., Li, L., Shonai, D., Gay, S., Basu Mallik, B., Chim, H. Y., Chen, L., Atienza Juanatey, M., Klein, H., Rieger, D., Schlegel, P., Macintyre, A. U., Secor, M., Granata, D., Cha, S., Cao, Z., Zhou, G., Geffner, T., Chen, X., Livne, M., Zhang, Z., Zhang, T., Gion, K., Bronstein, M. M., Steinegger, M., Deibler, K., Soderling, S., Schoeder, C. T., Khmelinskaia, A., Hollfelder, F., Dallago, C., Kucukbenli, E., Vahdat, A., Ogden, P., Kreis, K.
- DOI: 10.64898/2026.09.12.751118
- Source URL: <https://doi.org/10.64898/2026.09.12.751118>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.12.751118>

Abstract: De novo protein design has advanced rapidly, yet designing binders to polar, solvent-exposed epitopes and small, flexible ligands remains challenging. Such hydrated surfaces and flexible molecules, including carbohydrates, provide few of the hydrophobic contacts favoured by current methods and have largely resisted de novo binders. To address this challenge, here we introduce latent generative search for binder design, a novel framework that uses reward-guided search at inference time to steer the Proteina-Complexa generative model. The model codesigns sequence and structure - generating them together in a continuous latent space - and thereby removes the inverse-folding step on which current methods rely. In a screen of more than one million designs by multiplexed phage display, latent generative search produced more validated binders than every other method tested, its codesigned sequences surpassing post hoc redesign. It delivered high-affinity binders across therapeutic receptors, a viral attachment protein and intracellular signalling targets. Our approach also accessed previously untapped biology, generating the first de novo proteins that bind a free carbohydrate, including one that discriminates between blood-group antigens - a polar, flexible target class beyond the reach of current design methods.

## Learned Human Aortic Morphodynamics: A Dynamical-Systems Model of Post-EVAR Sac Remodeling
- Source: medRxiv (preprints)
- Date: 2026-09-18
- Authors: Pugar, J., Kim, J., Mansour, M., Davis, C., Nguyen, N., Lee, C. J., Babrowski, T., Verhagen, H., Milner, R., Klishin, A., Pocivavsek, L.
- DOI: 10.1101/2025.09.29.25336910
- Source URL: <https://doi.org/10.1101/2025.09.29.25336910>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1101%2F2025.09.29.25336910>

Abstract: Biological tissue reorganizes in response to changes in its mechanical environment, yet quantitative models of that reorganization are almost always built where observations are dense. At organ-scale in living humans the observational data is sparse and noisy: a few irregularly spaced images per individual, acquired with variable instrumentation, with loss to follow-up and patient attrition. We ask whether governing equations for organ-scale remodeling can be recovered from this regime, and pose what those equations may reveal about the underlying biology. We study the human abdominal aorta following endovascular aneurysm repair (EVAR), in which a stent graft excludes the aneurysm sac from arterial flow and thereby changes the mechanical loading of the sac wall. Each computed tomography (CT) scan is reduced to a two-dimensional state: surface area $\\widetilde\{A\}$ (size) and fluctuation in integrated Gaussian curvature $\\widetilde\{\\delta K\}$ (shape), both normalized to a non-pathological aortic population. Across 100 patients and 220 scans pooled from two international centers, we use Z--SINDy, a sparse identification method with statistical-mechanical uncertainty quantification, to infer affine systems of ordinary differential equations (ODEs) governing $(\\widetilde\{A\}, \\widetilde\{\\delta K\})$ for anatomy that remodels successfully (regressing sacs) and anatomy that does not (stable sacs). Both outcome classes approach stable fixed points, but at different locations in the morphological state space and on different timescales. Regressing anatomy approaches $(\\widetilde\{A\}^\\ast, \\widetilde\{\\delta K\}^\\ast) = (2.0, 1.0)$ with characteristic timescales of $1.3$ and $4.6$ years; stable anatomy approaches its fixed point of $(3.8, 2.5)$ with an order of magnitude slower timescale ($10.9$ and $16.6$ years). Embedding the learned equations within Bayesian classifiers demonstrates that trajectory-based information resolves patient outcome classification earlier than positional evidence alone. Toward model utilization in future studies, these advantages are also observed when empirically observed loss to follow-up is carried through the priors. Recovering the equations that govern remodeling, rather than tracking its instantaneous state, therefore provides a foundation for clinical models built from the sparse records that routine imaging surveillance actually produces.

## Linear antibody epitope prediction using AlphaFold2
- Source: eLife (journals)
- Date: 2026-09-18T00:00:00+00:00
- Categories: Proteins & structural biology, Tools & resources
- Authors: Jacob DeRoo, James S Terry, Ning Zhao, Timothy J Stasevich, Christopher Snow, Brian J Geiss
- Journal: eLife
- DOI: 10.7554/elife.98369
- Source URL: <https://doi.org/10.7554/elife.98369>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.7554%2Felife.98369>
- Code: <https://github.com/jbderoo/PAbFold>

Abstract: Defining the binding epitopes of antibodies is essential for understanding how they bind to their antigens and perform their molecular functions. However, while determining linear epitopes of monoclonal antibodies can be accomplished utilizing well-established empirical procedures, these approaches are generally labor- and time-intensive, and costly. To take advantage of the recent advances in protein structure prediction algorithms available to the scientific community, we developed a calculation pipeline based on the localColabFold implementation of AlphaFold2 that can predict linear antibody epitopes by predicting the structure of the complex between antibody heavy and light chains and target peptide sequences derived from antigens. We found that this AlphaFold2 pipeline, which we call PAbFold, was able to accurately flag known epitope sequences for several well-known antibody targets (HA/Myc) when the target sequence was broken into small overlapping linear peptides and antibody complementarity determining regions were grafted onto several different antibody framework regions in the single-chain antibody fragment format. To determine if this pipeline was able to identify the epitope of a novel antibody with no structural information publicly available, we determined the epitope of a novel anti-SARS-CoV-2 nucleocapsid-targeted antibody using our method and then experimentally validated our computational results using peptide competition ELISA assays. These results indicate that the AlphaFold2-based PAbFold pipeline we developed is capable of accurately identifying linear antibody epitopes in a short time using just antibody and target protein sequences. This emergent capability of the method is sensitive to methodological details such as peptide length, AlphaFold2 neural network versions, and multiple-sequence alignment databases. PAbFold is available at https://github.com/jbderoo/PAbFold .

## Linear antibody epitope prediction using AlphaFold2
- Source: eLife (journals)
- Date: 2026-09-18T00:00:00+00:00
- Categories: Proteins & structural biology, Tools & resources
- Authors: Jacob DeRoo, James S Terry, Ning Zhao, Timothy J Stasevich, Christopher Snow, Brian J Geiss
- Journal: eLife
- DOI: 10.7554/elife.98369.3
- Source URL: <https://doi.org/10.7554/elife.98369.3>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.7554%2Felife.98369.3>
- Code: <https://github.com/jbderoo/PAbFold>

Abstract: Defining the binding epitopes of antibodies is essential for understanding how they bind to their antigens and perform their molecular functions. However, while determining linear epitopes of monoclonal antibodies can be accomplished utilizing well-established empirical procedures, these approaches are generally labor- and time-intensive, and costly. To take advantage of the recent advances in protein structure prediction algorithms available to the scientific community, we developed a calculation pipeline based on the localColabFold implementation of AlphaFold2 that can predict linear antibody epitopes by predicting the structure of the complex between antibody heavy and light chains and target peptide sequences derived from antigens. We found that this AlphaFold2 pipeline, which we call PAbFold, was able to accurately flag known epitope sequences for several well-known antibody targets (HA/Myc) when the target sequence was broken into small overlapping linear peptides and antibody complementarity determining regions were grafted onto several different antibody framework regions in the single-chain antibody fragment format. To determine if this pipeline was able to identify the epitope of a novel antibody with no structural information publicly available, we determined the epitope of a novel anti-SARS-CoV-2 nucleocapsid-targeted antibody using our method and then experimentally validated our computational results using peptide competition ELISA assays. These results indicate that the AlphaFold2-based PAbFold pipeline we developed is capable of accurately identifying linear antibody epitopes in a short time using just antibody and target protein sequences. This emergent capability of the method is sensitive to methodological details such as peptide length, AlphaFold2 neural network versions, and multiple-sequence alignment databases. PAbFold is available at https://github.com/jbderoo/PAbFold .

## LLMsFold: Integrating Large Language Models and Biophysical Simulations for De Novo Drug Design
- Source: bioRxiv (preprints)
- Date: 2026-09-18
- Categories: Proteins & structural biology, Tools & resources
- Authors: Waththe Liyanage, W. W., Rigoni, D., Bove, F., Righelli, D., Romano, S., Visone, R., Iorio, M. V., Grassia, M., Mangioni, G., Lio, P., Taccioli, C.
- DOI: 10.64898/2026.03.02.709055
- Source URL: <https://doi.org/10.64898/2026.03.02.709055>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.03.02.709055>

Abstract: The discovery of novel small molecules is challenging because of the vastness of chemical space and the complexity of protein-ligand interactions, leading to low success rates and time-consuming workflows. Here, we present LLMsFold, a computational framework that combines Large Language Models (LLMs) and biophysical foundation tools to design and validate new small molecules targeting pathogenic proteins. The pipeline starts by identifying viable binding pockets on a target protein through geometry-based pocket detection. A 70-billion-parameter transformer model from the LlaMA family then generates candidate molecules as SMILES strings under prompt constraints that enforce drug-likeness. Each molecule is evaluated by Boltz-2, a diffusion-based model for protein-ligand co-folding that predicts bound 3D structure and binding affinity. Promising candidates are iteratively optimized through a reinforcement learning loop that prioritizes high predicted affinity and synthetic accessibility. We demonstrate the approach on two challenging targets: ACVR1 (Activin A Receptor Type 1), implicated in fibrodysplasia ossificans progressiva (FOP), and CD19, a surface antigen expressed on most B-cell lymphoma and leukemia cells. Top candidates show strong in silico binding predictions and favorable drug-like profiles. All code and models are made available to support reproducibility and further development.

## Machine Learning for Toxicity Prediction in Low-Sample Molecular Classes
- Source: bioRxiv (preprints)
- Date: 2026-09-18
- Authors: Barajas, C., Dunphy, L., Mullany, L., Tiburzi, O., Lloyd, E.
- DOI: 10.64898/2026.09.16.751426
- Source URL: <https://doi.org/10.64898/2026.09.16.751426>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.16.751426>

Abstract: Deep learning models such as Chemprop have advanced quantitative molecular property prediction, but their reliance on large training sets limits use in data-scarce domains. We propose a framework that fine-tunes a general baseline model trained on publicly available data on small, class-specific datasets. The resulting models retain the baseline's generalization ability while gaining class-specific accuracy and produce probabilistic outputs that capture uncertainty in the training data. We demonstrate the approach on three toxicity classes defined by a common core structure, target, or mode of action: (i) organophosphates, (ii) androgen receptor antagonists, and (iii) estrogen receptor beta antagonists. Each fine-tuned model outperforms classical machine-learning methods and the EPA TEST tool. The probabilistic nature of the predictions enables prioritization of compounds for experimental validation and seamless integration with data streams of varying quality, supporting iterative decision-making in chemical safety and drug discovery.

## Mapping Gene Expression to an Interpretable Semantic Space
- Source: bioRxiv (preprints)
- Date: 2026-09-18
- Categories: Genomics & sequence analysis, Single-cell & spatial, Tools & resources
- Authors: Duan, X., Aggarwal, M., Periwal, V.
- DOI: 10.64898/2026.09.11.750957
- Source URL: <https://doi.org/10.64898/2026.09.11.750957>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.11.750957>

Abstract: Cell embeddings organize single-cell expression data, but their dimensions have no biological meaning, so clusters are interpreted afterward. We present MESIC (Mapping Expression to Semantic space with Interpretable Components), which builds the written knowledge about genes held in curated databases into the dimensions themselves. A biomedical language model converts each gene's summary into a semantic embedding. MESIC compresses these embeddings into a small number of components, each concentrated on a small set of genes and explained by their annotations. The components are computed once from the summaries, so any expression dataset can be mapped onto them, and every cluster, outlier, or cell-type assignment is then characterized by named genes. In cardiomyocytes, outliers in the component space were enriched for hypertrophic cardiomyopathy. In a lung atlas, unsupervised clusters in that space matched the broad cell types that experts had annotated. In both, the components that separated the cells matched their known biology. For about half of the cells that the atlas itself had left unannotated, the same space gave a confident cluster assignment, and with it an interpretation through component-associated genes. Gene summaries thus give single-cell analysis a coordinate system in which every result is traced to genes and what is written about them.

## Mining association rules for targeted spatiotemporal aquatic environmental DNA (eDNA) sampling
- Source: bioRxiv (preprints)
- Date: 2026-09-18
- Categories: Tools & resources
- Authors: Toth, N., Antonie, L., Hanner, R. H., Gillis, D. J., Phillips, J. D.
- DOI: 10.64898/2026.09.16.752056
- Source URL: <https://doi.org/10.64898/2026.09.16.752056>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.16.752056>

Abstract: Environmental DNA (eDNA) offers a non-invasive alternative to traditional, more destructive sampling methods for determining species occupancy at ecological sites of interest. Aquatic eDNA sampling entails filtering a known volume of water to capture and detect genetic material shed by organisms. While the influence of individual environmental properties on the presence of target eDNA has been widely studied, it remains unclear how variables like temperature, pH, flow rate and conductivity correlate collectively with site electrofishing counts and eDNA concentrations. Resolving this question is important for two reasons: (1) typically, only eDNA, not physical specimens, is collected and measured, and (2) when methods like electrofishing and eDNA sampling are used in tandem, results often differ. Here unsupervised association rule-based machine learning is employed to discover interesting relationships among sampled covariates within a previously published case study of native brook trout (Salvelinus fontinalis) collected from Hanlon Creek (Guelph, Ontario, Canada) in September 2019. From a dataset of only 126 observations, the mining process revealed over 12000 plausible association rules linking covariates to eDNA concentrations (low/high) and electrofishing outcomes (absence/presence of brook trout). A strict pruning strategy reduced this ruleset to a manageable size of 153 associations, some of which were corroborated by existing literature, and some of which were novel (such as those potentially relating electrical conductivity to microbial and enzymatic activity). The entire workflow is included as a new R package called RulesTools. These results highlight the promise of association rule mining as a tool for guiding eDNA metadata collection, complementing statistical modelling, and informing conservation and management decision-making.

## MiRNA Atlas: A Literature-Derived Database of MicroRNAs Bridging Osteoarthritis and Appendage Regeneration
- Source: bioRxiv (preprints)
- Date: 2026-09-18
- Categories: Tools & resources
- Authors: Parker, A. T., Kraus, V. B.
- DOI: 10.64898/2026.09.17.748327
- Source URL: <https://doi.org/10.64898/2026.09.17.748327>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.17.748327>

Abstract: Many microRNAs (miRNAs) regulate tissue remodeling, cellular plasticity, and repair across evolutionarily distant vertebrate lineages that are capable of regenerating appendages such as amputated limbs, fins, and antlers, as well as in human articular cartilage responding to injury. These miRNAs often belong to the same families and exert conserved, though occasionally inverted, regulatory effects. This strong cross-species overlap motivates the present literature-derived analysis. Osteoarthritis (OA), the most prevalent joint disease, still lacks disease-modifying therapies, in part because of the longstanding assumption that adult mammalian cartilage demonstrates no intrinsic reparative capacity. Yet human cartilage retains a latent repair program activated by mechanical and inflammatory stress. Some injured or degenerating joints may never be clinically recognized as osteoarthritic because their intrinsic repair capacity is sufficient to restore tissue integrity; in others, where repair capacity is diminished or damage exceeds it, the repair program is insufficient and OA becomes clinically manifest. MiRNAs are established post-transcriptional regulators of cartilage homeostasis, degeneration, and appendage regeneration, yet because the OA and regeneration research fields have advanced largely independently, the insights available at their intersection have gone unrecognized. To close this gap, we systematically mined both literatures to construct an auto-updating, cross-referenced atlas of OA- and appendage regeneration-associated miRNAs. Integrating these datasets identified a core set of shared miRNA families, delineated miRNAs unique to each field, and mapped convergent families onto common pathways governing matrix remodeling, dedifferentiation, senescence, and inflammation. We propose that regeneration-competent species can inform the identification of therapeutic miRNAs, such as miR-133, miR-21, and let-7, capable of activating endogenous cartilage repair. Collectively, this synthesis and its accompanying web-based miRNA atlas (https://mirnaatlas.shinyapps.io/mirnaatlas/) establish a comparative framework for regenerative miRNA biology and provide a continually updated resource to accelerate discovery of disease-modifying, RNA-based therapies for OA.

## Modeling parasite clearance and transmission delays in American Cutaneous Leishmaniasis transmission dynamics
- Source: Journal of Mathematical Biology (journals)
- Date: 2026-09-18T00:00:00+00:00
- Categories: Mathematical biology & statistics
- Authors: Sharmin Sultana, Gilberto Gonzalez-Parra, Luis Fernando Chaves
- Journal: Journal of Mathematical Biology
- DOI: 10.1007/s00285-026-02460-9
- Source URL: <https://doi.org/10.1007/s00285-026-02460-9>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1007%2Fs00285-026-02460-9>

Abstract: American Cutaneous Leishmaniasis (ACL) is a multi-host vector-borne disease with complex transmission dynamics involving sand fly vectors, reservoirs, and incidental hosts. In this study, we develop and analyze two delay differential equation (DDE) models to explore the role of biological time delays in ACL transmission dynamics. The first model incorporates a delay in the parasite clearance term of the incidental host, reflecting the minimum parasite clearance time between infection and recovery. The second model introduces a delay that represents delays in skin lesion development following exposure to infective sand fly bites. For both models, we derived the basic reproduction number $$\\mathcal \{R\}\_0$$ R 0 and performed an elasticity analysis which revealed that vector-reservoir transmission and vector mortality are key drivers of outbreak potential. For both models we studied the stability of the disease-free and endemic equilibria. Using linearization, we identified critical delay thresholds that generate Hopf bifurcations. We show that time delays can destabilize the endemic equilibrium and induce periodic oscillations. Despite having the same $$\\mathcal \{R\}\_0$$ R 0 , the two models exhibit distinct qualitative behaviors due to differences in the delay structure. Simulations support the analytical findings, illustrating how increasing delays can lead to oscillations in transmission. These findings highlight the importance of incorporating biologically inspired delays when modeling ACL transmission.

## MorphCell learns reconstructable shape representations with explicit scale control for 3D cell morphology
- Source: bioRxiv (preprints)
- Date: 2026-09-18
- Categories: Biological imaging
- Authors: Zhang, G., Wang, G., Cao, R., Guo, J., Xu, S., Yu, Z., Zheng, Y., He, Y., Feng, X.
- DOI: 10.64898/2026.09.17.752374
- Source URL: <https://doi.org/10.64898/2026.09.17.752374>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.17.752374>

Abstract: Three-dimensional (3D) cell morphology provides a measurable phenotype of cellular state and function, but representing it across biological systems and imaging modalities remains difficult. Existing approaches do not jointly provide transferable representations, complete surface reconstruction, geometric interpretation and explicit control of physical scale. Here we introduce MorphCell, a self-supervised framework that uses cell-surface point clouds to learn shape-driven representations independently of physical scale. MorphCell combines cross-view reconstruction pretraining with spherical self-reconstruction. The former captures geometric relationships between surface regions, whereas the latter recovers complete 3D morphology from individual representations. Pretrained on non-biological object surfaces, MorphCell transfers to cellular datasets without biological task-specific training and outperforms existing point-cloud representations in morphological classification. The learned representations capture both global contour and local surface variation, enabling geometric interpretation through reconstruction, biophysical descriptors and saliency analysis. By retaining physical scale as a separate variable for controlled fusion, MorphCell further reveals that shape and scale contribute differently across biological distinctions. This framework provides a general approach for representing, reconstructing and interpreting 3D cellular morphology across imaging modalities and biological contexts.

## Multivalent lamin binding controls the meshwork structure in a self-assembled model of the nuclear lamina
- Source: bioRxiv (preprints)
- Date: 2026-09-18
- Authors: Hameed, H. A., Ozkan, A. U., Erbas, A.
- DOI: 10.64898/2026.03.14.711786
- Source URL: <https://doi.org/10.64898/2026.03.14.711786>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.03.14.711786>

Abstract: The nuclear lamina is a specialized two-dimensional filamentous polymer meshwork that provides structural integrity and elasticity to the nucleus while orchestrating diverse cellular processes. Composed of interacting A- and B-type lamin networks, this structure undergoes tightly regulated self-assembly that is frequently perturbed by disease-causing mutations such as those observed in laminopathies or cardiomyopathies. However, because filament assembly, peripheral adsorption of lamins, and network branching occur concurrently in vivo, isolating the specific biophysical parameters that dictate emergent lamina topology has remained a major challenge. Here, we present a polymer-physics approach that explicitly resolves the spontaneous self-assembly of lamin networks under nuclear confinement. By modeling lamin dimers as semiflexible filaments with distinct interactive domains, we demonstrate that the formation of continuous, high-aspect-ratio fibers strictly requires a coordination cascade of parallel lateral alignment sites and longitudinal head-to-tail interactions between lamins. We show that the thermodynamic affinity between lamin-A and the peripheral boundary (i.e., the inner nuclear membrane and lamin B network) acts as a kinetic switch: weak surface adsorption drives network phase separation and large lamin-free gaps, whereas robust substrate binding stabilizes highly branched networks with uniform lamin distribution. Finally, uniaxial compression simulations reveal that mutations altering these molecular binding interfaces severely compromise macroscale nuclear load-bearing capacity and induce structural vulnerabilities. Collectively, our model establishes a predictive, multi-scale view that directly bridges nanoscale lamin interactions with mesoscale topological remodeling of lamina and macroscale nuclear mechanopathology.

## NANHA - Neurostimulation in Atypical Neurodevelopment: a Harmonized Atlas
- Source: bioRxiv (preprints)
- Date: 2026-09-18
- Categories: Tools & resources
- Authors: Mondal, M., Guha, C., Suresh, S. A., Vashishth, S., Muralidharan, V., Samal, A.
- DOI: 10.64898/2026.09.12.751126
- Source URL: <https://doi.org/10.64898/2026.09.12.751126>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.12.751126>

Abstract: Non-invasive brain stimulation (NIBS) has become an important tool for investigating brain function and exploring therapeutic intervention in neurodevelopmental disorders (NDDs). Nevertheless, relevant studies are scattered across different NDDs, stimulation modalities, target regions, and experimental designs, which makes systematic exploration and comparative analysis challenging. Therefore, we present NANHA, a manually curated, harmonized and freely accessible resource of published NIBS studies in NDDs. The database was constructed using systematic PubMed searches covering 22 NDDs and five NIBS modalities. NANHA captures participant characteristics, study design, stimulation purpose and parameters, targeted brain regions, behavioral, cognitive, and neurophysiological outcomes, safety information, and follow-up assessments, whenever available. The NANHA web platform (https://cb.imsc.res.in/nanha/) provides searchable and downloadable records, advanced filtering options, and interactive visualizations to support data exploration. Overall, NANHA is a standardized resource for evidence synthesis, comparative analyses, identification of research gaps, and the development of disease-specific stimulation protocols for NDDs.

## Neural Architectures of Slow and Fast Dynamics in the Human Brain
- Source: bioRxiv (preprints)
- Date: 2026-09-18
- Categories: Computational neuroscience
- Authors: Lu, Y., Li, Z., Mao, H., Lyu, Q., Lu, Y., Yao, C., Chen, J., Tao, L., Xiao, Z., Tian, X.
- DOI: 10.64898/2025.12.31.696728
- Source URL: <https://doi.org/10.64898/2025.12.31.696728>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2025.12.31.696728>

Abstract: The human brain operates across a vast temporal range, from fast perception and action to slow physiological regulation. The capacity has attributed to a unitary cortical gradient of intrinsic timescales, yet such a unidimensional model cannot explain how local circuits simulateously support both rapid external behavior and slow internal body-coupled dynamics. Using SPLIT (spectral piecewise-linear inference of timescales) on a large-scale intracranial stereo-electroencephalography (8,619 contacts, 185 individuals), we identified dissociable fast (~10-100 Hz) and slow (~1-10 Hz) temporal components. Only the fast-component timescales followed the canonical sensorimotor-to-association cortical hierarchy. Slow-component timescales showed no hierarchical gradient, but were instead covaried with heart rate and enriched at site with heart-related neural responses. This dual temporal architecture persisted across wakefulness, rest, sleep, and anesthesia, challenging the unitary view of brain timescales and revealing an intrinsic bipartite organization that is associated with cortical hierarchy or neurophysiology.

## Novel two-stage deep learning-based approach applied to gene expression data pertaining to esophageal adenocarcinoma boosting biological knowledge discovery
- Source: bioRxiv (preprints)
- Date: 2026-09-18
- Categories: Genomics & sequence analysis
- Authors: Jamie, F., Turki, T., Alsolami, F., Taguchi, Y.-h.
- DOI: 10.64898/2026.09.12.751181
- Source URL: <https://doi.org/10.64898/2026.09.12.751181>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.12.751181>

Abstract: Esophageal cancer (EC) is characterized by complex transcriptional alterations and therapeutic resistance, posing challenges for traditional computational methods. In this study, we propose a deep learning (DL)-based computational framework to identify important genes and biologically relevant pathways in bulk cell RNA-seq data (GSE234304 and GSE273848), which comprise tumor and non-tumor esophageal tissue samples. A fully connected feedforward neural network was trained for binary classification, and two feature selection strategies were implemented: Neural Network followed by Support Vector Regression (NN+SVR) and Integrated Gradients combined with SVR (IG+SVR). The genes were then ranked according to their weights in SVR deriving the importance scores, and the top 100 genes were subjected to enrichment analysis using Enrichr and Metascape. The proposed DL-based approaches identified a greater number of expressed genes across established esophageal cancer cell lines than LIMMA, SAM, and the t-test did. Specifically, in the GSE234304 dataset, IG + SVR, our best method, identified a total of 9 expressed genes while the best baseline method, LIMMA, identified a total of 3 expressed genes. In terms of GSE273848 dataset, IG + SVR was also the best identifying a total of 11 expressed genes while the best baseline method, t-test, had a total of 7 expressed genes. The key genes identified included CEBPB, SUMO1, RORA, STAT1, GATA, OCT1, RUNX1, and NR3C1, as well as pathways related to nucleoprotein maturation, collagen fibril organization, the immunoglobulin-mediated immune response, immune regulation, and insulin signaling. These results show that combining neural networks and attribution-based regression creates an effective and interpretable framework for selecting genes in esophageal cancer research.

## On a new model of COVID-19 transmission by incorporating booster dose vaccination and suggested treatments in India
- Source: Scientific Reports (journals)
- Date: 2026-09-18T00:00:00+00:00
- Authors: Pankaj Singh Rana, Longjun Zhan
- Journal: Scientific Reports
- DOI: 10.1038/s41598-026-70688-y
- Source URL: <https://doi.org/10.1038/s41598-026-70688-y>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1038%2Fs41598-026-70688-y>

Abstract: In this study, a COVID-19 model that is driven by an eight-dimensional system of ordinary differential equations is developed and analyzed, incorporating the primary and booster dose vaccinated individual’s compartments. Initially, the basic properties of the model are examined, and the threshold quantity is obtained. Further, the stability of the equilibrium points of the model is investigated analytically. Moreover, a nonlinear least squares technique is used to calibrate the model parameters based on the cumulative number of COVID-19 reported cases in India. The best-fitted model parameters are found to interpret the consequences of various parameters. In addition, sensitivity analysis is investigated and awareness parameter is identified as the most influential parameter. Moreover, a numerical simulation of the model has been done to compare the consequences of vaccination. In addition, the essence of vaccine efficacy and awareness are examined by considering the different scenarios. It has been found that perceiving preventive measures (pharmaceutical or non-pharmaceutical) significantly reduces the spread of disease amongst the people. Particularly, an increment in the booster dose vaccination and awareness of the disease reduces the number of infected individuals and expands the recovery. Therefore, it is instructed that to protect India from another outbreak of COVID-19, the speed of the booster dose vaccination and awareness campaign should be encouraged.

## Online synaptic credit assignment in active dendrites
- Source: bioRxiv (preprints)
- Date: 2026-09-18
- Categories: Computational neuroscience
- Authors: He, G., Zhang, S., Du, K., Huang, T.
- DOI: 10.64898/2026.09.14.751357
- Source URL: <https://doi.org/10.64898/2026.09.14.751357>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.14.751357>

Abstract: Credit assignment in neural networks is usually formulated as the computation of an abstract error gradient. Whether such a gradient can take a physical, causal form in biophysically detailed multi-compartment neuron models, and enable online, supervised learning, remains unclear. Here we show that the gradient of a detailed neuron's voltage with respect to a synaptic weight is itself a voltage. Differentiating the discrete backward-Euler update solved by standard simulators yields equations of the same form as the original voltage dynamics, driven by gradient currents. Replaying weight-specific gradient currents forward in time reproduces the exact gradient with high fidelity in an L5 pyramidal neuron across diverse input regimes (R2 > 0.998). Pairing the replayed gradient voltage with a local learning signal yields a causal, online learning rule. A single L5 pyramidal neuron with active dendrites learns to reproduce target voltage trajectories containing calcium plateaus and bursts, and recurrent networks of detailed neurons learn to generate target temporal patterns, far outperforming a readout-only control. These results demonstrate that synaptic credit assignment can be implemented online by the voltage dynamics of detailed neurons, potentially suggesting a physical substrate for gradient-based supervised learning in the brain.

## OpenAntigens: a structure-aware database for antigen construct design across the human cell-surface and secreted proteome
- Source: bioRxiv (preprints)
- Date: 2026-09-18
- Categories: Proteins & structural biology, Tools & resources
- Authors: Teixeira, A. A. R., Zhu, H., Kothiwal, D., Cao, R., Mills, A.
- DOI: 10.64898/2026.07.30.741735
- Source URL: <https://doi.org/10.64898/2026.07.30.741735>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.07.30.741735>

Abstract: Choosing which region of a protein to express remains poorly standardized in antibody discovery, recombinant reagent generation, structural biology and computational binder design. For human cell-surface and secreted proteins, this requires reconciling topology, processing, predicted and experimental structure, modifications, interaction partners, orthologs, paralogs and cross-reactivity risk before ordering DNA. OpenAntigens is a free, no-login database of construct-design reports for 5328 human secreted, GPI-anchored, single-pass and multipass proteins. It integrates UniProt topology, AlphaFold pLDDT and PAE, PDB precedent, InterPro and Pfam domains, mouse and cynomolgus orthologs, paralog and family context, Open Targets disease associations, partner and assembly context, and BLAST searches. It provides 55 305 construct suggestions spanning full design regions, PDB-backed boundaries, annotated domains, pLDDT/PAE-derived regions and membrane-expression options, plus 148 722 sequence-similarity hits to help choose constructs and assess cross-reactivity. For targets with compatible AlphaFold models, the interactive designer links sequence, structure, pLDDT and PAE, allowing users to revise boundaries and export species-equivalent sequences with real-time cysteine and modification warnings. OpenAntigens places reproducible construct suggestions, comparative context and browser editing in one workflow, reducing manual reconciliation across resources. OpenAntigens is available at openantigens.org.

## OpenLipid: a large language model workflow for targeted analysis of DIA mass spectrometry data in lipidomics
- Source: bioRxiv (preprints)
- Date: 2026-09-18
- Categories: Tools & resources
- Authors: Li, J., Rost, H.
- DOI: 10.64898/2026.09.12.751192
- Source URL: <https://doi.org/10.64898/2026.09.12.751192>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.12.751192>

Abstract: Mass spectrometry has become a central technology for lipidomics, with data-independent acquisition (DIA) enabling broad and reproducible sampling of lipid signals. However, the multiplexed fragment-ion spectra in DIA data complicate lipid identification. Here, we introduce OpenLipid, a large language model (LLM)-based workflow for targeted DIA lipidomics. Using assay libraries built from data-dependent acquisition (DDA) results, OpenLipid directly evaluates extracted ion chromatograms (XICs) from DIA data in a zero-shot setting to identify target lipid peaks and generate human-readable rationales for individual lipid identification decisions. We benchmarked OpenLipid against manual annotations across four datasets comprising human plasma and mouse feces analyzed in positive and negative ionization modes. The plasma assay libraries contained 199 target lipids in positive mode and 147 in negative mode. The fecal assay libraries contained 181 target lipids in positive mode and 264 in negative mode. At a 5% false discovery rate (FDR) threshold, OpenLipid identified 110 (55.3%) and 28 (19.0%) library targets in plasma and 130 (71.8%) and 84 (31.8%) in feces, in positive and negative ionization modes, respectively. OpenLipid achieved an overall identification rate comparable to that of DIAMetAlyzer (57.8%, 19.7%, 89.0%, and 17.8% across the corresponding datasets) and substantially higher than that of untargeted MS-DIAL DIA analysis (12.6%, 0.0%, 47.5%, and 0.0%) on the same assay-library targets. LLM-derived chromatographic features also enabled supervised discrimination between correct and incorrect candidate peak groups for target lipids, with median cross-validation average precision values of 0.838-0.912. Together, these results demonstrate that OpenLipid is an effective LLM-based workflow for FDR-controlled targeted analysis of DIA lipidomics data.

## Optimal control and sensitivity analysis of an SEIHR TB-COVID-19 coinfection model with cost intervention strategies
- Source: Scientific Reports (journals)
- Date: 2026-09-18T00:00:00+00:00
- Authors: Anagandula Praveen Kumar, Poosan Muthu
- Journal: Scientific Reports
- DOI: 10.1038/s41598-026-70128-x
- Source URL: <https://doi.org/10.1038/s41598-026-70128-x>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1038%2Fs41598-026-70128-x>

Abstract: A compartmental extended SEIHR-based model is developed to simulate the transmission of TB and COVID-19. The positivity and boundedness of the model are demonstrated. The Basic Reproduction Numbers ( $$R\_\{0T\}$$ and $$R\_\{0C\}$$ ) are calculated for TB and COVID-19 models using the Next Generation Matrix (NGM) method. The individual models exhibit backward bifurcation when $$R\_\{0T\}$$ , $$R\_\{0C\}$$ < 1. The local stability analysis is performed by using Lienard-Chipart criteria. Sensitivity analysis is conducted using Latin Hypercube Sampling (LHS) and Partial Rank Correlation Coefficient (PRCC) methods. Optimal control analysis is performed to evaluate the effectiveness of interventions. The co-infection model showed that mask usage significantly reduces both infections. The LHS-PRCC analysis revealed that variables with low p-values strongly influenced the model. Key parameters, including inflow rate, COVID-19 transmission and hospitalization rate, and co-infection progression, exhibited strong correlations. Optimal control measures, including isolation, testing, and treatment, effectively reduced contact between COVID-19-exposed and TB-infected individuals. Post-isolation played a pivotal role in strengthening immunity and aiding recovery from the disease. Implementing all control strategies mitigated the growth of infections and accelerated their decline. The cost-effective analysis, by implementing all interventions, markedly reduces the disease burden (Infection Averted Ratio = 23.72%) and is the most cost-effective strategy, exhibiting a dominant ACER (Average Cost-Effectiveness Ratio) of 0.000359. Finally, we compared the proposed model with the existing coinfection model and validated it with real-time data of individual models.

## Pep-PU-GAN: Positive-Unlabeled Adversarial Learning for Peptide Function Prediction
- Source: bioRxiv (preprints)
- Date: 2026-09-18
- Categories: Proteins & structural biology
- Authors: Midjani, F., Hashemi, S., Keshtkar, F. Z., Malekpour, M., Saberzadeh Ardestani, B., Khosravi, B.
- DOI: 10.64898/2026.09.13.751209
- Source URL: <https://doi.org/10.64898/2026.09.13.751209>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.13.751209>

Abstract: Peptide classification remains challenging in bioinformatics because of limited labeled data, particularly the scarcity of verified negative examples, and the complex relationship between amino acid sequences and biological functions. This study introduces Pep-PU-GAN, a deep learning framework that combines positive-unlabeled (PU) learning, generative adversarial networks (GANs), and graph neural networks (GNNs) for peptide classification. Peptides are represented as sequence-derived residue graphs, with amino acids as nodes and edges connecting adjacent residues, enabling attention-based message passing over local neighborhoods. The architecture includes a generator that produces synthetic peptide embeddings in encoder space and a dual-function discriminator that distinguishes real from synthetic embeddings while performing PU classification. Training uses a custom loss integrating non-negative PU (nnPU) risk estimation with adversarial objectives. A self-training mechanism further incorporates high-confidence synthetic positive embeddings to augment the training set and improve performance. Evaluated on neuropeptide classification using 4,049 positive neuropeptides and 8,558 unlabeled peptides, Pep-PU-GAN outperformed baseline models, achieving an F1 score of 0.93 and an AUROC of 0.98 on an independent held-out benchmark. Pep-PU-GAN provides a promising approach for peptide classification tasks with scarce labeled and abundant unlabeled data, with potential applications in computational biology and drug discovery.

## Physical priors improve performance of structure-based binding affinity models
- Source: bioRxiv (preprints)
- Date: 2026-09-18
- Categories: Proteins & structural biology
- Authors: Kaminow, B., Payne, A. M., MacDermott-Opeskin, H. I., Chodera, J. D., Singh, S.
- DOI: 10.64898/2026.09.11.750982
- Source URL: <https://doi.org/10.64898/2026.09.11.750982>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.11.750982>

Abstract: Structure-based drug discovery is a widely used paradigm for the rational design of novel small molecule therapeutics. However, the benefits conferred by the use of structural information has seen limited adoption in machine learning, where ligand-only ("2D") models are still the industry standard for molecular property or binding affinity prediction. Structure-based ("3D") ML models for binding-affinity prediction promise to present a clear advantage, but have not yet overtaken existing 2D models. Here, we show that physics-based priors can improve predictive performance of structure-based models by comparing different model architectures with varying physical priors on several prediction tasks. We present the Modular Training and Evaluation of Neural Networks (mtenn) package, where we decompose affinity prediction into separate steps of embedding structure into learned representations and combining those embeddings into a predicted binding affinity. We consider both E(3)-invariant and E(3)-equivariant architectures to determine the importance of encoding roto-translational inductive biases, as well as different methods for combining learned embeddings. By first optimizing several aspects of model construction using the general purpose PDBBind dataset, we are able to improve the performance and data efficiency of structure-based models. When subsequently trained and evaluated on the COVID Moonshot small molecule drug discovery dataset, our tuned models perform on par with industry standard ligand-only models. Our decomposed model framework highlights that encoding some physical priors improves model performance, while more complex biases such as equivariance offer limited benefit. Additionally, structure-based models generalize better to an unseen target and display higher training efficiency. Overall, these results emphasize that structure-based models benefit from their ability to incorporate physics-informed constraints, giving promising directions for model architecture development. These results also suggest that the strength of these models may be in tasks specifically aimed at generalizability, providing guidelines for their use in early-stage drug discovery campaigns.

## Post-selection inference in testing for phenotypic differences with scRNA-Seq
- Source: bioRxiv (preprints)
- Date: 2026-09-18
- Categories: Genomics & sequence analysis
- Authors: Sanchez, N., Etourneau, L., Purdom, E.
- DOI: 10.64898/2026.09.17.752411
- Source URL: <https://doi.org/10.64898/2026.09.17.752411>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.17.752411>

Abstract: For the purpose of differential expression (DE) analysis in single-cell RNA-sequencing (scRNA-Seq), phenotype differences between samples are often tested within specific cell types. Cell types are regularly imputed by clustering the same gene expression data which is later used for phenotype testing. This creates the potential for a "double-dipping" or post-selection inference problem resulting in inflated rates of false discoveries. While this selection bias is known to inflate significance in cell-type marker identification, its effect on sample-level phenotype testing, e.g. in patient cohorts, has never been explored despite the growing preponderance of this type of analysis. To address this, we perform an extensive simulation study and demonstrate that naive clustering on uncorrected embeddings can severely inflate the False Discovery Rate (FDR) in the presence of strong phenotypic differences. However, we further show that applying batch-correction methods to remove phenotypic effects prior to clustering resolves the FDR inflation with no obvious loss of power. Finally, we provide measures of phenotypic imbalance that can be applied to real datasets which closely track the false discovery proportion and thus can be used to as part of data exploration to gauge the risk of post-selection inflation of p-values in a particular dataset.

## Predicting Neoantigen Immunogenicity from In Vivo Immune Editing
- Source: bioRxiv (preprints)
- Date: 2026-09-18
- Categories: Proteins & structural biology
- Authors: Sears, T. J., Lee, K.-h., Munoz Perez, M., Rasmussen, R., Pagadala, M. S., Tanaka, K., Subramanian, A., Moding, E. J., Zanetti, M., Carter, H.
- DOI: 10.64898/2026.09.17.752425
- Source URL: <https://doi.org/10.64898/2026.09.17.752425>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.17.752425>

Abstract: Neoantigen immunogenicity prediction is fundamental to personalized cancer vaccines, tumor-infiltrating lymphocyte (TIL) therapy, and TCR-T cell engineering. Existing computational predictors rely primarily on in-vitro correlates of peptide presentation or models trained against assay-based reactivity, and they are typically validated within a single therapeutic setting. We reasoned that the most direct evidence of neoantigen immunogenicity is longitudinal in-vivo elimination: under immune checkpoint blockade (ICB), subclones bearing recognized neoantigens are selectively depleted over time. Here, we present the Neoantigen Elimination Model (NEMo), a two-compartment (CD8 and CD4) machine learning classifier trained on the in-vivo editing (IVE) of neoantigens across serially sequenced, ICB-treated tumors. By using mechanistically inspired NeoPrecis features designed to capture determinants of immunogenicity beyond MHC binding affinity, NEMo recovered assay-confirmed immunogenic neoantigens across four independent, unseen clinical settings -- pre-existing immunogenicity screening, personalized cancer vaccines, TIL therapy, and a radiotherapy +/- ICB ctDNA cohort -- and stratified progression-free survival more strongly than ELISPOT-confirmed reactivity. The editing signal further revealed an immune-evasion architecture in which oncogenic drivers and neoantigens restricted to lost or silenced HLA alleles are systematically spared from editing.

## Pretrained gene representations transfer mean expression more broadly than spatial patterns in virtual spatial transcriptomics
- Source: bioRxiv (preprints)
- Date: 2026-09-18
- Categories: Genomics & sequence analysis, Biological imaging
- Authors: Chen, T., Hicks, S. C.
- DOI: 10.64898/2026.09.15.751768
- Source URL: <https://doi.org/10.64898/2026.09.15.751768>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.15.751768>

Abstract: Models that combine tissue images with pretrained gene representations aim to predict spatial expression for genes not used to fit the downstream predictor. Yet success on held-out genes can reflect two capabilities: estimating a gene's mean expression across tissue locations and recovering its spatial variation. Across four cohorts spanning three human brain regions and HER2-positive breast cancer, we evaluated held-out genes in held-out individuals and separated these components. For spatial predictors using fixed gene representations from Decima or scGPT, reductions in gene-mean error accounted for more than 91% of the reduction in mean squared error relative to matched random vectors. Independently fitted mean-only models using the same representations but no tissue images retained 90-99% of the corresponding gain in full-matrix correlation. Spatial gains were smaller on average, increased with expression variation in training tissue and differed across cohorts and representations. Across these settings, pretrained gene representations broadly transferred mean expression but selectively improved spatial recovery, showing that cross-gene generalization in virtual spatial transcriptomics is not a single capability.

## Prospecting the protein design landscape.
- Source: FEBS letters (journals)
- Date: 2026-09-18T00:00:00Z
- Categories: Proteins & structural biology
- Authors: Jakob R. Riccabona, Katharina T. Stonig, J. Meiler, Clara T. Schoeder, Monica L. Fernández-Quintero
- Journal: FEBS letters
- DOI: 10.1002/1873-3468.70459
- External ID: 6dff162463966e086bfd7907f0a9afb8c72bdd62
- Keywords: peptide, antibody, peptides, protein design
- Source URL: <https://doi.org/10.1002/1873-3468.70459>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1002%2F1873-3468.70459>

Abstract: Generative AI has driven remarkable breakthroughs in protein design, enabling the rapid, computationally guided creation of high-affinity binders against diverse targets. While remarkable experimental success has been demonstrated, the confidence metrics used to filter and evaluate designs remain optimized for static protein interfaces and can fail when applied to underrepresented or conformationally complex targets. In this review, we outline the current landscape of deep learning-driven protein design pipelines, discuss tailored applications in peptide, small molecule, binder, vaccine, and antibody design, and argue that the integration of ensemble-based methods represents a promising avenue for improving design success rates. Beyond single target binder design, we further highlight emerging strategies that expand the functional scope of designed proteins, including fold-switching scaffolds and molecular glues realized through engineered cyclic peptides, which enable context-dependent control of protein-protein interaction networks. Together, these advances position de novo protein design as a broadly applicable technology platform at the intersection of structural biology, biophysics, and molecular medicine.

## QBayMic: Quantum-coupled variational Bayes for clustering and feature selection in low-signal microbiome data
- Source: bioRxiv (preprints)
- Date: 2026-09-18
- Categories: Evolution & metagenomics
- Authors: Dang, T., Lysenko, A., Tsunoda, T.
- DOI: 10.64898/2026.09.14.751634
- Source URL: <https://doi.org/10.64898/2026.09.14.751634>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.14.751634>
- Code: <https://github.com/tungtokyo1108/QBayMic>

Abstract: Clustering microbiome samples into community types is central to cohort stratification and biomarker discovery, yet the resulting inference becomes unstable when the between group signal is small compared with sampling noise: variational Bayes yields different partitions across initialisations, and common fixes do not solve the problem. Simple restarts are ineffective because the variational free energy is anti-correlated with clustering accuracy; deterministic annealing collapses to the same solution as greedy ascent, with the operator staying diagonal at every temperature; and parallel tempering replicas remain too similar to permit configuration exchanges. We propose QBayMic, which replaces the assignment step of a Dirichlet-multinomial mixture with sparse variable selection via a quantum Gibbs state under an annealed Hamiltonian, coupling competing assignments through a transverse-field term that cannot be reproduced by temperature scaling alone. We present two gate-based circuit designs for this step, evaluating on noiseless qubit-register simulations, and we derive a signal fraction, computable prior to clustering, that predicts the expected strength of quantum coupling. With matched compute in the predicted regime, the three classical methods recovered the reference partition (ARI > 0.4) in 0/100 seeds, while QBayMic recovered it in 47-64/100; when the number of clusters exceeded three, only QBayMic recovered the correct cluster count. For a soil pH dataset, the diagnostic indicates a narrow separation margin; for a human-derived dataset tuned into the predicted band via controlled dilution, classical methods recovered the cluster count in 0/100 seeds, compared with 61-76% for QBayMic. The implementation is publicly available at https://github.com/tungtokyo1108/QBayMic.

## Real-time, cross-modal genotype mapping of free-moving Drosophila larvae via simultaneous mechano-electrophysiological recording
- Source: Science Advances (journals)
- Date: 2026-09-18T00:00:00+00:00
- Categories: Evolution & metagenomics
- Authors: Kairu Dong, Hao Song, Qianhui Zhao, Zhiying Song, Fu Lv, Siouwen Wan, Tianyu Zheng, Yunlong Fan, Wen-Che Liu, Shaomin Zhang, Yongjun Wu, Yuhui Huang, Jizhou Song, Zhefeng Gong, Nenggan Zheng, Kewang Nan
- Journal: Science Advances
- DOI: 10.1126/sciadv.aef7492
- Keywords: genotyping
- Source URL: <https://doi.org/10.1126/sciadv.aef7492>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1126%2Fsciadv.aef7492>

Abstract: Drosophila larvae provide a powerful model for interrogating genes associated with human muscle and neurological disorders; however, existing genotyping and phenotyping approaches remain low-throughput and often rely on destructive, invasive, or toxic procedures. Here, we present a scalable bioelectronic platform that enables real-time, simultaneous mechano-electrophysiological recording from freely moving Drosophila larvae in an open three-dimensional (3D) space, allowing high-throughput cross-modal genotype mapping (CMGM). The system integrates conductive and piezoelectric microneedle electrodes into a flexible sensory array that achieves stable, long-term signal acquisition during unrestricted and complex 3D locomotion. By coupling dual-modal signal acquisition with machine-learning-assisted classification, we directly identify muscle defects in unlabeled RNAi-knockdown larvae within 30 minutes, without invasive manipulation or time-consuming sample preparation. Incorporation of both electrophysiological and mechanical waveform features improves overall classification accuracy to 96%, outperforming single-modality approaches. This non-destructive, high-throughput CMGM strategy establishes a generalizable framework for bridging genotype and phenotype in intact, freely behaving organisms, with broad implications for functional genetics and disease modeling.

## Regulatory-prior-guided attention preserves biological structure during unpaired single-cell RNA–ATAC integration
- Source: Bioinformatics (journals)
- Date: 2026-09-18T00:00:00+00:00
- Categories: Genomics & sequence analysis, Single-cell & spatial, Tools & resources
- Authors: Zhenglong Cheng, Jiao Zhang, Shixiong Zhang
- Journal: Bioinformatics
- DOI: 10.1093/bioinformatics/btag696
- Source URL: <https://doi.org/10.1093/bioinformatics/btag696>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1093%2Fbioinformatics%2Fbtag696>
- Code: <https://github.com/zlCreator/scHPGT>

Abstract: Motivation Single-cell RNA sequencing and single-cell ATAC sequencing provide complementary views of transcriptional output and chromatin regulatory potential, but integrating unpaired profiles remains challenging because the modalities differ in feature space, sparsity and noise. Existing approaches often frame integration as distribution matching, which can over-align biologically distinct, condition-specific or modality-specific cell states. We present scHPGT, a single-cell Heterogeneous Prior-Guided Transformer for regulatory-prior-guided integration of unpaired RNA and chromatin accessibility profiles. scHPGT uses modality-specific encoders to model RNA and ATAC signals, a prior-guided cross-modal Transformer to constrain gene–peak attention using regulatory links, and a domain-adversarial objective to reduce modality-specific discrepancies in a shared latent space. Results Across PBMC3k, mouse spleen, CITE-seq/ASAP-seq PBMC and PBMC10k benchmarks, scHPGT improves clustering agreement, label transfer and biological structure preservation while maintaining effective modality alignment. In partial-overlap and condition-shift settings, scHPGT aligns shared populations without forcing unmatched or condition-specific states into inappropriate correspondence. Attention-derived links recover regulatory relationships, highlight marker-gene regulatory regions, recover transcription factor programs and produce regulatory activity profiles consistent with cell-type-specific transcriptional programs. Availability and Implementation Code and datasets are released at https://github.com/zlCreator/scHPGT. Supplementary Information Supplementary data are available at Bioinformatics online.

## Representing Sex in Cardiovascular Models: Calibrating Reference Parameters from Healthy Cohorts
- Source: bioRxiv (preprints)
- Date: 2026-09-18
- Authors: Lakshmikanthan, A., Plappert, F., Shen, W., Oomen, P. J. A.
- DOI: 10.64898/2026.09.17.752488
- Source URL: <https://doi.org/10.64898/2026.09.17.752488>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.17.752488>

Abstract: Reduced-order models are increasingly used to study cardiac physiology and inform patient-specific therapies. However, a model's prediction is only as reliable as its underlying parameters: representative model parameterization is essential to reflect the physiology of the populations these models are meant to represent, including biological sex. Most current models are parameterized from male or sex-agnostic data and/or focus on specific pathologies. Therefore, the goal of this work is to establish a formal parameter estimation pipeline for deriving reduced-order cardiovascular model parameter ranges that are physiologically representative of healthy women and men. We calibrated a closed-loop reduced order model of the heart and circulation separately for healthy female and male populations, using data pooled from eleven healthy cohorts. To account for parameter sensitivity and identifiability, we employed a three-stage parameter subset reduction pipeline: global sensitivity analysis (Sobol's method), collinearity screening (Fisher information matrix), and profile-likelihood identifiability analysis. Sex-specific distributions of parameters that were deemed sensitive and identifiable for each sex, ten for women and 9 for men, were obtained by Hamiltonian Monte Carlo. All the calibrated parameters showed less than 80% overlap between sexes, with the smallest overlap observed in some of the most influential parameters, such as stressed blood volume. Comparing simulations of the calibrated models against allometrically size-matched simulations showed that body size explained some, but not all, of the sex differences. The resulting parameter distributions provide reference ranges usable in future mechanistic and patient-specific simulations to contribute to more inclusive cardiovascular modeling.

## ReScale4DL: balancing pixel and contextual information for enhanced bioimage segmentation
- Source: Nature Communications (journals)
- Date: 2026-09-18T00:00:00+00:00
- Categories: Biological imaging
- Authors: Mariana G. Ferreira, Bruno M. Saraiva, António D. Brito, Mariana G. Pinho, Ricardo Henriques, Estibaliz Gómez-de-Mariscal
- Journal: Nature Communications
- DOI: 10.1038/s41467-026-77930-1
- Source URL: <https://doi.org/10.1038/s41467-026-77930-1>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1038%2Fs41467-026-77930-1>

Abstract: Deep learning is the state-of-the-art approach for bioimage segmentation. However, it presents a paradox regarding image resolution: counterintuitively, deep learning segmentation performance can improve with lower image resolutions. This phenomenon is particularly significant in microscopy, where high-resolution acquisitions come with substantial costs in throughput, storage requirements and potential photodamage. We systematically evaluate how image resolution impacts segmentation by training popular architectures on datasets downsampled to 6-50% of their original resolution, mimicking lower-magnification acquisitions. Compared with models trained on native-resolution images, segmentation accuracy either improves (by up to 25% of mean Intersection over Union (IoU)) or degrades minimally (< 5% of mean IoU) when using images downsampled by up to fourfold (25% of the original resolution). Downsampling proportionally increases information throughput while reducing storage requirements and inference time. These findings provide practical guidelines for creating efficient, sustainable and cost-effective bioimaging pipelines that reduce data and computing needs while optimising microscopy techniques.

## Resource limitation rewires chromosome instability and ploidy evolution across in vitro and in vivo cancer models
- Source: bioRxiv (preprints)
- Date: 2026-09-18
- Categories: Evolution & metagenomics
- Authors: Li, T., Beck, R., Tagal, V., Yu, X., Andor, N.
- DOI: 10.64898/2026.09.16.752131
- Source URL: <https://doi.org/10.64898/2026.09.16.752131>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.16.752131>

Abstract: Whole-genome doubling (WGD) and elevated ploidy are pervasive features of cancer that shape chromosomal instability (CIN), therapeutic response, and metastatic fitness. Yet ploidy is strikingly context dependent: many tumors remain near diploid in vivo despite the frequent emergence of highly polyploid states in vitro. Here, we estimate ploidy dependent chromosome missegregation tolerance using a mathematical model calibrated to growth and chromosome number data from matched near diploid and near tetraploid breast cancer cultures and xenografts. The model architecture reflects a tug of war between two opposing selective forces acting on ploidy--resource limitation that caps high ploidy by imposing energetic and biosynthetic costs, and CIN that can favor higher ploidy by buffering the fitness impact of chromosome gains and losses. The fitted model reproduced chromosome losses in 4N cultures and WGD followed by chromosome losses in 2N cultures. In vivo, the strongest determinants of ploidy shifted from stress associated death at low oxygen to baseline missegregation at higher oxygen. Joint calibration to both in vivo and in vitro contexts assigned tumors lower proliferation, an approximately tenfold higher stress-associated death scale, and an 11-16 fold larger maximal stress induced missegregation increment than cultures. Simulating populations across combinations of constant oxygen levels and missegregation settings showed that ploidy increased with oxygen under conditions supporting population growth. In both culture and tumors, mean ploidy above tetraploidy was associated with population decline. Together, the framework predicts when resource constraints favor chromosome loss and when missegregation tolerance permits high ploidy expansion, providing testable expectations for how CIN perturbations reshape ploidy evolution in different resource environments.

## Role of foundation models in data-driven tissue diagnostics.
- Source: Molecular aspects of medicine (journals)
- Date: 2026-09-18T00:00:00Z
- Categories: Genomics & sequence analysis, Biological imaging
- Authors: Mohsin Bilal, Aadam, Manahil Raza, Anas Alsuhaibani, Youssef N. Altherwy, Abdulrahman Alabduljabbar, Fahdah A. Almarshad, P. Golding, N. Rajpoot
- Journal: Molecular aspects of medicine
- DOI: 10.1016/j.mam.2026.101504
- External ID: fc55fe015219e1de4f0046e0fb94d8f1a26e3670
- Keywords: genomics, histopathology, foundation models
- Source URL: <https://doi.org/10.1016/j.mam.2026.101504>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1016%2Fj.mam.2026.101504>

Abstract: Artificial intelligence (AI)-based tissue diagnostics is entering a new phase driven by pathology foundation models: large-scale encoders and multimodal systems pretrained with self-supervised and vision-language objectives. Yet the evidence base has expanded faster than the methods used to evaluate and interpret claimed capabilities. The central clinical question is no longer whether these models can achieve strong retrospective performance, but what new capabilities they provide, under what conditions those capabilities are demonstrated, and whether they can translate into reliable diagnostics in clinical settings. This review makes three contributions. First, we define histopathology-centric foundation models and distill the technical factors that shape their behavior: pretraining data regime, learning objective, and downstream adaptation to clinical endpoints. Second, we introduce a practical capability framing that distinguishes general-purpose capabilities (coverage across tissues, scales, stains, scanners, and institutions) from functional breadth (task primitives, adaptability, and analysis level), while accounting for modality scope spanning vision, language, and genomics. Third, we synthesize reported results as an evidence map rather than a leaderboard, clarifying where capability is supported by reported evidence, where reproducibility is constrained by access or reporting limits, and where further validation is needed. We then analyze current benchmarking practice and identify common confounders, including heterogeneous adaptation protocols and pretraining-evaluation overlap, and propose deployment-aware recommendations built around broad-and-deep benchmarks, standardized adaptation "budgets," overlap auditing, and domain-agnostic evaluation. Finally, we review emerging clinical-utility evidence and argue that foundation models are most compelling when tied to workflow-defined endpoints, calibrated operating points, and measurable operational benefit. We conclude with actionable recommendations for converting capability demonstrations into clinically reliable data-driven diagnostic systems.

## sabinaMBM: An R package for Multiscale Bayesian species distribution Modelling using INLA
- Source: bioRxiv (preprints)
- Date: 2026-09-18
- Categories: Tools & resources
- Authors: Morales-Barbero, J., Gomez-Rubio, V., Seoane, J., Adde, A., Goicolea, T., Mateo, R. G.
- DOI: 10.64898/2026.09.17.752384
- Source URL: <https://doi.org/10.64898/2026.09.17.752384>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.17.752384>

Abstract: 1. Regional species distribution models (SDM) calibrated over spatially restricted extents tend to truncate species' ecological niches. Existing nested SDM workflows integrate multi-scale information through sequential combination, which limits formal uncertainty propagation across scales and prevents regional predictions from being explicitly constrained within globally-informed niche boundaries. 2. We introduce sabinaMBM, an R package implementing joint multiscale Bayesian SDM within a unified probabilistic framework built on inlabru/R-INLA. It propagates uncertainty across scales without the computational bottlenecks of MCMC-based approaches. The framework offers multiple coupling architectures ranging from complete independence to hierarchical constraint that can be configured independently for intercepts and covariates. 3. In a range-margin population, hierarchical constraint most improves out-of-sample discrimination where regional data were scarcest, while leaving predictions unchanged where they already suffice, delivering gains precisely where sequential approaches are expected to struggle most. Applied to Quercus petraea across its Iberian trailing-edge, including a spatial field produced the largest single performance gain, consistent across every coupling configuration, and covariate responses diverged by scale for at least one climatic predictor. Under future climate, the constrained model yielded lower suitable habitat estimates and redistributed uncertainty in proportion to cross-scale agreement rather than uniformly. 4. sabinaMBM makes multiscale Bayesian inference accessible without specialist programming. This framework provides robust value for trailing-edge populations and spatial (invasive species) or temporal (climate change) projections where niche truncation risks ecologically implausible outcomes, while simultaneously delivering fine-resolution predictions with properly propagated uncertainty whenever global and regional covariates offer complementary information.

## Setting the SCENE for Interpretable Cell-Gene Embeddings in Single-Cell RNA-seq
- Source: bioRxiv (preprints)
- Date: 2026-09-18
- Categories: Genomics & sequence analysis
- Authors: Moberg, O. L., Petersen, M. B., Herlau, T., Kristensen, L. E., Jessen, L. E., Morup, M.
- DOI: 10.64898/2026.09.12.750699
- Source URL: <https://doi.org/10.64898/2026.09.12.750699>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.12.750699>

Abstract: Single-cell RNA sequencing measures cellular states at high resolution, but sparse high-dimensional count data remain difficult to model interpretably. We introduce the Single-Cell Euclidean Network Embedding (SCENE), a probabilistic latent-distance model that jointly embeds cells and genes from Unique Molecular Identifier (UMI) counts. SCENE treats the count matrix as a weighted bipartite cell-gene graph, where Euclidean distances represent transcriptional affinity, and combines this geometry with a zero-inflated count likelihood that separates gene detection from expression magnitude. Across real and simulated scRNA-seq datasets, SCENE recovers biologically structured cell and gene embeddings with state-of-the-art performance. Surprisingly, major biological structure is preserved in native two- and three-dimensional latent spaces, enabling directly interpretable visualization. Perturbation analyses show that SCENE organizes glucocorticoid-response genes and T-cell receptor regulatory programs coherently in gene space, capturing biology beyond cell-type separation. SCENE provides a transparent representation learning framework in which low-dimensional Euclidean geometry supports accurate modeling and biological interpretation.

## Several multiple sequence alignment-perturbing methods enhance AlphaFold3 sampling of alternative protein states.
- Source: Communications chemistry (journals)
- Date: 2026-09-18
- Categories: Proteins & structural biology
- Authors: Samuel Eriksson Lidbrink, Ivan Nissen, Rebecca J Howard, Jonathan Kenichi Ahrlind, Erik Lindahl
- Journal: Communications chemistry
- DOI: 10.1038/s42004-026-02198-x
- External ID: 42760289
- Source URL: <https://doi.org/10.1038/s42004-026-02198-x>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1038%2Fs42004-026-02198-x>

Abstract: Protein function often involves multiple conformational states. Several multiple sequence alignment-perturbing strategies, including stochastic subsampling, clustering, and column masking, have been shown to enhance AlphaFold2 (AF2) sampling of alternative protein states. Here, we evaluate these strategies on AlphaFold3 (AF3) and compare their performance with the BioEmu Boltzmann sampling model on 107 proteins with multiple experimentally solved conformational states. We find that unperturbed AF3 samples alternative states with significantly higher TM-scores compared to AF2 and comparable to BioEmu. In particular, all MSA perturbation methods improve AF3 sampling at a statistically significant level, improving the top 1% TM-score by at least 0.05 in approximately 20% of cases each, while rarely worsening the performance. Furthermore, we find that different choices of amino acid masks can improve column-masked AF3 sampling for specific targets. Our results highlight how MSA perturbations remain relevant in AF3, providing a useful tool for understanding dynamic biological processes.

## SimpleMicrobiome: An integrated web-based platform for streamlined microbiome data analysis and visualization.
- Source: Journal of microbiology (journals)
- Date: 2026-09-18T00:00:00Z
- Categories: Tools & resources
- Authors: Seong-In Na, Juhee Kim, So-Yeon Kim, Jin Park, Yong-Joon Cho
- Journal: Journal of microbiology
- DOI: 10.71150/jm.2606011
- External ID: 6875c1bcead7baf28d8f43ee6fc38baa9a56675b
- Source URL: <https://doi.org/10.71150/jm.2606011>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.71150%2Fjm.2606011>
- Code: <https://github.com/yjcho2252/SimpleMicrobiome>

Abstract: Microbiome studies require multiple analytical steps after initial sequence processing. These steps commonly include data harmonization, preprocessing, taxonomic profiling, diversity analysis, differential abundance testing, predictive modeling, network inference, and preparation of publication-ready outputs. Although robust packages are available for many of these tasks, routine use often depends on command-line workflows, repeated data reformatting, and method-specific scripting. These requirements can limit accessibility for experimental researchers and complicate consistent analysis across interdisciplinary teams. We developed SimpleMicrobiome, a web-based R Shiny platform that integrates established microbiome analysis methods into a single interactive downstream workflow. The application accepts standard abundance, taxonomy, and metadata tables, supports interactive preprocessing and sample filtering, and provides modules for taxa profile visualization, alpha and beta diversity analysis, ANCOM-BC2 and MaAsLin2 differential abundance testing, Random Forest modeling with SHAP-based interpretation, microbial association network inference using SparCC and SPIEC-EASI through NetCoMi, correlation heatmaps, and dbRDA/CAP-style association biplots. The platform is implemented as a modular Shiny application so that preprocessing choices are propagated across downstream analyses, results can be exported as figures and tables, and the same application can be run through the public server, source-code installation, or a Docker image. SimpleMicrobiome consolidates major downstream microbiome analysis tasks in an accessible browser-based environment while retaining links to established analytical frameworks. The platform may reduce technical barriers for non-programming users, improve consistency across exploratory and reporting-oriented analyses, and support collaborative microbiome research. The public application is available at https://simplemicrobiome.mglab.org, the source code is available at https://github.com/yjcho2252/SimpleMicrobiome, and a Docker image for local deployment is available at https://hub.docker.com/r/mglab2252/simplemicrobiome.

## SpaHDSRL: hierarchical dual-graph self-supervised representation learning for integrating spatially resolved multi-omics data
- Source: Briefings in Bioinformatics (journals)
- Date: 2026-09-18T00:00:00+00:00
- Categories: Genomics & sequence analysis, Single-cell & spatial, Tools & resources
- Authors: Xiang Li, Kangkang Zhang, Yifei Li, Fangrong Yan, Bosheng Li, Qian Ding
- Journal: Briefings in Bioinformatics
- DOI: 10.1093/bib/bbag517
- Source URL: <https://doi.org/10.1093/bib/bbag517>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1093%2Fbib%2Fbbag517>
- Code: <https://github.com/Lisa62103/SpaHDSRL>

Abstract: Spatial multi-omics technologies facilitate simultaneous measurement of multiple molecular modalities within their native spatial context, offering opportunities to characterize tissue organization and cellular heterogeneity. However, effective integration remains challenging because such data concurrently encode spatial adjacency and molecular similarity, while also being limited by high dimensionality, sparsity, noise, and cross-modality heterogeneity. Here, we propose SpaHDSRL, a hierarchical dual-graph self-supervised representation learning framework for spatial multi-omics integration. SpaHDSRL jointly models a shared spatial graph and modality-specific feature graphs, which are integrated through an adaptive gated hierarchical fusion strategy to learn coherent and informative latent representation. To further enhance representation quality, SpaHDSRL combines a Deep Graph Infomax-based objective with spatial regularization, preserving both global informativeness and local spatial consistency. Experiments on simulated and real datasets demonstrate that SpaHDSRL consistently achieves superior performance over existing methods in both the accuracy and robustness of spatial domain identification. Downstream analyses further highlight its utility in marker discovery, functional enrichment, second-modality-associated analysis, and cell–cell communication inference, underscoring its value for dissecting tissue architecture, developmental programs, and multicellular interactions in complex biological systems. The source code of SpaHDSRL is available at https://github.com/Lisa62103/SpaHDSRL.

## Sparse Machine Learning Pipeline with Stabl Identifies Cord Blood Multi-Omic Signatures of Bronchopulmonary Dysplasia
- Source: bioRxiv (preprints)
- Date: 2026-09-18
- Categories: Single-cell & spatial, Proteins & structural biology, Systems & networks
- Authors: Mestan, K., Newar, J., Zhao, J., Chakraborty, A., Reiss, J., Funk, W., Stelzer, I., Waked, B., Bellan, G., Durand, X., Hedou, J.
- DOI: 10.64898/2026.09.12.748996
- Keywords: multi omic, proteomics, metabolomics, pathways, pipeline
- Source URL: <https://doi.org/10.64898/2026.09.12.748996>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.12.748996>

Abstract: Background: Several omics studies have been completed in recent years, with the goal of identifying biomarkers of complex multifactorial diseases, such as bronchopulmonary dysplasia (BPD). Objective: To evaluate the performance of 3 distinct omics platforms, using a machine learning pipeline with integration of sparse, reliable and adaptive biomarker identification (Stabl). Methods: Using a well-characterized birth cohort, cord blood metabolomics, proteomics and adductomics data were integrated with Least Absolute Shrinkage and Selection Operator (LASSO) regression and Stabl, to evaluate predictive performance for BPD. Results: Sparse multivariable modeling of 45,000 features measured in 217 infants (52 term, 165 extremely preterm <28 weeks; 82 with BPD and 35 with severe BPD/death) identified a perfect signature for preterm birth with both LASSO and Stabl (AUROC=1.0; p<0.001). Analysis of the preterm group yielded excellent predictive power for severe BPD (AUROC=0.83; p=0.005). Stabl identified a set of 12 biomarkers (2 adducts, 3 proteins and 7 metabolites) with good performance for predicting grade III BPD (AUROC=0.76; P=0.03). Biomarkers across the 3 omics platforms revealed dysregulated pathways of innate/adaptive immune responses, metabolic programming and oxidative stress. Conclusions: The sparse machine learning pipeline is a complementary approach for identifying novel pathways and biomarkers of multifactorial BPD and its endotypes.

## Spatially Constrained Monte Carlo Permutation Test Reveals Diffusion Changes Near Stress Granules
- Source: bioRxiv (preprints)
- Date: 2026-09-18
- Categories: Mathematical biology & statistics
- Authors: Korunova, E., Sikirzhytski, V., Twiss, J. L., Shtutman, M., Vasquez, P.
- DOI: 10.64898/2026.09.17.752460
- Source URL: <https://doi.org/10.64898/2026.09.17.752460>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.17.752460>

Abstract: Intracellular diffusion is inherently heterogeneous, yet single-particle tracking (SPT) analyses are often summarized using cell-wide average parameters that can obscure localized effects. Here, we tracked 40-nm genetically encoded multimeric (GEM) nanoparticles during stress granule (SG) formation and developed SPaCe-MC (Spatially Constrained Monte Carlo permutation test), a statistical framework that generates cytoplasm-specific null models to test whether diffusion associated with a specific cellular structure differs from that expected in the surrounding heterogeneous cytoplasm. Across three SG-inducing conditions, including oxidative stress, DDX3 inhibition, and combined treatment, bulk cytoplasmic analyses revealed distinct responses ranging from increased nanoparticle mobility to increased subdiffusive behavior. In contrast, SPaCe-MC consistently detected a local diffusion constraint in SG-associated regions relative to their treatment-matched cytoplasmic background, revealing a conserved local diffusion effect despite divergent global cytoplasmic responses. Together, our findings establish SPaCe-MC as a framework for identifying compartment-specific diffusion changes in heterogeneous cellular environments.

## Spatially resolved multimodal hallmarks of response to neoadjuvant immunotherapies in the melanoma ecosystem in 2D and 3D
- Source: bioRxiv (preprints)
- Date: 2026-09-18
- Categories: Single-cell & spatial, Biological imaging, Tools & resources
- Authors: Liu, Z., Song, X., Chen, W.-S., Dhingra, S., He, J., He, L., Chen, C., Balasi, J. A., Sayegh, Z., Moran Segura, C. M., Lopez-Blanco, N., Alleyne, A., Nguyen, J. V., Johnson, J. O., Marchion, D., Yoder, S. J., Messina, J. L., Sondak, V. K., Markowitz, J., Reder, N. P., Hwu, P., Mule, J. J., Chuang, J. H., Chen, P.-L.
- DOI: 10.64898/2026.09.12.751186
- Source URL: <https://doi.org/10.64898/2026.09.12.751186>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.12.751186>

Abstract: Neoadjuvant immunotherapy has transformed cancer treatment, yet the spatial molecular architecture governing response and resistance across distinct immune checkpoint blockade (ICB) regimens remains poorly defined. We assembled the largest neoadjuvant ICB (NICB) spatial multi-omics cohort to date, profiling over 112 million cells at single-cell resolution across three melanoma ICB regimens using MERFISH spatial transcriptomics, multiplexed immunofluorescence, and scRNA-sequencing. These analyses revealed the full multicellular spatial architecture of the NICB tumor microenvironment, including mature TLS with germinal centers, TCF7+ stem-like T cells, myeloid cells organized into spatially distinct cellular neighborhoods with unique intercellular signaling circuits, and CCL19/CCL21-expressing fibroblasts as a previously unrecognized stromal scaffold sustaining these immune hubs. We developed three purpose-built computational tools that together enabled comprehensive quantification of this microenvironment for the first time: SCIRA for whole-slide single-cell receptor-ligand quantification, GCSCAN for molecularly grounded TLS and germinal center structural delineation, and PathNet-TLS for automated TLS detection on H&E images. Applying these tools across the cohort, we defined the immune and stromal composition and cellular neighborhood organization distinguishing responders from non-responders. We also quantified cell-cell interactions and regimen-specific immune architectures, including a markedly stronger mature TLS/germinal center response with IPI-NIVO than NIVO-RELA. Importantly, GCSCAN-quantified TLS and germinal center density each stratified disease-free survival, with responders that lack germinal centers having an elevated risk of relapse. Open-top light-sheet imaging and CODA-based 3D reconstruction further uncovered interconnected germinal center-TLS tunnels invisible to standard 2D histopathology. These findings establish a discovery-to-tool paradigm linking single-cell tumor microenvironment interrogation to clinically deployable computational pathology for biomarker-driven NICB assessment across cancer types.

## SPIRAL: A versatile online single time-point circadian analysis platform for rice
- Source: Science Advances (journals)
- Date: 2026-09-18T00:00:00+00:00
- Categories: Genomics & sequence analysis, Tools & resources
- Authors: Yabo Shi, Li Yao, Zhenxian Han, Yingke Ma, Yufeng Xu, Xingwei Wang, Zhaoxiong Jiang, Shuyu Wang, Mian Zhou, Dong Zou, Zhang Zhang, Wei Wang
- Journal: Science Advances
- DOI: 10.1126/sciadv.aec9727
- Source URL: <https://doi.org/10.1126/sciadv.aec9727>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1126%2Fsciadv.aec9727>

Abstract: The circadian clock synchronises plant physiology with environmental oscillations to promote plant fitness. The commonly-used methods for rhythm monitoring in dicots include rhythmic leaf movement tracking and luciferase-based imaging. For monocots, however, the leaf erectness makes these methods ineffective. Leveraging over 11,000 transcriptome samples, the circadian time-course profiling, the simulation-based algorithm optimisation, and the experimental validation, we developed SPIRAL, an online single time-point circadian analysis platform for rice and unexpectedly revealed a ∼28-hour endogenous rhythm in V4-stage Nipponbare leaves, making period-matched or long-day photoperiods comparatively more permissive growth conditions for the assayed experimental system. We demonstrated the versatility of SPIRAL by quantifying global rhythm sensitivity to abiotic stresses, pinpointing when nitrogen deficiency starts to perturb rhythms, a temporal resolution surpassing that of the state-of-the-art methods, and identifying candidate components connecting the clock to stresses through factorial analyses. Online deployment of SPIRAL enables platform-independent analysis of public and user-supplied rice transcriptomes to accelerate discoveries in crop adaptation and chronoculture.

## ssJSD: A fusion of sparsity and spatial information for HiC single-cell clustering
- Source: bioRxiv (preprints)
- Date: 2026-09-18
- Categories: Genomics & sequence analysis, Single-cell & spatial
- Authors: Lee, S. W., Lin, S.
- DOI: 10.64898/2026.09.14.751457
- Source URL: <https://doi.org/10.64898/2026.09.14.751457>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.14.751457>

Abstract: Single-cell high-throughput chromatin conformation capture (scHiC) enables profiling three dimensional genome architecture at cellular resolution, providing insights into cell-to-cell variability and cellular functions. Recent frameworks utilize spatial interaction patterns to derive dissimilarity measures for downstream tasks such as cell clustering. However, the inherent sparsity and ultra-high dimensionality of scHiC contact matrices pose significant challenges. A central hurdle is that existing measures typically treat all zeros without distinction, failing to differentiate biologically meaningful structural zeros (SZs) from technical dropouts. Here, we introduce ssJSD (spatial and sparsity informed Jensen-Shannon Divergence), a computational framework designed to explicitly account for scHiC-specific sparsity patterns. By integrating band-wise contact frequency profiles with SZ-induced sparsity matrices, ssJSD leverages both spatial interaction patterns and biological absence of contacts. We adopted two complementary integration strategies: an early fusion approach that concatenates information into a single representation, and a late fusion approach that integrates JSD-based dissimilarities through diverse averaging methods. Through simulations and applications to human cell lines and prefrontal cortex data, we demonstrate that ssJSD improves clustering accuracy and effectively distinguishes cell types. Our study indicates that integrating SZ patterns is important for accurately quantifying cell-to-cell variability in 3D genomics.

## Stacked enviromic-genomic models improve prediction of genotype performance in new environments.
- Source: G3 (journals)
- Date: 2026-09-18T00:00:00Z
- Categories: Genomics & sequence analysis
- Authors: Marcos Antonio de Godoy, Maurício dos Santos Araújo, J. T. B. Chagas, J. B. Pinheiro
- Journal: G3
- DOI: 10.1093/g3journal/jkag257
- External ID: df5cc024b486ee8fa81a30be8dbaa77d85e1cb61
- Keywords: genomic
- Source URL: <https://doi.org/10.1093/g3journal/jkag257>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1093%2Fg3journal%2Fjkag257>

Abstract: Genotype-by-environment interaction is a major challenge for breeding programs, limiting the predictive ability of genomic selection in untested environments. We propose a Stacked Generalization framework that integrates linear mixed models (factor analytic and genomic best linear unbiased prediction), enviromic reaction norms, and machine learning (Extreme Gradient Boosting) to predict phenotypic plasticity. The framework was evaluated on large multi-environment trials of maize and rice, under scenarios that simulate new environments and seasons. The genetic covariance structures differed between crops, requiring a factor analytic model of order k=6 for continental maize and order k = 2 for the local rice network. Across all scenarios, the Stacking ensemble improved on the genomic baseline (M1), with gains in predictive ability from 10% (r = 0.45$ vs. 0.41 for M1) to 27.5% (r = 0.51 vs. 0.40), and reduced the root mean squared error by 30% to 43% relative to the Enviromic Reaction Norm (M2) when ensembles were selected to minimize error. These gains relied on careful feature engineering. Latent variables from genomic and environmental dimensionality reduction (principal component analysis and PaCMAP) and their interactions were the most important features for the machine learning models, and the first genomic PaCMAP component ranked first for both crops. These results indicate that non-linear dimensionality reduction is a promising tool for genomic and enviromic prediction. By combining the stability of mixed models with the flexibility of machine learning, the framework improves robustness, reduces dependence on any single model, and enhances prediction in new environments, supporting cultivation zone expansion and recommending superior genotypes.

## Stroke onset time estimation from NCCT with censoring-aware learning and robustness to infarct segmentation uncertainty
- Source: Scientific Reports (journals)
- Date: 2026-09-18T00:00:00+00:00
- Categories: Biological imaging
- Authors: Linda Vorberg, Leonhard Rist, Hendrik Ditt, Andreas Maier, Oliver Taubmann
- Journal: Scientific Reports
- DOI: 10.1038/s41598-026-71771-0
- Source URL: <https://doi.org/10.1038/s41598-026-71771-0>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1038%2Fs41598-026-71771-0>

Abstract: In patients with acute ischemic stroke and unknown symptom onset, reliable estimation of time since stroke onset is important for guiding reperfusion treatment decisions, particularly in settings where advanced imaging is unavailable. In this study, we propose a fully automated onset time estimation pipeline based on non-contrast CT (NCCT) using radiomics features extracted from automatically segmented infarct regions. A segmentation network is used to localize the infarct core, after which radiomics features are derived from both the lesion and the corresponding contralateral region. These features are used to train machine learning regressors, including kernel-based models and a multilayer perceptron, and are compared against pure attenuation-based baselines reflecting net water uptake (NWU). As approximately 25% of patients present without exact onset time and only a last-known-well time (LKWT) is available, these cases are commonly omitted from model development, further reducing already limited training cohorts. To address this problem, we investigate multiple strategies for incorporating LKWT, including surrogate target assignment, stochastic target sampling, and censoring-aware learning formulations. In addition, we assess the robustness of the deployed models to variations in infarct delineation by simulating multiple plausible segmentation boundaries at inference. The literature-parameterized NWU model remained a competitive baseline, while radiomics-based censoring-aware models achieved the lowest errors. Incorporating LKWT through censoring-aware formulations reduced the error compared with known-onset-only training, although this difference was not statistically significant. On the external test set of 32 patients with documented onset time, the censoring-aware support vector regression formulation achieved the lowest median absolute error of 1.12 h (95% BCa CI: 0.91–1.42 h). Censoring-aware formulations also showed low variability under the investigated infarct-boundary perturbations. These findings highlight the potential of NCCT-based regression models for continuous and interpretable onset time estimation. By incorporating LKWT cases, the proposed framework can expand otherwise limited training cohorts, while providing flexible predictions independent of fixed decision thresholds and demonstrating robustness to segmentation uncertainty. Together, these properties suggest that the proposed framework may have future value for clinical decision support, although validation in larger cohorts and prospective studies remains necessary before clinical translation.

## Systems biology framework for the rational design of operational conditions for in vitro/in vivo translation of tissue models
- Source: Science Advances (journals)
- Date: 2026-09-18T00:00:00+00:00
- Categories: Systems & networks
- Authors: Jose L. Cadavid, Nikolaos Meimetis, Tyler Matsuzaki, Erin N. Tevonian, Linda G. Griffith, Douglas A. Lauffenburger
- Journal: Science Advances
- DOI: 10.1126/sciadv.aef7756
- Keywords: systems biology, pathways, framework
- Source URL: <https://doi.org/10.1126/sciadv.aef7756>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1126%2Fsciadv.aef7756>

Abstract: Preclinical models are used extensively to study diseases and therapies. In vitro monoculture or microphysiological system (MPS) platforms incorporating multiple different human cell types can emulate diseased tissues, but determining experimental conditions (e.g., media supplements) that provide the most effective translatability to humans (in vivo) is a major challenge. Using metabolic dysfunction–associated steatotic liver disease (MASLD) as a case study, we developed a machine learning framework \[called LIV2TRANS (Latent In Vitro to In Vivo Translation)\] that first maps MPS onto in vivo data, then elucidates translation insights, and lastly nominates experimental conditions that increase translatability. Our findings highlight TGFβ (transforming growth factor–β) as a crucial cue for MPS translatability and indicate that adding interferon-mediated JAK (Janus kinase)-STAT (signal transducer and activator of transcription) signaling perturbations could increase the predictive performance of MPS for MASLD. Last, an optimization algorithm highlights key signaling pathways to maximize germane human-relevant information captured by this MPS. This work establishes a mathematically principled approach for identifying experimental conditions that most beneficially capture in vivo–relevant molecular processes, generalizable to a wide range of diseases where suitable molecular data exist.

## The causal integration ladder: a multilevel evidence framework for therapeutic target evaluation in cervical cancer
- Source: Frontiers in Systems Biology (journals)
- Date: 2026-09-18T00:00:00Z
- Categories: Genomics & sequence analysis, Proteins & structural biology
- Authors: Shishir Singh, M. Srivastava, Pragathi Uppada, Diksha Ameta, Aesha Singh, Monisha Banerjee, Atar Singh Kushwah
- Journal: Frontiers in Systems Biology
- DOI: 10.3389/fsysb.2026.1874355
- External ID: d793f463e4c31046018cec16fe91a37172d82775
- Keywords: genomic, framework
- Source URL: <https://doi.org/10.3389/fsysb.2026.1874355>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.3389%2Ffsysb.2026.1874355>

Abstract: Cervical-cancer genomic studies nominate many altered genes, but the evidence needed to distinguish disease association from a causal, therapeutically tractable mechanism is rarely stated explicitly. We present the Causal Integration Ladder (CIL), an evidence-accounting framework that distinguishes Level I association, inherited Level II-G evidence, acquired Level II-S evidence, Level III computational and direct functional evidence, and Level IV translational validation, while treating viral etiology as an explicit context modifier. An exploratory TCGA-CESC screen (306 tumours, 3 normal samples) supplied 25 Level I candidates. HPV annotations were available for 291 primary tumours (280 positive, 9 negative, 2 indeterminate). Restriction to HPV-positive tumours preserved the direction of all 25 Level I effects; HPV-positive versus HPV-negative comparisons were exploratory because the negative group was small and histologically heterogeneous. Somatic analysis used 194 mutation-evaluable and 295 copy-number-evaluable tumours. Ten candidates showed false-discovery-rate-significant copy-number–expression associations, including CDKN2A, whereas recurrent protein-altering mutation was uncommon. Eight genes were additionally audited using public cis-eQTL, GWAS, dependency, pharmacogenomic, and cell-compartment resources. The revised CIL reports inherited, somatic, etiological, and functional evidence independently; absence of germline support is not interpreted as evidence against somatic, viral, or functional relevance. No observational result is presented as experimental validation, and Level III direct perturbation and Level IV translational claims remain prospective.

## The PSInet Plant Water Potential Database: advancing new perspectives on plant water status, traits, and hydraulic processes
- Source: bioRxiv (preprints)
- Date: 2026-09-18
- Categories: Tools & resources
- Authors: Guo, J., Restrepo Acevedo, A. M., Browne, M., Johnson, D. M., McCulloh, K. A., Nippert, J. B., Poyatos, R., Kannenberg, S. A., Beverly, D. P., Endsley, A., Feldman, A. F., Konings, A. G., Liu, Y., Martinez-Vilalta, J., Hammond, W. M., Hultine, K. R., Lowman, L. E. L., Dukes, J. S., Green, J. K., Sack, L., Vinod, N., Hu, J., Allen, J., Adams, C. E., Adams, H. D., Adet, L. L. A., Ambrose, A., Anderegg, L. D. L., Anderegg, W. R. L., Aparecido, L. M. T., Aranda, I., Arcoverde de Mattos, E., Avila-Lovera, E., Bailey, K. C., Baldocchi, D. D., Bassiouni, M., Batllori, E., Batterton, B. E., Baugh, T.
- DOI: 10.64898/2026.09.17.752364
- Source URL: <https://doi.org/10.64898/2026.09.17.752364>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.17.752364>

Abstract: Water potential gradients drive water flow within and between soils and plants, and the internal plant water potential controls a wide range of physiological processes including photosynthesis, growth, and mortality. Notwithstanding this clear relevance for many critical aspects of ecosystem function, water potential data have historically been relatively inaccessible and unnetworked. The absence of a centralized repository for plant water potential time series limits our ability to integrate a wealth of ecophysiological information from other networks and from remote sensing. Closing this gap is necessary to address unresolved questions about plant responses to drought and heat stress, and to make confident predictions about plant and ecosystem function in a warming world. Here, we introduce the PSInet database -- a global collection of plant water potential time series from 285 datasets representing 523 species. We present the workflow that guided database development and evaluate its key features. Through a series of preliminary analyses, we then highlight the potential of the PSInet database for applications including: a) advancing plant water use strategy frameworks; b) disentangling the impacts of soil versus atmospheric drought stress; c) assessing the long-held assumption of pre-dawn equilibration of ecosystem water potential; d) understanding the risk of drought-driven mortality; and e) benchmarking remote-sensing data products and land-surface models.

## This must be the place: deep learning local adaptation
- Source: bioRxiv (preprints)
- Date: 2026-09-18
- Categories: Genomics & sequence analysis, Evolution & metagenomics
- Authors: Rodriguez, J., Cronn, R. C., Tittes, S., Kern, A. D.
- DOI: 10.64898/2026.09.16.752190
- Source URL: <https://doi.org/10.64898/2026.09.16.752190>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.16.752190>

Abstract: Climate change is increasingly disrupting the relationship between locally adapted populations and the environments in which they evolved, creating an urgent need for tools that connect genomic variation to climate. Common-garden and provenance trials remain the gold standard for characterizing local adaptation, but their time and resource requirements limit how broadly they can be applied. Genomic approaches provide a complementary path. Genotype--environment association (GEA) methods identify environmentally associated loci. Machine-learning models have also shown that geographic origin can be predicted directly from genotypes. Here we introduce EcoLocator, a supervised deep neural network that jointly predicts geographic location and climate of origin from genotypes. Through extensive simulations we demonstrate that EcoLocator accurately recovers geographic location and environment of origin from genotype data, and, with SHAP-based feature attribution, identifies adaptive loci more reliably than benchmark GEA methods. We apply our method to coastal Douglas-fir (Pseudotsuga menziesii var. menziesii), where EcoLocator predicts geographic origin (R2=0.75--0.83) and climate of origin (R2=0.52--0.72) under leave-one-out cross-validation. Notably, we find that climate predicted directly from genotypes outperforms climate inferred by first predicting geographic origin, showing that EcoLocator captures genotype--climate signal that cannot be recovered from geography alone. Our prediction errors fall within the tolerances used in existing seed-transfer guidelines, demonstrating that EcoLocator's predictions are ready for practical application, and our approach is readily extendable to other species and conservation contexts.

## TopoAdapter: a plug-and-play multi-hop topology adapter for MI-EEG decoding.
- Source: Journal of neuroscience methods (journals)
- Date: 2026-09-18
- Categories: Computational neuroscience
- Authors: Yu Cai, Qianjin Guo, Xiaozhu Lin
- Journal: Journal of neuroscience methods
- DOI: 10.1016/j.jneumeth.2026.110909
- External ID: 42759576
- Source URL: <https://doi.org/10.1016/j.jneumeth.2026.110909>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1016%2Fj.jneumeth.2026.110909>

Abstract: BACKGROUND: Motor imagery electroencephalography (MI-EEG) decoding is limited by low signal-to-noise ratio, non-stationarity, inter-subject variability, and small calibration sets. Lightweight decoders are attractive for online BCI but often learn channel relations only from limited training data. NEW METHOD: We introduce TopoAdapter, a plug-and-play input module that injects a fixed electrode-layout prior into existing EEG backbones. It builds a physical electrode graph from the montage, computes cumulative multi-hop channel-to-neighborhood contrasts, and adds a learnable low-amplitude residual while preserving input shape. RESULTS: Under an aligned 500-epoch, five-seed ATCNet protocol, mean accuracy changes were +1.17 and +0.13 percentage points on the BCI Competition IV-2a and Zhou2016 motor-imagery datasets, respectively, and +0.36 points on the High-Gamma executed-movement dataset. All three dataset-level means were positive; the BCI IV-2a and High-Gamma bootstrap intervals excluded zero, although no paired test remained significant after three-dataset Holm correction. In a separate BCI IV-2a compatibility study, all seven selected backbones improved on average and three retained Holm-adjusted Wilcoxon evidence. COMPARISON WITH EXISTING METHODS: Unlike graph neural decoders that redesign the backbone, TopoAdapter keeps downstream components unchanged. In a matched comparison with the open-source Adaptive Channel Mixing Layer (ACML), TopoAdapter attained 60.52% versus 60.42% accuracy with 70 versus 506 added parameters; the direct paired difference was unresolved. CONCLUSIONS: TopoAdapter provides an explicit, ultra-lightweight spatial prior for motor EEG decoding. Positive mean changes on two motor-imagery datasets, one executed-movement dataset, and seven heterogeneous BCI IV-2a backbones support portability and a favorable cost-benefit profile, while effect magnitude remains dataset- and subject-dependent.

## TRACEDD: A Tool-grounded Reasoning and Agentic Coordination for Explainable Drug Design
- Source: bioRxiv (preprints)
- Date: 2026-09-18
- Categories: Tools & resources
- Authors: Vangala, S. R., Kasturi, V. V., Bung, N., Roy, A.
- DOI: 10.64898/2026.09.12.751167
- Source URL: <https://doi.org/10.64898/2026.09.12.751167>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.12.751167>

Abstract: Drug discovery depends on coordinated decisions across target validation, structure analysis, molecular design, developability assessment and synthetic feasibility, but current computational methods often operate as disconnected tools. Here, we introduce TRACEDD (Tool-grounded Reasoning and Agentic Coordination for Explainable Drug Design), a framework that makes three primary contributions: (1) It establishes a 'tool-first' multi agentic architecture where LLMs orchestrate validated computational tools rather than replace them, ensuring scientific rigor. (2) It implements a multi-agent system that mirrors expert discovery teams, enabling transparent and traceable decision-making through a Reason-Act-Observe loop. (3) It demonstrates an end-to-end workflow, from target validation to synthesis planning, that adaptively handles real-world data variability, such as the absence of experimental structures. The framework decomposes discovery into specialized agents for target validation, druggability assessment, molecular generation, lead optimization, ADMET evaluation, literature evidence integration and retrosynthesis, all operating through a Reason Act Observe workflow. Using JAK2 as a representative case, we show that the system can retrieve experimental protein structures, invoke AlphaFold when structures are unavailable, identify druggable pockets and perform de novo molecular generation. Known JAK2 inhibitors are used to define design hypotheses and guide reinforcement learning-based molecular generation, with docking scores/predicted pIC50 and other physicochemical/ADMET properties serving as reward and prioritization signals. The framework demonstrates a tool-first, reasoning-driven approach in which each major decision is linked to explicit tool invocation, intermediate evidence. By combining agentic orchestration with domain-specific computational tools, the system supports transparent, adaptable and human-verifiable molecular design workflows, providing a foundation for more reliable AI-assisted drug discovery.

## Transmission of mutated SARS-CoV-2 variants is favored by relatively prolonged infections due to delayed immunity
- Source: bioRxiv (preprints)
- Date: 2026-09-18
- Categories: Evolution & metagenomics
- Authors: Owens, K., Radecki, P., Tempia, S., von Gottberg, A., Cohen, C., Boritz, E., Schiffer, J. T., Reeves, D. B.
- DOI: 10.64898/2026.09.15.751875
- Source URL: <https://doi.org/10.64898/2026.09.15.751875>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.15.751875>

Abstract: SARS-CoV-2 evolution enhanced viral fitness and immune evasion, extending the COVID-19 pandemic and resulting in millions of excess deaths. Viral diversity is generated within infected individuals, yet the timing and interplay of viral and immunological forces that drive transmissible evolution are incompletely understood. We developed a multi-scale within host phylodynamic (WiPhy) model of SARS-CoV-2 infection which couples viral replication, innate and acquired immune responses, and viral mutation. We then validated the model against quantitative viral and phylodynamic metrics. Model output predicts that typical acute infections rapidly generate genetic diversity due to accumulation of minor variants which in most cases do not achieve sufficient concentrations for transmission. Delayed innate immune responses correlate with higher peak viral load and diversification, allowing higher transmission risk of the founder virus or with a novel variant that is equally or less fit. In contrast, the risk of transmitting a fitter variant is highest during the ~10% of infections in which viral loads remain sufficiently high for transmission after 10-14 days. In these cases, non-sustained innate and/or weak acquired immune responses allow sufficient time for selection of a variant with one or more fitness enhancing non-synonymous mutations. Across a simulated cohort of ~1500 individuals, 5% of transmission risk came from variants with enhanced fitness from nonsynonymous mutations, and 13% of simulated infections accounted for 90% of fitter variant transmission risk. Our results highlight how the timing and interplay of viral and immunological forces within a host create bottlenecks that severely limit between host evolution.

## Two distinct excitability types delineate the partition between normal brain function, engram encoding, and the two phases of hyperexcitability/epileptic susceptibility
- Source: Scientific Reports (journals)
- Date: 2026-09-18T00:00:00+00:00
- Authors: A. Rabinovitch, R. Rabinovitch, D. Braunstein, E. Smolik, Y. Biton
- Journal: Scientific Reports
- DOI: 10.1038/s41598-026-70663-7
- Source URL: <https://doi.org/10.1038/s41598-026-70663-7>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1038%2Fs41598-026-70663-7>

Abstract: The conventional conceptualization of neuronal excitability as a unitary phenomenon obscures critical distinctions between synaptic and ephaptic mechanisms of neural activation. In the present investigation, we separate excitability into two independent parameters synaptic ( p ) and ephaptic ( b ) within a cellular automata framework. This separation facilitates the precise demarcation of operational regimes across the (p, b) parameter space, encompassing normal brain function (with and without engram encoding), and hyperexcitability/epileptic susceptibility phases (HEPS), including tonic and clonic manifestations. Note that hyperexcitability (HEPS) as defined here does not distinguish between cases of non-epileptic episodes and actual epileptic seizures. Simulations reveal possible contiguous HEPS domains intrinsically linked to memory (normal / encoding) processes, situated (p < 0.90) beneath the elevated synaptic excitabilities traditionally associated with epileptogenesis. Notably, this low-p HEPS region(s) could emerge within the hippocampus during engram formation, suggesting a mechanistic overlap between physiological memory encoding and possible pathological hyperexcitability. Implications for pharmacotherapy are explored, emphasizing targeted modulation of p and b to mitigate epileptic risk in individuals with varying baseline excitabilities, while preserving cognitive faculties. These findings underscore the necessity of disentangling excitability subtypes to refine diagnostic and therapeutic paradigms in neurology and cognitive science.

## UCOD: A Near-Field Benthic Organism Dataset for Underwater Visual Camouflage and Multi-Task Analysis
- Source: Scientific Data (journals)
- Date: 2026-09-18T00:00:00+00:00
- Categories: Tools & resources
- Authors: Ruixue Wang, Xuanhe Chu, Xinyu Zhao, Ximan Zhao, Shuhao Zhang, Miaoxin Lu, Ruyi Chen, Chunlei Zhan, Zhuo Chen, Junwen Tian, Jie An, Minyi Xu, Zhiying Jiang, Xianping Fu, Yongjun Gong, Siyuan Liu
- Journal: Scientific Data
- DOI: 10.1038/s41597-026-08254-4
- Keywords: dataset
- Source URL: <https://doi.org/10.1038/s41597-026-08254-4>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1038%2Fs41597-026-08254-4>

Abstract: Accurate near-field underwater visual perception plays a crucial role in marine ecological monitoring and benthic resource exploration. However, the superimposition of the biomimetic characteristics of benthic organisms and the optical degradation caused by the water medium frequently induces significant underwater visual camouflage phenomena. This causes the target foreground and the background to become highly fused in terms of color, texture, and structure, thereby severely limiting the performance of underwater vision algorithms in core tasks such as object detection, image segmentation, and 3D reconstruction. Existing public datasets predominantly focus on salient targets in clear water or under simple backgrounds, lacking comprehensive multi-task benchmarks explicitly tailored for visual camouflage scenarios. To address this gap, we constructed an Underwater Camouflaged Object Dataset (UCOD) for near-field benthic organisms, designed for multi-task analysis to jointly support image enhancement, object detection, pixel-level segmentation, and 3D scene reconstruction. The dataset comprises 7,000 high-resolution RGB images, including 3,500 images with detection annotations, 3,500 images with segmentation masks, and 16 reconstruction sequence folders for 3D reconstruction. It covers six categories of benthic organisms exhibiting typical camouflage characteristics: scallops, fish, conches, abalones, starfish, and sea cucumbers. During data acquisition, calibrated underwater imaging equipment was employed, and the kinematic parameters of the acquisition platform were controlled to improve the stability and spatial consistency of the collected data. The dataset provides useful training data and an evaluation basis for underwater multi-task perception research under visual camouflage conditions, while offering complementary data support for studies on underwater optical sensing and intelligent exploration.

## Unsupervised learning of mapping between brain lesions and behavior
- Source: bioRxiv (preprints)
- Date: 2026-09-18
- Authors: Wahle, I. A., Griffis, J., Adolphs, R., Grafman, J., Tranel, D., Boes, A., Eberhardt, F.
- DOI: 10.1101/2023.12.22.573110
- Source URL: <https://doi.org/10.1101/2023.12.22.573110>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1101%2F2023.12.22.573110>

Abstract: Human lesion studies offer one of the most direct routes to investigating the relations between brain regions and behavioral outcomes in circumstances where experimental interventions are highly restricted. However, these studies face a major challenge in identifying the right level of granularity at which brain regions and behavioral outcomes should be analyzed to identify the relation between lesions to specific brain regions and specific behavioral outcomes. Here we showcase a novel data-driven approach, Causal Feature Learning (CFL), that learns the appropriate level of analysis and the relation between lesion and cognitive impairment at the same time. The method avoids specifying brain regions and specific outcome measures a priori, allowing for the discovery of new cross-cutting lesion-behavior maps. We show that CFL robustly recovers lesion behavior maps in a simulated dataset where Canonical Correlation Analysis fails to provide interpretable results. We then show that CFL recovers known lesion-behavior maps for language deficits and visuospatial processing using a large dataset of lesion subjects, and we illustrate how CFL can be used to identify new groupings of outcomes when mapping lesions to depression symptoms.

## VCCV: conservative transcriptomic corroboration for measurement prioritization of computational drug–target hypotheses
- Source: Bioinformatics (journals)
- Date: 2026-09-18T00:00:00+00:00
- Categories: Proteins & structural biology, Tools & resources
- Authors: Haihui Huang, Yanan Zhou, Dingkui Kang, Yong Liang
- Journal: Bioinformatics
- DOI: 10.1093/bioinformatics/btag694
- Source URL: <https://doi.org/10.1093/bioinformatics/btag694>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1093%2Fbioinformatics%2Fbtag694>
- Code: <https://github.com/bio-ai-source/VCCV>

Abstract: Motivation Computational drug-target interaction (DTI) models nominate plausible binders but cannot determine which candidate best accounts for an observed cellular response. Perturbational transcriptomics offers orthogonal mechanistic evidence, yet pharmacology-to-genetics mismatch, incomplete reference coverage, and non-specific stress programs make simple signature matching unreliable. This motivates a principled integration layer that corroborates hypotheses conservatively, abstains under global mismatch, and prioritizes informative follow-up measurements when the evidence remains ambiguous. Results We present Virtual-to-Cellular Corroboration for Validation (VCCV), a model-agnostic posterior-triage layer for pre-trained DTI models. VCCV updates calibrated DTI working weights with context-aligned perturbational evidence using a near-identity affine map and exact covariance transport. Within the stated Gaussian class, exact transport is the unique uncertainty update that preserves posterior-odds comparisons under invertible affine changes of measurement coordinates. VCCV also introduces an empirical warning branch for abstention and converts residual ambiguity into compact follow-up gene panels, using a submodular objective for deep near-ties. Each query is assigned one of three actionable states: a target-resolved hypothesis, an abstention, or a prioritized panel. Across five DTI models, VCCV improved discrimination (paired ROC-AUC gains 0.021-0.037) and reduced negative log-likelihood in every case. Additional evaluations showed improved ranking on the same-cell-supported endpoint, warning-score discrimination of strong-response profiles (ROC-AUC 0.876), and better recovery of full-coordinate leading hypotheses by selected panels than by random panels. Across these retrospective kinase-focused evaluations, VCCV provides a principled bridge from computational nomination to conservative, measurement-directed cellular corroboration. Availability and implementation Source code of VCCV is publicly available at https://github.com/bio-ai-source/VCCV. Supplementary information Supplementary data are available online.

## Viral Burden: New insights into Estimating IgG Recognition Across Human Populations
- Source: bioRxiv (preprints)
- Date: 2026-09-18
- Categories: Proteins & structural biology
- Authors: Harhala, M. A., Gembara, K., Nelson, D. C., Konieczny, A., Jedruchniewicz, N., Rybicka, I., Dabrowska, K.
- DOI: 10.64898/2026.09.16.751894
- Keywords: epitope, peptides, proteome, antibodies, epitopes
- Source URL: <https://doi.org/10.64898/2026.09.16.751894>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.16.751894>

Abstract: Assessment of viral burden in populations is essential for understanding virus epidemiology and herd immunity potential. While serological profiling is straightforward at the individual level, it remains challenging at scale. We propose analysis advancing in serological technologies for broader population-level analysis and comparison. We used an epitope library in Phage Display ImmunoPrecipitation (PhIP) technology (VirScan type library) to assess IgG recognition of 49,630 representative viral oligopeptides in 134 serum samples from two populations (Poland and the US). Only 5.9% of oligopeptides were immunogenic, yet IgG recognition of viruses was consistent across populations - over 90% of virus species, 99% of genera, and 97% of families were detected. Shannon Diversity Index analysis supported these findings. Among immunogenic peptides, 9.1% were significantly more frequently recognized, though not correlated with recognition strength. We further proposed a normalization method to account for differences in viral proteome representation when assessing immune burden: the burden score. Finally, we demonstrated how this approach facilitates serological comparisons between populations. These observations show that while people are exposed to similar viruses, they produce antibodies against different viral epitopes. This individual variability, combined with broad virus recognition, likely strengthens population-level protection. Epitope recognition frequency seems to be shaped more by population exposure than by magnitude of response that an epitope can induce. Accurately measuring viral burden can inform healthcare planning, and antiviral technologies development.

## VirPLM: Antigenic prediction of influenza A/H3N2 viruses with a fine-tuned protein language model
- Source: Bioinformatics (journals)
- Date: 2026-09-18T00:00:00+00:00
- Categories: Genomics & sequence analysis, Proteins & structural biology, Tools & resources
- Authors: Xingyi Li, Kexin Xiao, Chunyan Zhou, Xiangting Jia, Dongmin Zhao, Jialuo Xu, Xianying Zeng, Jianzhong Shi, Xuequn Shang, Junnan Zhu, Huihui Kong
- Journal: Bioinformatics
- DOI: 10.1093/bioinformatics/btag692
- Source URL: <https://doi.org/10.1093/bioinformatics/btag692>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1093%2Fbioinformatics%2Fbtag692>
- Code: <https://github.com/xingyili/VirPLM>

Abstract: Motivation Human influenza A/H3N2 viruses undergo rapid antigenic evolution primarily driven by the hemagglutinin subunit 1 (HA1). Within HA1, amino acid substitutions under immune pressure cause antigenic drift, necessitating frequent updates to vaccine strains. While hemagglutination inhibition (HI) assays remain the gold standard for assessing antigenic relationships, their labor-intensive and low-throughput nature limits scalability. Fortunately, the rapid accumulation of HA1 sequences enables sequence-based antigenic prediction, yet effectively extracting informative representations from these viral sequences remains challenging. Results In this study, we present VirPLM, a two-stage framework that adapts the ESM-2 protein language model to H3N2 HA1 sequences for antigenic prediction. VirPLM significantly outperforms representative methods and maintains robust performance under both cross-validation and retrospective time-split evaluations. Moreover, VirPLM identifies highly critical sites enriched in known regions related to antigenic evolution. In the season-specific coverage analysis, VirPLM-prioritized strains achieve higher estimated coverage rates than the corresponding historical strains recommended by the World Health Organization in most evaluated seasons, suggesting that VirPLM can provide complementary sequence-based evidence for candidate strain prioritization. Availability and Implementation The source code is available at https://github.com/xingyili/VirPLM, and the version used in this study is archived in Zenodo (DOI: 10.5281/zenodo.21650323). Supplementary Information Supplementary information is available at Bioinformatics online.

## Virtual experiments bridge sequence and microscopy with generative models
- Source: bioRxiv (preprints)
- Date: 2026-09-18
- Categories: Proteins & structural biology, Biological imaging
- Authors: Zheng, D., Hong, K., Huang, B.
- DOI: 10.64898/2026.09.13.751243
- Source URL: <https://doi.org/10.64898/2026.09.13.751243>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.13.751243>

Abstract: Large-scale screening and mapping efforts have produced vast libraries of perturbation-readout data. Converting these measurements into mechanistic insights requires models that link perturbation and genetic input to phenotypes, i.e., labels from experimental readouts, which are usually task specific. We propose a different, virtual experiment modeling approach: train generative models to recreate readouts conditioned on the experimental context, and then let established downstream models extract phenotypes from the synthetic data. As an illustrative case, we develop a bidirectional sequence-image generative framework, CELL-FM, that maps protein sequence and cellular context to fluorescence microscopy images and back, enabling in silico localization prediction, image-conditioned functional motif analysis and generation, and large-scale virtual mutagenesis revealing the amino acid features controlling condensate formation of intrinsically disordered peptides. This approach decouples representation learning from task-specific annotation, reuses rich experimental modalities across many downstream tasks, and preserves the spatial and organizational detail that hand-crafted labels often discard.

## Widespread detection of Polyethylene glycol reveals chronic human exposure through pharmaceutical drugs
- Source: bioRxiv (preprints)
- Date: 2026-09-18
- Categories: Tools & resources
- Authors: Gouda, H., Kelly, P., Whiley, L., Gomez-Perez, D., Mascellani Bergo, A., Goncalves Nunes, W. D., Chekmeneva, E., Yuen, A. H. Y., David, M., McKirdy, S., Nelson, A., Smith, D. L., Zhao, H. N., Seo, J. I., Mathias, E. J., Farrell, G., Scheurink, T., Strobel, M., Mannochio-Russo, H., Gkikas, K., Wang, M., Havlik, J., Quince, C., Takats, Z., Lewis, M., Gerasimidis, K., Rattray, N. J. W., Dorrestein, P. C.
- DOI: 10.64898/2026.09.11.750999
- Source URL: <https://doi.org/10.64898/2026.09.11.750999>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.11.750999>

Abstract: Polyethylene glycol (PEG) is a synthetic polymer ubiquitous in pharmaceuticals, personal care products, food additives, and industrial manufacturing. Despite its widespread use and potential importance as an exposure chemical, the prevalence of PEG exposure and its excretion in human populations remain largely uncharacterized. Moreover, PEG detected in human biofluids is frequently assumed to arise from analytical contamination during sample preparation, potentially obscuring its contribution to the human exposome and confounding metabolic phenotyping studies. Here, we show that PEG is as component of the human xenobiotic exposome and is further metabolized into PEG hydoxy acid and diacid metabolites in humans. We further find that PEG exposure is associated with alterations in microbiome composition and short-chain fatty acid metabolism, suggesting its biological impact of its exposure. We identify PEG exposure in approximately 2.3% of publicly available metabolomics data files and provide a reusable 85,484 candidate PEG and PEGylated MS/MS spectral library for future use for the metabolomic community. Together, these findings establish PEG signal in human biofluids can reflect genuine exposure, and that PEG exposure is neither metabolically inert nor biologically silent.

## Will there be a warning for the next pandemic?
- Source: Science Advances (journals)
- Date: 2026-09-18T00:00:00+00:00
- Authors: Justin Lessler, C. Jessica E. Metcalf
- Journal: Science Advances
- DOI: 10.1126/sciadv.aef2194
- Source URL: <https://doi.org/10.1126/sciadv.aef2194>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1126%2Fsciadv.aef2194>

Abstract: How to best allocate resources to combat the threat of pathogen emergence remains an important open question. Using archetypes that characterize the link from genotype to fitness in both zoonotic reservoirs and humans, we show that, across a range of plausible conditions, emergences of pathogens with pandemic potential in humans are unlikely to be preceded by detectable, failed, attempts. Yet, the number of “failed” emergence events contains information about the emergence potential of zoonotic pathogens, and should modify our beliefs about the underlying fitness landscape. Our work suggests that the most important, modifiable, risk factors for emergence may be phenomena that alter fitness landscapes, such as viral ecology in bridge species and human immunological landscapes.

## Zero-inflated Joint Species Distribution Models for improved partial-correlation network inference from community data
- Source: bioRxiv (preprints)
- Date: 2026-09-18
- Categories: Evolution & metagenomics
- Authors: Tous, J., Chiquet, J., Deacon, A. E., Fontrodona-Eslava, A., Fraser, D. F., Magurran, A. E.
- DOI: 10.1101/2025.07.24.666553
- Source URL: <https://doi.org/10.1101/2025.07.24.666553>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1101%2F2025.07.24.666553>

Abstract: 1. A long-term goal of community ecology has been to decipher the mechanisms that shape the spatio-temporal organization of species communities. Understanding these processes is critical to predicting the responses of ecological communities to environmental change. To this end, Joint Species Distribution Models (JSDMs) offer statistical tools to analyze community data, identify the impact of abiotic factors on them and study inter-species correlations in their distributions. In particular, the JSDM-inferred partial-correlation networks allow one to identify direct links between species that can help decipher the mechanisms that shape their joint distribution. 2. Community data based on species counts often contains numerous zeros. However, not accounting for these zeros in a data set is known to hinder parameter inference. We investigate this issue in the context of JSDMs, and ask what impact it can have on the inference of partial-correlation networks. 3. We propose a novel JSDM, the ZIPLN-network model, based on the PLN-network (Poisson log-normal network) and ZIPLN (Zero-Inflated Poisson log-normal) model, which models count data while including zero-inflation and infers a partial-correlation network. Using simulated data, we compare the results obtained by this model with existing JSDMs in terms of association network inference from abundance data containing structural zeros. We then illustrate the ZIPLN-network approach using real data from tropical freshwater fish communities. 4. Simulations show that zero-inflation can significantly bias the inference of partial-correlation networks from community data and that the ZIPLN-network model efficiently counterbalances these effects. The ZIPLN-network approach is widely applicable to community data, delivers ecologically-insightful analyses, helps distinguish amongst potential mechanisms, and aids better understanding of community assembly rules. We provide guidance for getting started with our approach.

## Neural Learning as a Game Induced by Spike-Timing-Dependent Plasticity
- Source: arXiv (preprints)
- Date: 2026-09-17T22:31:51Z
- Categories: Computational neuroscience
- Authors: Xinhao Fan, Shreesh P. Mysore
- External ID: 2609.21131v1
- Source URL: <https://arxiv.org/abs/2609.21131v1>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Farxiv.org%2Fabs%2F2609.21131v1>
- PDF: <https://arxiv.org/pdf/2609.21131v1>

Abstract: A general framework for inferring the computational role of spike-timing-dependent plasticity (STDP) does not currently exist. Here, we develop a game-theoretic description for a canonical, convergent neural circuit with general STDP and postsynaptic-potential (PSP) kernels. Parity-matched STDP-PSP interactions induce a potential game among presynaptic neurons, whereas parity-mismatched interactions induce a zero-sum game. These components, respectively, implement contrastive PCA and flow selection; their weighted interplay determines learning dynamics and computation for arbitrary STDP rules.

## Hybrid quantum-classical attention for histopathology-based molecular profiling in data-limited cancers
- Source: arXiv (preprints)
- Date: 2026-09-17T22:00:55Z
- Categories: Biological imaging
- Authors: Kahn Rhrissorrakrai, Aritra Bose, Aldo Guzman-Saenz, Filippo Utro, Laxmi Pardia
- External ID: 2609.21115v1
- Source URL: <https://arxiv.org/abs/2609.21115v1>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Farxiv.org%2Fabs%2F2609.21115v1>
- PDF: <https://arxiv.org/pdf/2609.21115v1>

Abstract: Molecular profiling from routine histopathology could expand access to precision oncology when sequencing is unavailable, tissue is limited, or training cohorts are small. We developed a hybrid quantum-classical strategy that replaces softmax attention in a transformer for histopathology-based gene expression prediction with a quantum-derived doubly stochastic matrix (QDSM). Across 29 cancer cohorts from The Cancer Genome Atlas and an independent pancreatic cancer cohort from the Clinical Proteomic Tumor Analysis Consortium, QDSM attention produced selective gains, with the largest relative improvements in smaller, data-limited cohorts, including adrenocortical carcinoma and uveal melanoma. Rather than improving transcriptome-wide performance uniformly, QDSM redistributed predictive accuracy across genes and pathways, improving biologically relevant targets in some tumor contexts while worsening others. In adrenocortical carcinoma, preferentially improved genes were enriched for adverse overall-survival associations, linking enhanced molecular inference to prognostically relevant biology. In pancreatic cancer transfer experiments, QDSM improved selected metabolic and lineage-associated genes but did not consistently improve performance under cross-cohort shift. Leave-one-cancer-out mixed-effects analysis showed that baseline molecular features predicted part of the gene-level benefit, while residuals identified cancer-specific programs that improved more or less than expected. Separate experiments on IBM quantum processors recovered the doubly stochastic matrix primitive underlying the attention mechanism. These findings position QDSM attention as a context- and target-dependent strategy for image-based molecular profiling and molecular triage when direct testing is unavailable, incomplete, or impractical.

## Retention-Constrained Post-Training Quantization of Cellpose-SAM for Stem Cell Microscopy
- Source: arXiv (preprints)
- Date: 2026-09-17T19:55:15Z
- Categories: Biological imaging
- Authors: Sebastián A. Cruz Romero
- External ID: 2609.21038v1
- Source URL: <https://arxiv.org/abs/2609.21038v1>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Farxiv.org%2Fabs%2F2609.21038v1>
- PDF: <https://arxiv.org/pdf/2609.21038v1>

Abstract: Induced pluripotent stem cell (iPSC) culture increasingly relies on segmentation foundation models, yet deployment on laboratory CPUs and edge hardware requires compression schemes that are both efficient and auditable. We present a deployment-oriented evaluation of compressed Cellpose-SAM using a pre-specified retention criterion: the 95% cluster-bootstrap interval of mean change from FP32 must remain above a fixed -0.02 margin for every imaging modality. On a stratified 176-field panel spanning BBBC038 nuclei, BBBC039 U2OS fluorescence, and NIST iPSC images across density regimes, weight-only W8A16 preserves instance F1 across all modalities. A sensitivity-guided mixed W4/W8 scheme, using four INT8 exceptions, achieves a 6.76x reduction in weight storage with no observed catastrophic failures (0/176 fields), matching W8A16 at this sample size. In contrast, ternary weight-only quantization achieves 12.08x compression but fails catastrophically on 169/176 fields. These results demonstrate that compression should be evaluated by modality-stratified downstream retention rather than single-number accuracy, and establish a reproducible protocol for auditing compressed foundation models in regulated stem-cell imaging.

## Flow, dynamics and active fracture in hydraulic multicellular systems
- Source: arXiv (preprints)
- Date: 2026-09-17T19:16:37Z
- Categories: Mathematical biology & statistics
- Authors: John D. Treado, Arthur Boutillon, Frank Jülicher, Otger Campàs
- External ID: 2609.21021v1
- Source URL: <https://arxiv.org/abs/2609.21021v1>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Farxiv.org%2Fabs%2F2609.21021v1>
- PDF: <https://arxiv.org/pdf/2609.21021v1>

Abstract: From interstitial space to luminal cavities, fluid pressure and flow can remodel, reshape and even redefine a biological tissue. Fluids can either govern or react to mechanical interactions between cells. However, measuring flows at cellular scales is difficult, which makes it challenging to understand tissue hydraulics. Here, we develop a theoretical approach that captures cellular mechanics and fluid flow in one framework. We find that hydraulics can drastically influence tissue behavior. Hydraulic coupling between cell shape and size governs a tissue's response to osmotic shock, while tuning a tissue's permeabilities can channel fluid either between or across cell membranes. In active tissues, hydraulics can suppress cell mobility to the point of fracture, where we discover a hydraulic ratchet that drives fluid out of cells to generate small luminal spaces. We find experimental evidence that hydraulics can suppress cell motion in early stage zebrafish embryos injected with a thickening agent, which indicates that hydraulics may generally govern the behaviors of many multicellular systems.

## ERCPMP-Gx: Endoscopic Image and Video Dataset for Morphological, Histopathological, and Genomic Characterization of Colorectal Polyposis
- Source: arXiv (preprints)
- Date: 2026-09-17T17:59:28Z
- Categories: Genomics & sequence analysis, Biological imaging, Tools & resources
- Authors: Zahra Ghaffari, Massih Bahar, Mojgan Forootan, Ali Darvishi, Hamidreza Bolhasani
- External ID: 2609.20815v1
- Keywords: genomic, histopathological, histopathology, dataset
- Source URL: <https://arxiv.org/abs/2609.20815v1>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Farxiv.org%2Fabs%2F2609.20815v1>
- PDF: <https://arxiv.org/pdf/2609.20815v1>

Abstract: Hereditary polyposis syndromes can be precursor lesions to colorectal cancer and are associated with a broad spectrum of extracolonic tumors. Early identification and accurate classification of these syndromes are essential for timely diagnosis, individualized patient management, and targeted surveillance strategies for affected families. However, public endoscopic datasets are largely organized around the individual sporadic polyp, and none links the polyposis phenotype to histopathology and germline findings at the patient level. Here, we present ERCPMP-Gx, an endoscopic, histopathological, and genomic dataset developed to support the application of artificial intelligence (AI) in the recognition, characterization, and classification of colorectal polyposis. Most procedures were performed using the Olympus EVIS X1 system with white-light endoscopy (WLE), narrow-band imaging (NBI), magnifying NBI (M-NBI), and NBI with near focus modes, yielding 160 images and accompanying video clips. Approximately eighty percent of cases represent clinically and/or genetically confirmed hereditary polyposis syndromes (PG), including familial adenomatous polyposis (FAP), Peutz-Jeghers syndrome (PJS), juvenile polyposis syndrome (JPS), and ganglioneuroma syndrome (GNS), while the remaining twenty percent comprise non-hereditary polyps and polyp-mimicking lesions with overlapping morphological features (Non-PG), included to support differential classification. Each released record is linked, where available, to standardized endoscopic annotations, representative histopathology, and clinically reported germline findings, forming an AI-ready, patient-level annotation framework. The dataset is publicly accessible at Mendeley (https://doi.org/10.17632/nzyfc544bx.2). For the latest updates and further information, readers are referred to the DataBioX website: https://databiox.com.

## Three biological data bills, with OpenAI support
- Source: Stephen Turner (feeds)
- Date: 2026-09-17T15:57:10+00:00
- Categories: Blog
- Source URL: <https://blog.stephenturner.us/p/biodata-bills-nsceb>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fblog.stephenturner.us%2Fp%2Fbiodata-bills-nsceb>
- Abstract: not stored for this record.

## The Motile-Units model: Interacting spins model of cell polarization and motility
- Source: arXiv (preprints)
- Date: 2026-09-17T15:40:18Z
- Categories: Mathematical biology & statistics
- Authors: Jonathan E. Ron, Nir S. Gov
- External ID: 2609.20587v1
- Source URL: <https://arxiv.org/abs/2609.20587v1>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Farxiv.org%2Fabs%2F2609.20587v1>
- PDF: <https://arxiv.org/pdf/2609.20587v1>

Abstract: We introduce a coarse-grained interacting-spin model for two-dimensional cell motility, in which the cell perimeter is discretized into stochastic binary spins that switch between active and inactive states. Each perimeter spin represents a "motile-unit" that is a source of protrusive force and retrograde flow when active. Long-range interactions between the motile-units arise through a polarity cue advected by the collective actin retrograde flow, providing a minimal realization of spontaneous symmetry breaking and self-propulsion. The model exhibits three dynamical phases, a random walk phase, persistent random walk phase, and an intermittent bistable phase characterized by run-and-tumble migration. Additional nearest-neighbor interactions modulate speed and persistence without altering the overall phase structure. Owing to its simplicity, the framework naturally incorporates external cues, reproducing chemotactic migration, steering by localized optogenetic activation, and directional decision-making (symmetry breaking) under competing stimuli. The model introduces a new class of active-particle model in which both speed and polarity emerge from internal stochastic spin dynamics, rather than being imposed as particle-level variables, offering a framework for the study of cell migration and extends the scope of active-matter physics.

## Dynamic Generalized Gromov-Wasserstein Optimal Transport
- Source: arXiv (preprints)
- Date: 2026-09-17T10:17:33Z
- Categories: Genomics & sequence analysis, Single-cell & spatial, Biological imaging
- Authors: Junda Ying, Zhiwei Zeng, Peijie Zhou, Lei Zhang
- External ID: 2609.20008v1
- Source URL: <https://arxiv.org/abs/2609.20008v1>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Farxiv.org%2Fabs%2F2609.20008v1>
- PDF: <https://arxiv.org/pdf/2609.20008v1>

Abstract: Gromov--Wasserstein optimal transport (GW-OT) extends classical optimal transport by introducing structure-aware transport cost. This is particularly relevant for spatial transcriptomics, where dynamical reconstruction should preserve tissue structure in addition to matching expression patterns. While static formulations have been widely used for such structure-aware alignment, a general dynamic formulation for reconstructing continuous trajectories is still missing. We introduce Travelling Pair Dynamical Alignment and Trajectory Estimation (TP-DATE), a theoretical and computational framework to generalize GW-OT dynamically in a simulation-free manner. We formulate a broad class of static and dynamic Quadratic-form OT (QOT) through path actions and prove the static dynamic equivalence. We further develop travelling-pair flow matching, which allows interacting conditional paths and marginalizes their interactions into a single vector field. On synthetic and real spatial transcriptomics data, TP-DATE better preserves spatial structure and improves continuous 3D dynamics reconstruction.

## CellRFT: Reinforcement Fine-Tuning for Single-Cell Perturbation Modeling
- Source: arXiv (preprints)
- Date: 2026-09-17T09:43:12Z
- Categories: Genomics & sequence analysis, Single-cell & spatial
- Authors: Jie Yan, Li Liu, Hanze Guo, Jiaxin Hu, Houxin He, Xiaoning Qi, Haoran Wang, Cong Li, Zhong-Yuan Zhang, Yong Wang
- External ID: 2609.19970v1
- Source URL: <https://arxiv.org/abs/2609.19970v1>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Farxiv.org%2Fabs%2F2609.19970v1>
- PDF: <https://arxiv.org/pdf/2609.19970v1>

Abstract: Predicting cellular responses to perturbations supports the study of gene function, disease mechanisms, and therapeutic strategies. Despite advances in single-cell perturbation modeling, existing models typically optimize surrogate losses that do not directly reflect the biological criteria used for evaluation, so better data fitting need not yield better biological predictions. To address this mismatch, we introduce \\textbf\{CellRFT\}, a reinforcement fine-tuning framework that uses biological evaluation as direct training feedback. CellRFT uses policy-gradient optimization to learn from non-differentiable evaluations of generated cell populations and integrates multiple biological rewards through hierarchical reward aggregation. Comprehensive experiments demonstrate CellRFT's applicability across different pretrained models and effectiveness in improving perturbation prediction, reveal that optimizing one biological criterion can help or hinder others, and show that complementary rewards can improve criteria beyond those directly optimized, offering a way to probe how biological metrics shape model behavior, with the potential to inform evaluation design. Code will be made available.

## Benchmarking: the foundation of reliable AI in biomedicine
- Source: EMBL (feeds)
- Date: 2026-09-17T08:56:04+00:00
- Categories: Blog
- Source URL: <https://www.embl.org/news/science-technology/benchmarking-the-foundation-of-reliable-ai-in-biomedicine/>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fwww.embl.org%2Fnews%2Fscience-technology%2Fbenchmarking-the-foundation-of-reliable-ai-in-biomedicine%2F>
- Abstract: not stored for this record.

## Identifying Damage Pathways Linking Sequence Composition to Storage Failure in DNA Data Storage via High-Dimensional Mediation Analysis
- Source: arXiv (preprints)
- Date: 2026-09-17T07:31:05Z
- Categories: Mathematical biology & statistics
- Authors: Jingyi Li, Huaming Wu, Haixiang Zhang
- External ID: 2609.19822v1
- Source URL: <https://arxiv.org/abs/2609.19822v1>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Farxiv.org%2Fabs%2F2609.19822v1>
- PDF: <https://arxiv.org/pdf/2609.19822v1>

Abstract: DNA data storage offers extraordinary information density and long-term durability, but its reliability is limited by sequence-dependent errors introduced during synthesis and accumulated during storage. It remains unclear how sequence composition is associated with storage failure through specific molecular damage components. We develop a high-dimensional semiparametric mediation framework for survival outcomes. GC content is treated as the exposure, a high-dimensional baseline damage spectrum (a vector of per-read damage counts stratified by trinucleotide context and error type) as the mediator, and storage-quality failure as the outcome. Nonlinear covariate effects in both the mediator and survival models are approximated using deep neural networks. A three-step procedure combining product-of-coefficients screening, Smoothly Clipped Absolute Deviation (SCAD) penalized estimation, and joint significance testing is developed for mediator selection and inference. Applied to an aging experiment on electrochemically synthesized DNA, the method identifies 14 significant mediators, all corresponding to single-base deletions, with estimated mediated effects concentrated in trinucleotide contexts ending in C. These results reveal deletion-type damage as a major pathway linking sequence composition to reduced archival reliability and suggest candidate sequence features for future optimization and error-control strategies. The proposed framework thus offers a mechanism-oriented statistical approach for understanding and improving the reliability of DNA data storage.

## Learning-Based Reconstruction of Optical Properties in Bilayered Media from Single-distance Time-Resolved Reflectance Measurements
- Source: arXiv (preprints)
- Date: 2026-09-17T06:52:51Z
- Categories: Biological imaging
- Authors: Caterina Amendola, Giulia Maffeis, Lorenzo Buffoni, Lorenzo Chicchi, Francesco Coghi, Duccio Fanelli, Raffaele Marino, Fabrizio Martelli, Riccardo Paoli, Lorenzo Pattelli, Lorenzo Spinelli
- External ID: 2609.19786v1
- Source URL: <https://arxiv.org/abs/2609.19786v1>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Farxiv.org%2Fabs%2F2609.19786v1>
- PDF: <https://arxiv.org/pdf/2609.19786v1>

Abstract: The inverse problem of reconstructing optical properties, specifically absorption and scattering coefficients, in layered biological media from time-domain reflectance measurements remains a significant challenge for traditional analytical models. Inverse solvers based on the diffusion equation often struggle with structural heterogeneity, frequently yielding poor accuracy for superficial absorption and deep-layers scattering. In this work, we propose a machine learning framework as an alternative approach to reconstruct the optical properties of a bilayered medium, benchmarking its efficiency and accuracy against model-based algorithms. To overcome the intrinsic approximations of diffusion theory and inverse reconstruction, we generated a robust synthetic dataset of forward DTOF using exact Monte Carlo simulations at multiple source-detector distances. A machine learning pipeline was then trained on this dataset and validated against state-of-the-art model-based reconstruction methods. Besides the significant reconstruction speed-up, the machine learning approach achieves higher accuracy than model-based inverse solvers, further providing an estimate of the parameter space dimensionality without requiring any a priori information about the number of layers in the investigated geometry. Further enhancements in the reconstruction accuracy can be expected in future extensions of this work, by training the pipeline over multiple DTOF curves from the same medium, in a joint multi-distance reconstruction approach.

## TorchCraft: Unified binder design by inverting an all-atom structure predictor
- Source: arXiv (preprints)
- Date: 2026-09-17T06:37:20Z
- Categories: Proteins & structural biology, Tools & resources
- Authors: TorchCraft Team, Yu Liu, Zhouhanyu Shen, Zhengyi Li, Xikun Huang, Jiaqi Liu, Shuxian Gao, Qilin Yu, Xiayan Qin, Yucheng Zhang, Mingchen Chen
- External ID: 2609.19770v1
- Source URL: <https://arxiv.org/abs/2609.19770v1>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Farxiv.org%2Fabs%2F2609.19770v1>
- PDF: <https://arxiv.org/pdf/2609.19770v1>

Abstract: All-atom structure predictors model diverse molecular interactions, but using their learned structural priors for binder design remains challenging. Here we present TorchCraft, a unified binder-design framework that optimizes sequence logits through a frozen all-atom predictor. Implemented in TorchFold, TorchCraft combines confidence, contact, geometric, and sequence-prior objectives within a shared optimization procedure for minibinders, framework-conditioned VHHs, cyclic peptides, and ligand-binding proteins. Using pretrained AlphaFold 3 weights, TorchCraft generated representative minibinders and VHHs with experimentally measured binding across four targets in each format, without post hoc sequence redesign. Computational benchmarks further demonstrated the framework's applicability to cyclic peptides and ligand-conditioned pocket design. TorchCraft extends predictor inversion to multiple binder formats and molecular contexts, providing a common framework for reusing all-atom structural priors in design.

## The segmentation ceiling: why explicit left-ventricular masks do not improve learned ejection-fraction regression
- Source: arXiv (preprints)
- Date: 2026-09-17T05:41:28Z
- Authors: Farshid Farhadi Khouzani, Paul La Plante, Bryar Mustafa Shareef, Laxmi Gewali
- External ID: 2609.19730v1
- Source URL: <https://arxiv.org/abs/2609.19730v1>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Farxiv.org%2Fabs%2F2609.19730v1>
- PDF: <https://arxiv.org/pdf/2609.19730v1>

Abstract: Accurate estimation of left ventricular ejection fraction (EF) from echocardiography is central to cardiovascular care, and deep learning enables automated EF prediction from echocardiographic video. Because EF is clinically derived from left-ventricular (LV) volumes, a widely held intuition is that explicit LV segmentation should improve prediction. We introduce a quantitative criterion, the segmentation ceiling, that makes this testable: from EF as a normalized difference of end-diastolic and end-systolic volumes, we derive in closed form how per-frame segmentation area error propagates into EF error, and thus the accuracy a mask must reach before it can improve on direct regression. Using EchoNet-Dynamic, a UniFormer-S backbone, and the empirically measured within-patient error correlation, the criterion places the break-even near 10% per-frame area error, whereas a representative segmenter operates at roughly 14%, above the ceiling. Consistent with this, four strategies for injecting segmentation or area information (a predicted-mask channel, end-diastolic/end-systolic clip sampling, and per-bin and amplitude area-consistency objectives) fail to beat a raw-video baseline; ground-truth masks help only through label leakage. Input representation thus not being the limit, we identify generalization as the practical lever: weight averaging with strong augmentation attains a test R^2 of 0.806 (MAE 4.08) under a matched dense-clip protocol, comparable to an R(2+1)D baseline (0.811) while tightening the validation-to-test gap. Finally, a heteroscedastic beta-NLL formulation yields informative, well-calibrated per-prediction uncertainty, larger for clinically harder low-EF cases, where Monte-Carlo dropout does not. The segmentation ceiling gives a concrete design criterion for when mask-guided EF estimation is worthwhile, plus a simple, uncertainty-aware recipe for EF regression.

## Matrix Graphical Model Via Joint Estimation of Partial Correlations
- Source: arXiv (preprints)
- Date: 2026-09-17T05:22:52Z
- Categories: Proteins & structural biology
- Authors: Hyewon Kim, Seongoh Park
- External ID: 2609.19718v1
- Source URL: <https://arxiv.org/abs/2609.19718v1>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Farxiv.org%2Fabs%2F2609.19718v1>
- PDF: <https://arxiv.org/pdf/2609.19718v1>

Abstract: Matrix graphical models aim to characterize conditional dependence structures in matrix-variate data under a separable covariance assumption. In this framework, the precision matrix is decomposed as a Kronecker product, enabling separate modeling of undirected graphs across row and column domains. Existing methods have been developed for this problem, including likelihood-based approaches and regression-based procedures for graph estimation. Likelihood-based methods estimate precision matrices directly and recover graph structures indirectly, whereas regression-based approaches directly target estimating edges among variables, thus outperforming the former. However, existing regression-based methods are based on multiple penalized regression problems, which naturally yields asymmetry in estimated graphs and computational difficulty in selecting tuning parameters. To address the limitations, we propose a joint estimation of partial correlations in matrix graphical models. The proposed method estimates all partial correlations simultaneously within a unified optimization framework, thereby preserving symmetry and easing the pain of selecting the best models. Numerical studies demonstrate that the proposed method improves graph recovery performance compared to existing approaches. We also analyze protein expression data collected from patients with pulmonary tuberculosis, measured repeatedly at multiple time points, where the proposed method compares protein networks between two groups of patients and recovers the temporal dependence structure.

## Improving Sample Efficiency in Peptide-HLA Binding Prediction with Hybrid Quantum-Classical Neural Networks
- Source: arXiv (preprints)
- Date: 2026-09-17T03:33:06Z
- Categories: Proteins & structural biology
- Authors: Chenyan Jia, Cong Guo, Siyue Chen, Pengpeng Ye, Xiaochun Chen
- External ID: 2609.19642v1
- Source URL: <https://arxiv.org/abs/2609.19642v1>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Farxiv.org%2Fabs%2F2609.19642v1>
- PDF: <https://arxiv.org/pdf/2609.19642v1>

Abstract: Peptide-HLA binding prediction is a critical step in neoantigen identification for personalized cancer immunotherapy and holds significant clinical value. However, the training data available for many HLA alleles are extremely limited, which severely constrains the performance of conventional methods on this task. Parameterized quantum circuits are hypothesized to induce inductive biases beneficial for learning from small datasets, yet their application to biological sequence prediction remains underexplored. To address this, we propose a hybrid quantum-classical neural network (HQNN) specifically designed for peptide-HLA binding prediction. HQNN integrates multi-source biological feature encoding with parallel quantum feature extractors and a quantum-enhanced classifier. On two HLA alleles (A\*02:01 and B\*07:02), HQNN outperforms a parameter-matched classical CNN baseline across all training sizes, with the performance gap widening as training data decreases. Ablation studies confirm the respective contributions of the quantum feature extraction module and the quantum classifier. In noise-aware simulations, performance degrades only mildly, and such degradation is reasonable and acceptable under realistic quantum hardware noise levels. These results suggest that hybrid quantum-classical architectures can provide practical sample-efficiency gains for immunoinformatics tasks in low-data regimes.

## Large Language Model Agents for Evidence Based Genetic Disease Severity Classification
- Source: arXiv (preprints)
- Date: 2026-09-17T01:56:38Z
- Categories: Genomics & sequence analysis
- Authors: Tohid Ghasemnejad, Ahmadreza Argha, Mark Grosser, John Wang, Min Yang, Thantrira Porntaveetus, Tony Roscioli, Nigel H. Lovell, Mahmoud Aarabi, Hamid Alinejad-Rokny
- External ID: 2609.19569v2
- Keywords: genomic, language model
- Source URL: <https://arxiv.org/abs/2609.19569v2>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Farxiv.org%2Fabs%2F2609.19569v2>
- PDF: <https://arxiv.org/pdf/2609.19569v2>

Abstract: Disease severity classification for genetic conditions is subjective and labor-intensive, creating bottlenecks in genomic screening, where commercial panels vary widely in size and overlap. We developed an autonomous AI agent integrating Reasoning and Acting (ReAct) with Retrieval-Augmented Generation (RAG) to classify 10,211 Human Phenotype Ontology terms. It uses American College of Medical Genetics (ACMG)-endorsed severity guidelines and American College of Obstetricians and Gynecologists (ACOG) quality-of-life criteria to retrieve PubMed literature, generate interpretable reasoning chains, and independently verify claims. At the phenotype level, using expert-curated cohorts, the agent achieved 93.55% accuracy (MCC 0.9237) with 82.6% to 91.4% of claims supported by direct evidence or valid inferences. Gene-level severity was aggregated across 8,738 pairs, identifying 3,283 autosomal recessive pairs with severe or profound presentations. External validation showed 95.2% concordance with Mackenzie's Mission gene list. This system enables standardized panel design by providing reliable, automated classification supported by direct evidence.

## 3D Segmentation of Pathological Muscle with a Physics-Informed Latent-Regularized U-Net
- Source: medRxiv (preprints)
- Date: 2026-09-17
- Categories: Biological imaging
- Authors: Mehrabi, N., Pegard, N. C. R., Handsfield, G. G.
- DOI: 10.64898/2026.09.15.26363155
- Source URL: <https://doi.org/10.64898/2026.09.15.26363155>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.15.26363155>

Abstract: ObjectiveWe developed a data-efficient deep learning framework for three-dimensional segmentation of pathological musculoskeletal anatomy from magnetic resonance imaging (MRI) when only limited manual annotations are available. MethodsWe developed a Physics-Informed Latent-Regularized U-Net (PILR-U-Net) that combines transfer learning from healthy MRI, latent-space anatomical regularization, and elasticity-based physics-informed constraints. The physics-informed loss enforces mechanical equilibrium, near-incompressibility, and spatial smoothness of predicted deformation fields. The framework was evaluated on MRI datasets from 50 participants with cerebral palsy across 15 lower-limb musculoskeletal structures using sparse manual annotations. Performance was assessed using volumetric overlap, boundary accuracy, volume error, sensitivity, and precision. Ablation experiments evaluated the individual contributions of latent-space and physics-informed regularization. ResultsOur experimental results show accurate segmentations with three-dimensional Dice coefficients ranging from 0.750 to 0.943 across evaluated structures, while most structures exhibited low surface-distance errors. Predicted deformation fields maintained positive Jacobian determinants near unity and smooth strain-energy distributions. Ablation analysis showed that both regularization components improved performance, with removal of physics-informed regularization producing the largest reductions in Dice accuracy and increases in boundary error. ConclusionPILR-U-Net enables accurate and anatomically plausible segmentation of pathological musculoskeletal MRI under sparse supervision. SignificanceIncorporating anatomical priors and biomechanical constraints into deep segmentation networks may reduce dependence on extensive pathological annotations and support patient-specific musculoskeletal modeling and clinical analysis.

## A directed TM->JM coupling in receptor tyrosine kinase dimers, set by activating mutations and the membrane environment
- Source: bioRxiv (preprints)
- Date: 2026-09-17
- Categories: Proteins & structural biology
- Authors: Sato, T., Tamagaki-Asahina, H.
- DOI: 10.64898/2026.05.22.727318
- Source URL: <https://doi.org/10.64898/2026.05.22.727318>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.05.22.727318>

Abstract: The direction of conformational coupling in a membrane protein, that is, which domain drives which, has been inaccessible to experiment. We recover this directivity from molecular dynamics (MD) of transmembrane-juxtamembrane (TM-JM) dimers of receptor tyrosine kinases EGFR and FGFR3. Coupling is detected with a Bayesian-network framework (CASCADE); its direction is measured with PERI (Phase-plane Estimation of Rotational Irreversibility), the net phase-plane circulation, validated on synthetic data and resolved at 0.1 ns. Direction is summarized as the TM\[->\]JM directed-mass fraction f+ (0.5 = balanced) via a hierarchical Bayesian model. The activating TM mutants EGFR L658Q and FGFR3 A391E are TM-JM (posterior probability 0.95 and 0.99); fluid wild-type EGFR leans the same way (0.93), in agreement with its experimentally reported constitutive activity in fluid but not ordered bilayers; the ligand-dependent ordered wild type is balanced (0.45); and an activating mutation raises the TM-JM bias above the ordered wild type with probability 0.94. The directivity thus tracks the measured activity state of the receptor, distinguishing signaling-competent from ligand-dependent RTK dimers by a property not apparent from structure alone.

## A Drug-Specific, Half-Life-Adjusted Framework for Classifying CNS-Active Systemic Therapy Exposure During and After Radiotherapy
- Source: medRxiv (preprints)
- Date: 2026-09-17
- Authors: Pari Mitre, L., Drapkin, B., Dohopolski, M.
- DOI: 10.64898/2026.06.11.26354463
- Source URL: <https://doi.org/10.64898/2026.06.11.26354463>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.06.11.26354463>

Abstract: Clinical oncology datasets often store systemic therapy as a regimen label with a start date and an end date. Those records are clinically recognizable but can be analytically incomplete when the research question concerns whether a patient was exposed to a concurrent CNS-active drug (cCNS-aD) or an adjuvant CNS-active drug (aCNS-aD) around radiotherapy. Contemporary CNS-oncology studies usually define CNS activity by empiric drug lists and define concurrency by fixed calendar windows, although the literature shows substantial heterogeneity across both concepts. This paper proposes a generalizable framework for converting raw systemic therapy records into reproducible cCNS-aD and aCNS-aD variables, useful in subgrouping for clinical studies. The framework uses a transparent CNS scoring model based on three clinical evidence components: intracranial objective response rate, consensus CNS endorsement, and intrathecal route of administration. It then defines a pharmacokinetic exposure proxy as the recorded end date plus five half-lives plus a drug-specific steady-state accumulation term, to account for repeated dosing. Concurrent exposure is classified by overlap with the radiotherapy interval, while post-radiotherapy exposure is classified by overlap with a prespecified post-RT attribution window. The framework separately identifies post-RT pharmacokinetic persistence and post-RT treatment initiation, allowing investigators to distinguish continued exposure from true adjuvant initiation. This is a methodological framework and reference implementation. Implementation audits and endpoint-specific sensitivity analyses remain necessary before use as a definitive exposure classifier.

## A FAIR layer for the INHERENT haemoglobinopathy patient registry
- Source: medRxiv (preprints)
- Date: 2026-09-17
- Categories: Tools & resources
- Authors: Tamana, S., Yiangou, C., Orphanou, K., Xenophontos, M., Papasavva, P. L., Bernabe, C., Roos, M., Wijnbergen, D., Kersloot, M. G., Cornet, R., Minaidou, A., Stephanou, C., Chatzimatthaiou, S., Landi, A., Giannuzzi, V., Bonifazi, F., Lederer, C. W., Kountouris, P.
- DOI: 10.64898/2026.09.16.26363187
- Source URL: <https://doi.org/10.64898/2026.09.16.26363187>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.16.26363187>

Abstract: Haemoglobinopathy registries support research and outcome monitoring. Still, reuse is limited due to heterogeneous structures, registry-specific coding, and incomplete semantic representation. We designed and implemented a FAIRification workflow for the INHERENT haemoglobinopathy platform, an international genotype-phenotype registry, as part of the HemaFAIR project. The workflow was extended to the Cyprus Haemoglobinopathy Patient Registry to demonstrate its applicability across a second registry. Source data and metadata were transformed through two independent but complementary harmonisation branches executed in parallel: one producing an OMOP CDM representation, and the other generating a CARE-SM representation. The workflow generated graph-based semantic resources, predefined query services, public aggregate dashboards, and application programming interface (API) endpoints. Registry metadata were published through the European Rare Disease Registry Infrastructure and a FAIR Data Point. These outputs support findability, interoperable analysis and controlled reuse while preserving existing governance over patient-level data.

## A Gene-Program Architecture of Mouse T cells
- Source: bioRxiv (preprints)
- Date: 2026-09-17
- Categories: Genomics & sequence analysis, Single-cell & spatial, Tools & resources
- Authors: Zhang, Z., Wang, T., Panigrahi, S. S., Carbonetto, P., Stephens, M., Benoist, C., Mostafavi, S., Brbic, M., Zemmour, D., the immgenT Project,
- DOI: 10.64898/2026.09.15.751918
- Source URL: <https://doi.org/10.64898/2026.09.15.751918>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.15.751918>

Abstract: We present immgenT-GP, a gene-program framework for resolving mouse T cell heterogeneity across the immgenT atlas. On ~ 800,000 T cells spanning lineages, organs, and immune challenges, we defined 200 reproducible gene programs, discovered by empirical Bayes matrix factorization approach and validated through a new deep learning approach, that capture major axes of T cell variation, including lineage identity, activation states and tissue location. Gene-program analysis complemented cluster-based annotation by decomposing T cell states into molecular modules, revealing quantitative, shared, modules not represented with discrete labels alone. Across tissues, gene programs reflected both tissue-imposed programs and changes in cluster composition. Integrating GP activity with cell-surface marker expression from the CITE-seq data, revealed that markers can report different programs depending on lineage and context. Together, immgenT-GP extends the atlas from a map of T cell states to a molecular reference of the programs that underlie them.

## A hundred years of dinosaur research in Mexico: skeletal completeness and phylogenetic information of the dinosaur fossil record in Mexico
- Source: Frontiers in Ecology and Evolution (journals)
- Date: 2026-09-17T00:00:00Z
- Authors: O. R. Regalado Fernández
- Journal: Frontiers in Ecology and Evolution
- DOI: 10.3389/fevo.2026.1882629
- External ID: d1742e813c72126fd77e9ba049166a4ef20de4b2
- Source URL: <https://doi.org/10.3389/fevo.2026.1882629>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.3389%2Ffevo.2026.1882629>

Abstract: The dinosaur fossil record from Mexico is sparsely documented since most of it is fragmentary remains with little taxonomic information and a lot of it continues undescribed in institutions and university collections. There are several large-scale analyses that have attributed the low alpha-diversity to paleobiogeographic processes or to low research activity. Like most large-scale analyses based on big data, the material from the collection (physical fossil record) becomes dissociated from difficult to code contextual information as an occurrence is transformed into a taxonomic name (abstracted fossil record). Although several studies have attempted to evaluate biases in the fossil record — preservation, sampling, research interest, identification effort — information processing disconnects the initial observer from the extracted data. In this study, this problem is approached from the perspective of database management by developing a function (estimator) that can represent the observed material and the research effort invested on it. The estimator is developed as a cross product between a skeletal completeness metric and a character coverage metric that can produce absolute values or be ranked relative to every element of the database. This estimator, named here as phylogenetic estimator (pê), reflects the status of distribution of hypotheses of homologies assigned to specific groups. Although it would be expected that two random variables (skeletal completeness and character coverage) behave in a chaotic and non-linear manner, they reflect more covariance with each other and show strong correlation. High skeletal completeness values produce a leverage effect making the linear regression consistent with the intuition that more complete specimens are more informative. After removing the high skeletal completeness, the two variables become more disjointed with low skeletal completeness scores, suggesting that fragmentary remains can still be potentially informative. The estimator pê applied to the Mexican dinosaur fossil record behaves according to the geological framework and shows potential in capturing information related to the local depositional environment. When performing data wrangling, there is a general intuition that fragmentary remains are less informative than complete specimens, and data tends to be disregarded based on this assumption. From the point of view of database management in the age of big data, the estimator pê enables communicating first-hand observers, those in contact with the material, with the users who will work on the data in the future, finding gaps in the literature and having more objective criteria to discard occurrence data, connecting the physical with the abstracted fossil record.

## A Latent Inflammatory Tissue-State Variable Mechanistically Links Radiotherapy-Induced Immune Remodeling to Recurrent Tumor Permissiveness
- Source: bioRxiv (preprints)
- Date: 2026-09-17
- Categories: Systems & networks
- Authors: Mayeaux, M. A., Zhou, X. M., Rafat, M.
- DOI: 10.64898/2026.09.14.751444
- Source URL: <https://doi.org/10.64898/2026.09.14.751444>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.14.751444>

Abstract: Triple-negative breast cancer recurrence following radiotherapy is associated with a microenvironment characterized by immune dysfunction and persistent inflammation. We developed an experimentally constrained agent-based model to investigate how transient immune remodeling becomes a persistent recurrence-permissive tissue state. The model reproduced experimentally observed macrophage recruitment and phenotype dynamics but demonstrated that recurrent recruitment, impaired inflammatory resolution, and adaptive immune bias were insufficient to reproduce the recurrent macrophage ecology. We therefore introduced recurrence-associated microenvironmental inflammation (RAMI), a latent tissue-state variable representing accumulated unresolved inflammatory remodeling. Coupling RAMI to the emergence of experimentally constrained interleukin-6 signaling generated tissue-to-cell feedback that reinforced recurrence-associated macrophage phenotypes and increased tumor establishment, supporting inflammatory tissue memory as a mechanistic intermediary between transient immune perturbation and persistent recurrent tumor permissiveness.

## A Molecularly Anchored Spatial Transcriptomic Framework for Precise CA1–Subiculum Parcellation and Region-Resolved Analysis in Alzheimer’s Disease
- Source: GigaScience (journals)
- Date: 2026-09-17T00:00:00+00:00
- Categories: Single-cell & spatial
- Authors: Yuyang Liu, Youzhe He, Yanrong Wei, Pan Wang, Chunyu Huang, Quyuan Tao, Langjian Zhu, Xun Xu, Longqi Liu, Shiping Liu, Lei Han, Jing Zhang, Lifang Wang
- Journal: GigaScience
- DOI: 10.1093/gigascience/giag094
- Source URL: <https://doi.org/10.1093/gigascience/giag094>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1093%2Fgigascience%2Fgiag094>

Abstract: Background The precise molecular delineation of the interface between the Subiculum (Sub) and cornu ammonis 1 (CA1) is a challenge in hippocampal research, as conventional cytoarchitectural boundaries are often ambiguous and limit reproducible regional annotation. Here, we developed a molecularly anchored spatial transcriptomic framework to define CA1-Sub regional identities using high-definition spatial transcriptomics (Stereo-seq) and single-nucleus RNA sequencing (snRNA-seq) references. Findings Using a human hippocampal Stereo-seq dataset from 12 donors, we established a data-driven parcellation framework that defines reproducible molecular features distinguishing CA1 and Sub while capturing the transition between these regions. FN1 was identified as a Sub-enriched marker in a subset of EX\_Sub and, together with ETV1 and additional regional markers, enabled molecular assignment of CA1 and Sub identities across datasets. The Sub association of FN1 and ETV1 was further supported by human 10X Genomics spatial transcriptomics, mouse in situ hybridization data, and a mouse spatial transcriptomic dataset. Applying this framework to Alzheimer’s disease (AD) tissues revealed region-specific transcriptional alterations across CA1 and Sub, including enrichment of mitochondrial energy metabolism-related transcripts in the Sub, suggesting exploratory transcriptional associations of altered metabolic function. Conclusions This study provides a molecularly anchored framework for human CA1–Sub parcellation that complements conventional annotation. By defining regional molecular states while preserving the biological continuum across CA1–Sub interface, this approach enables more consistent regional analysis of human hippocampus tissue across donors, datasets, and disease conditions.

## A nanopore-based next-generation sequencing workflow for comprehensive adventitious virus testing
- Source: Scientific Reports (journals)
- Date: 2026-09-17T00:00:00+00:00
- Authors: Michael Karbiener, Tim Walker, Stephen Rudd, Natalia Garcia-Garcia, Jens Modrof, W. Paul Duprex, Veronica L. Fowler, Thomas R. Kreil
- Journal: Scientific Reports
- DOI: 10.1038/s41598-026-72037-5
- Source URL: <https://doi.org/10.1038/s41598-026-72037-5>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1038%2Fs41598-026-72037-5>

Abstract: Traditional adventitious virus testing of biological drugs has missed the presence of viral contaminants in the past. Next-generation sequencing (NGS) has emerged as a promising additional tool with unprecedented capability for breadth of detection. This study introduces a Nanopore-based NGS workflow for detection of all classes of replicating viruses in banked cells employed in medicinal biotechnology (e.g., production cell lines, cell lines employed for adventitious agent testing \[AAT\], allogeneic cell therapies). Starting with isolation of total cellular RNA, cDNA library preparation was designed to detect also viruses which do not poly-adenylate their transcripts while retaining strand-specific information. The newly generated bioinformatic analysis pipeline provides versatility with respect to the use of databases for host sequence depletion (publicly available, customized) and convenience features (direct comparison of sample to run-specific negative control, implemented visualization of sample read alignment to virus database entries). Via deliberate infection of typical biotechnological cell lines, the workflow was found to specifically identify diverse classes of viruses. Preliminary results are available within 26 h, a feature particularly useful for emergency situations requiring fast decision-making, e.g., to investigate potentially positive signals in traditional AAT.

## A Real‐Data‐Driven Framework for Evaluating Differential Transcript Usage Methods Across Long‐Read Bulk, Single‐Cell, and Spatial Transcriptomics
- Source: Advanced Science (journals)
- Date: 2026-09-17T00:00:00Z
- Categories: Genomics & sequence analysis, Tools & resources
- Authors: Chenxing Zhang, Jun Liu, Qi Zhao, Hui-Long Yin, Ang-Ang Yang, Min-Hua Zheng, Rui Zhang
- Journal: Advanced Science
- DOI: 10.1002/advs.77756
- External ID: ef5a2c6289ca024ddb5bfce0842a6eb69256e21f
- Source URL: <https://doi.org/10.1002/advs.77756>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1002%2Fadvs.77756>

Abstract: Differential transcript usage (DTU) analysis reveals transcript‐level regulation in alternative splicing. With the rapid adoption of long‐read sequencing in bulk, single‐cell, and spatial transcriptomics, reliable evaluation of DTU methods under real biological conditions becomes essential. Current evaluation frameworks mainly rely on simulated data, which can introduce bias and may not reflect true regulatory mechanisms. A real‐data‐driven framework is constructed to evaluate DTU methods across long‐read bulk, single‐cell, and spatial transcriptomics. The framework includes two key components. First, a biologically grounded reference transcript set is defined using RNA‐binding protein (RBP) knockout or knockdown RNA‐seq data together with experimentally validated RBP‐transcript interactions. Second, a DTU‐specific evaluation metric, the transcript set enrichment score, is introduced to quantify how effectively a method prioritizes reference transcripts in ranked results. The framework is systematically validated for reliability, unbiasedness, stability, effectiveness, and robustness using multiple real RNA‐seq datasets. Supported by this validation, ten representative DTU methods are evaluated across long‐read and short‐read data, revealing performance differences across data types. Beyond evaluating DTU methods, the framework is further extended to predict transcript‐level RBP activity, recovering perturbed RBPs more consistently than gene‐level differential expression strategies. Together, this study establishes a biologically interpretable and data‐driven standard for DTU method evaluation.

## A segmentation-guided CNN–Vision transformer feature fusion framework for multi-class breast ultrasound image classification
- Source: PLOS One (journals)
- Date: 2026-09-17T00:00:00+00:00
- Categories: Biological imaging
- Authors: Meiru Wu, Jian Wang
- Journal: PLOS One
- DOI: 10.1371/journal.pone.0353628
- Source URL: <https://doi.org/10.1371/journal.pone.0353628>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1371%2Fjournal.pone.0353628>

Abstract: Breast ultrasound (BU) imaging is widely used for detecting breast abnormalities because it is cost-effective, non-invasive, and suitable for dense breast tissue. However, multi-class classification of BU images is considered a challenging task due to low contrast, speckle noise, and overlapping visual patterns between benign and malignant tumours. To address this issue, we develop a segmentation-guided CNN and Vision Transformer based feature fusion framework for efficient multi-class BU image classification. The framework first applies a lesion segmentation model to identify the region of interest by using Fusion-Enhanced Transformer (FET) Unet model. The FET Unet model uses CNNs and Swin Transformers integrated and constructs a UNet-like architecture. In the second step, CNN-based and Vision Transformer-based features are extracted from the segmented lesion regions. In the third step, the CNN and Vision Transformer features are fused to generate a robust feature representation. A deep neural network classifier comprising four dense layers with 1,024, 512, 256, and 128 neurons, respectively, followed by a softmax output layer, is then developed using the fused features to classify breast ultrasound images into benign, malignant, and normal categories. The proposed framework was evaluated using the publicly available Breast Ultrasound Images (BUSI) dataset, which contains 780 ultrasound images collected from women aged 25–75 years, including 437 benign, 210 malignant, and 133 normal cases. Numerical results show that the proposed framework showed 95.11% of accuracy, 95.66% of sensitivity and 97.63% of specificity and F1 score of 0.944644. Based on the classification accuracy, the effectiveness of the proposed hybrid framework is demonstrated in comparison with previously reported methods for multi-class breast ultrasound image classification.

## A spatially and temporally aligned contrast-non-contrast cardiac CT dataset of pigs
- Source: Scientific Data (journals)
- Date: 2026-09-17T00:00:00+00:00
- Authors: Øyvind Nordbø, Muhammad Fahad, Rune Sagevik, Kevin Mikkelsen, Eli Grindflek, Mohib Ullah, Faouzi Alaya Cheikh, Frode Johannesen, Kim Samuel Sollie, Marianne Oropeza-Moe
- Journal: Scientific Data
- DOI: 10.1038/s41597-026-08293-x
- Source URL: <https://doi.org/10.1038/s41597-026-08293-x>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1038%2Fs41597-026-08293-x>

Abstract: Traditionally, cardiac CT imaging relies on the use of contrast agents. However, in some cases such agents cannot be used, and there is a need to develop novel AI-based algorithms to accurately segment cardiac structures from CT images, recorded without the use of contrast agents. To address this challenging problem, we have collected a spatially and temporally aligned cardiac CT dataset of pigs with and without the use of contrast fluid. This dataset consists of 10 different pigs, CT-scanned multiple times between 15 and 130 kg to cover a wide set of body weights. A subset of the dataset is manually labelled, to enable development of supervised learning-based segmentation techniques.

## A Thermodynamic Framework Linking Black Box Growth models with Genome Scale Metabolite models
- Source: bioRxiv (preprints)
- Date: 2026-09-17
- Categories: Systems & networks, Mathematical biology & statistics
- Authors: PUJIANG, J., Wang, D., Shi, H.
- DOI: 10.64898/2026.09.14.751451
- Source URL: <https://doi.org/10.64898/2026.09.14.751451>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.14.751451>

Abstract: Genome scale metabolic models provide detailed mechanistic descriptions of cellular metabolism, whereas black box growth models capture physiological behaviors using a small number of effective parameters. However, the quantitative relationship between these two modeling scales remains unclear. In this study, we develop a thermodynamic framework that connects black box growth models with thermodynamically constrained genome scale metabolic models. By linking black box model parameters with Gibbs energy dissipation rates derived from genome scale metabolic models, we demonstrate that coarse grained physiological descriptions can be obtained from detailed metabolic networks while preserving their thermodynamic foundation. We validate this framework in both Escherichia coli and yeast, showing that the resulting black box models reproduce key physiological behaviors, including growth dynamics, biomass yield, and overflow metabolism observed experimentally. Our results indicate that black box growth models and genome scale metabolic models are connected through shared thermodynamic constraints, revealing a consistent thermodynamic basis across different levels of metabolic description. This framework provides a general approach for integrating detailed metabolic networks with simple black box growth models and enables efficient multiscale modeling of cellular metabolism.

## A trainable language model with potential to modulate translation rates in non-model organisms by generating upstream untranslated region sequence libraries
- Source: PLOS One (journals)
- Date: 2026-09-17T00:00:00+00:00
- Categories: Genomics & sequence analysis
- Authors: Alexander D. Duggan, Matthew P. Newman, David R. McMillen
- Journal: PLOS One
- DOI: 10.1371/journal.pone.0348455
- Source URL: <https://doi.org/10.1371/journal.pone.0348455>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1371%2Fjournal.pone.0348455>

Abstract: Tuning protein expression in non-model organisms is often constrained by the lack of validated genetic parts and predictive design tools. Translational tuning through the modulation of upstream untranslated regions (5′-UTRs) offers a potentially organism-agnostic route, but existing methods typically rely on mechanistic assumptions, prior knowledge that may not be available in non-model contexts, or the screening of sequence libraries. Here, we present a simple generative approach for creating synthetic 5′-UTR libraries based solely on the genomic sequence statistics of any desired organism. The method uses a sliding-window n-gram language model applied to native 5′-UTR sequences to produce novel sequences that preserve organism-specific base distributions and motifs without hard-coding specific motifs or mechanistic rules into inflexible statistical templates. We have applied this approach to the model bacterium Escherichia coli and the non-model probiotic Limosilactobacillus reuteri . Libraries of approximately 1,000 sequences were generated for each organism, from which about 100 unique sequences were experimentally tested for translation of a fluorescent reporter protein. In both organisms, the synthetic libraries yielded a broad range of translation levels from this relatively small number of tested variants. Sequences derived from an organism’s own genomic statistics provided a more uniformly distributed range of translation rates in that organism than sequences derived from the other species. Correlations of individual sequence performance across the two species were weak, and thermodynamic predictions of ribosome binding strength showed very little predictive power, especially in the non-model L. reuteri . The results demonstrate that simple statistical language model approaches applied to genomic data can generate functional translational regulatory sequence libraries without detailed mechanistic knowledge or explicit reference to consensus motifs. The approach requires minimal computational resources, avoids reproducing native sequences, and can be readily applied to any organism with a sequenced genome. This strategy may lower technical barriers to expression tuning in non-model organisms.

## Ambient, real-time digitization and datafication of glass slide microscopy towards AI-at-the-microscope
- Source: Nature Communications (journals)
- Date: 2026-09-17T00:00:00+00:00
- Categories: Biological imaging
- Authors: Cooper Maira, Max S. Cooper, Kimberly L. Ashman, Andrew B. Sholl, Sharon E. Fox, Shams Halat, David Manthey, Roni Choudhury, Jonathan Sears, Carola Wenk, J. Quincy Brown, Brian Summa
- Journal: Nature Communications
- DOI: 10.1038/s41467-026-77887-1
- Keywords: microscopy, microscope
- Source URL: <https://doi.org/10.1038/s41467-026-77887-1>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1038%2Fs41467-026-77887-1>

Abstract: Pathology remains central to clinical diagnosis, yet adoption of digital pathology is constrained by financial, operational, and workflow burdens of fully digital infrastructure. We introduce HistoCAM, a platform for ambient, real-time datafication and digitization of glass-slide microscopy that preserves microscope workflows. A 31-megapixel, high space-bandwidth-time-product camera and custom software application stream and composite the pathologist’s eyepiece view, passively generating multi-resolution images from 2X to 40X while recording magnification use, search paths, and dwell times. These outputs provide immediate workflow uplift through digital annotation, measurement, quality assurance, and real-time integration of configurable AI tools. Simultaneously, HistoCAM links image content with expert interaction data and supports rapid generation of annotated, pre-embedded training data during routine slide review. By converting routine microscopy into an AI-ready data stream without requiring additional acquisition steps, HistoCAM provides a practical bridge to computational pathology while creating process-aware datasets that capture how pathologists examine and interpret tissue.

## An archaic reference-free method to jointly infer Neanderthal and Denisovan introgressed segments in modern human genomes
- Source: Molecular Biology and Evolution (journals)
- Date: 2026-09-17T00:00:00+00:00
- Categories: Genomics & sequence analysis, Evolution & metagenomics, Tools & resources
- Authors: Léo Planche, Anna Ilina, María C Ávila-Arcos, Flora Jay, Emilia Huerta-Sanchez, Vladimir Shchur
- Journal: Molecular Biology and Evolution
- DOI: 10.1093/molbev/msag235
- Source URL: <https://doi.org/10.1093/molbev/msag235>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1093%2Fmolbev%2Fmsag235>

Abstract: Admixture between populations is a common feature of human history. Admixture events introduce new genetic variation that can fuel evolution. Characterizing the significance of admixture events on the evolution of populations across various species is of great interest to evolutionary geneticists. Local Ancestry Inference (LAI) methods infer genetic ancestry of an individual at a particular chromosomal location. Certain methods specialize in detecting archaic introgression, which consists of interbreeding between modern and archaic humans like Neanderthals and Denisovans. Most current LAI methods allow the detection of a single archaic ancestry, and post-processing may distinguish between multiple waves of introgression. These methods vary in how they choose archaic or modern reference genomes for the inference. Here, we present a new HMM-based method (DAIseg), which has the advantage of simultaneously distinguishing between multiple waves of ancient and recent admixture, using only modern human reference genomes. Simulations demonstrate that DAIseg achieves higher overall performance than state-of-the-art methods. We also apply DAIseg to Papuan populations to jointly detect Denisovan and Neanderthal introgressed segments, and identify a higher number of archaic segments than previous methods. Analysis of inferred introgressed segments, shows that we can identify evidence for two Denisovan introgression events in Papuans. Overall, on top of being able to deal with both Archaic and recent admixture, DAIseg provides a more principled approach for detecting and classifying Denisovan and Neanderthal segments which will improve downstream analysis of introgressed segments to infer the impact of archaic introgression in humans.

## An explicit birth-death-reticulation model for studying the diversification of phylogenetic networks
- Source: bioRxiv (preprints)
- Date: 2026-09-17
- Categories: Evolution & metagenomics, Mathematical biology & statistics
- Authors: May, M. R., Rothfels, C. J.
- DOI: 10.64898/2026.09.11.750987
- Source URL: <https://doi.org/10.64898/2026.09.11.750987>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.11.750987>

Abstract: Much of the history of life is reticulate and better represented by phylogenetic networks than by strictly bifurcating trees. Understanding the processes that generated that history thus requires models of diversification (speciation and extinction) that incorporate reticulation. However, we currently lack tractable reticulate diversification models. Here we develop a simple birth-death-reticulation model that includes unidirectional and bidirectional gene flow, homoploid hybrid speciation, and allopolyploidization, and derive a practical probability density function for networks under this model. We demonstrate that the model can extract information about reticulation processes from known phylogenetic networks. We also explore the empirical utility of the model using an allopolyploid network of ferns of the family Cystopteridaceae, revealing evidence in favor of the controversial hypothesis that polyploids have lower diversification rates than their diploid relatives. While the model represents an advance in our ability to learn about the diversification of reticulate lineages, we also identify significant statistical, computational, and empirical challenges that face this nascent framework.

## Attachment site, not linker length, bounds tethered base editor windows
- Source: bioRxiv (preprints)
- Date: 2026-09-17
- Authors: Mahaboob Ali, A. A., Nelson, E. J. R.
- DOI: 10.64898/2026.09.16.751980
- Source URL: <https://doi.org/10.64898/2026.09.16.751980>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.16.751980>

Abstract: Base editors act on the DNA strand displaced within an R-loop, and the set of positions they convert, the activity window, has been engineered for a decade on the assumption that linker length and attachment geometry determine it. We built a geometric model that predicts where a tethered deaminase acts from linker statistics, steric exclusion, and R-loop geometry alone; fitted three parameters to two previously reported profiles, held them fixed, and scored predictions against 50 architectures from seven studies, with each substrate coordinate withheld. Varying the contour length by a factor of 16 does not shift the predicted window at all, whereas changing the attachment site does: across 185 buildable single-linker designs at thirteen attachment sites, the predicted peak never leaves protospacer positions 3 to 12, positions 1, 2 and 13 to 20 are reached by no design, and transfer to Cas12a fails by five to six nucleotides in a way that localizes to the fusion junction rather than to reach, sterics or substrate. Tether geometry, therefore, bounds where a fused deaminase can act without predicting where it does so, making attachment site rather than linker length the effective design variable.

## Auditing bacterial dark-gene screens for superimposed open reading frame artefacts: A multi-layer analysis of Rv2438A in Mycobacterium tuberculosis.
- Source: Journal of microbiological methods (journals)
- Date: 2026-09-17T00:00:00Z
- Authors: C. Guyeux
- Journal: Journal of microbiological methods
- DOI: 10.1016/j.mimet.2026.107715
- External ID: 32fc7b63f8ae6cc25fb49ae062b6dda923d67fb8
- Source URL: <https://doi.org/10.1016/j.mimet.2026.107715>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1016%2Fj.mimet.2026.107715>

Abstract: Essentiality and knockdown-vulnerability screens can promote spurious bacterial open reading frames when those frames overlap essential genes, because such a frame inherits its neighbour's signals undiluted and therefore satisfies the screen's criteria better than a genuine small gene. We present a multi-layer audit that tests this failure mode across genome annotation, transposon mutagenesis, CRISPR interference, homology, transcript mapping, proteomics, and population variation. We apply it to Rv2438A, a 92-codon conserved hypothetical open reading frame of Mycobacterium tuberculosis ranked first by our own dark-gene target screen. Rv2438A is superimposed on the essential NAD synthetase locus nadE: 44% lies within its coding sequence on the opposite strand, and the remainder covers its promoter and transcription start site. Consequently, three of five Himar1 sites lie within nadE, no CRISPRi guide can target Rv2438A without binding nadE, and the cross-species hit maps to the same nadE start junction. Rv2438A lacks its own transcription start site and is absent from every proteomic dataset that detects nadE. A genome-wide scan identifies six short, overlapping, uncharacterised loci among 3907 annotated genes, but only Rv2438A combines overlap and essentiality with non-detection across all proteomic datasets; rare genome-wide, it ranked first among screen hits. We provide an implementable audit workflow and a codon-position control, but measure the control's sensitivity as only two of five genes with attested protein, limiting it to confirmatory use. Overlap coordinates and neighbour-specific experimental resolution should therefore be reported before bacterial dark genes are prioritised.

## Automatic segmentation and modeling of the aortic vessel tree: Overview of the SEG.A 2023 aorta segmentation challenge.
- Source: Medical image analysis (journals)
- Date: 2026-09-17
- Categories: Biological imaging, Tools & resources
- Authors: Yuan Jin, Antonio Pepe, Gian Marco Melito, Yuxuan Chen, Gege Ma, Yunsu Byeon, Hyeseong Kim, Kyungwon Kim, Doohyun Park, Euijoon Choi, Dosik Hwang, Andriy Myronenko, Dong Yang, Yufan He, Daguang Xu, Ayman El-Ghotni, Mohamed Nabil, Hossam El-Kady, Ahmed Ayyad, Amr Nasr, Marek Wodzinski, Henning Müller, Hyeongyu Kim, Yejee Shin, Abbas Khan, Muhammad Asad, Alexander Zolotarev, Caroline Roney, Anthony Mathur, Martin Benning, Gregory Slabaugh, Theodoros Panagiotis Vagenas, Konstantinos Georgas, George K Matsopoulos, Jihan Zhang, Zhen Zhang, Liqin Huang, Christian Mayer, Heinrich Mächler, Jan Egger
- Journal: Medical image analysis
- DOI: 10.1016/j.media.2026.104324
- External ID: 42762598
- Source URL: <https://doi.org/10.1016/j.media.2026.104324>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1016%2Fj.media.2026.104324>

Abstract: The automated analysis of the aortic vessel tree (AVT) from computed tomography angiography (CTA) is crucial for clinical applications but lacks shared, high-quality data. To address this, we launched the SEG.A. challenge, introducing a large, public, multi-institutional dataset for AVT segmentation and benchmarking automated algorithms. The challenge results showed a strong trend toward deep learning, with 3D U-Net architectures being most effective. The winning solution used an ensemble-based strategy, highlighting the value of model ensembling for robust AVT segmentation. Performance strongly correlated with algorithmic design, notably the use of customized post-processing and training data characteristics. This initiative establishes a new performance benchmark and provides a lasting resource to drive future innovation toward robust, clinically translatable AVT analysis tools.

## Axiomatic Community Ecology, Topology, and Dynamic Distance
- Source: bioRxiv (preprints)
- Date: 2026-09-17
- Categories: Evolution & metagenomics
- Authors: Wontner, N. J., Spencer, M.
- DOI: 10.1101/2025.08.21.671550
- Source URL: <https://doi.org/10.1101/2025.08.21.671550>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1101%2F2025.08.21.671550>

Abstract: The super-organismal view of ecosystems has largely been superseded by the individualistic view. One consequence of the dominance of the individualistic view is that many modern ecologists treat ecosystems as nothing more than vectors of relative abundances, ignoring the potentially important idea that an understanding of ecosystems should be based on dynamics rather than abundances or species identities. We develop a mathematical framework in which we compare dynamical properties of ecosystems with different sets of species, using ideas from functional analysis, metric spaces, and topology. We give two proof-of-principle applications of our framework to marine sessile communities and to a large database of ecosystem models. We show that under a set of biologically-motivated axioms designed to capture the properties of predator-prey systems, there is only one natural kind of ecosystem.

## Benchmarking methods for inferring single-cell transcription factor activity using large-scale perturbation sequencing data
- Source: Briefings in Bioinformatics (journals)
- Date: 2026-09-17T00:00:00+00:00
- Categories: Genomics & sequence analysis, Single-cell & spatial, Systems & networks
- Authors: Yuehui Zhu, Dongmei Han, Zhen Wang
- Journal: Briefings in Bioinformatics
- DOI: 10.1093/bib/bbag513
- Source URL: <https://doi.org/10.1093/bib/bbag513>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1093%2Fbib%2Fbbag513>

Abstract: Transcription factor activity (TFA) is determined not solely by the expression level of the transcription factor (TF) gene itself, but is also modulated by a series of post-transcriptional regulatory processes. Although numerous computational methods have been developed to infer TFA from single-cell transcriptomic data by constructing gene regulatory networks (GRNs), a systematic and unified evaluation of these methods using high-quality experimental data remains lacking in the field. In this study, we conducted a comprehensive evaluation of eight mainstream TFA inference methods spanning three categories—prior GRN-based, de novo GRN-based, and integrated GRN-based approaches—using large-scale, high-quality single-cell perturbation sequencing (Perturb-seq) datasets. Our results demonstrate that metaTF, which employs an integrated GRN, achieves the best performance across multiple metrics, including TF coverage, predictive accuracy for perturbed cells, and accuracy for perturbed TFs. Among de novo GRN-based methods, pySCENIC exhibits predictive accuracy second only to metaTF but with lower TF coverage; meanwhile, decoupleR, a prior GRN-based method, ranks highly across all evaluated metrics. Further investigation reveals that the enrichment of reconstructed regulons within differentially expressed genes, the selection of prior GRNs and TFA scoring algorithms, and the perturbation types of target TFs are all critical factors influencing the accuracy of TFA inference. This study provides practical recommendations for the application and development of TFA inference methods.

## Benchmarking of bulk transcriptomic harmonization tools in a multi-platform B-cell lymphoma cohort identifies feature-specific quantile normalization and surrogate variable analysis as top-performing methods
- Source: bioRxiv (preprints)
- Date: 2026-09-17
- Categories: Genomics & sequence analysis, Tools & resources
- Authors: Nikitin, D., Borisov, N. M., Savchenko, M., Bobe, A., Meerson, M., Nesmelov, A., Harutyunyan, N., Paponova, S., Kravets, A., Zaitsev, A., Bagaev, A., Arakelyan, A.
- DOI: 10.64898/2026.09.15.751825
- Source URL: <https://doi.org/10.64898/2026.09.15.751825>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.15.751825>

Abstract: Cross-platform harmonization of bulk transcriptomic datasets remains a fundamental challenge for developing cancer biomarkers because of persistent unresolved batch effects. Most harmonization tools are benchmarked on datasets with large inter-group biological differences (for example TCGA tumor types), whereas actionable biomarker mining requires preserving subtle transcriptional distinctions between closely related diagnoses. Here we present ComboBatch, a benchmarking pipeline that evaluates the full cross-product of 14 batch-removal strategies, 3 imputation methods, 33 harmonization algorithms and 2 post-removal conditions across 7,174 samples from 88 germinal-center B-cell lymphoma cohorts spanning four transcriptomic platforms. Scoring 87 quality metrics across 2,234 harmonization approaches, we show that method choice (R2 0.36) and batch-removal strategy (0.26) are the principal determinants of harmonization quality, whereas imputation (0.016) and post-removal (<0.01) are secondary. Feature Specific Quantile Normalization and Surrogate Variable Analysis were the top methods, jointly resolving follicular lymphoma, diffuse large B-cell lymphoma and normal germinal-center B-cell differences in multi-platform and RNA-seq-only compositions, respectively. We provide a data-driven five-scenario decision tree for harmonization method selection, applicable to any retrospective multi-platform transcriptomic study. The ComboBatch pipeline is available on GitHub and can be used for harmonization, allowing bioinformaticians to utilize 33 harmonization and 3 imputation methods according to their needs.

## Benchmarking of tools for resolving the plasmidome from short-read assemblies for Klebsiella pneumoniae
- Source: bioRxiv (preprints)
- Date: 2026-09-17
- Categories: Genomics & sequence analysis
- Authors: Connor, C. H., Wick, R. R., Gorrie, C. L., Winkler, M. A., Lohr, I. H., Ingle, D. J., Lam, M. M.
- DOI: 10.1101/2025.07.24.666686
- Source URL: <https://doi.org/10.1101/2025.07.24.666686>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1101%2F2025.07.24.666686>

Abstract: Plasmids play a critical role in the dissemination of antimicrobial resistance genes and virulence factors in healthcare associated pathogens, such as Klebsiella pneumoniae. Surveillance of these plasmids relies on whole genome sequencing data often generated in clinical and public health settings which frequently use short-read platforms. Therefore, there is a need for robust, scalable tools that can identify and/or reconstruct plasmid sequences from short read data. A myriad of tools already exist to address this problem, however the optimum tool for plasmid identification in K. pneumoniae remains unclear. From a comprehensive search of the literature and code repositories we identified 44 plasmid identification tools, highlighting the uncertainty around best practices. Here, we sought to evaluate these 44 tools to determine which is best suited for reconstructing the plasmidome of K. pneumoniae and related species from the species complex (KpSC). We used a publicly available dataset of 568 diverse KpSC isolates that had both short-read Illumina data and closed hybrid assemblies available. This allowed us to investigate which tools perform best at recovering plasmid sequences when only short-read data is available, whilst knowing the ground truth. From the 44 tools, 34 were excluded as they: were intended for plasmid typing / characterisation (n=3), were not intended for KpSC (n=1), required metagenomic data (n=7), required long read data (n=1), could not be installed (n=13) or could not be run on the command line (n=10). The remaining nine tools had their precision and recall metrics calculated and combined into an overall F1 score. Each individual tool displayed the full range of F1 scores (0 to 1) across our collection of genomes, overall, the best performing was PlaScope followed closely by MOB-suite. Future tools developed in this crowded space should offer meaningful advancements over existing tools and be rigorously benchmarked using standardised datasets that reflect plasmid diversity.

## Beyond two alleles: Multiallelic genotypic selection and its estimation from time-series data
- Source: bioRxiv (preprints)
- Date: 2026-09-17
- Categories: Evolution & metagenomics, Mathematical biology & statistics
- Authors: Vellnow, N., Gossmann, T. I., Waxman, D.
- DOI: 10.1101/2024.11.08.622587
- Source URL: <https://doi.org/10.1101/2024.11.08.622587>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1101%2F2024.11.08.622587>

Abstract: Genetic diversity is central to evolutionary change, with both natural selection and random genetic drift depending on variation within a population. An individual in a diploid population carries two alleles per locus, yet the population as a whole can harbour many alleles, giving rise to a rich spectrum of homozygous and heterozygous genotypes. Such multiallelic variation is common at biologically and medically important loci such as the major histocompatibility complex, the ABO blood group system, and genes underlying monogenic diseases. However, much of population genetic theory and data analysis has focussed on biallelic loci. Here, we introduce a matrix representation of the genotypic selection acting at a multiallelic locus. This exploits the common mathematical structure underlying selection and drift, and separates the effects of genetic diversity and fitness. The representation accommodates diverse selection regimes, including additive, multiplicative, frequency-dependent, and temporally varying selection, as well as heterozygote advantage. We show how, under specific assumptions, genotype-specific fitness-effects can be estimated from allele frequency trajectories over microevolutionary timescales. Applying this estimation procedure to time-series data from experimental yeast evolution illustrates how multiallelic fitness interactions, including heterozygote advantage, may be characterised from haplotype frequency data. More broadly, this work provides a practical foundation for analysing evolutionary dynamics at multiallelic loci in experimental and natural populations.

## Building dynamical models of multi-step state transitions from single cell gene expression trajectories
- Source: bioRxiv (preprints)
- Date: 2026-09-17
- Categories: Genomics & sequence analysis, Single-cell & spatial, Systems & networks, Mathematical biology & statistics, Tools & resources
- Authors: You, Y., Caranica, C., Dai, G., Lu, M.
- DOI: 10.64898/2025.12.08.693064
- Source URL: <https://doi.org/10.64898/2025.12.08.693064>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2025.12.08.693064>

Abstract: Multi-step cell state transitions occur across biological processes, such as development and disease progression, yet the underlying gene regulation remains unclear. We introduce NetDes, a computational systems-biology method that infers core transcription factor (TF) regulatory networks and builds ODE-based dynamical models from single-cell gene expression trajectories. In benchmarks on synthetic trajectories with decoys and BEELINE scRNA-seq datasets, NetDes identifies regulatory interactions competitively with existing methods, and reconstructs a simulated cell-fate circuit. We applied it to time-series scRNA-seq data of iPSC-to-definitive-endoderm differentiation, epithelial-mesenchymal transition, erythropoiesis, and dendritic cell differentiation. NetDes has advantages over existing approaches in reconstructing a minimal network with a single model that reproduces observed expression dynamics and captures sequential state transitions. Network simulations predict TFs and their combinations driving each transition, recovering known master regulators, while network coarse-graining reveals the circuit logic of iPSC-to-DE differentiation. NetDes provides a general framework for mechanistic modeling of complex cell state transitions.

## CancerGeneHub: An Evidence-Linked Bidirectional Web Portal for Centralised Exploration of Cancer–Gene Associations
- Source: International Journal of Creative and Open Research in Engineering and Management (journals)
- Date: 2026-09-17T00:00:00Z
- Categories: Tools & resources
- Authors: Gulnaaz Parveen, Alam Zia, S. Alam, Ayushi Sharma
- Journal: International Journal of Creative and Open Research in Engineering and Management
- DOI: 10.55041/ijcope.v2i9.127
- External ID: 5cd6f240e09cfd8c736c4a3e7bd7864ccd2b7497
- Source URL: <https://doi.org/10.55041/ijcope.v2i9.127>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.55041%2Fijcope.v2i9.127>

Abstract: Cancer is a major global health challenge characterised by genetic and molecular alterations that disrupt normal cellular growth, regulation, and survival. The rapid expansion of cancer genomics has generated extensive information concerning cancer-associated genes; however, this information remains distributed across scientific publications and specialised biological databases, making efficient retrieval difficult, particularly for students and early-stage researchers. This study presents CancerGeneHub, a web-based Cancer Gene Information Portal developed to provide centralised and evidence-linked exploration of gene–cancer associations. The portal organises 8,077 gene–cancer associations involving 3,177 genes and 130 cancer types distributed across 11 body-system classifications. A bidirectional search mechanism allows users to investigate genes associated with a selected cancer type and, conversely, cancer types associated with a selected gene. Each association is linked to a PubMed Identifier, enabling users to trace the reported relationship to supporting scientific literature. The system was implemented using React and Tailwind CSS for the frontend, Express for backend services, and MongoDB for data storage and management. The resulting platform combines structured biological information, searchable many-to-many relationships, and literature traceability within a single web interface. The developed system provides an accessible resource for cancer-genetics education and exploratory research. It establishes a scalable foundation for future integration of genomic databases, evidence scoring, mutation information, functional annotations, and interactive analytical capabilities. Keywords— Cancer genomics; cancer-associated genes; gene–cancer association; bioinformatics; PubMed; biomedical information retrieval.

## CircExor enables interpretable prediction of circRNA localization into extracellular vesicles
- Source: Genome Research (journals)
- Date: 2026-09-17T00:00:00+00:00
- Categories: Genomics & sequence analysis, Proteins & structural biology, Tools & resources
- Authors: Yusa Zhang, Hanbo Lu, Pengfei Bao, Anhao Wang, Xiaohong Lyu, Yidong Zhou, Songjie Shen, Zhi John Lu
- Journal: Genome Research
- DOI: 10.1101/gr.281656.125
- Source URL: <https://doi.org/10.1101/gr.281656.125>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1101%2Fgr.281656.125>

Abstract: Certain circular RNAs (circRNAs) are selectively enriched in extracellular vesicles (EVs), in which they contribute to intercellular communication and represent promising biomarkers, yet the sequence determinants of their sorting remain unclear. Existing computational predictors are optimized mainly for linear RNAs and rarely address circRNA localization into EVs. Here we introduce circExor, the first framework specifically designed for circRNA EV localization. We curate a dedicated benchmark data set of 2102 circRNAs and implement a variable-length end-to-end concatenation strategy together with k -mer frequency encoding to accommodate circular topology, long sequence length, and length heterogeneity. Using a tree-based classifier, circExor achieves superior performance compared with RNAlocate-v3 and ExoGRU, reaching an AUROC of 0.743 on the internal test set and an average AUROC of 0.680 on the held-out test set. SHAP-based analysis, sequence perturbation analysis, motif mapping, and cell-based experimental validation support the predicted EV tendency and identify YBX1, HNRNPK, HNRNPL, and NOVA2 as candidate RBPs potentially associated with circRNA sorting. CircExor therefore provides a predictive and interpretable framework that links in silico modeling to mechanistic hypotheses, and supports biomarker discovery and candidate prioritization for downstream studies of EV-associated circRNAs.

## Climate change, infectious disease, and the spread of microblades across the Qinling-Huaihe line
- Source: bioRxiv (preprints)
- Date: 2026-09-17
- Authors: Elgart, S., Aoki, K., Feldman, M. W.
- DOI: 10.64898/2026.09.14.751440
- Source URL: <https://doi.org/10.64898/2026.09.14.751440>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.14.751440>

Abstract: Aoki et al. (2023) proposed that the sharp difference in microblade distribution between the north and south of the Qinling-Huaihe (H-Q) line before the early Holocene could have been due to a higher frequency of pathogens in the south. Greater susceptibility to pathogen-induced disease among northerners could have effectively prevented migration across the H-Q line. Here, we explore the possibility that migration between the north and south could have been stalled by a vector-borne disease whose vector distribution was climate dependent. The original wave equation approach is extended to include climate variation in the forms of (i) a period of linear warming; (ii) a period of linear cooling; and (iii) a period of temperature oscillations. Simulations of this extended model are carried out using published estimates of the climate record for the Northern Hemisphere since the Upper Paleolithic. This analysis suggests possible time intervals during which microblades could have reached the south.

## CMIGAT: Joint learning via Cyclic Modality-Interaction Graph attention for multi-omics integration.
- Source: Journal of biomedical informatics (journals)
- Date: 2026-09-17
- Categories: Single-cell & spatial
- Authors: Kai Wang, Jiang Xie, Mengfei Zhang, Haoyang Zhang
- Journal: Journal of biomedical informatics
- DOI: 10.1016/j.jbi.2026.105098
- External ID: 42753938
- Source URL: <https://doi.org/10.1016/j.jbi.2026.105098>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1016%2Fj.jbi.2026.105098>

Abstract: OBJECTIVE: With the rapid development of high-throughput sequencing technology, integrating multi-omics data has become a necessary means to elucidate complex disease mechanisms and achieve precision diagnosis. However, existing methods still face two major challenges: (1) the difficulty of effectively and accurately extracting cross-omics shared representations; and (2) the lack of effective strategies to combine specific and shared representations. To address these challenges, we propose the Cyclic Modality-Interaction Graph Attention Network (CMIGAT), which unifies specificity extraction, shared alignment, and topological fusion in an end-to-end framework. METHODS: CMIGAT comprises three coupled modules. Omics-specific features are first extracted via graph convolutional encoders with reconstruction regularization and confidence learning. We then extract shared features directly from the raw omics inputs through a lightweight linear alignment that bypasses the deep modality-specific encoders, and apply dual-alignment constraints (Maximum Mean Discrepancy and semantic consistency) to ensure cross-modal distributional and semantic agreement. For multi-omics integration, we propose the Cyclic Modality-Interaction Graph Integration Module (CMIGM). In this module, a Cyclic Modality-Interaction Graph (CMIG) is designed to integrate the shared and specific features of each omics, and a Graph Attention Network (GAT) is used to execute cross-modal information propagation, whereby effective information interaction and robust feature aggregation are achieved. RESULTS: Extensive experiments on six public benchmarks (ROSMAP, BRCA, LGG, KIPAN, GBM, and OV) show that CMIGAT achieves the best or competitive performance, ranking first on the large majority of metrics across the benchmarks. The two four-omics datasets (GBM and OV) further show that the framework scales naturally to more modalities. Ablation studies confirm the necessity and complementarity of each module. Shapley-based biomarker analysis on BRCA, together with KEGG and GO enrichment analyses, identifies biologically meaningful features closely associated with cancer-related pathways. CONCLUSION: CMIGAT effectively addresses the challenges of cross-omics shared representation extraction and specific-shared feature combination, achieving superior classification and interpretable biomarker identification. It provides a useful computational tool for multi-omics tasks such as cancer subtype classification and biomarker screening.

## Comparative Genomics-Guided Epitope Prioritization and in Silico Design of a Multi-Epitope DNA Vaccine Candidate Against Megalocytivirus pagrus 1.
- Source: Marine biotechnology (New York, N.Y.) (journals)
- Date: 2026-09-17
- Categories: Genomics & sequence analysis, Proteins & structural biology
- Authors: Sung-Bin Moon, Min-Young Sohn, Gyoungsik Kang, HyeongJin Roh, Yoonhang Lee, Min Jae Kim, Kwang Il Kim, Seong Don Hwang, Chan-Il Park, Kyung-Ho Kim
- Journal: Marine biotechnology (New York, N.Y.)
- DOI: 10.1007/s10126-026-10707-1
- External ID: 42753003
- Keywords: genomics, dna, genomes, epitope, peptide, epitopes, molecular dynamics
- Source URL: <https://doi.org/10.1007/s10126-026-10707-1>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1007%2Fs10126-026-10707-1>

Abstract: Megalocytivirus pagrus 1 infection is a World Organisation for Animal Health-listed aquatic animal disease caused by a virus species comprising the RSIV, ISKNV, and TRBIV genogroups. Here, we integrated comparative genomics and immunoinformatics to prioritize a multi-epitope protein construct, pMEV, and to design a DNA vaccine candidate encoding it, with emphasis on RSIV-type infection relevant to rock bream aquaculture. Analysis of 61 complete genomes identified 28 core gene clusters, from which myristoylated membrane protein (MMP) and major capsid protein (MCP) were prioritized as source antigens for epitope screening. Four cytotoxic T-cell, five helper T-cell, and five linear B-cell epitope candidates were selected based on sequence-based screening and exploratory peptide-MHC docking. The selected epitopes were assembled with rock bream beta-defensin-3, PADRE, and peptide linkers to generate the 283-aa pMEV construct. Sequence-based physicochemical analyses indicated properties relevant to subsequent structural and expression-based evaluation, while computationally refined structural modeling identified nine putative conformational B-cell epitope regions. TLR3 docking, normal mode analysis, and a 200-ns molecular dynamics simulation characterized the structural behavior of the selected computational complex without inferring receptor activation. C-ImmSim further generated model-dependent generic humoral and helper T-cell-associated response patterns within a mammalian-based simulation framework. Finally, the pMEV coding sequence was codon-optimized and incorporated into an in silico pcDNA3.1(+)-based DNA vaccine design. Collectively, this study provides a comparative genomics-guided framework for prioritizing an experimentally testable multi-epitope DNA vaccine candidate against M. pagrus 1, while construct expression, immunogenicity, and protective efficacy remain to be evaluated experimentally.

## Constructing Gene Regulatory Network using Chatterjee's Rank Correlation with Single-cell Transcriptomic Data
- Source: bioRxiv (preprints)
- Date: 2026-09-17
- Categories: Genomics & sequence analysis, Systems & networks
- Authors: Gupta, S., Chaudhuri, A., Raghuraman, V., Ni, Y., Cai, J. J.
- DOI: 10.1101/2025.09.17.676530
- Source URL: <https://doi.org/10.1101/2025.09.17.676530>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1101%2F2025.09.17.676530>

Abstract: Discovering gene regulatory networks (GRNs) from single-cell RNA sequencing (scRNA-seq) data is critical for understanding cellular function. Still, existing methods are limited by strong theoretical assumptions or high computational complexity. We introduce a multiple testing framework for GRN construction using Chatterjee's rank correlation coefficient, a nonparametric measure of dependence. Our approach overcomes the limitations of traditional methods while offering a transparent, scalable, and computationally efficient alternative to recent black-box machine learning models. Crucially, to address the non-independence of cellular observations inherent to scRNA-seq, we develop a data-driven algorithm for estimating robust testing cutoffs. Furthermore, we exploit the asymmetric nature of Chatterjee's correlation to propose a new test for active regulation, enabling the construction of biologically meaningful and directionally informed GRNs. We demonstrate that our method matches or outperforms state-of-the-art approaches in recovering true gene-gene dependencies and directed regulatory interactions from both simulated and real datasets, particularly for complex, non-linear dependencies, providing a powerful tool for dissecting complex GRNs.

## Construction of a Standardized Time-Lapse Imaging Database and a Gradient Boosting Ensemble Framework for Integrating Zygote Morphokinetic Parameters with Conventional Embryo Assessment
- Source: medRxiv (preprints)
- Date: 2026-09-17
- Categories: Biological imaging, Tools & resources
- Authors: ZHAO, M., LIU, J., HAN, D., ZHANG, C., ZHOU, Y., CHEN, S., Appiah, K., LIU, C.
- DOI: 10.64898/2026.08.20.26359523
- Source URL: <https://doi.org/10.64898/2026.08.20.26359523>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.08.20.26359523>

Abstract: In vitro fertilization (IVF) laboratories equipped with time-lapse incubators generate vast quantities of sequential embryo images, yet the absence of standardized, annotated databases impedes the development of reproducible computational tools for embryo assessment. Here we describe the standardized time-lapse imaging database built upon prior research ground-work comprising 631 two-pronuclear (2PN) zygotes from 218 treatment cycles performed at Guangdong Provincial Peoples Hospital (2020-2023), together with a gradient boosting decision tree (GBDT) ensemble framework designed to fuse heterogeneous data types for blastocyst outcome prediction. Each embryo record integrates 84 zygote-stage morphokinetic parameters extracted from EmbryoScope time-lapse sequences via a previously validated convolutional neural network segmentation pipeline with 8 conventional embryo assessment features recorded at cleavage and blastocyst stages according to the Istanbul consensus. The fusion framework employs LightGBM with equal-weight initialization and iterative residual-decreasing training, augmented by recursive feature elimination and nested five-fold cross-validation. Ablation experiments demonstrate that the full model (AUC = 0.78) outperforms morphokinetics-only (AUC = 0.71) and conventional-only (AUC = 0.65) configurations, confirming that zygote-stage temporal dynamics carry complementary information beyond standard morphological grading. SHAP analysis identifies cytoplasmic area slope, zona pellucida grayscale trend, and pronuclear fading time as the three most influential predictors. The database and fusion methodology provide a reproducible framework for integrating time-series imaging features with categorical clinical assessments in reproductive medicine.

## Convex approaches to isolate the shared and distinct genetic components of complex traits
- Source: Bioinformatics (journals)
- Date: 2026-09-17T00:00:00+00:00
- Categories: Genomics & sequence analysis, Tools & resources
- Authors: Saikat Banerjee, Shane O’Connell, Sarah M C Colbert, Niamh Mullins, David A Knowles
- Journal: Bioinformatics
- DOI: 10.1093/bioinformatics/btag670
- Source URL: <https://doi.org/10.1093/bioinformatics/btag670>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1093%2Fbioinformatics%2Fbtag670>
- Code: <https://github.com/daklab/clorinn>

Abstract: Motivation Groups of complex diseases, such as coronary heart disease, neuropsychiatric disorders, and cancers, often display overlapping clinical symptoms and pharmacological responses. Genetic variants with shared associations across diseases have the potential to help explain their underlying biological processes, but this sharing remains poorly understood. Results We model the matrix of summary statistics of trait-associated genetic variants as the sum of a low-rank component—representing shared biological processes—and a sparse component representing disease-unique processes and arbitrarily corrupted or contaminated components. We introduce Clorinn, an open-source Python library that uses convex optimization algorithms to recover these components by minimizing a weighted combination of nuclear norm and L1 terms. Clorinn provides two significant benefits: (a) convex optimization guarantees reproducibility of the components, and (b) the low-rank “uncorrupted” matrix allows robust singular value decomposition (SVD) and principal component analysis (PCA), which are otherwise highly sensitive to outliers and noise in the input matrix. In extensive simulations, we observe that Clorinn is uniquely able to recover the disease-group structure while remaining competitive on factor-level reconstruction error. We apply Clorinn to estimate 200 latent factors from GWAS summary statistics for 2,110 phenotypes from the Pan-UK Biobank (N = 420,531 European-ancestry individuals) and 10 latent factors from 14 psychiatric disorders. Availability Clorinn is available at https://github.com/daklab/clorinn. Supplementary information Supplementary data are available at Bioinformatics online.

## Coolsecture: an easy-to-use and improved framework for cross-species Hi-C contact map comparison
- Source: Bioinformatics (journals)
- Date: 2026-09-17T00:00:00+00:00
- Categories: Genomics & sequence analysis, Tools & resources
- Authors: Peng-Kai Zhu, Jiang-Qi Pan, Zhan-Chao Cheng, Jian Gao
- Journal: Bioinformatics
- DOI: 10.1093/bioinformatics/btag683
- Source URL: <https://doi.org/10.1093/bioinformatics/btag683>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1093%2Fbioinformatics%2Fbtag683>
- Code: <https://github.com/pk-zhu/Coolsecture>

Abstract: Summary Cross-species Hi-C comparison remains challenging because existing workflows often rely on multiple scripts, heterogeneous I/O formats, and limited diagnostic support. We present coolsecture, a Python 3 command-line toolkit that integrates multiple contact-matrix and synteny formats with bidirectional lift-over and reciprocal consistency assessment, multi-resolution percentile-based comparison, diagnostic visualization, and cross-sample similarity analysis, providing a streamlined and reproducible framework for comparative Hi-C analysis. Availability and implementation coolsecture is distributed under the GPL-3 license. Source code, Snakemake workflows, and documentation are freely available at https://github.com/pk-zhu/Coolsecture and are archived on Figshare at https://doi.org/10.6084/m9.figshare.30158440

## Cophylogeny simulators are not interchangeable: similarities, differences and structural biases in synthetic host-symbiont
- Source: bioRxiv (preprints)
- Date: 2026-09-17
- Categories: Evolution & metagenomics
- Authors: Di Palma, G., Matias, C., Sinaimeri, B.
- DOI: 10.64898/2026.09.13.751162
- Source URL: <https://doi.org/10.64898/2026.09.13.751162>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.13.751162>

Abstract: Synthetic data are becoming increasingly important for computational studies of cophylogeny, including machine learning inference, benchmarking, and method testing. Several generators have been proposed to produce such data, but each relies on different assumptions about host-symbiont coevolution. These assumptions are often implicit and rarely examined, even though results can depend strongly on the synthetic model being used. In this article, we present a systematic structural analysis of representative cophylogeny generators under controlled scenarios. The goal is to make their assumptions explicit and to understand how these choices shape the synthetic data they produce as well as the conclusions that may be drawn from them.

## CoSAG-nf: A Scalable Nextflow Pipeline for Co-assembly, Optimization, and Interactive Visualization of High-Throughput Single-Cell Genomes
- Source: Bioinformatics (journals)
- Date: 2026-09-17T00:00:00+00:00
- Categories: Genomics & sequence analysis, Tools & resources
- Authors: Linfeng Xu, Zhe-Xue Quan
- Journal: Bioinformatics
- DOI: 10.1093/bioinformatics/btag671
- Source URL: <https://doi.org/10.1093/bioinformatics/btag671>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1093%2Fbioinformatics%2Fbtag671>
- Code: <https://github.com/linfengxu/CoSAG-nf>

Abstract: Motivation Single-cell amplified genomes (SAGs) are crucial for resolving intra-population microbial heterogeneity and accurately understanding the metabolic potential of microbial dark matter populations. However, SAGs generated through multiple displacement amplification (MDA) of genomic DNA from single cells with single-copy chromosomes are highly fragmented and prone to contamination, severely hindering high-quality genome reconstruction and functional analysis, which greatly limits their scientific utility. Co-assembly of related SAGs can substantially improve genome quality, but to our knowledge no automated pipeline exists for high-throughput processing, forcing manual implementation of complex workflows that scale poorly to modern dataset sizes. Results We present CoSAG-nf, an automated high-throughput co-assembly and optimization pipeline for SAGs, implemented following the nf-core framework standards. The pipeline performs alignment-free clustering using sourmash MinHash signatures, then employs iterative tetranucleotide frequency profiling to identify and exclude outlier SAGs from co-assembly groups. CheckM2 quality assessment guides dynamic selection of optimal SAG combinations to optimize genome completeness and minimize contamination. Fully containerized, CoSAG-nf ensures reproducibility and scalability for the high-throughput processing of large-scale SAG datasets across diverse computing environments, including HPC and cloud platforms. The pipeline generates comprehensive HTML reports with quality metrics and taxonomic annotations, providing an end-to-end solution for automated high-throughput single-cell genome reconstruction. Availability CoSAG-nf is freely available under the MIT License at: https://github.com/linfengxu/CoSAG-nf. Archival code repository snapshots are published at zenodo with doi: https://doi.org/10.5281/zenodo.21525244. Supplementary information Supplementary data are available at Bioinformatics online.

## CrossBranch: cross-domain cell-type deconvolution with dual-branch representation learning
- Source: BMC Genomics (journals)
- Date: 2026-09-17T00:00:00+00:00
- Categories: Genomics & sequence analysis, Single-cell & spatial
- Authors: Qianbei Yi, Jiaqi Yuan, Peng Xu, Wenbin Liu
- Journal: BMC Genomics
- DOI: 10.1186/s12864-026-13360-z
- Source URL: <https://doi.org/10.1186/s12864-026-13360-z>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1186%2Fs12864-026-13360-z>

Abstract: Accurate estimation of cell-type composition from mixed omics data is essential for understanding tissue heterogeneity and disease mechanisms. However, existing deconvolution methods are often affected by discrepancies between reference single-cell data and target bulk, proteomic, or spatial omics measurements. This study aims to develop a robust and biologically informed framework for cross-domain cell-type deconvolution. We present CrossBranch, a dual-branch representation learning framework that integrates gene-level and pathway-level information. CrossBranch generates labeled simulated mixtures from single-cell references and jointly encodes simulated and target data through a gene-expression branch and a pathway-informed branch. A prediction head is trained using simulated mixtures with known cell-type proportions, while latent-space alignment reduces distribution discrepancies between simulated and target data. For spatial transcriptomics data, a neighboring-spot-based spatial consistency loss is further incorporated. Across bulk RNA-seq, proteomics, and spatial transcriptomics benchmarks, CrossBranch consistently achieves competitive deconvolution performance compared with existing statistical and deep learning methods. Ablation analyses confirm the contributions of pathway-level representation, cross-domain alignment, and spatial neighborhood modeling. Applications to prostate, colorectal, and pancreatic cancers further demonstrate that CrossBranch can identify tumor-associated cellular changes, survival-associated cell-type patterns, malignant epithelial localization, fibroblast–endothelial co-localization, and compartment-specific spatial organization in tumor microenvironments. CrossBranch provides a unified cross-domain deconvolution framework that improves cell-type composition inference across diverse omics modalities and supports biologically meaningful interpretation of disease microenvironments.

## Deep learning-based assessment of ulcerative colitis activity from full-length endoscopic videos with spatial characterisation and histological correlation
- Source: medRxiv (preprints)
- Date: 2026-09-17
- Categories: Biological imaging
- Authors: Bogush, A., Toskas, A., Ralli, G., Windell, D., Aljabar, P., DeLegge, M., Walsh, A., Thomas, J. P., Wakefield, P., Langford, C., Fryer, E., Goldin, R., Suzuki, N., Landy, J.
- DOI: 10.64898/2026.09.16.26363201
- Source URL: <https://doi.org/10.64898/2026.09.16.26363201>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.16.26363201>

Abstract: Background and study aims: Endoscopic assessment of ulcerative colitis (UC) is central to clinical decision-making however remains subjective and limited in characterisation of disease distribution. Artificial intelligence (AI) enables analysis of entire endoscopic examination, rather than relying on selected views. We aimed to develop and validate a deep learning model for automated assessment of UC severity from full-length endoscopic videos, introducing spatial representation of inflammation (Continuous Disease Score, CDS), and assess histological correlation. Patients and methods: Full-length endoscopy videos from adult patients with UC undergoing colonoscopy or flexible sigmoidoscopy were analysed; isolated proctitis was excluded. Videos were segmented and annotated using Mayo Endoscopic Score (MES) and Ulcerative Colitis Endoscopic Index of Activity (UCEIS). A deep learning model was trained for frame-level quality control and severity prediction, enabling analysis of full-length videos. Performance was evaluated using quadratic weighted kappa (QWK) and Cohen's kappa, with patient-level separation between datasets. CDS was derived from UCEIS predictions to quantify cumulative inflammatory burden, spatial extent of disease and histological prediction. Results: A total of 67 videos from 59 patients were included. The model demonstrated agreement for remission classification (MES=0 \{kappa\} 0.76; UCEIS\[≤\]1 \{kappa\} 0.84). CDS enabled quantification of inflammatory burden and revealed spatial heterogeneity not reflected in categorical scores. Agreement with histology was strong (AUROC 0.84-0.87). Conclusions: AI-based analysis enables automated assessment of UC activity from full-length endoscopic videos, including remission detection and estimation of histological healing. CDS provides continuous characterisation of inflammatory burden and disease extent beyond conventional categorical scores, with potential to support more standardised assessment in clinical trials and practice.

## DELPHAI predicts heterogeneous perturbation responses with learned single-cell fitness
- Source: bioRxiv (preprints)
- Date: 2026-09-17
- Categories: Genomics & sequence analysis, Single-cell & spatial
- Authors: Zhang, X., Wu, H., Liu, H.
- DOI: 10.64898/2026.07.01.735965
- Source URL: <https://doi.org/10.64898/2026.07.01.735965>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.07.01.735965>

Abstract: Current perturbation response modelling in single-cell transcriptomics assumes conserved cell mass and loses gene expression information to latent-space decoding. We propose DELPHAI, training a fitness network and an optimal transport network jointly without biological priors, and during inference applying a fitness-gated transport with a direct gene-space retrieval. Demonstrated across two benchmark frameworks, DELPHAI ranks first in predicting differentially expressed genes, while revealing which cell lineages a perturbation depletes.

## DepoCat: Interactive database of experimentally verified phage depolymerases
- Source: bioRxiv (preprints)
- Date: 2026-09-17
- Categories: Tools & resources
- Authors: Olejniczak, S., Otwinowska, A., Pozniak, M., Drulis-Kawa, Z.
- DOI: 10.64898/2026.09.11.750917
- Source URL: <https://doi.org/10.64898/2026.09.11.750917>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.11.750917>

Abstract: Klebsiella phage depolymerases degrade polysaccharide capsules and exhibit narrow substrate specificity for particular capsular types. Despite a growing number of experimentally characterized enzymes, these data remain scattered throughout the scientific literature, while existing protein sequence repositories are dominated by entries with computationally assigned, unverified functional annotations. Here we present DepoCat, the first interactive database of phage depolymerases with experimentally verified function and specificity, available at http://depocat.uwr.edu.pl. The database currently contains 131 proteins meeting rigorous inclusion criteria, spanning 75 distinct capsular types. Each entry integrates experimental and computational resources. The web interface provides an integrated Classifier tool with two search modes: sequence-based search and structure-based search - enabling preliminary structural classification and inference of putative substrate specificity of newly identified depolymerases. We demonstrated the utility of both modes on a set of 17 experimentally verified non-Klebsiella phage depolymerases, for which structural analysis enabled unambiguous class assignment in almost all cases despite low or undetectable sequence similarity to the database reference dataset. DepoCat constitutes a publicly accessible resource supporting research into the structural diversity and sequence-structure-specificity relationships of phage depolymerases, while also facilitating the identification of candidates for therapeutic and diagnostic applications.

## Development of a YOLOv9 model with angle loss for automatic detection of landmarks and cephalometric analysis
- Source: Scientific Reports (journals)
- Date: 2026-09-17T00:00:00+00:00
- Categories: Biological imaging
- Authors: Shinta Amini Prativi, Andriyan Bayu Suksmono, Tati Latifah Erawati Rajab, Donny Danudirdjo, Akira Hirose, Stefanie Mueller
- Journal: Scientific Reports
- DOI: 10.1038/s41598-026-71977-2
- Source URL: <https://doi.org/10.1038/s41598-026-71977-2>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1038%2Fs41598-026-71977-2>

Abstract: Artificial intelligence has been widely applied to identify anatomical landmarks on lateral cephalometric radiographs, reducing localization errors and improving the efficiency of cephalometric analysis. This study aimed to develop and evaluate an automated cephalometric landmark detection system based on You Only Look Once version 9 (YOLOv9) with an angle-based loss function and automated Steiner cephalometric analysis using radiographs from an Indonesian population. The proposed angle-based loss was incorporated into YOLOv9 to enforce geometric consistency among anatomically related landmarks. Model performance was evaluated using mean radial error (MRE), successful detection rate (SDR), and mean average precision (mAP). Automated Steiner measurements derived from the predicted landmarks were compared with expert annotations using the mean absolute error (MAE). The proposed model achieved an overall MRE of 0.99 mm, an SDR of 86.8% at a 2-mm threshold, and a mAP of 0.754. Automated Steiner analysis yielded a mean absolute error of 1.38° compared with expert measurements. Compared with the baseline YOLOv9 model, the proposed angle-based loss modestly improved overall landmark localization performance while preserving anatomical relationships among predicted landmarks. These findings suggest that the proposed framework can support automated cephalometric analysis. However, further validation using independent datasets is required to confirm its generalizability and clinical applicability.

## DiffDomain-Spectrum identifies structurally reorganized TADs from sparse aggregated single-cell Hi-C contact maps
- Source: bioRxiv (preprints)
- Date: 2026-09-17
- Categories: Genomics & sequence analysis, Mathematical biology & statistics, Tools & resources
- Authors: Zhu, J., Zhang, H., Du, Y., Zhang, X., Zhou, Y., Tian, D.
- DOI: 10.64898/2026.09.15.751728
- Source URL: <https://doi.org/10.64898/2026.09.15.751728>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.15.751728>

Abstract: Structurally reorganized topologically associating domains (TADs) capture condition- or cell-type-specific remodeling of chromatin contacts and are important for understanding genome organization in health and disease. Emerging single-cell Hi-C (scHi-C) technologies enable such comparisons across heterogeneous cell populations, but aggregated scHi-C contact maps remain sparse at biologically meaningful 25 kb resolution, limiting reliable TAD reorganization detection. Here we present DiffDomain-Spectrum, a spectral statistical framework for identifying reorganized TADs between conditions or cell types from aggregated raw scHi-C contact maps. It tests normalized TAD-level difference matrices without separately normalizing sparse maps or enhancing individual scHi-C contact maps. Comparison with a semicircle-law null integrates evidence across the full eigenvalue spectrum. Across multiple scHi-C platforms, DiffDomain-Spectrum balances false positive control and detection sensitivity relative to alternative bulk callers, and detects a substantially higher proportion of reference TADs as reorganized than the boundary-focused single-cell method scHiCluster. Detected TADs show coherent aggregate contact patterns and CTCF binding changes and are enriched for differentially expressed genes, supporting biological relevance. Together, these results establish DiffDomain-Spectrum as a statistically principled framework for comparative domain-level analysis of sparse aggregated scHi-C contact maps without single-cell map enhancement.

## Discovery of microbial intergenic features with genomic language modeling and multimodal search
- Source: bioRxiv (preprints)
- Date: 2026-09-17
- Categories: Genomics & sequence analysis, Tools & resources
- Authors: Zulaybar, N., Tranzillo, M., Silverstein, R., Hwang, Y., Cornman, A.
- DOI: 10.64898/2026.09.15.751765
- Source URL: <https://doi.org/10.64898/2026.09.15.751765>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.15.751765>

Abstract: Systematic characterization of microbial noncoding regions is limited by two distinct challenges: discovery of conserved sequence features without predefined motifs and functional interpretation of newly identified elements. We address these challenges by training a sparse autoencoder on genomic language model (gLM2) representations to identify intergenic sequence features without prior annotation, and by implementing multimodal search to generate functional hypotheses from conserved associations with neighboring proteins, RNA families, and genomic organization. This framework uncovered divergent, previously uncharacterized noncoding elements, including candidate regulatory DNA sequences and structured RNAs not captured by existing annotation models. gLM2-derived intergenic features can be explored through SeqHub's multimodal search, freely available for academic use at seqhub.org.

## Do papers tell the whole story? A benchmark and framework for uncovering hidden implementation gaps in bioinformatics
- Source: Briefings in Bioinformatics (journals)
- Date: 2026-09-17T00:00:00+00:00
- Categories: Tools & resources
- Authors: Tianxiang Xu, Xiaoyan Zhu, Xin Lai, Xin Lian, Sizhe Dang, Hangyu Cheng, Jiayin Wang
- Journal: Briefings in Bioinformatics
- DOI: 10.1093/bib/bbag509
- Keywords: benchmark
- Source URL: <https://doi.org/10.1093/bib/bbag509>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1093%2Fbib%2Fbbag509>

Abstract: As bioinformatics software is increasingly applied across a broader range of scenarios and the rapid development of large language models (LLMs) further lowers the barriers to software use and development, the composition of the bioinformatics research community is undergoing substantial change. Consequently, a growing number of researchers require a deeper understanding of methodological details and software behavior. In this context, systematically analyzing the relationship between paper descriptions and code implementations is emerging as an important new challenge in the field. To address this challenge, we introduce paper-code consistency analysis as a new research perspective and construct BioCon, the first benchmark dataset for paper-code consistency analysis in bioinformatics. Furthermore, we develop a unified cross-modal analysis framework to systematically investigate this problem from three perspectives: sentence-level detection, cross-modal retrieval, and project-level assessment. Experimental results demonstrate that the proposed framework can effectively model the semantic relationships between scientific publications and software implementations. Further case studies reveal that paper-code inconsistency is not a single phenomenon but arises from multiple underlying causes, among which Author-Perceived Non-Essential Details represents the most prevalent category. These findings suggest that paper-code consistency analysis is not merely a technical problem but also raises broader discussions regarding knowledge dissemination, the boundaries of code disclosure, and community norms. We hope that this work will encourage the bioinformatics community to re-examine the relationship between scientific publications and software implementations while providing a foundation for future research in paper-code consistency analysis.

## Enhancer Activity-informed Gap GEne Regulatory Network (EAGER) to model Drosophila gap gene expression on the entire anterior-posterior (A-P) axis
- Source: bioRxiv (preprints)
- Date: 2026-09-17
- Categories: Systems & networks, Mathematical biology & statistics
- Authors: Shaikh, R., Busato, S., Dima, S. S., Williams, C., Reeves, G. T.
- DOI: 10.64898/2026.09.15.751584
- Source URL: <https://doi.org/10.64898/2026.09.15.751584>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.15.751584>

Abstract: Across metazoa, morphogen gradients differentially regulate gene expression and activate a spatially distinct program to specify body axis development. The Drosophila gap gene network, initiated by maternal morphogen Bicoid, is one of the most well-studied systems. Several regulatory interactions act synergistically to produce distinct gap gene expression patterns along the anterior-posterior (AP) axis of the blastoderm stage Drosophila embryo to ensure the proper segmentation of the larval and, eventually, adult stage fly. Several mathematical models have been proposed to summarize the interconnectivity of gap gene regulatory elements and predict expression in mutant systems. However, these models have not successfully predicted the gap gene expression profile over the entire AP axis. Here, we present an Enhancer Activity-informed Gap GEne Regulatory network (EAGER) model that incorporates enhancer activity-driven differential regulation along the AP axis, which successfully summarizes gap gene expression patterns over the entire AP axis. We validated the predictions of EAGER on Kr mutants and performed a comprehensive parametric sensitivity and identifiability analysis to evaluate the robustness of the EAGER model fits and predictions. We also propose a reduced version of the model, rEAGER, which identifies a minimal set of regulatory interactions and successfully summarizes the gap gene interactions over the entire AP axis. Our results suggest that the expression driven by individual enhancers must be accounted for in models of developmental pattern formation.

## European ash pangenome reveals widespread structural variation and genetic basis of low ash dieback susceptibility
- Source: Nature Communications (journals)
- Date: 2026-09-17T00:00:00+00:00
- Categories: Genomics & sequence analysis
- Authors: Daniel P. Wood, Mohammad Vatanparast, Dario Galanti, Catherine Gudgeon, Katherine Wheeler, Emma Curran, Levi Yant, Richard Whittet, Richard A. Nichols, Richard J. A. Buggs, Laura J. Kelly
- Journal: Nature Communications
- DOI: 10.1038/s41467-026-77194-9
- Source URL: <https://doi.org/10.1038/s41467-026-77194-9>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1038%2Fs41467-026-77194-9>

Abstract: European Ash ( Fraxinus excelsior ) is a keystone tree species, whose populations are being decimated by ash dieback disease (ADB) – better characterisation of genetic variants associated with low susceptibility to the disease is needed. Here, we develop a F. excelsior pangenome to more fully capture sequence variability within this species compared with a linear reference genome, using a geographically diverse set of fifty F. excelsior samples. We identify 362,965 structural variants (SVs), including 174 Mb of sequence absent from the linear reference genome (22% of the linear reference size), and identify 3,412 high-confidence dispensable genes (those present only in some individuals). We use the pangenome to analyse existing genomic data from over 1,200 individuals, revealing 220 single nucleotide polymorphisms (SNPs) showing consistent allele frequency shifts between healthy individuals and those highly damaged by ADB, across UK seed sources, explicitly demonstrating the existence of a shared genetic component to low ADB susceptibility.

## Facilitating genome annotation using ANNEXA and long-read RNA sequencing
- Source: bioRxiv (preprints)
- Date: 2026-09-17
- Categories: Genomics & sequence analysis, Tools & resources
- Authors: Hoffmann, N., Besson, A., Cadieu, E., Lorthiois, M., Le Bars, V., Houel, A., Hitte, C., Andre, C., Hedan, B., Derrien, T.
- DOI: 10.1101/2025.04.16.648718
- Source URL: <https://doi.org/10.1101/2025.04.16.648718>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1101%2F2025.04.16.648718>
- Code: <https://github.com/IGDRion/ANNEXA>

Abstract: With the advent of complete genome assemblies, genome annotation has become essential for the functional interpretation of genomic data. Long-read RNA sequencing (LR-RNAseq) technologies have significantly improved transcriptome annotation by enabling full-length transcript reconstruction for both coding and non-coding RNAs. However, challenges such as transcript fragmentation and incomplete isoform representation persist, highlighting the need for robust quality control (QC) strategies. This study presents ANNEXA, a pipeline designed to enhance genome annotation using LR-RNAseq data while also providing QC for reconstructed genes and transcripts. ANNEXA integrates two transcriptome reconstruction tools, StringTie2 and Bambu, applying stringent filtering criteria to improve annotation accuracy. It also incorporates deep learning models to evaluate transcription start sites (TSSs) and employs the tool FEELnc for the systematic annotation of long non-coding RNAs (lncRNAs). Additionally, the pipeline offers intuitive visualisations for comparative analyses of coding and non-coding repertoires. Benchmarking against multiple reference annotations revealed distinct patterns of sensitivity and precision for both known and novel genes and transcripts and mRNAs and lncRNAs. To demonstrate its utility, ANNEXA was applied in a comparative oncology study involving LR-RNAseq of two human and eight canine cancer cell lines. The pipeline successfully identified novel genes and transcripts across species, expanding the catalog of protein-coding and lncRNA annotations in both species. Implemented in Nextflow for scalability and reproducibility, ANNEXA is available as an open-source tool: https://github.com/IGDRion/ANNEXA.

## Fast Diffusion of Bound Ca: Analytical and Experimental Characterization of One- and Two-Dimensional Traveling Waves
- Source: bioRxiv (preprints)
- Date: 2026-09-17
- Categories: Mathematical biology & statistics
- Authors: Mironov, S.
- DOI: 10.64898/2026.07.06.735233
- Source URL: <https://doi.org/10.64898/2026.07.06.735233>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.07.06.735233>

Abstract: Reaction diffusion (RD) systems play a fundamental role in numerous biochemical and biophysical processes. Here, we present a novel analytical framework for solving RD equations by applying the Wentzel Kramers Brillouin Jeffreys (WKBJ) formalism to Ca nanodomains generated by individual membrane channels, a widely used paradigm for intracellular Ca signaling. Previous models have primarily focused on stationary Ca nanodomains while neglecting diffusion and saturation of intracellular Ca buffers and sensors. In contrast, we derive analytical solutions without these simplifying assumptions. Our analysis demonstrates that sustained Ca influx generates continuously expanding distributions of free Ca, whereas Ca bound buffers and sensors propagate as traveling waves. These predictions are supported experimentally by measurements of one-dimensional fluorescence profiles produced by single-channel activity and two-dimensional profiles generated by whole cell Ca currents. The analytical framework developed here readily extends Michaelis Menten type kinetics to reaction diffusion systems and may therefore be broadly applicable to biochemical and biophysical processes in which diffusion cannot be neglected.

## FISHFINDER: Catching fish name mistakes in text
- Source: Fisheries (journals)
- Date: 2026-09-17T00:00:00Z
- Categories: Tools & resources
- Authors: Zachery D. Zbinden
- Journal: Fisheries
- DOI: 10.1093/fshmag/vuag058
- External ID: 64a3b866218a968442b5c559b66c416d88d7920c
- Source URL: <https://doi.org/10.1093/fshmag/vuag058>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1093%2Ffshmag%2Fvuag058>

Abstract: In a sea of text, fish name mistakes are easy to miss. Latin binomials are easy to misspell, and taxonomic nomenclature is continuously revised as new phylogenetic evidence accumulates, creating a persistent disconnect between current taxonomy and the names used in the literature. For example, the eighth edition of Common and Scientific Names of Fishes from the United States, Canada, and Mexico introduced 850 changes to fish nomenclature. But no existing tool validates fish names in unstructured manuscript text against this authority. Therefore, I developed FISHFINDER, a free, easy-to-use web application that classifies scientific and common fish names by comparing them against the 5,086 species in the Names of Fishes, 8th edition and 8,729 synonyms compiled from Eschmeyer's Catalog of Fishes (https://researcharchive.calacademy.org/research/ichthyology/catalog/fishcatmain.asp). An analysis of synonym data revealed that species accumulate a mean of 1.7 synonyms (median = 1, max = 38), with 61% of species having at least one historical synonym that could appear in the literature. To quantify the prevalence of naming errors in recent literature, I applied the FISHFINDER classification engine to the body text of 38 papers whose study systems were verified to lie within the Names of Fishes area (the United States, Canada, and Mexico) and that used at least one scientific name. Of these, 42% contained at least one outdated synonym or misspelled species name: 29% used outdated synonyms and 29% contained misspellings. An additional 53% of papers referenced species whose names changed between the seventh and eighth editions. These findings demonstrate that nomenclatural errors are common and that automated validation tools can help authors, reviewers, and editors ensure taxonomic accuracy before publication. FISHFINDER is freely available at fishnames.net.

## FoldaVirus, a knowledge-based icosahedral capsid prediction tool using AlphaFold
- Source: Proceedings of the National Academy of Sciences (journals)
- Date: 2026-09-17T00:00:00+00:00
- Categories: Proteins & structural biology, Tools & resources
- Authors: Oscar Rojas Labra, David S. Montoya-Munoz, Nelly Santoyo-Rivera, Jeffrey McDonald, Daniel Montiel-Garcia, David A. Case, Vijay S. Reddy
- Journal: Proceedings of the National Academy of Sciences
- DOI: 10.1073/pnas.2613741123
- Source URL: <https://doi.org/10.1073/pnas.2613741123>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1073%2Fpnas.2613741123>

Abstract: Coat protein (CP) tertiary structures and their capsid organization of spherical viruses are generally well conserved within each virus family. While AlphaFold successfully predicts the tertiary structures of individual CPs, their association to form proper quaternary assemblies cannot be easily accomplished. Here, we report a generalized methodology and an associated web-based utility ( https://foldavirus.org ) that combines AlphaFold predictions of CPs with the knowledge on corresponding icosahedral architectures (e.g., T = 1, 3, 4…) based on the known structures from the related virus family to generate associated capsids. The resulting assemblies are relaxed using Amber energy minimization to relieve any steric clashes at the intersubunit interfaces. Significantly, the capsid models are validated by calculating robust Mahalanobis distance using the residue annotations categorized as interface, core, and surface amino acids with respect to those observed in the experimentally determined structures from the corresponding virus family. Given the amino acid sequence of CP(s), we successfully generated capsids up to T = 9 icosahedral symmetry, including those of picornaviruses that display pseudo- T = 3 symmetry comprising VP1-VP4. As the number of currently available CP sequences are 2 to 3 orders of magnitude larger than the experimentally determined 3D-structures, this approach bridges the huge gap that exists between the sequence and structural space of viruses.

## FORGE audits residue-level information encoded in RNA tertiary structure geometry
- Source: bioRxiv (preprints)
- Date: 2026-09-17
- Categories: Proteins & structural biology, Tools & resources
- Authors: Gow, L., Li, J., Tan, X., Liang, K., Gui, N., Luo, B.
- DOI: 10.64898/2026.07.05.736550
- Source URL: <https://doi.org/10.64898/2026.07.05.736550>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.07.05.736550>

Abstract: Coarse RNA coordinate representations are widely used, yet the biological information they encode remains unquantified. We introduce FORGE, which converts a seven-atom RNA geometry representation into 935 interpretable descriptors and reports which residue-level annotations this geometry supports. On 4,135 post-2025 RNA chains, FORGE recovered 64.6% of native nucleotides; a six-atom control lacking the glycosidic nitrogen retained 58.5%, locating most of this signal in phosphate-sugar geometry. Confidence was sharply graded: abstaining from the least-confident half of positions raised accuracy to 94.4%, yet many chains remained only partially identifiable. The same descriptors predicted base-pair state far better than a DMS-like proxy or protein-proximal context. Native-decoy, OpenKnot and solved-pseudoknot analyses showed that nucleotide identifiability, foldability and experimental design score are separable: AlphaFold3 reproduced the experimental fold for one of four AI-designed constructs and none of the sequences FORGE read from their geometry. FORGE provides a reproducible audit layer for RNA structural interpretation.

## Fragmented criticality of infectious disease epidemics
- Source: medRxiv (preprints)
- Date: 2026-09-17
- Categories: Mathematical biology & statistics
- Authors: Wang, B., Valdano, E.
- DOI: 10.64898/2026.09.08.26362507
- Source URL: <https://doi.org/10.64898/2026.09.08.26362507>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.08.26362507>

Abstract: Population vulnerability to an infectious disease epidemic is commonly summarized by system-level indicators such as the epidemic threshold: a single critical boundary. Yet transmission is heterogeneous across communities, host groups and transmission pathways, and public-health decisions often require identifying which parts of the system become vulnerable, and under which conditions. Here we show that epidemic criticality can itself be fragmented across population structure. Using multitype branching processes, we identify singularities governing the expected size of outbreaks that ultimately become extinct and show that coupling between population strata transforms their individual thresholds into complex-valued critical points. Their real parts locate critical changes along the transmissibility axis, their imaginary parts determine their strength and smearing, and their modes identify the subpopulations involved. In Italy, this framework reveals localized vulnerability to respiratory-pathogen emergence and improves vaccine allocation over importation-based strategies. For measles in Texas, it identifies spatial units more homogeneous in observed outbreak burden than standard administrative or metropolitan partitions. In a One Health model of livestock-associated MRSA, it separates occupational and human--animal transmission pathways. Epidemic vulnerability is therefore organized by a structured critical landscape rather than a single threshold, and resolving this landscape can directly inform public-health risk assessment and intervention.

## From Nagelkerke's R2 to Liability-Scale Variance Explained for Polygenic Scores
- Source: bioRxiv (preprints)
- Date: 2026-09-17
- Categories: Mathematical biology & statistics
- Authors: Uffelmann, E., Visscher, P. M.
- DOI: 10.64898/2026.09.12.751172
- Source URL: <https://doi.org/10.64898/2026.09.12.751172>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.12.751172>

Abstract: It is desirable to quantify the prediction accuracy of polygenic scores (PGS) for disease on the scale of liability and adjusted for case-control ascertainment in the test sample, because that allows comparison across prevalence and ascertainment. Previous expressions have focused in their derivation and implementation on linear regression on the observed 0-1 scale followed by a transformation of the coefficient of determination (R2) to the scale of liability, adjusted for ascertainment. Yet most statistical analyses with empirical data use logistic regression. The differences in scale have led to confusion and incorrect transformations in the literature. Here we provide a new derivation and simple equation, validated by simulation, that allows a direct transformation from the empirical results from logistic regression to the scale of liability.

## From Prompt to Pipeline: A Comparative Evaluation of Large Language Model Coding Agents for Reproducible Bioinformatics Pipeline Construction
- Source: bioRxiv (preprints)
- Date: 2026-09-17
- Categories: Tools & resources
- Authors: Munn, P. R., Grenier, J. K.
- DOI: 10.64898/2026.09.11.751004
- Keywords: pipeline
- Source URL: <https://doi.org/10.64898/2026.09.11.751004>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.11.751004>

Abstract: Agentic coding systems are increasingly presented as a way to reduce the engineering burden of scientific software development. Bioinformatics is a strong test case for this claim because useful pipelines must combine domain-specific analysis choices, command-line software, sample metadata, workflow orchestration, container or HPC execution, and interpretable quality-control reporting. We evaluated three agentic systems - Biomni, Claude Code, and Codex - on the same task: constructing a Nextflow DSL2 pipeline for paired-end CUT&Tag data that included read QC, trimming, alignment, filtering, duplicate removal, signal track generation, per-sample and group-level peak calling, control-aware group merging, annotation, FRiP calculation, deepTools visualizations, and final MultiQC reporting. Each system received the same detailed CRAFT-style prompt and was assessed against a hand-coded reference pipeline developed by the authors. All three systems produced pipeline implementations that appeared plausible at the level of documentation and file structure, but none fully satisfied the requested analysis. The most consequential failure was shared: the agent-generated pipelines performed some form of group-level merging but did not produce the requested merged-group reporting outputs. Sample-level MultiQC reports also disagreed with the reference report. Codex was closest to the reference for primary mapped-read counts, although its total-read accounting and report structure still differed. Claude Code produced the broadest final report, but its mapping summary mixed stages and therefore could not be treated as numerically correct. Biomni produced the strongest subjective documentation, but its read-count agreement with the reference report was poor and several failures required substantial Nextflow expertise to diagnose. These results suggest that current coding agents can accelerate scaffolding, documentation, and routine implementation, but they do not eliminate the need for expert review in bioinformatics workflow construction. For complex sequencing workflows, prompts must specify not only the biological intent, but also the exact stage semantics, acceptance tests, metadata contracts, expected report sections, resource propagation rules, and failure criteria needed to distinguish a plausible pipeline from a correct one.

## From surveillance to intelligence: a scoping review of machine learning for antimicrobial resistance surveillance intelligence across One Health
- Source: Frontiers in Public Health (journals)
- Date: 2026-09-17T00:00:00Z
- Categories: Genomics & sequence analysis, Evolution & metagenomics
- Authors: S. A. Rabbani, Mohamed El-Tanani, I. Matalka, Shrestha Sharma, Manita Saini, Rakesh Kumar
- Journal: Frontiers in Public Health
- DOI: 10.3389/fpubh.2026.1922265
- External ID: a80721bd52c426188b40362bc1dbd2bcfc469018
- Keywords: genomic, genome, metagenomic
- Source URL: <https://doi.org/10.3389/fpubh.2026.1922265>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.3389%2Ffpubh.2026.1922265>

Abstract: Antimicrobial resistance (AMR) is a leading global health threat requiring coordinated surveillance across human, animal, environmental, and genomic systems. Machine learning is increasingly applied to AMR data, yet its contribution to actionable surveillance intelligence, rather than prediction alone, remains poorly defined. To map how machine-learning approaches generate AMR surveillance intelligence, to characterise their validation and implementation maturity, and to propose a framework distinguishing technical prediction from actionable surveillance intelligence. We conducted a scoping review following JBI methodology and PRISMA-ScR reporting. PubMed/MEDLINE, Scopus, and Web of Science were searched from January 2015 to May 2026 for studies applying machine learning or related methods to AMR surveillance intelligence. Two reviewers independently screened and charted records. Of 1,985 records, 66 met eligibility and formed the working evidence base; 41 studies (40 core empirical and one supporting preprint) were appraised against TRIPOD+AI- and PROBAST-aligned reporting, validation, and implementation-readiness domains. Machine learning was applied across five clusters: clinical and electronic-health-record risk prediction and decision support; genomic and whole-genome-sequencing prediction; MALDI-TOF-based rapid resistance prediction; wastewater and metagenomic surveillance; and environmental, animal, food-chain, and One Health early warning. Prediction and risk stratification predominated, but validation maturity was limited: most studies were retrospective or internally validated, with few using external, cross-country, temporal, prospective, or drift-focused evaluation. On appraisal, discrimination was reported in 31 of 41 studies (76%) and explainability in 26 (63%); by contrast, external or temporal validation was present in only 15 (37%), calibration in 5 (12%), prospective evaluation in 1 (2%), and operational deployment with measured clinical or public-health impact in a single study (2%). Machine learning can support AMR surveillance intelligence across clinical, genomic, diagnostic, environmental, and One Health settings, but the evidence demonstrates technical feasibility far more convincingly than operational readiness. Realising this transition will require external and prospective validation, calibration and drift monitoring, transparent and equitable reporting, workflow integration, and explicit linkage of model outputs to clinical and public-health action. We propose a One Health AMR Surveillance Intelligence Framework to organise this shift from data generation toward actionable, adaptive surveillance intelligence.

## Gelato streamlines reproducible and auditable assembly of Western blot figures
- Source: bioRxiv (preprints)
- Date: 2026-09-17
- Categories: Tools & resources
- Authors: Pirozhkova, M. A., Babitz, E., Benhalevy, D.
- DOI: 10.64898/2026.09.11.750878
- Source URL: <https://doi.org/10.64898/2026.09.11.750878>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.11.750878>

Abstract: Western blotting has been widely used for decades, yet figure preparation remains unstandardized, error-prone and difficult to audit. Gelato is a free Fiji/ImageJ plugin that streamlines intuitive figure preparation, logs processing coordinates, provides a straightforward audit platform andenables reproduction from raw images. By making these capabilities simple and accessible, Gelato facilitates robust Western blot reporting, reducing a substantial burden on authors and reviewers while strengthening scientific rigor.

## Generative Design of New-to-nature Biosynthetic Assembly Lines with Genomic Language Modeling
- Source: bioRxiv (preprints)
- Date: 2026-09-17
- Categories: Genomics & sequence analysis, Proteins & structural biology
- Authors: Lanclos, N., Ibrahim, K., Cornman, A., Huang, M., Gill, V., Jiang, A., Abraham, J., Gin, J., Chen, Y., Petzold, C., Baerwald, J., Kortemme, T., Keasling, J., Hwang, Y.
- DOI: 10.64898/2026.09.11.750945
- Source URL: <https://doi.org/10.64898/2026.09.11.750945>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.11.750945>

Abstract: Reprogramming biosynthetic assembly lines can extend biosynthesis beyond the chemical space explored by nature. However, this remains difficult because assembly-line function depends on coordinated interactions across large multidomain enzymes. Here, we couple gLM2, a genomic language model trained on metagenomic sequences, with discrete diffusion and domain-level conditioning to enable generative design and optimization of biosynthetic gene clusters. We apply this approach to a chimeric type I polyketide synthase (PKS) engineered to produce \{delta\}-valerolactam, a molecule not naturally synthesized by PKSs. Through iterative redesign of two multi-domain regions in the context of the full PKS sequence, gLM2 progressively improved \{delta\}-valerolactam production, yielding variants with up to 9.4-fold higher titer than the starting enzyme. Together, these results demonstrate that evolutionary sequence information can be learned and applied to complex, multi-domain enzyme design problems, expanding biosynthetic assembly lines to produce molecules outside their natural biosynthetic repertoire.

## GeneSIS: enhancing transferability of polygenic scores with variant-level gene-by-sex interaction effects
- Source: bioRxiv (preprints)
- Date: 2026-09-17
- Categories: Systems & networks
- Authors: Tanigawa, Y., Kellis, M.
- DOI: 10.64898/2026.09.14.751452
- Keywords: pathway
- Source URL: <https://doi.org/10.64898/2026.09.14.751452>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.14.751452>

Abstract: Advancing precision medicine requires accurate prediction of disease liability across populations and contexts. A major challenge is the limited transferability of polygenic scores (PGS) across genetic ancestry groups. We present GeneSIS (GENE and Sex Interaction Score), a supervised statistical learning framework for jointly modeling additive and context-dependent genetic effects at single-variant resolution directly from individual-level data. We analyze 406,659 individuals, including admixed individuals, in the UK Biobank and 1.3 million variants to develop predictive models for 99 complex traits. We report that ~8% of selected variables capture gene-by-sex (GxS) effects, validated by sex-stratified analyses. Modeling GxS effects improves prediction across 32 traits in non-European individuals. For predicting hip circumference in Africans, GeneSIS achieves a 3.7-fold improvement (p=8.0x10-7) over linear-only PGS and highlights biologically plausible hypotheses, such as pleiotropic GxS effects of GCKR (rs1260326) on anthropometry and menopause, as well as GxS pathway enrichments for interleukin-4 regulation. Overall, our results highlight the benefits of integrating context-dependent effects in human genetics studies.

## Global immuno-epidemiology and the persistence of mpox clade IIb in MSM
- Source: medRxiv (preprints)
- Date: 2026-09-17
- Categories: Evolution & metagenomics
- Authors: Gubela, N., Berghold, R., Bartel, A., Kühnert, D., von Kleist, M.
- DOI: 10.64898/2026.09.16.26363229
- Keywords: phylogenetic
- Source URL: <https://doi.org/10.64898/2026.09.16.26363229>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.16.26363229>

Abstract: In 2022, mpox clade IIb spread globally in men who have sex with men (MSM) via sexual contact. By summer 2022, case numbers declined in Europe and Norther America due to behavior change and immunization, with almost no reported cases in 2023, followed by low-level transmission. By autumn 2022, many susceptible individuals at high risk of transmission may have become immunized. Moreover, because individuals with mpox are infectious for only 2--4 weeks, sustaining transmission chains requires considerable infection throughput. This raises the question: How did mpox persist and transition to endemic circulation? We developed an agent-based model of mpox transmission on the Berlin MSM sexual contact network that incorporates viral shedding kinetics, vaccination, infection-derived immunity, immune waning, and case importations. After calibrating the model, we estimated that the probability of mpox extinction in Berlin exceeded $80\\%$ in 2023, when acquired immunity fragmented the transmission network. Extending the analysis to a European MSM metapopulation reduced the extinction probability to less than $60\\%$, whereas incorporating global metapopulation dynamics reduced it to nearly zero. We find that mpox persistence was enabled by epidemic asynchrony: When transmission declined in Europe/the Americas, mpox transmission continued in Asia, long enough for immune waning to partially restore transmission potential in Europe, thereby enabling subsequent re-importation and endemic circulation. These model-based findings are supported by phylogenetic evidence indicating that all active German transmission clusters are either linked to (i) re-importation of clade F from the Americas, or (ii) the emergence of clades C and E linked to Asia. Our findings on asynchronous transmission across connected global MSM networks highlight the role of global immuno-epidemiology in enabling mpox persistence and emphasize the importance of coordinated international surveillance and control efforts.

## Halo: a pretrained model for whole-cell segmentation from nuclei images in spatial transcriptomics
- Source: Briefings in Bioinformatics (journals)
- Date: 2026-09-17T00:00:00+00:00
- Categories: Single-cell & spatial, Biological imaging, Tools & resources
- Authors: Xingyuan Zhang, Haotian Zhuang, Zhicheng Ji
- Journal: Briefings in Bioinformatics
- DOI: 10.1093/bib/bbag515
- Source URL: <https://doi.org/10.1093/bib/bbag515>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1093%2Fbib%2Fbbag515>

Abstract: Spatial transcriptomics (ST) enables measurement of gene expression while preserving spatial organization within tissues. Accurate reconstruction of single-cell transcriptomes requires precise whole-cell segmentation, yet many ST experiments provide only nuclear staining images, making reliable inference of cell boundaries difficult. Here we introduce Halo, a pretrained segmentation model that reconstructs whole-cell boundaries by integrating nuclear morphology with the spatial distribution of RNA transcripts. Halo converts transcript coordinates into molecular density maps that are processed jointly with DAPI images using a Cellpose-SAM segmentation architecture. Halo is pretrained on multimodal Xenium data from 12 tissue types and can be directly applied to new datasets without additional training, providing a ready-to-use alternative to supervised approaches that require dataset-specific fitting. Across diverse tissues, Halo achieves substantially higher agreement with the 10$\\times$ multimodal reference segmentation than existing methods, in terms of both cell boundaries and RNA-to-cell assignments, while requiring only DAPI staining and transcript spatial information. Improved segmentation leads to more reliable cell-type identification and more accurate estimation of cell morphological features. By providing a pretrained, generalizable model for whole-cell reconstruction, Halo enables scalable and reproducible cell segmentation for image-based ST.

## High-throughput physics-based enzyme engineering
- Source: bioRxiv (preprints)
- Date: 2026-09-17
- Categories: Proteins & structural biology
- Authors: Wang, X., Qiu, X., Wu, Y., Niu, T., Zhang, S., Gao, R., Cho, I., Tang, H., Hu, K., Lei, X., Isayev, O., Wang, J.
- DOI: 10.64898/2026.09.15.751901
- Source URL: <https://doi.org/10.64898/2026.09.15.751901>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.15.751901>

Abstract: Enzyme engineering aims to tailor natural enzymes for industrial and therapeutic applications, yet physically grounded rational design has been limited by a trade-off between accuracy and cost, leaving the field heavily dependent on expert intuition. Here we present a scalable physics-based framework that combines field-aware machine learning with molecular mechanics to capture enzyme electrostatics at quantum-mechanical accuracy while enabling efficient, atomistic exploration of reaction free-energy landscapes. Coupled with microkinetic modelling, the framework translates molecular free-energy landscapes into catalytic rates and selectivity across competing, multistep reaction pathways. Applied to a newly engineered oxidative amidase (OxiAm), the framework predicts catalytic rate constants with near-experimental accuracy, quantitatively resolves the selectivity between hydrolysis and aminolysis, and generalizes across substrates, mutations and enzyme homologues. Transition-state ensemble analysis further reveals the reaction mechanism and guides the design of enzyme variants for pharmaceutical synthesis. By bringing chemical accuracy and high-throughput sampling to enzyme catalysis, this approach shifts rational design from static, empirical practice toward dynamic, free-energy-driven design, and should accelerate the engineering of biocatalysts.

## Host type governs influenza evolutionary strategy across reservoir and spillover hosts
- Source: bioRxiv (preprints)
- Date: 2026-09-17
- Categories: Evolution & metagenomics
- Authors: Maltepes, M. A., Markin, A., Anderson, T. K., Kistler, K., Park, G., Damodaran, L., Ort, J., Sabre, J., Shank, S., Moncla, L. H.
- DOI: 10.64898/2026.09.15.751820
- Source URL: <https://doi.org/10.64898/2026.09.15.751820>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.15.751820>

Abstract: Despite its high propensity for host switching, the evolutionary mechanisms underlying influenza host adaptation remain unclear. H3Nx influenza viruses are uniquely generalist, with long-term lineages that circulate in avian, human, swine, equine, and canine hosts. Using 13,295 H3Nx sequences, we quantified host-specific adaptive evolution and developed a pipeline to map reassortment events onto trees with measures of statistical uncertainty. We find that while H3Nx viruses in mammals undergo adaptive evolution in HA and NA, viruses in birds experience very little directional selection. Instead, avian lineages exhibit high rates of reassortment, frequently generating novel reassortant lineages that persist transiently and turn over rapidly. 29.8-47.4% of all avian reassortant lineages are purged within the first year of circulation, and reassortment shows no fitness benefit in birds. In contrast, reassorted lineages in swine are more likely to persist long-term, suggesting that reassortment in swine may be broadly beneficial. Segment-specific reassortment patterns were also distinct between avian and mammalian viruses, with NA reassorting more frequently than expected in birds, but less frequently than expected in swine. Reassortment events are enriched between mammalian, but not avian, host switches, suggesting that reassortment may be most beneficial for mediating host switches among mammalian species. Together, our data suggest that host differences drive fundamentally different evolutionary outcomes for influenza viruses, transitioning from reassortment-dominant evolution in their avian reservoir, to varying degrees of adaptation upon establishment in mammals.

## iDriver: A patient-centric framework for genome-wide cancer driver discovery
- Source: bioRxiv (preprints)
- Date: 2026-09-17
- Categories: Genomics & sequence analysis, Tools & resources
- Authors: Bahari, F., Montazeri, H.
- DOI: 10.64898/2026.02.16.706129
- Source URL: <https://doi.org/10.64898/2026.02.16.706129>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.02.16.706129>

Abstract: Tumor genomes harbor a mixture of neutral and positively selected mutations, yet distinguishing true cancer drivers remains a major challenge. Several factors can obscure the detection of selection signals, among which patient-specific variation in mutational burden plays a significant role. Current approaches often fail to account for the heterogeneity in mutation burden across different patients; in particular, no existing method explicitly accounts for it when integrating both mutation recurrence and functional impact. Here we present iDriver, a probabilistic graphical model that integrates both mutation recurrence and functional impact at the individual-patient level, enabling an enhanced estimation of positive selection across functional genomic elements. Applying iDriver to 29 cancer types, we identify both known and previously unrecognized drivers spanning coding and noncoding regions, and provide evidence for their clinical and biological relevance. In comprehensive benchmarks against 12 established driver discovery methods, iDriver consistently outperformed all competitors, achieving the highest rankings for known cancer drivers across both coding and noncoding elements.

## iMTSS: an integrated framework for biology- and patient-driven prognosis in myelofibrosis undergoing transplantation.
- Source: Transplantation and cellular therapy (journals)
- Date: 2026-09-17T00:00:00Z
- Categories: Genomics & sequence analysis, Systems & networks
- Authors: N. Gagelmann, R. Salit, T. Schroeder, P. Chiusolo, M. Finazzi, C. Gurnari, S. Pagliuca, C. Rautenberg, M. Rubio, J. Maciejewski, A. Vannucchi, Paola Guglielmelli, Chiara Nozzoli, N. Leimkühler, E. Angelucci, M. Gambella, A. Rambaldi, H. C. Reinhardt, A. Bacigalupo, Bart L. Scott, F. Heidel, U. Popat, N. Kröger
- Journal: Transplantation and cellular therapy
- DOI: 10.1016/j.jtct.2026.09.029
- External ID: 1d68b3c36991986878fecf26c663d705a63a29bc
- Keywords: genomically, pathway, framework
- Source URL: <https://doi.org/10.1016/j.jtct.2026.09.029>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1016%2Fj.jtct.2026.09.029>

Abstract: BACKGROUND Allogeneic hematopoietic cell transplantation is the only curative treatment for myelofibrosis, but failure occurs by two mechanistically distinct routes: relapse of the neoplasm, which reflects its underlying genetics, and non-relapse mortality, which reflects whether the patient and graft tolerate the procedure. Established prognostic systems either lack molecular granularity or were derived in the non-transplant setting, and all collapse these two routes into a single survival estimate. None can indicate why an individual patient is at risk, or which class of intervention might reduce that risk. OBJECTIVE To determine why an individual patient is at risk and to develop and validate an integrated framework that quantifies biology- and patient-driven prognosis. STUDY DESIGN We analyzed 1,550 adults undergoing first allogeneic transplantation for primary or secondary myelofibrosis across international centers, the largest genomically annotated transplant cohort in this disease. The cohort was split into development (n=930) and validation (n=620) sets. Overall survival was modeled by Cox regression; relapse and non-relapse mortality were modeled as competing events by Fine-Gray subdistribution-hazard regression at 2 years. Discrimination was assessed by the concordance index with bootstrap confidence intervals. The molecular contribution was quantified by variance decomposition of, and robustness to the analytic choices was examined by resampling. RESULTS A genetically defined disease-intrinsic axis, including TP53 allelic state, RAS pathway mutations, ASXL1 and driver genotype, blasts and blood counts, predicted 2 year relapse incidence (validation concordance 0.69, 95% CI 0.63 to 0.74), whereas a non-overlapping host and structural axis, including portal vein thrombosis, donor type, patients' performance status, and age predicted 2-year non-relapse mortality (0.63, 95% CI 0.59 to 0.68). The two scores shared only 3.4% of their variance, indicating that a patient's disease genetics carried almost no information about non-relapse mortality. Variance decomposition showed that TP53 allelic state alone accounted for 30% of the relapse score. Recombined, the framework discriminated overall survival (concordance 0.640, 95% CI 0.616 to 0.662) better than every established prognostic system. For proof of concept, 3 risk groups separated in the validation cohort, with 5 year survival of 72%, 58%, and 39% (P<0.001), and the models were well calibrated. CONCLUSIONS Relapse and non-relapse mortality after transplantation for myelofibrosis are governed by distinct dimensions. Estimating both outcomes independently with genetic and clinical information, in addition to overall survival, establishes an individualized basis for transplant decision-making. The calculator is openly available (https://imtss-calculator.com).

## Ingrams for engrams: co-active inhibitory-inhibitory plasticity shapes inhibitory assemblies that stabilize and recall embedded engrams through disinhibition
- Source: bioRxiv (preprints)
- Date: 2026-09-17
- Categories: Computational neuroscience
- Authors: Kania, M., Confavreux, B., Vogels, T. P.
- DOI: 10.64898/2026.09.15.751748
- Source URL: <https://doi.org/10.64898/2026.09.15.751748>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.15.751748>

Abstract: Inhibitory neurons have been widely understood to play a supporting role in neural function and memory formation, but recent advances highlight that memories are encoded by both excitatory and inhibitory neurons (EI assemblies), carving out a bigger role for inhibition beyond mere stabilization. However, the computations enabled by such EI assemblies are still unclear, and the synaptic plasticity rules that could sustain and retrieve memories are unknown. Here, we construct a computational model for EI assembly recall and investigate the computational benefits of such mixed engrams over classical, excitatory-only engrams. Towards this end, we consider large recurrent spiking networks with symmetrical Hebbian synaptic plasticity at both inhibitory-to-inhibitory (I-to-I) and inhibitory-to-excitatory (I-to-E) synapses. The conjunction of these rules can robustly stabilize embedded EI assemblies in the asynchronous irregular regime. Assemblies can then be reactivated by two distinct mechanisms: direct stimulation of the engram or disinhibition through the ingram. Both mechanisms of recall lead to reliable pattern completion and separation. Crucially, we show that networks that lack I-to-I plasticity show weak recall and cannot discriminate between overlapping engrams via disinhibitory activation. This suggests that inhibitory plasticity can facilitate recall of overlapping engrams, with inhibitory neurons controlling multiple excitatory engrams. Furthermore, we show that ingrams can emerge from pre-existing E-I-E loops in the network's random connectivity. Our work proposes that inhibitory plasticity is a plausible mechanism for high-quality disinhibitory recall, allowing inhibitory neurons to selectively control multiple excitatory engrams through synapse-specific plasticity rules.

## Integrating Causal Inference into Pharmacovigilance: Target Trial Emulations for Proactive Signal Detection of Atorvastatin Initiation in Medicare Beneficiaries
- Source: medRxiv (preprints)
- Date: 2026-09-17
- Categories: Biological imaging
- Authors: Rowan, C. G., Tazare, J., Tran, M., Srivastava, S., Dreyer, N. A., Maringe, C.
- DOI: 10.64898/2026.07.01.26356874
- Keywords: cell counts, inference
- Source URL: <https://doi.org/10.64898/2026.07.01.26356874>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.07.01.26356874>

Abstract: ImportanceAdverse drug events (ADEs) in older adults are a substantial public health burden, yet spontaneous reporting systems detect them poorly owing to underreporting and the lack of a defined population. These limitations are of particular concern for older adults, who are underrepresented in pre-approval trials yet at elevated ADE risk owing to polypharmacy, multimorbidity, and age-related changes in drug metabolism. ObjectiveTo develop and apply an active, claims-based pharmacovigilance framework using sequential target trial emulation to detect ADE signals in older adults, with atorvastatin as the initial application. MethodsUsing Medicare fee-for-service claims (2017-2019), we studied statin-naive beneficiaries aged 65 years or older following hospitalization for myocardial or cerebral infarction. We emulated up to 14 daily sequential trials from the discharge date, classifying patients as initiating atorvastatin (A1), initiating a different medication (A2), or no new medication (A0); the primary contrast was A1 versus A2. For each trial, incident outcomes were ascertained and classified into 552 outcomes based on the Clinical Classifications Software Refined categories. Per-protocol effects were estimated over a 6-month follow-up period using Fine-Gray regression weighted by the inverse probability of treatment and censoring, treating death as a competing risk, with the false discovery rate controlled via the Benjamini-Hochberg procedure. A signal was declared when the q-value was \[≤\] 0.10 and the subdistribution hazard ratio (sHR) was \[≥\] 1.20 in any prespecified analytic stratum (sensitivity analyses used thresholds of q \[≤\] 0.20 and sHR \[≥\] 1.20). ResultsOf 70,130 eligible patients, 39,948 initiated atorvastatin (A1), 19,182 initiated another new medication (A2); after weighting, baseline characteristics were closely balanced. After excluding outcomes with sparse cell counts, 295 outcomes were analyzed; five met the primary signal detection criteria: valve disorders (sHR 1.71, 1.20-2.43); sprains and strains (sHR 1.79, 1.26-2.54); general sensation/perception symptoms (sHR 1.23, 95% CI 1.11-1.36); abnormal findings without diagnosis (sHR 1.55, 1.18-2.05); and prediabetes (sHR 1.71, 1.24-2.36). In the sensitivity analysis, we additionally detected: posthemorrhagic anemia, hemorrhagic stroke, varicose veins, other circulatory and skin conditions. ConclusionsAn active, claims-based framework using sequential target trial emulation detected both expected and previously unrecognized ADE signals following atorvastatin initiation in older adults, offering a systematic alternative to passive surveillance that can be extended to other commonly prescribed medications. Confirmatory analyses with disaggregated outcomes and narrower comparators are underway and will be reported separately. KEY POINTSO\_ST\_ABSQuestionC\_ST\_ABSCan an active, claims-based pharmacovigilance framework using sequential target trial emulation detect adverse drug event signals among older Medicare beneficiaries who initiated atorvastatin? FindingsIn this cohort study of 59,130 Medicare beneficiaries discharged after myocardial or cerebral infarction, a hypothesis-free scan across hundreds of prespecified outcomes detected five primary signals (valve disorders, sprains and strains, sensory symptoms including dizziness, abnormal laboratory findings, and prediabetes) and additional signals in sensitivity analyses (including hemorrhagic stroke), several of which align with known or labeled statin adverse effects (serving as positive controls that validate the framework) while others remain uncertain. MeaningThis framework offers a systematic, quantitative, and proactive alternative to spontaneous reporting for detecting medication safety signals in high-risk older adults. PLAIN LANGUAGE SUMMARYOlder adults use more prescription drugs than any other age group, yet they are often left out of clinical trials conducted before a drug is approved. As a result, the safety of many medications in older patients is poorly understood until the drugs are already in wide use. The main system for detecting new drug safety problems in the United States, the FDA Adverse Event Reporting System, relies on voluntary reports and captures only a small fraction of harms, which limits its usefulness. Using Medicare records, the researchers developed an active monitoring method and tested it on a commonly used cholesterol-lowering medication (atorvastatin) - that is frequently prescribed after a heart attack or stroke. They followed thousands of patients for six months, statistically balancing those who started atorvastatin with those who started a different new drug, and screened for hundreds of possible side effects. Evidence of potential medication harm emerged along a clear range of how likely the drug caused each problem. This ranged from well-known side effects (such as diabetes and sprains and strains) and outcomes already noted in the product labeling or prior research (such as hemorrhagic stroke and dizziness) to more uncertain associations (such as cardiac valve disorders) that are difficult to interpret. Overall, atorvastatin appeared largely safe in this older population, and the same method can now be used to study other widely used drugs.

## Integrating complementary biological information for multi-objective enzyme engineering
- Source: bioRxiv (preprints)
- Date: 2026-09-17
- Categories: Proteins & structural biology
- Authors: Blalock, N., Sosa, Y., Heuschkel, J., Li, R., Ma, X., Radomkit, S., Wu, H., Buono, F., Song, J., Pefaur, N., Kingsley, L. J., Romero, P. A.
- DOI: 10.64898/2026.09.15.751856
- Source URL: <https://doi.org/10.64898/2026.09.15.751856>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.15.751856>

Abstract: Enzyme catalysts are increasingly used for sustainable pharmaceutical manufacturing, but engineering industrial biocatalysts remains challenging as multiple catalytic and developability properties must be optimized simultaneously from limited experimental data. Here, we develop a machine learning-guided multi-objective design framework that integrates sparse functional measurements with complementary evolutionary and structural information to engineer the ketoreductase Gre2. The resulting designs achieved simultaneous improvements in catalytic performance, protein yield, and thermal stability, with selected designs retaining improved performance under process-relevant conditions. More broadly, our results demonstrate that integrating complementary biological information enables efficient multi-objective enzyme engineering from sparse experimental data, providing a general strategy for accelerating industrial biocatalyst development.

## Integrating Genomic Annotations and Traits Dependencies for single-nucleotide polymorphisms Prioritization with Causal Concept Bottleneck Models
- Source: bioRxiv (preprints)
- Date: 2026-09-17
- Categories: Genomics & sequence analysis
- Authors: De Santis, F., Malpetti, D., Gualdi, F., Mangili, F.
- DOI: 10.64898/2026.09.04.749112
- Source URL: <https://doi.org/10.64898/2026.09.04.749112>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.04.749112>

Abstract: Predicting common traits from single-nucleotide polymorphism (SNPs) data is challenging due to polygenicity, small effect sizes, and the presence of potentially mediated or spurious cross-trait associations. We propose a modeling approach that combines genomic annotations with known cross-trait relations by leveraging Causally Reliable Concept Bottleneck Models (C2BM), a deep learning architecture that factors the joint trait distribution over a graph of interpretable concepts. This design allows trait predictions to leverage information from other observed traits in addition to genomic inputs. Furthermore, the interpretable architecture of the model enables us to investigate how specific trait-trait relationships influence SNP-level predictions. We evaluate the approach on a multi-trait GWAS dataset covering five traits and show that C2BM improves predictions when ground-truth labels for related traits are available. Moreover, by analyzing variations in how trait-trait relationships influence predictions, we postulate that such differences may reflect the presence or absence of shared genetic mechanisms or indirect effects. Accepted at the CIBB 2026 conference (https://cibb2026.teralab.ai/)

## Integrating structural and biological evidence to rerank ESMFold2 protein-protein interactions
- Source: bioRxiv (preprints)
- Date: 2026-09-17
- Categories: Proteins & structural biology
- Authors: Xie, J., Li, M., Chai, Y., Ou, G., Li, W., Guo, Z.
- DOI: 10.64898/2026.09.16.751194
- Source URL: <https://doi.org/10.64898/2026.09.16.751194>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.16.751194>

Abstract: Large-scale protein structure prediction enables proteome-wide protein-protein interaction (PPI) screening, but distinguishing biologically meaningful interactions from spurious interfaces remains challenging. Here, we develop a scalable framework combining fast, MSA-free ESMFold2 prediction with PAE-guided domain parsing and evidence-based reranking. Three-recycle ESMFold2-Fast achieved 57% acceptable-or-better DockQ scores on FoldBench, comparable to AlphaFold2-Multimer while substantially reducing computation. PAE-guided parsing preserved 98.1% of XL-MS cross-links within parsed domain pairs. We then developed the Structure Prediction and Omics informed Classifier (SPOC)-ESMFold2, which integrates structural features with independent biological evidence to prioritize predicted interactions. SPOC-ESMFold2 achieved an AUROC of 0.93 and AUPR of 0.90, compared with 0.87 and 0.79 for a structural-only classifier. Under a 1:128 positive-to-negative ratio, SPOC-ESMFold2 achieved 17.2% recall at 5% false-discovery rate, substantially outperforming structural confidence metrics alone. This framework enables scalable PPI screening by integrating structural plausibility with orthogonal biological evidence to prioritize candidates for experimental investigation.

## inteRelate: flexible and thorough relating of genomic interval datasets through comparative overlap analysis
- Source: bioRxiv (preprints)
- Date: 2026-09-17
- Categories: Genomics & sequence analysis, Tools & resources
- Authors: Mamane-Logsdon, A., Kalemera, M. D., Maertens, G. N.
- DOI: 10.64898/2026.09.14.751391
- Source URL: <https://doi.org/10.64898/2026.09.14.751391>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.14.751391>
- Code: <https://github.com/loggy01/interelate>

Abstract: Summary Testing spatial relationships between genome-mapped features is both a common source of hypothesis generation and an additional layer of supporting evidence for experimental findings. Many computational tools automate the statistical association procedures used to assess overlap between pairs of genomic features. However, none provides a dedicated workflow that directly compares multiple features by assessing their overlap with a separate, common feature, first testing for overall heterogeneity and then identifying which overlap rates differ. Such comparative overlap analysis could directly facilitate comparisons of features within the same class across disease states, cell types, experimental perturbations, and other biological contexts. Here, we describe inteRelate, a software package that uses genomic interval datasets to test spatial relationships between genome-mapped features through comparative overlap analysis. The package functions as an end-to-end pipeline that is highly tunable and thorough in its statistical association procedures. We use experimental data to demonstrate the automation, accuracy, and insight inteRelate provides. Availability and implementation inteRelate is available at https://github.com/loggy01/interelate and archived at https://zenodo.org/records/21891072. Example uses are available in the online supplement. Additionally, the example datasets and results are available at https://zenodo.org/records/22012767.

## Intrinsic and circuit mechanisms of predictive coding in a grid cell network model
- Source: bioRxiv (preprints)
- Date: 2026-09-17
- Categories: Computational neuroscience
- Authors: Shaikh, I., Assisi, C.
- DOI: 10.1101/2025.04.11.648301
- Source URL: <https://doi.org/10.1101/2025.04.11.648301>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1101%2F2025.04.11.648301>

Abstract: Grid cells in the medial entorhinal cortex (MEC) fire at the vertices of a hexagonal lattice, forming an allocentric code for the animal's current position. Recent studies have identified a class of grid cells that represent locations ahead of the animal. How do these predictive representations emerge from the wetware of the MEC? We developed a detailed conductance-based model of the MEC network, constrained by empirical data on the biophysical properties of stellate cells and the topology of the MEC network. The model revealed two mechanisms by which grid cells can signal future locations. First, hyperpolarizing inhibition from interneurons activates HCN channels in stellate cells, whose slow kinetics maintain a depolarizing influence after inhibition ends, advancing spike timing and shifting the inferred position forward by ~5% of a grid field diameter. Second, introducing asymmetry into the inhibitory connectivity, by skewing the Gaussian profile of interneuron-to-stellate connections, causes inhibition to rise more steeply and enables earlier spiking, advancing the inferred position by up to ~25%. A corollary of our model is that the extent of the predictive code changes monotonically along the dorsoventral axis of the MEC, following the experimentally measured dorsoventral gradient in HCN time constants.

## KSTAR v1.2: A faster and more and accessible KSTAR for kinase activity inference
- Source: bioRxiv (preprints)
- Date: 2026-09-17
- Categories: Proteins & structural biology, Tools & resources
- Authors: Crowl, S., Custer, J.-L., Salazar Lopez, G., Lei-Dadey, C., Shimpi, A. A., Naegle, K. M.
- DOI: 10.64898/2026.09.14.751568
- Source URL: <https://doi.org/10.64898/2026.09.14.751568>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.14.751568>
- Code: <https://github.com/NaegleLab/KSTAR>

Abstract: Motivation: KSTAR is an algorithm with high flexibility for inferring kinase activity from any phosphoproteomic pipeline. However, in its first instantiation (v0.1) it requires Python programming and lots of memory and computational resources. Hence, we wished to improve speed and accessibility for broader uptake by researchers. Results: Here, we provide an updated algorithm that improves speed and memory, without affecting accuracy, along with some new features for increased usability and insight. KSTAR v1.2 has also been integrated into Galaxy for programming-free activity analysis and ProteomeScout for dataset preparation and interactive plotting. Availability and implementation: KSTAR is available at https://github.com/NaegleLab/KSTAR or on Galaxy on https://usegalaxy.org/. KSTAR Network resource assets are managed on Figshare at: https://doi.org/10.6084/m9.figshare.14944305.

## Lacuna: Cryptic Binding Pocket Discovery via Conformational Ensemble Analysis
- Source: bioRxiv (preprints)
- Date: 2026-09-17
- Categories: Proteins & structural biology, Tools & resources
- Authors: Moore, C.
- DOI: 10.64898/2026.08.14.744956
- Source URL: <https://doi.org/10.64898/2026.08.14.744956>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.08.14.744956>
- Code: <https://github.com/mooreneural/lacuna>

Abstract: Lacuna, an open-source Python tool for discovering cryptic binding pockets: sites that are absent or too small to detect in a protein's unbound structure and open only during conformational fluctuation. Most binding-site predictors score a single static structure, which is precisely the structure in which a cryptic site is invisible. Lacuna instead generates a conformational ensemble from any input structure, detects pockets independently in every conformer, clusters the detections into persistent sites across the ensemble, and ranks those sites with a model fitted on within-structure pairs. Ensemble generation is pluggable: normal mode analysis by default, with implicit-solvent molecular dynamics, Boltz-2 diffusion sampling, or a user-supplied ensemble as alternatives. On the designated test fold of CryptoBench, Lacuna recovers 55.6% of cryptic sites in its top five predictions with the zero-dependency default and 66.1% with an optional PLM-assisted ranker; pooling the geometric detector with an optional learned surface detector recovers 73.9% while raising the fraction of sites found from 68.5% to 86.4%, measured on the held-out fold at five conformers. It recovers 73%, 45% and 87% on the PocketMiner set, a curated set of literature apo/holo pairs, and COACH420 respectively. The default backend completes in a median of 2.6 seconds per chain on one CPU core, so ensemble-based pocket finding does not require a simulation budget. Every site carries a continuous crypticity score, and outputs are emitted as docking-ready Boltz YAML constraints, AutoDock Vina boxes, pseudoatom PDB files, and the generated conformational ensemble as a multi-model PDB. Lacuna is MIT licensed and available at https://github.com/mooreneural/lacuna and on PyPI as lacuna-pockets.

## Local ancestry inference identifies robust evidence of selection in Neolithic Europe
- Source: Molecular Biology and Evolution (journals)
- Date: 2026-09-17T00:00:00+00:00
- Categories: Genomics & sequence analysis, Evolution & metagenomics
- Authors: Georgia Mies, Iain Mathieson
- Journal: Molecular Biology and Evolution
- DOI: 10.1093/molbev/msag236
- Source URL: <https://doi.org/10.1093/molbev/msag236>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1093%2Fmolbev%2Fmsag236>

Abstract: During the European Neolithic, migrating Anatolian farmers admixed with local hunter-gatherers, coinciding with major shifts in diet, environment, and lifestyle that imposed strong selective pressures. Local ancestry inference is widely used to detect selection following admixture, but most methods were developed and validated on present-day populations. Their performance in ancient DNA – where reference panels are smaller, data are sparser, and admixture is more ancient – remains unresolved. We benchmark eight local ancestry inference methods on 176 imputed Neolithic genomes. While individual-level ancestry estimates are highly correlated across methods, inferred tract lengths and admixture time estimates vary by an order of magnitude. Overall, we recommend Gnomix or RFMix for general use. We also investigated our ability to detect natural selection using LAI. Integrating results across methods and replicating across methods and in two independent datasets (n=378 and 1,121) we identify a robust ancestry deviation at FADS1/2, consistent with adaptation on metabolism. We also identify IRAK4 (innate immunity) as a candidate locus, but with less consistent signal across methods. Finally, we replicate previous reports of excess hunter-gatherer ancestry at the HLA, but these results are inconsistent across methods and suggest that they may be affected by bias in local ancestry inference. Our findings demonstrate that while local ancestry inference recovers biologically meaningful signals in ancient genomes, results can be sensitive to the methods used for inference, particularly in complex regions like the HLA. Method choice critically influences inferred ancestry patterns and selection signals, underscoring the importance of multi-method validation.

## Local interaction networks reconstructed from global biodiversity data improve pollinator restoration decision making
- Source: bioRxiv (preprints)
- Date: 2026-09-17
- Categories: Tools & resources
- Authors: Baiotto, T., Cosma, C., Cheung, Y. Y. J., Narango, D., Woodard, J., McCarville, P., Echeverri, A., Horne, G., Wood, E., Williams, N. M., Seltmann, K. C., Fleri, J. R., Owens, A., Lequerica Tamara, M., Boren, A., Doneski, S., Guralnick, R. P., Li, D., Guzman, L. M.
- DOI: 10.64898/2026.03.30.715389
- Source URL: <https://doi.org/10.64898/2026.03.30.715389>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.03.30.715389>

Abstract: Global pollinator declines threaten the health of ecosystems and food systems, underscoring the urgency of conservation actions such as habitat restoration. However, data gaps on plant use among pollinators continue to limit reliable design of restoration plant mixes. To address this, we present NECTAR (Network-Enhanced Conservation Tool for Analysis and Recommendation), a new modular framework that integrates multiple data modalities - including species distributions, phenological metrics, and phylogenetic data - to infer flower visitation and host plant interactions from spatial, temporal, and phylogenetic overlap, generating spatially explicit plant-insect interaction networks that guide planting recommendations for pollinator habitat restoration. We demonstrate the utility of NECTAR by generating a large plant-insect metaweb across California, comprising 2,473,729 spatially explicit interactions that included 3,792 pollinator species and 4,363 native plant species. NECTAR achieved high interaction recall across withheld interactions and independent datasets, substantially outperforming null models and matching or exceeding values reported in comparable studies. NECTAR's data-driven plant mix recommendations are predicted to support up to 2.4 times more pollinator species compared to existing resources and random selection of plants. This optimization facilitates the inclusion of multiple goals and constraints, and provides complementary decision-making information to existing resources. NECTAR offers a scalable, evidence-based framework for translating increasingly available global biodiversity data into locally actionable restoration guidance, with broad potential to improve pollinator habitat restoration worldwide.

## Locat: Joint enrichment and depletion testing identifies localized marker genes in single-cell transcriptomics
- Source: bioRxiv (preprints)
- Date: 2026-09-17
- Categories: Genomics & sequence analysis, Single-cell & spatial, Tools & resources
- Authors: Lewis, W. R., Aizenbud, Y., Strino, F., Kluger, Y., Parisi, F.
- DOI: 10.64898/2026.04.03.716370
- Source URL: <https://doi.org/10.64898/2026.04.03.716370>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.04.03.716370>

Abstract: Several methods identify marker genes that delineate cell populations in single-cell transcriptomic data, yet most emphasize enrichment within candidate populations without testing whether expression is significantly reduced elsewhere. We present Locat, a framework for identifying highly specific localized genes by testing whether expression is concentrated within compact regions of a cellular embedding and depleted outside them. For each gene, Locat fits weighted Gaussian mixture models to gene-specific and background densities, computes concentration and depletion statistics, and integrates them into a unified localization score. Across synthetic benchmarks with controlled ground truth, Locat detects uni-modal, multi-modal, and sparse localized patterns and loses significance when expression becomes indistinguishable from background structure. In developmental, perturbation, and differentiation datasets, Locat identifies compact marker sets that capture lineage organization, condition-specific programs, and temporal dynamics. These sets are often smaller than highly variable gene selections, while embeddings built from them preserve major cell populations and developmental programs in several cases. In murine dermis, interferon-treated PBMCs, and retinoic acid-induced embryonic stem cell differentiation, localized genes recover differentiation trajectories, stimulus-responsive programs, and reproducible stage-specific patterns. Together, these results show that jointly assessing concentration and depletion yields specific, interpretable marker genes.

## LRP2: A proteogenomics pipeline for long-read informed protein isoform analysis and discovery
- Source: bioRxiv (preprints)
- Date: 2026-09-17
- Categories: Genomics & sequence analysis, Proteins & structural biology, Tools & resources
- Authors: Schertzer, M. D., Lewandowski, J. T., Watts, E. F., Rosenow, W., Mehlferber, M. M., Jeffery, E. D., Adamson, S. I., Bruand, J., Tseng, E., Neelamraju, Y., Garrett-Bakelman, F. E., Dolzhenko, E., Knowles, D. A., Sheynkman, G.
- DOI: 10.64898/2026.05.27.728216
- Source URL: <https://doi.org/10.64898/2026.05.27.728216>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.05.27.728216>

Abstract: Most human genes produce multiple RNA isoforms, yet it remains unclear which isoforms are translated into stable, functional proteins. Long-read RNA sequencing resolves full-length transcript structures and, when paired with mass spectrometry, can provide empirical evidence of isoform translation. Despite this opportunity, comprehensive workflows integrating isoform discovery, open reading frame prediction, peptide identification, and protein inference remain limited, leaving users to handle these steps piecemeal. Here, we present LRP2, a modular, end-to-end long-read proteogenomics pipeline built in Nextflow. LRP2 scales transcript discovery to hundreds of samples via PacBio's latest Isocall tool, removes technical artifacts with SQANTI QC, generates and classifies predicted proteomes via CPAT and SQANTI Protein, performs multi-group differential expression and usage analysis via edgeR, DRIMSeq, and a long-read adaptation of LeafCutter, and integrates protein-level evidence from DDA and DIA MS data through FragPipe. For cross-dataset comparison of novel isoforms, LRP2 employs deterministic splice-junction, coordinate-based isoform identifiers. Used as an integrated pipeline, LRP2 enables the detection of novel peptides and improves the protein isoform inference to confirm protein isoform translation.

## Mapping immune targets in peste des petits ruminants virus hemagglutinin: An integrated computational framework for vaccine candidate prioritization.
- Source: Veterinary immunology and immunopathology (journals)
- Date: 2026-09-17
- Categories: Genomics & sequence analysis, Proteins & structural biology
- Authors: Abubakar Garba
- Journal: Veterinary immunology and immunopathology
- DOI: 10.1016/j.vetimm.2026.111213
- External ID: 42759151
- Keywords: sequence alignment, epitope, epitopes, peptide, epitope score, framework
- Source URL: <https://doi.org/10.1016/j.vetimm.2026.111213>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1016%2Fj.vetimm.2026.111213>

Abstract: BACKGROUND: Peste des petits ruminants virus (PPRV) is a major transboundary viral pathogen of small ruminants and causes substantial economic losses in endemic regions. The hemagglutinin (H) protein mediates receptor recognition and host-cell attachment and is an important target for vaccine development. This study applied an integrated computational framework to identify conserved immunogenic regions within the PPRV H protein. METHODS: A total of 64 unique PPRV H protein sequences were analyzed using multiple sequence alignment, entropy-based conservation profiling, conservation-aware epitope prediction, comparison with experimentally characterized immune determinants, structural mapping, candidate-region ranking, and exploratory neural-network attribution analysis. Predicted epitopes were compared with reported immune determinants and contextualized using conservation and structural data. RESULTS: The workflow identified 9 predicted B-cell epitope candidates and 151 predicted T-cell peptide candidates distributed throughout the H protein sequence. The highest-scoring predicted B-cell epitope candidate was localized within residues 399-407 (SGPWSEGRI, Epitope\_Score: 1.0000, length: 9 aa), whereas the highest-scoring T-cell peptide candidate corresponded to residues 36-44 (YILLGVLLV; score: 0.889). Conservation analysis identified 412 residues with conservation scores greater than 0.9. Comparison with experimentally characterized immune determinants showed literature-based correspondence with selected predicted regions. Structural mapping provided three-dimensional context for conserved and predicted regions. Exploratory neural-network attribution scores were generated descriptively, without biological interpretation as validated antigenicity measures. CONCLUSIONS: Integrated computational analysis combining conservation profiling, immune epitope prediction, comparison with experimentally characterized immune determinants, structural interpretation, candidate-region ranking, and exploratory neural-network attribution analysis identified conserved and computationally predicted immune candidate regions within the PPRV H protein. Residues 399-407 (SGPWSEGRI) represented the highest-scoring computationally predicted B-cell epitope region and warrant further investigation and experimental confirmation.

## MMAD-Risk: Multivariate Mixed Survival Analysis for the Prediction of Age-Dependent Disease Risks from Plasma Proteomes
- Source: bioRxiv (preprints)
- Date: 2026-09-17
- Categories: Mathematical biology & statistics
- Authors: Hilger, A. M., Soeding, J.
- DOI: 10.64898/2026.09.16.752074
- Source URL: <https://doi.org/10.64898/2026.09.16.752074>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.16.752074>

Abstract: Motivation: Multivariate survival analysis with hundreds of correlated outcomes is computationally challenging. Established approaches either ignore correlations between response variables, rely on black-box deep learning or are limited to small-scale outcomes. Results: We introduce MMAD-Risk, a novel multivariate mixed accelerated failure time (AFT) model that enables scalable analysis of high-dimensional survival analysis data. We train MMAD-Risk using amortized variational inference where we design the variational distribution such that it factorizes across diseases, allowing us to decompose multivariate disease risk prediction into a series of tractable, one-dimensional problems. This allows us to calculate the ELBO analytically, enabling fast computation. The model employs a low-rank decomposition of the effect size matrix B = VW to capture shared disease mechanisms and latent random effects Vz to model comorbidity. MMAD-Risk is trained on the UK Biobank Pharma Proteomics cohort (N \{approx\} 55000, P \{approx\} 3000 proteins, D = 271 diseases). Using the full 3,000-protein dataset, MMAD-Risk achieved a mean concordance index (c-index) of 0.744 for diagnoses occurring \[>=\] 10 years after blood sample collection, outperforming a Cox proportional hazards model (mean c-index = 0.709). Greedy backward selection identified a 10-protein panel that preserved > 99% of the full-model performance. On this reduced panel MMAD-Risk still outperformed Cox regression (0.738 vs. 0.679).

## Modeling Protein Sequence Evolution as an Ornstein-Uhlenbeck Process in a Latent Space
- Source: bioRxiv (preprints)
- Date: 2026-09-17
- Categories: Genomics & sequence analysis, Proteins & structural biology, Evolution & metagenomics
- Authors: De Leonardis, M., Pagnani, A.
- DOI: 10.64898/2026.09.16.751972
- Source URL: <https://doi.org/10.64898/2026.09.16.751972>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.16.751972>

Abstract: High-throughput directed evolution produces longitudinal sequence libraries that are ideal for probing local fitness neighborhoods but often underpowered for global inference tasks such as contact prediction. We present an unsupervised inference model that integrates directed-evolution sequencing time series with natural homologs. We project sequences into a low-dimensional latent space learned from the natural multiple sequence alignment and model the experimental process as an Ornstein-Uhlenbeck dynamics in that space. Maximum-likelihood estimation of the latent drift and noise parameters determines a stationary Gaussian distribution, which induces an effective Potts model in sequence space. The inferred couplings improve structural contact prediction by combining global evolutionary constraints from nature with local, experiment-specific signals. Experiments on PSE1 \{beta\}-lactamase and dihydrofolate reductase demonstrate the ability to identify correct complementary contacts not recovered by methods using either natural or experimental data alone, with gains concentrated in intermediate- and long-range contacts.

## Monte Carlo modeling of the formation and organization of ion channel clustering
- Source: PLOS Computational Biology (journals)
- Date: 2026-09-17T00:00:00+00:00
- Categories: Mathematical biology & statistics
- Authors: Nicolae Moise, Seth H. Weinberg
- Journal: PLOS Computational Biology
- DOI: 10.1371/journal.pcbi.1014791
- Source URL: <https://doi.org/10.1371/journal.pcbi.1014791>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1371%2Fjournal.pcbi.1014791>

Abstract: The spatial organization of ion channels on cell membranes critically influences many key physiological processes, such as cardiac and neuronal excitability and cellular signaling, yet the mechanisms governing channel clustering remain poorly understood. In this study, we present a stochastic computational framework that models the dynamic organization of ion channels through Monte Carlo simulations incorporating membrane insertion, removal, channel-channel interactions, and diffusion processes. Our model reveals several fundamental principles of membrane domain formation. In single-channel systems, we demonstrate a biphasic relationship between interaction energy and cluster size, with optimal clustering occurring at intermediate interaction strengths, suggesting that excessively strong interactions can impede cluster growth by restricting channel mobility. In two-channel systems, we find that the interplay between homotypic and heterotypic interactions determines whether channels form mixed or segregated clusters, with asymmetric clustering behaviors emerging when homotypic interaction strengths differ between channel types. Simulations of three-channel systems demonstrate emergent organizational principles leading to hierarchical clustering patterns and specialized domain formation. These findings generate testable predictions about how channel density, trafficking dynamics, and interaction energies collectively alter ion channel spatial organization, in the setting of both physiological function and pathophysiological conditions.

## Multi-cohort machine learning identifies a ferroptosis-linked prognostic signature in lung adenocarcinoma
- Source: Frontiers in Bioinformatics (journals)
- Date: 2026-09-17T00:00:00Z
- Categories: Genomics & sequence analysis
- Authors: Rana Salihoğlu
- Journal: Frontiers in Bioinformatics
- DOI: 10.3389/fbinf.2026.1921468
- External ID: 042b1e6b47d25032b2c969b2062ce767aef4cf0c
- Keywords: transcriptomic
- Source URL: <https://doi.org/10.3389/fbinf.2026.1921468>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.3389%2Ffbinf.2026.1921468>

Abstract: Lung adenocarcinoma (LUAD) shows marked outcome heterogeneity within clinicopathological stage groups. This study developed and externally validated a ferroptosis-linked transcriptomic risk model using a leakage-controlled multi- cohort survival-learning framework and characterized its immune context. Candidate genes were defined by combining FerrDb V3-annotated genes within a ferroptosis-associated weighted gene co-expression network analysis (WGCNA) module with a filtered WGCNA discovery branch derived from the independent GSE81089 cohort. Model development used TCGA-LUAD, GSE31210, GSE72094, and GSE136961 (1,149 patients; 344 overall-survival events) with bagged cross-cohort Cox screening, leave-one-cohort-out (LOCO) stability locking, and benchmarking of 150 survival-learning configurations. External evaluation used the GSE50081, GSE68465, and GSE30219 cohorts (912 patients; 507 events). The development-selected extra survival trees configuration reached a mean leave-one-cohort-out Uno’s C-index of 0.723. The highest cohort-specific C-indices among the prespecified candidate configurations were 0.611, 0.691, and 0.699 and arose from different model configurations. Applying the same development-selected configuration to all three external cohorts yielded Harrell’s C-indices of 0.605, 0.675, and 0.682 and 5-year Uno’s C-indices of 0.605, 0.678, and 0.696. In TCGA-LUAD, the stored out-of-fold molecular score remained prognostic after TNM stage adjustment (hazard ratio per standard deviation 2.32, 95% confidence interval: 1.39−3.87; p = 0.0013). At 3 years and 5 years, the combined TNM-plus-score models showed close calibration and modest gains in discrimination, while decision-curve analysis identified positive incremental net benefit only over restricted threshold ranges. Low-risk tumors were enriched for interferon, complement, and immune-cell programs. These computational findings require further prospective assay-level and experimental validation before clinical implementation.

## Multi-omics insights into growth impairment mechanisms in children with persistent diarrhea
- Source: Microbiology Spectrum (journals)
- Date: 2026-09-17T00:00:00Z
- Categories: Single-cell & spatial, Systems & networks
- Authors: Jian Shen, Jun-Le Yan, Ying Yuan, Li-Lin Le, Juan Xu, Bai-Lu Chen, Xiao-Ying Liu, Hui-Jie Chen, Li-Jun Chen, Mei-Xiang Yi, Jiajia Lyu, Jun Diao, Xin-Lin Zhang, Yin-Qiu Zhao, Jing-Ru Chen
- Journal: Microbiology Spectrum
- DOI: 10.1128/spectrum.03905-25
- External ID: 9c6368dbabe7226b165e8e43bda4555da78176dd
- Keywords: multi omics, pathways
- Source URL: <https://doi.org/10.1128/spectrum.03905-25>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1128%2Fspectrum.03905-25>

Abstract: This study aimed to elucidate the key mechanisms underlying short stature with pediatric persistent diarrhea in children (PDC) and to propose an integrated model linking the butyrate-producing gut microbial niche, short-chain fatty acids (SCFAs), Th17/Treg balance, the growth hormone-insulin-like growth factor 1 (GH-IGF-1) axis, and growth regulation. Samples from healthy controls (HCs), PDC patients without short stature (PDC-NS), and PDC patients with short stature (PDC-S) were analyzed using multi-omics profiling and machine learning-based predictive modeling. The results showed that PDC-S patients exhibited disruption of core butyrate-producing bacterial communities and related metabolic pathways, markedly reduced fecal butyrate levels, systemic Th17/Treg imbalance characterized by a pro-inflammatory state, and dual suppression of receptor- and ligand-level components of the GH-IGF-1 axis. Multi-omics machine learning identified a five-factor risk prediction panel composed of metabolic, immune, and endocrine markers, while causal inference further established butyrate as a central regulator of growth. Mouse experiments further validated that combined intervention with butyrate, Faecalibacterium prausnitzii, and recombinant human growth hormone (rhGH) ameliorated growth retardation. Overall, this study proposes a precision therapeutic strategy based on butyrate and GH co-intervention, providing new mechanistic insights and translational tools for PDC-associated short stature. IMPORTANCE The research conducted in this study holds significant implications for improving the understanding and treatment of stunted growth with persistent diarrhea in children (PDC). By unraveling the intricate pathways linking gut microbiota, immune responses, and growth hormone regulation, the study sheds light on the underlying mechanisms contributing to stunting in these vulnerable populations. The identification of key factors, such as butyrate-producing bacteria and the Th17/Treg balance, not only enhances our comprehension of PDC-related growth impairment but also offers a promising avenue for targeted interventions. The proposed integrated mechanistic model and intervention strategy pave the way for precision therapies tailored to address the specific biological mechanisms at play, potentially leading to more effective and personalized treatments for children suffering from PDC-associated stunting. The research conducted in this study holds significant implications for improving the understanding and treatment of stunted growth with persistent diarrhea in children (PDC). By unraveling the intricate pathways linking gut microbiota, immune responses, and growth hormone regulation, the study sheds light on the underlying mechanisms contributing to stunting in these vulnerable populations. The identification of key factors, such as butyrate-producing bacteria and the Th17/Treg balance, not only enhances our comprehension of PDC-related growth impairment but also offers a promising avenue for targeted interventions. The proposed integrated mechanistic model and intervention strategy pave the way for precision therapies tailored to address the specific biological mechanisms at play, potentially leading to more effective and personalized treatments for children suffering from PDC-associated stunting.

## Multi-organelle signatures map cell-state diversity and metabolic adaptation in tissues
- Source: Science (journals)
- Date: 2026-09-17T00:00:00+00:00
- Categories: Biological imaging
- Authors: Raghabendra Adhikari, Alexander Hillsley, Alana Dowdell Johnson, Shihong Max Gao, Isabel Espinosa-Medina, Jan Funke, Daniel Feliciano
- Journal: Science
- DOI: 10.1126/science.ady6372
- Source URL: <https://doi.org/10.1126/science.ady6372>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1126%2Fscience.ady6372>

Abstract: Cell-state diversity drives tissue adaptability, repair, and disease resilience, but capturing this complexity is a challenge. Current approaches rely on transcriptional profiling and overlook organelle structure, a key indicator of metabolism and stress. We developed spatial Organellomics (sOrganellomics), an imaging workflow that integrates automated segmentation with machine learning to classify and spatially map cell states from multi-organelle signatures. In liver and pancreas, these signatures distinguished broad cellular classes. In liver, sOrganellomics revealed that zonal position did not fully explain organelle-defined hepatocyte categories. Instead, hepatocytes formed intermixed communities within canonical zones, supporting a refined subzonal diversity model. Nutritional stress reshaped this organization. Intravital imaging linked fasting-induced organelle remodeling with altered mitochondrial membrane potential in vivo, supporting multi-organelle architecture as a structural readout of tissue adaptation.

## Multiplexed embryo profiling links cellular state to zygotic genome activation in single cells
- Source: Nature Communications (journals)
- Date: 2026-09-17T00:00:00+00:00
- Categories: Single-cell & spatial, Biological imaging
- Authors: Max Hess, Marvin F. Wyss, Edlyn Wu, Gian-Marco Schaniel, Joel Lüthi, Chiara Rebagliati, Daniel Hannuschke, Nadine L. Vastenhouw, Darren Gilmour, Shayan Shamipour, Lucas Pelkmans
- Journal: Nature Communications
- DOI: 10.1038/s41467-026-77784-7
- Source URL: <https://doi.org/10.1038/s41467-026-77784-7>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1038%2Fs41467-026-77784-7>

Abstract: Multicellular self-organization depends on interactions across multiple length scales, yet mapping protein states at high spatial resolution in whole-mount embryos remains challenging. Here, we introduce high-throughput 3D in toto iterative immunofluorescence imaging (3D-4i) and a dedicated computer vision pipeline to quantify morphological and molecular features from subcellular to whole-embryo scales across hundreds of samples. Applying this pipeline to early zebrafish embryos undergoing mid-blastula transition, we determine the cell cycle phase for each cell across the embryo, and uncover the spatiotemporal dynamics by which global meta-synchronous mitotic waves transition to cell cycle desynchronization. Using statistical analysis, we find that the cell cycle phase is the major source of variability in transcription within a division cycle, and combining this with the analysis of key transcription factors and chromatin modifier state, included in our multiplexed dataset, enables accurate prediction of transcriptional output during zygotic genome activation in individual cells. Together, these findings establish 3D-4i as a powerful approach for quantifying multimodal, multiscale biological processes underlying multicellular self-organization.

## Neural representational geometry of a joint code for stimulus category and category-independent features
- Source: bioRxiv (preprints)
- Date: 2026-09-17
- Authors: Tiberi, L., Sompolinsky, H.
- DOI: 10.64898/2026.03.23.713692
- Source URL: <https://doi.org/10.64898/2026.03.23.713692>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.03.23.713692>

Abstract: A central question in neuroscience and machine learning is how a single neural representation can support linear access to multiple kinds of information about the same stimulus. While this ability is widespread in biological and artificial systems, the representational principles that make such joint coding possible remain poorly understood. We address this question for an important class of representations: those jointly encoding discrete stimulus categories and continuous features that vary independently of category. Using the framework of category manifolds (the sets of neural representations elicited by stimuli from the same category) we extend existing theories of manifold classification to the equally essential task of regressing category-independent features, determining which aspects of manifold geometry govern regression performance. This provides a unified framework for understanding how classification- and regression-relevant geometry can be jointly optimized to implement an effective joint code. Applying this framework to convolutional neural networks (CNNs), we find that regression-relevant geometry can be optimized through subtle changes that largely preserve classification-relevant geometry. This explains why common representational-similarity measures previously failed to distinguish joint codes from codes optimized exclusively for classification. Motivated by prior work in visual neuroscience suggesting that macaque inferotemporal cortex may jointly encode object category and category-independent features such as object position and size, we use our framework to identify principled geometric signatures that distinguish joint codes from classification-only codes in CNNs and can be tested in future neural recordings. Finally, we characterize how these signatures are affected by common experimental constraints: limited stimulus categories and neural-population subsampling.

## Neurological and hematological safety profiles of GSK-3β inhibitors: insights from multi-database pharmacovigilance and experimental validation
- Source: Frontiers in Immunology (journals)
- Date: 2026-09-17T00:00:00Z
- Categories: Genomics & sequence analysis, Systems & networks, Computational neuroscience, Tools & resources
- Authors: Xin-Chi Luan, Bing-Cheng Fan, Xue-Zhe Wang, Xiao-Xuan Li, Yuhui Song, Xiao-Lei Zhang, Huhu Zhang, Ruo-Lan Chen, Yi Li, Ze-Ling Yang, Ning Liu, Wei-Wei Qi, Wen-Sheng Qiu, Jing Guo
- Journal: Frontiers in Immunology
- DOI: 10.3389/fimmu.2026.1915614
- External ID: 1d3687f224182306018861025c53437a758b7d55
- Keywords: neuronal, transcriptomic, pathway, database
- Source URL: <https://doi.org/10.3389/fimmu.2026.1915614>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.3389%2Ffimmu.2026.1915614>

Abstract: Glycogen synthase kinase 3 beta (GSK-3β) inhibitors have received substantial attention for their therapeutic potential; however, their systemic safety profile remains incompletely characterized. This study characterized the research landscape and exploratory safety-reporting signals associated with agents with reported GSK-3β activity by integrating bibliometric analysis, pharmacovigilance, and preliminary experimental assessment. A multi-layered framework included bibliometric analysis, disproportionality analyses of FAERS, JADER, and CVARD reports, and transcriptomic profiling. SH-SY5Y cells were treated with 9-ING-41 (1 μM, 24 h); cell viability, qRT-PCR, and DCFH-DA-based intracellular oxidative-stress measurements were assessed. Bibliometric analysis showed sustained growth in GSK-3β-related research. Pharmacovigilance identified neurological and hematologic disproportionality signals across 27 System Organ Classes. In SH-SY5Y cells, 9-ING-41 was associated with modest changes in neuronal-function, inflammatory-response, and GSK-3β/Wnt-pathway transcripts, increased DCFH-DA fluorescence, and high cell viability. These cell-based observations are preliminary and do not establish clinical causality. The literature-derived study-drug panel showed exploratory neurological and hematologic reporting signals. Cross-database recurrence can prioritize hypotheses, whereas pharmacological heterogeneity and the limitations of spontaneous reporting require cautious interpretation. The SH-SY5Y experiments provide preliminary biological plausibility only.

## On the feasibility of temporal interference stimulation of human brains using two arrays of electrodes
- Source: bioRxiv (preprints)
- Date: 2026-09-17
- Authors: Huang, Y.
- DOI: 10.64898/2026.03.31.715653
- Source URL: <https://doi.org/10.64898/2026.03.31.715653>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.03.31.715653>

Abstract: Background: Conventional temporal interference stimulation (TI, TIS, or tTIS) leverages two pairs of electrodes to induce an interfering electrical field in the brain. Both computational and experimental studies show that TI can stimulate deep brain regions without significantly affecting shallow areas. While promising, optimization of the locations and dosages on these two pairs of electrodes for maximal focal modulation remains computationally challenging. We are the first to propose two arrays of electrodes instead of two or multiple pairs of electrodes to boost modulation focality. However, the optimization algorithm outputs too many electrodes with overlaps across two frequencies, making it difficult to implement in practice. Objective: Based on recent progress in developing multi-channel TI devices and computational work on TI optimization, here we again advocate two-array TI with feasibility data. Methods & Results: We give a review on major algorithms for TI optimization, and compare these algorithms over 25 individual heads across six brain targets. We show that the latest optimization algorithm for two-pair TI innately works for two-array TI with the fastest speed (under 30s) and a similar amount of electrodes as in multi-pair TI. At four of the six targets, this fastest algorithm for two-array TI achieves similar or even better focality than TI using up to 16 pairs of electrodes that takes days to optimize. We also show a hardware implementation of two-array TI using 10 electrodes on our 8-channel TI device. We argue that two-pair TI is only preferred when one does not care about modulation focality or when hardware is limited to only two current sources. We restate the focality-intensity tradeoff but in the context of TI and provide a first voxel-level map (at 4 mm resolution) of achievable focality and modulation strength by TI in the MNI-152 head template. Conclusions: Compared to two-pair or multi-pair TI, we promote two-array TI for its similar performance in focality and lower cost in terms of both optimization time and electrodes needed. We hope this work will pave the way for future adoptions of two-array TI for more focal non-invasive deep brain stimulation.

## Optimising digital volume correlation across materials: a practical framework for accuracy and spatial resolution
- Source: bioRxiv (preprints)
- Date: 2026-09-17
- Categories: Biological imaging
- Authors: Parmenter, A. L., Sharma, A., Bay, B. K., Lee, P. D.
- DOI: 10.64898/2026.09.15.751389
- Source URL: <https://doi.org/10.64898/2026.09.15.751389>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.15.751389>

Abstract: Digital volume correlation (DVC) provides full-field three-dimensional measurements of internal deformation, but its accuracy and effective spatial resolution depend on image-processing and analysis choices. Here, we use fibrous, cartilaginous and mineralised tissues within a rat intervertebral disc (IVD) as a controlled multi-material case study, combining virtual compression with experimentally loaded synchrotron computed tomography images. We systematically evaluate phase retrieval and image filtering, bit-depth conversion, point-cloud design, subvolume size and strain-field smoothing. Stronger phase retrieval reduced correlation residuals while increasing displacement and strain errors, showing that residual minimisation alone can select poorer parameters. Image filtering sensitivity was greatest in fibrous tissue, which had the smallest characteristic image feature size, whereas inappropriate intensity mapping during 16-bit to 8-bit conversion preferentially degraded low-contrast cartilage. Increasing subvolume size, point spacing or strain-window size improved measurement robustness but progressively smoothed local strain heterogeneity. We demonstrate that the spatial resolution of strain measurement must be matched across tissue types in order to compare strain magnitude; in the IVD, matching DVC spatial resolution changed the apparent ratio of compressive strain among fibrous, cartilaginous and mineralised tissues from 3.1:1.9:1 to 14:7.7:1. These findings establish a sequential, deformation-based optimisation framework in which image characteristics guide processing, known deformations validate accuracy and strain fields are compared at matched measurement scales. The framework supports more reproducible and mechanically interpretable DVC analyses in heterogeneous biological and engineered materials.

## Pan-Genome-Scale Metabolic Reconstruction Reveals Conserved Metabolic Functions in Candida albicans
- Source: Journal of Fungi (journals)
- Date: 2026-09-17T00:00:00Z
- Categories: Genomics & sequence analysis, Systems & networks, Tools & resources
- Authors: Ya Meng, Yi-Ming Zhang, Lei Zhang
- Journal: Journal of Fungi
- DOI: 10.3390/jof12090697
- External ID: cb1a9e5ff49283b1af0526aef1d00e8c40aa5c73
- Source URL: <https://doi.org/10.3390/jof12090697>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.3390%2Fjof12090697>

Abstract: Candida albicans is a major cause of human mucosal and invasive fungal infections, but the relationship between its intraspecific genomic diversity and metabolic variation remains poorly understood. Here, we integrated 80 public C. albicans genome assemblies, published fungal genome-scale metabolic models (GEMs), public reaction databases, and orthogroup-linked gene–protein–reaction (GPR) evidence to construct a species-level C. albicans pan-GEM and derived 80 strain-specific GEMs (ssGEMs) through genome projection. The pan-genome comprised 10,308 orthogroups, including 4215 core, 5947 accessory, and 146 singleton orthogroups. The final pan-GEM contained 1986 reactions, 1777 metabolites, and 865 genes. After feasibility rescue, all 80 ssGEMs met the feasibility criterion for predicted growth and passed the closed-uptake energy-generating-cycle test. Among experimentally essential genes with resolvable GPR associations, 23 were consistently predicted as model-essential across all final ssGEMs. As an application of the ssGEM collection, nutrient-boundary simulations showed that increasing D-glucose uptake markedly increased predicted growth across 79 feasible ssGEMs. This framework provides a reusable resource for comparing conserved metabolic functions and genome-projected reaction differences across C. albicans strains.

## PanDelos-plus: A parallel algorithm for computing sequence homology in pangenomic analysis
- Source: PLOS Computational Biology (journals)
- Date: 2026-09-17T00:00:00+00:00
- Categories: Genomics & sequence analysis, Tools & resources
- Authors: Simone Colli, Emiliano Maresi, Vincenzo Bonnici
- Journal: PLOS Computational Biology
- DOI: 10.1371/journal.pcbi.1014724
- Source URL: <https://doi.org/10.1371/journal.pcbi.1014724>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1371%2Fjournal.pcbi.1014724>

Abstract: The identification of homologous gene families across multiple genomes is a central task in bacterial pangenomics traditionally requiring computationally demanding all-against-all comparisons. PanDelos addresses this challenge with an alignment-free and parameter-free approach based on k-mer profiles, combining high speed, ease of use, and competitive accuracy with state-of-the-art methods. However, the increasing availability of genomic data requires tools that can scale efficiently to larger datasets. To address this need, we present PanDelos-plus, a fully parallel, gene-centric redesign of PanDelos. The algorithm parallelizes the most computationally intensive phases (Best Hit detection and Bidirectional Best Hit extraction) through data decomposition and a thread pool strategy, while employing lightweight data structures to reduce memory usage. Benchmarks on synthetic datasets show that PanDelos-plus achieves up to 14x faster execution and reduces memory usage by up to 96%, while maintaining consistency with the original algorithm. These improvements allow the PanDelos methodology to be applied to population-scale comparative genomics, thus enabling more precise characterisation of pangenome structure and dynamics. PanDelos-plus is available at github.com/synbionics/PanDelos-plus .

## Pareto Suboptimal Resource Allocation and Microbial Growth
- Source: bioRxiv (preprints)
- Date: 2026-09-17
- Authors: Baldazzi, V., Mairet, F., Gedeon, T., de Jong, H.
- DOI: 10.64898/2026.09.08.749853
- Source URL: <https://doi.org/10.64898/2026.09.08.749853>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.08.749853>

Abstract: Microbial growth has often been analyzed under the assumption that microorganisms have evolved to optimize phenotypic characteristics of interest, such as growth rate and growth yield. This assumption has been useful, for example, for the genome-scale modeling of metabolism and the study of the allocation of cellular resources to physiological processes. In many experimental situations of interest, however, microorganisms are found to be suboptimal with respect to phenotypic characteristics that are thought to be favorable in that situation. Whereas a strong theoretical framework exists to mathematically relate cellular resource allocation strategies to Pareto optimality of microbial growth and other biological processes, much less is known about the consequences of Pareto suboptimality. We extend the framework to the latter case and show that a given Pareto suboptimal phenotype can be explained by a range of underlying resource allocation strategies, each corresponding to a different growth physiology and biomass composition. We test the predictions with the help of a coarse-grained model of microbial growth and published experimental data, which relate Pareto suboptimal rate-yield phenotypes of Escherichia coli and the microalga Tisochrysis lutea to the macromolecular composition of the cells (storage metabolite and total protein contents). Changing the focus from Pareto optimality to suboptimality provides interesting leads to exploring the diversity of growth strategies that can support a given phenotype. This change of perspective is of practical interest, because some of the growth physiologies within this range may be important for biotechnological applications

## Per- and polyfluoroalkyl substances and kidney disease: Genetic associations and computational prioritization of candidate toxicogenomic pathways
- Source: PLOS Computational Biology (journals)
- Date: 2026-09-17T00:00:00+00:00
- Categories: Systems & networks
- Authors: Dianjie Zeng, Yuxi Wang, Yinhuai Wang, Guoqiang Li, Wenpeng Wang, Zhongkun Zuo
- Journal: PLOS Computational Biology
- DOI: 10.1371/journal.pcbi.1014665
- Keywords: pathways, pathway
- Source URL: <https://doi.org/10.1371/journal.pcbi.1014665>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1371%2Fjournal.pcbi.1014665>

Abstract: Per- and polyfluoroalkyl substances (PFAS) are persistent environmental pollutants with bioaccumulation potential, but their associations with kidney diseases remain incompletely understood. This study integrated Mendelian randomization and computational toxicology to examine associations between genetically predicted circulating PFAS levels and kidney disease outcomes and to prioritize candidate toxicogenomic pathway themes. Genetically predicted higher PFOA levels were inversely associated with IgA nephropathy (OR = 0.21, P = 0.004), but positively associated with hypertensive nephropathy (OR = 1.20, P < 0.001) and calculus of kidney (OR = 1.24, P = 0.016). Genetically predicted higher PFOS levels were inversely associated with IgA nephropathy (OR = 0.27, P = 0.046) and urinary tract infection (OR = 0.94, P = 0.003). A primary association was also observed between PFOA and membranous nephropathy (OR = 1.56, P = 0.028), but this association was not retained after targeted SNP-exclusion analyses and was therefore not interpreted as a robust or established causal association. Computational toxicology analyses prioritized database-derived candidate targets and pathway themes related to immune response, inflammation, oxidative stress, and apoptosis. Network-prioritized candidate nodes included CTNNB1, TP53, and EGFR in the PFOA–calculus of kidney network, IGF1 in the PFOA–hypertensive nephropathy network, and CCL2, TLR4, MMP9, and IFNG in the PFOA/PFOS–IgA nephropathy networks. In contrast, IL1B, TNF, and IL6 were observed as shared inflammatory nodes across multiple nephropathy-related networks. These candidates should be interpreted as database-derived and network-prioritized targets rather than experimentally validated causal mediators of kidney disease. Overall, this study provides a hypothesis-generating framework for exploring associations among genetically predicted PFAS-related traits, kidney disease outcomes, and candidate toxicogenomic pathway themes that require experimental validation.

## Phenotype-stratified computational convergence of osteoarthritis loci across human knee cell programs.
- Source: Computational biology and chemistry (journals)
- Date: 2026-09-17T00:00:00Z
- Categories: Genomics & sequence analysis, Single-cell & spatial
- Authors: Hao-Nan Zhang, Jin-Mei Ye, Min-Cong Wang, Cheng-Long Pan, Yong Hu
- Journal: Computational biology and chemistry
- DOI: 10.1016/j.compbiolchem.2026.109424
- External ID: 9fc359263052ca58dff6e8a3570535ded4dbd700
- Keywords: genome, single cell, single nucleus
- Source URL: <https://doi.org/10.1016/j.compbiolchem.2026.109424>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1016%2Fj.compbiolchem.2026.109424>

Abstract: Genome-wide association studies of osteoarthritis use related but non-identical endpoints, including knee osteoarthritis, broader hip-and-knee osteoarthritis, and total knee replacement. Whether these endpoint-specific genetic signals map to distinct human joint cell programs remains unclear. We developed and applied a phenotype-stratified computational convergence framework that integrates osteoarthritis genetic loci with a unified human knee single-cell/single-nucleus atlas. Three predefined locus groups were analyzed: knee osteoarthritis-enriched, total knee replacement-enriched, and shared broader-osteoarthritis loci. Candidate effectors were assigned using a refined locus-to-gene evidence layer and scored against knee cell programs using gene-wise program z-scores. Convergence was evaluated using one-sided upper-tail gene-set permutation testing, with Benjamini-Hochberg correction across the complete family of three phenotype groups × eight cell programs. Knee osteoarthritis-enriched loci showed their strongest discovery-atlas alignment with a fibro-inflammatory synovial program (mean program z = 1.107; nominal permutation p = 0.008; BH q = 0.192), whereas total knee replacement-enriched loci aligned most strongly with a cartilage ossification-like structural program (mean program z = 0.597; nominal permutation p = 0.018; BH q = 0.216). No primary convergence test remained significant after correction across the 24-test family. In donor-aware cartilage analysis, the cartilage ossification-like program ranked first in 28 of 31 cartilage sample units. External public-data stress testing identified transferability boundaries: synovial marker-level signals were partly directionally consistent, whereas small candidate-effector modules were not consistently reproduced across external bulk or single-cell datasets. These results provide phenotype-linked computational prioritization of human knee cell programs while defining clear statistical and biological limits on their interpretation.

## PlantRegMoD: An integrative and AI-driven multi-omics database for plant regeneration research
- Source: bioRxiv (preprints)
- Date: 2026-09-17
- Categories: Tools & resources
- Authors: Li, Y.-M., Ma, Y.-B., Gui, L.-Y., Tang, Y., Meng, C., Li, Y., Yao, L., Zhang, J., Xia, S., Peng, Y., Song, S., Zeng, Z., He, J.-B., Zhang, N., Xiao, P.-X., Xu, Y., Tan, L., Iwase, A., Chen, C., Jiao, W.-B.
- DOI: 10.64898/2026.09.15.751352
- Source URL: <https://doi.org/10.64898/2026.09.15.751352>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.15.751352>

Abstract: Plant regeneration underpins plant developmental plasticity, tissue culture, genetic transformation and crop improvement. Although numerous related omics datasets have been generated, specialized multi-omics databases for this field are still scarce. Here, we constructed PlantRegMoD, an AI-powered integrated database dedicated to plant regeneration. This platform hosts 20.54 TB standardized multi-omics data from 147 projects across 32 plant species and 2,593 samples. We established a unified hierarchical classification system covering five major categories and nine regeneration models, and curated 236 regeneration genes as well as their 28,190 homologs across 58 representative plant species. It contains extensive transcriptomic resources across all regeneration models, together with 196,423 single cells and over 8.81 million epigenetic peaks to dissect cellular heterogeneity and multi-layered epigenetic regulation. Equipped with nine online omics-related tools and a RAG-based intelligent Q&A system, PlantRegMoD greatly reduces bioinformatic barriers and serves as a robust resource for mechanistic, functional and evolutionary studies of plant regeneration.

## Population-Scale Precision Safety in Oncology Reveals Clinical and Genetic Determinants of Systemic Therapy Toxicity
- Source: medRxiv (preprints)
- Date: 2026-09-17
- Categories: Tools & resources
- Authors: Bakouny, Z., Guo, X. A., Huang, F., Mohan, S., Walser, R., Lu, Z., Perea-Chamblee, T., Khan, L., Ahmed, N., Ocejo, A., Hakimi, A. A., Shah, N., Voss, M. H., Arbour, K. C., Cheung, Y.-M. M., Azhari, H., Faleck, D. M., Niec, R., Donoghue, M. T. A., Orgera, J. J., Syed, A., Berger, M. F., Waters, M., Pichotta, K., Fong, C., Jee, J., Schultz, N., Schrag, D., Yarmus, L., Schoenfeld, A. J., Gusev, A., Kotecha, R. R., Motzer, R. J., Thompson, C. B., Tansey, W., Carrot-Zhang, J., Reznik, E.
- DOI: 10.64898/2026.09.16.26363259
- Source URL: <https://doi.org/10.64898/2026.09.16.26363259>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.16.26363259>

Abstract: Treatment toxicity constrains the use of effective cancer therapies, but its clinical and genetic determinants remain poorly defined. We developed a large language model-based approach to produce MSK-Tox, a pan-cancer resource capturing the incidence, temporality, and grade of toxicity to anti-cancer therapy across more than 50,000 patients. Analysis of six representative adverse events - pneumonitis, adrenal insufficiency, liver toxicity, colitis, hyperthyroidism and hypothyroidism - revealed distinct toxicity landscapes shaped by cancer type, treatment regimen, and clinical context. Pretreatment clinical features enabled individualized prediction of toxicity risk across adverse events, supporting risk assessment before therapy initiation. Beyond clinical predictors, we identified two modes of germline susceptibility to treatment toxicity: an organ-intrinsic mode, in which germline variation confers risk across systemic therapies, exemplified by a regulatory variant near FOXE1 associated with hypothyroidism; and an immune-mediated mode, confined to immune checkpoint inhibitor-treated patients, in which HLA-DRB1\*15 was a major determinant of adrenal insufficiency. Notably, the same allele predisposes to multiple sclerosis in individuals without cancer, indicating that immune checkpoint inhibition unmasks a latent autoimmune predisposition. These findings provide an empirical basis for a new precision safety paradigm for predicting who will be harmed by a therapy on the same principles that guide prediction of therapeutic benefit.

## Post-Hoc Long-Read Sequencing Links Leukemic Mutation Status to Single-Cell Transcriptomes.
- Source: European journal of haematology (journals)
- Date: 2026-09-17
- Categories: Genomics & sequence analysis, Single-cell & spatial, Evolution & metagenomics
- Authors: Sofia Papavasileiou, Chenyan Wu, Daryl Boey, Lucille Margerie, Jiezhen Mo, Ulla Olsson-Strömberg, Stina Söderlund, Gunnar Nilsson, Joakim S Dahlin
- Journal: European journal of haematology
- DOI: 10.1111/ejh.70322
- External ID: 42752856
- Keywords: transcriptomes, rna, gene expression, genomics, single cell, genotyping
- Source URL: <https://doi.org/10.1111/ejh.70322>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1111%2Fejh.70322>

Abstract: Single-cell RNA-sequencing-based characterization of cells that belong to the neoplastic clone is a major challenge in hematologic neoplasms, where malignant and normal cells coexist. Confident molecular profiling requires simultaneous analysis of gene expression and genetic mutations in individual cells, an ability that is not supported by the standard 10X Genomics workflow. Here, we systematically evaluated the potential and limitations of repurposing amplified cDNA generated during the 10X Genomics 3' workflow for post hoc genotyping of individual cells. We first established a mixed leukemic cell line system comprising one cell line with KIT point mutations and another with the BCR::ABL1 fusion gene. Targeted long-read PacBio sequencing enabled post hoc assignment of mutation data to transcriptionally profiled cells, but recovery differed between targets. Consistent with ambient RNA in microfluidics-based single-cell workflows, mutation-associated transcripts were detected in cells not expected to carry the corresponding mutations, illustrating how transcript recovery complicates cell-level genotype assignment. Target-specific thresholds mitigated this source of misclassification. In primary chronic myeloid leukemia samples, the post hoc approach detected BCR::ABL1-positive cells at diagnosis, but not during imatinib treatment. Together, we present a framework for adding mutation status to cells already profiled using the 10X Genomics workflow and highlight broader considerations for transcript-based single-cell genotyping.

## Predictive control of human pancreatic cell fate using a digital model of in vitro differentiation
- Source: bioRxiv (preprints)
- Date: 2026-09-17
- Categories: Genomics & sequence analysis, Single-cell & spatial, Systems & networks, Tools & resources
- Authors: Sanchez-Castro, E. E., Ishahak, M., Le, T., Maestas, M. M., Hernandez-Rincon, D. C., Mukherjee, N., Bradley, K., Lu, J., Gale, S. E., Millman, J. R.
- DOI: 10.64898/2026.04.27.721124
- Source URL: <https://doi.org/10.64898/2026.04.27.721124>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.04.27.721124>

Abstract: The controlled generation of mature stem cell-derived islets (SC-islets) remains a barrier to scalable cell therapy for diabetes. Here, we develop a predictive digital model defining the cell-state-specific regulatory logic governing fate specification during human SC-islet differentiation. We integrate 400,603 cells from 9 original single-cell multiomic datasets and 52 public single-cell RNA-seq and ATAC-seq datasets across 4 cell lines and 7 differentiation protocols. This model resolves transcriptional and chromatin accessibility dynamics while enabling time-resolved inference and in silico perturbation of cell-state-specific gene regulatory networks. We identify regulators across trajectories from endoderm progenitors to pancreatic exocrine and endocrine lineages, nominating new candidate regulators. Among these candidates, we validate previously unreported roles for STAT1 as an exocrine driver and ZEB1 as a dynamic regulator of early endocrine specification and later off-target serotonergic islet cell fate. This work provides an experimentally supported predictive framework and an interactive resource comprising 1,116 simulations to prioritize transcription factors and intervention windows for refining SC-islet differentiation.

## pykarambola: Minkowski tensor morphometry of 3D structures
- Source: Bioinformatics Advances (journals)
- Date: 2026-09-17T00:00:00+00:00
- Categories: Biological imaging, Tools & resources
- Authors: Yajushi Khurana, Keisuke Ishihara
- Journal: Bioinformatics Advances
- DOI: 10.1093/bioadv/vbag273
- Source URL: <https://doi.org/10.1093/bioadv/vbag273>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1093%2Fbioadv%2Fvbag273>
- Code: <https://github.com/Ishihara-SynthMorph/pykarambola>

Abstract: Three-dimensional biological morphologies encode functional and physiological state, yet their directional, orientational, and topological properties are rarely captured by morphometric tools in bioimage analysis. Minkowski tensors encode surface curvature and directionality for arbitrary topologies; their eigensystems directly quantify elongation axes and anisotropy. A C ++ implementation, karambola, computes Minkowski tensors for triangulated surfaces but is inaccessible within Python-based bioimage workflows. We present pykarambola, a Python package that accepts NumPy arrays and standard mesh formats and returns Minkowski tensors, including derived anisotropy and orientation quantities. A high-level label-image interface converts three-dimensional integer arrays into per-object Minkowski tensors in a single call, making pykarambola directly compatible with the output of segmentation tools. An optional Cython extension accelerates graph-traversal steps of mesh initialization for large-scale analyses. Validated on synthetic meshes spanning three topologies and benchmarked on 1,584 adrenal gland meshes, pykarambola reproduces all 121 karambola features to near-floating-point agreement and is 2.8-fold faster, with speedups primarily attributable to per-object file input/output. pykarambola is freely available as an open-source software package. Availability and implementation: pykarambola is distributed under the BSD 3-Clause License and is available on GitHub at https://github.com/Ishihara-SynthMorph/pykarambola. It can be installed via pip.

## RBApy: Extending resource allocation modeling to eukaryotes in complex environments
- Source: Bioinformatics Advances (journals)
- Date: 2026-09-17T00:00:00+00:00
- Categories: Systems & networks, Tools & resources
- Authors: Oliver Bodeit, Nadia Bessoltane, Delphine Charif, Anaghim Temtem, Olivier Inizan, Anne Goelzer
- Journal: Bioinformatics Advances
- DOI: 10.1093/bioadv/vbag276
- Source URL: <https://doi.org/10.1093/bioadv/vbag276>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1093%2Fbioadv%2Fvbag276>
- Code: <https://github.com/RBAgroup>

Abstract: Motivation Resource allocation modeling—as the Resource Balance Analysis (RBA) framework— provides a way to understand and predict how limited cellular resources (e.g., energy, proteins, etc.) in cells are distributed among competing cell processes within a limited cellular space. Currently, resource allocation modeling for any eukaryotes remains limited due to the lack of software capable of generating calibrated RBA models for these types of cells, unlike prokaryotes, which benefit from the software tools RBApy, RBAtools and the RBAml format for model encoding. Results Here we extended the RBA toolkit (RBApy, RBAtools and RBAml) to account for specific aspects of eukaryotic cells growing in complex environments such as varying temperature, light or nutritional conditions. We used them to generate and simulate RBA models of both prokaryotic (Escherichia coli) and eukaryotic (Arabidopsis thaliana) cells for varying temperatures. The resulting models show excellent prediction capabilities when benchmarked against published experimental datasets. The upgraded RBA toolkit will pave the way to creating, calibrating and running resource allocation models for crops, livestock or humans for a wide range of medical, biotechnological or agricultural applications in the future. Availability and implementation RBApy and RBAtools are available via PyPI, and at https://github.com/RBAgroup.

## Redescription of two Allocreadium species and molecular dating of the family Allocreadiidae (Trematoda: Gorgoderoidea)
- Source: Journal of Helminthology (journals)
- Date: 2026-09-17T00:00:00Z
- Categories: Evolution & metagenomics
- Authors: K. S. Vainutis, A. Zhokhov, M. Urabe, A. Aydogdu
- Journal: Journal of Helminthology
- DOI: 10.1017/S0022149X26102156
- External ID: a9bf99549ed647e7ec8d0a12aaafbe320ff15d61
- Keywords: phylogenetic
- Source URL: <https://doi.org/10.1017/S0022149X26102156>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1017%2FS0022149X26102156>

Abstract: The family Allocreadiidae comprises diverse parasites of freshwater fishes, but their evolutionary timescale remains poorly understood. We provide a molecular dating analysis based on an expanded 28S rRNA dataset and redescriptions of two East Asian species, Allocreadium hasu and A. pseudaspii, with new morphological and genetic data. Allocreadium hasu from Lake Biwa (Japan) is genetically close to Far Eastern A. khankaiense and A. anastasii (0.09–0.97% divergence in 28S) but has not been recorded in the Russian Far East. Allocreadium pseudaspii, first recorded in the Bolshaya Ussurka River, possesses eyespot remnants and occupies a unique position among European species, suggesting secondary eastward dispersal. Phylogenetic analyses confirm family monophyly and resolve relationships among Palaearctic genera. Divergence time estimates indicate the most recent common ancestor of Allocreadiidae originated in East Asia during the Lower Cretaceous (~110 Ma). Diversification of major genera coincides with radiation of primary fish hosts: Allocreadium with Cypriniformes (~93 Ma), Bunodera with Perciformes (~75 Ma), and Crepidostomum s. str. with Nemacheilidae (~49 Ma). The basal genus Acrolichanus, parasitic on sturgeons, diverged earlier (40–80 Ma). The split between the Neotropical genus Margotrema and its Palaearctic relatives is estimated at ~6.5 Ma, consistent with closure of the Panama Isthmus. Our results provide a temporal framework for Allocreadiidae evolution, linking diversification to biogeographic history of freshwater fish hosts in East Asia and subsequent dispersal to North America via the Bering land bridge and to South America through the Panama Isthmus.

## Resolving Allopolyploid Origins Within the Genus Clarkia Using a Novel Read-Mapping and Modeling Approach
- Source: bioRxiv (preprints)
- Date: 2026-09-17
- Categories: Genomics & sequence analysis, Evolution & metagenomics
- Authors: Stanton, K., Rausher, M. D.
- DOI: 10.64898/2026.09.13.751269
- Source URL: <https://doi.org/10.64898/2026.09.13.751269>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.13.751269>

Abstract: Whole genome duplications are a common occurrence in plants, but this creates challenges for reconstructing the evolutionary history between species, especially when polyploidy is a result of hybridization. While multiple methods have been developed to try to tackle these issues, most are computationally intensive, restrictive on the number of taxa that can be evaluated, and benefit immensely from a priori hypotheses about the allopolyploid progenitors, rendering these methods unfeasible for many understudied polyploids. We present a rapid, low-cost, and computationally light method for determining the relative time of hybridization as well as the most likely progenitor species of a given allopolyploid species, including progenitors that are extinct, ancestral, or unknown. The method utilizes a combined approach of first mapping sequencing reads from the polyploid against a diploid pantranscriptome to generate hypotheses about possible progenitor pairs and then modeling various hybridization scenarios to estimate the likelihood of each hypothesis. We demonstrate the utility of our methods by identifying likely progenitors and times of origin for six allotetraploid species from the genus Clarkia. While the methods outlined here do not conclusively confirm the origins of these allopolyploids, they provide well-supported working hypotheses for further intensive exploration.

## RNAbridge: a database of extended and non-canonical helices
- Source: bioRxiv (preprints)
- Date: 2026-09-17
- Categories: Proteins & structural biology, Tools & resources
- Authors: Zakrzewski, D., Antczak, M., Zok, T.
- DOI: 10.64898/2026.09.15.751775
- Source URL: <https://doi.org/10.64898/2026.09.15.751775>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.15.751775>

Abstract: RNA function is dictated by 3D architecture. Although 2D structural models based on canonical Watson-Crick base pairs are widely used, they often fail to capture the non-canonical interactions, tertiary contacts, and coaxial stacking important for biological activity. We have developed RNAbridge, a comprehensive database and web application that systematically identifies, quantifies, and visualizes extended non-canonical helices and multi-way junctions. Using a geometry- and stacking-based pipeline, we analyzed the Protein Data Bank and compiled a catalog of 135,541 structural motifs. RNAbridge includes a user-friendly interface with interactive filters, synchronized 2D and 3D visualizations, and detailed structural data, including helical bend angles and stacking paths. By connecting simplified 2D topologies with complex 3D structures, RNAbridge serves as a valuable resource for structural biologists and lays the foundation for future machine learning applications in RNA structure prediction. The database, available at https://rnabridge.cs.put.poznan.pl/, is automatically updated once a week.

## Seasonal influenza vaccine strain selection by quantifying viral fitness
- Source: bioRxiv (preprints)
- Date: 2026-09-17
- Categories: Evolution & metagenomics, Tools & resources
- Authors: Qin, L., Zhang, M., Liu, X., Ding, X., Li, Q., Yang, L., Qi, Y., Liu, J., Zhou, H., Li, Z., Xie, W., Li, Z., Ma, Y., Yang, J., Wang, H., Wang, J., Jiang, T., Wang, D., Wang, Y., Wu, A.
- DOI: 10.64898/2026.09.15.751745
- Source URL: <https://doi.org/10.64898/2026.09.15.751745>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.15.751745>

Abstract: The long-standing challenge of seasonal influenza vaccines providing poor protection is attributed to the rapid and continuous evolution of the virus. In theory, the recommendation of vaccine strain is to predict the dominant strain with best fitness in the upcoming season. Here, we develop FutureFlu, a biologically-grounded framework that quantifies the fitness score of influenza variants in upcoming seasons, by integrating metrics on three levels: molecular genetic divergence, individual immune escape, and population-scale transmission dynamics. The fitness scores of influenza variants present strong positive correlation with their actual observed frequencies in the next seasons. Validation across 24 seasons of three influenza subtypes shows FutureFlu recommends antigenically matched vaccine strains more frequently than annual recommendations, particularly for challenging subtypes: 83.3% versus 45.8% seasons for A/H3N2, and 75.0% versus 33.3% seasons for B/Victoria. Furthermore, FutureFlu significantly outperforms the currently used methods in vaccine strain selection whether or not there is available antigenic data from hemagglutination inhibition (HI) assay, which provides a valuable supplement for WHO vaccine recommendation. To support global public health implementation, an open online platform (futureflu.com.cn) has been established to offer real-time viral fitness predictions and vaccine recommendations.

## Seasonal Light and Temperature Timing in a Stoichiometric Food Web
- Source: bioRxiv (preprints)
- Date: 2026-09-17
- Categories: Evolution & metagenomics, Mathematical biology & statistics
- Authors: Ramirez, R.
- DOI: 10.64898/2026.09.13.751257
- Source URL: <https://doi.org/10.64898/2026.09.13.751257>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.13.751257>

Abstract: Seasonal food-web interactions can depend on whether consumer performance is high when food quantity and elemental quality are favorable. We extend a closed-phosphorus model containing pelagic and benthic producers, variable producer phosphorus quotas, and a shared Daphnia grazer by allowing annual light and temperature cycles to have an adjustable phase difference. The previous light-seasonality preprint has a nonisolated grazer-free boundary. We parameterize its annual periodic extension by the fraction of producer phosphorus in phytoplankton and derive the unique positive annual producer orbit for every fixed allocation. Linearization in the rare-grazer direction then gives an exact conditional Floquet exponent. Its decomposition into a mean-rate term and a covariance term identifies the effect of seasonal timing. A reconstructed descriptive thermal proxy uses quasi-acclimated filtration-capacity means from the official Muller et al. dataset; it supplies only a relative response shape over 15-25 degrees C, not an absolute ingestion calibration. In a representative configuration, changing phase while preserving the annual light and temperature distributions changes the invasion exponent from -0.00382 to 0.01486 day^-1, with annual multipliers 0.248 and 227, respectively. The constant-mean-ingestion exponent is positive, so the negative case is generated by adverse timing covariance. Sign reversal persists across a range of phosphorus allocations, but not across the entire boundary family. The result is a local invasion criterion for specified grazer-free cycles, not a theorem of global persistence or extinction. Within this parameterized boundary problem, relative seasonal timing can change the sign of infinitesimal consumer growth.

## Selection bias in Mendelian randomization studies with adjustment for medication use
- Source: medRxiv (preprints)
- Date: 2026-09-17
- Authors: Shi, J., Swanson, S. A., Diemer, E. W., Gerlovin, H., Posner, D. C., Wilson, P. W. F., Gaziano, J. M., Cho, K., Hernan, M. A., on behalf of the VA Million Veteran Program,
- DOI: 10.64898/2026.09.16.26363225
- Source URL: <https://doi.org/10.64898/2026.09.16.26363225>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.16.26363225>

Abstract: Background. Mendelian randomization (MR) studies often evaluate exposures such as LDL cholesterol (LDL-C). The widespread use of lipid lowering medications complicates the interpretation and the validity of MR estimates. Methods. We describe two causal estimands in populations with medication use: a lifetime effect, and a lifetime effect under no medication use. In simulations, we compared the common approach of excluding medication users with estimation based on inverse probability (IP) weighting to adjust for medication use. We applied both approaches to a MR analysis of LDL-C and coronary artery disease in the Million Veteran Program (MVP), a large prospective cohort of U.S. veterans with linked electronic health record and genetic data. Results. In simulations, MR analyses that did not adjust for medication use estimated a lifetime effect that reflected a valid estimate of a total effect that included both the harms of higher LDL-C and the benefits of statins. To estimate the effect under no medication use, excluding statin users introduced selection bias. Alternatively, IP weighting could address bias from incident statin users, but could not address bias related to prevalent medication use. Estimates from MVP data varied considerably, reflecting the importance of these analytic choices. Conclusions. Different medication adjustment strategies in MR studies implicitly target different causal estimands and are subject to distinct biases. Transparent analytical choices and careful interpretation are essential for informative MR results in the context of widespread medication use.

## Self-organized Regulation of Group Size and Number in Natural and Artificial Collectives
- Source: bioRxiv (preprints)
- Date: 2026-09-17
- Authors: Zhang, T., Lee, S., Hamann, H.
- DOI: 10.64898/2026.08.25.746978
- Source URL: <https://doi.org/10.64898/2026.08.25.746978>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.08.25.746978>

Abstract: From animal societies to self-organizing multi-agent systems, collectives adapt their group structure to tasks and environments. However, how they determine appropriate group sizes and the number of subgroups to form remains unclear. We formulate the Group Size and Number Regulation Problem (GSNRP), which asks how individuals regulate group sizes and numbers using only local information. In a first step, we establish a graph-theoretic model demonstrating that simple following behavior suffices to form group structures that match theoretical expectations, but is insufficient for active regulation of group size and number. In a second step, we operationalize individual group-size preferences in a decentralized fission-fusion mechanism based on perceived group size. Through multi-agent simulations, we validate that this mechanism achieves stable convergence across three signaling regimes, from position-only sensing to continuous group-size communication. Using tracking data from wild white-nosed coatis (mammals in the raccoon family), we calibrate individual group-size preferences and show that the controller recovers selected group-size, subgroup-count, and transition statistics. This in-sample case study demonstrates descriptive consistency with natural fission-fusion dynamics without establishing the underlying behavioral mechanism. These results suggest that natural and engineered collectives may share local principles of perception, preference, and response for regulating group structure.

## Shallow recurrent decoders for neural and behavioural dynamics.
- Source: Philosophical transactions of the Royal Society of London. Series B, Biological sciences (journals)
- Date: 2026-09-17
- Categories: Computational neuroscience
- Authors: Amy Rude, J Nathan Kutz
- Journal: Philosophical transactions of the Royal Society of London. Series B, Biological sciences
- DOI: 10.1098/rstb.2024.0461
- External ID: 42750452
- Source URL: <https://doi.org/10.1098/rstb.2024.0461>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1098%2Frstb.2024.0461>

Abstract: Machine learning algorithms are affording new opportunities for building bio-inspired and data-driven models characterizing neural activity. Critical to understanding decision-making and behaviour is quantifying the relationship between the activity of neuronal population codes and individual neurons. We leverage a SHallow REcurrent Decoder (SHRED) architecture for mapping the dynamics of population codes to individual neurons and other proxy measures of neural activity and behaviour. SHRED is constructed from a temporal sequence model, which encodes the temporal dynamics of limited sensor data in multiple scenarios, and a shallow decoder, which reconstructs the corresponding high-dimensional neuronal and/or behavioural states. It is a robust and flexible sensing strategy which allows for decoding the diversity of neural measurements with only a few sensor measurements. Thus, estimates of whole-brain activity, behaviour and individual neurons can be constructed with only a few neural time-series recordings. Several examples in this article further highlight the potential of leveraging non-invasive or minimally invasive measurements to estimate large-scale brain dynamics. We empirically demonstrate the capabilities of the method on a number of model organisms including Caenorhabditis elegans, mouse, zebrafish and human biolocomotion. This article is part of the discussion meeting issue 'Digital healthcare for the management of functional neurological disorders'.

## Simple a posteriori insertion of fossil tips in molecular phylogenies can improve inferences of continuous trait evolution.
- Source: Evolution; international journal of organic evolution (journals)
- Date: 2026-09-17T00:00:00Z
- Categories: Evolution & metagenomics
- Authors: Lindsey M. DeHaan, Graham J. Slater, M. Friedman
- Journal: Evolution; international journal of organic evolution
- DOI: 10.1093/evolut/qpag166
- External ID: d232e03951ce798336546c7b1b692424687f5dbb
- Source URL: <https://doi.org/10.1093/evolut/qpag166>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1093%2Fevolut%2Fqpag166>

Abstract: Integrating neontological and paleontological information improves inferences about the tempo and mode of phenotypic evolution over long timescales. However, outside a few exemplar clades, morphological matrices establishing formal phylogenetic placements of fossils within molecular phylogenies are often lacking, limiting the inclusion of fossil data in macroevolutionary inferences. Informal fossil placements are commonly established through morphological assessment (e.g., as is often the case for node-age calibrations) but determining branch length durations is less straightforward. Using simulations, we compared four contrasting methods of a posteriori fossil insertion and assess their impacts on estimating the tempo and mode of continuous trait evolution. We find that arbitrarily assigning branch lengths to fossil taxa or analytically estimating them using continuous trait data had negligible biases on phenotypic inferences relative to true branch lengths. However, use of minimum fossil branch lengths biased model selection toward an Ornstein-Uhlenbeck process and inflated rate estimates. We find that the inclusion of fossil taxa using our preferred approaches improves support for an early burst when it is the generating model. Applying these methods to a comparative dataset of carangarian fishes (flatfishes, jacks, barracudas, and billfishes) reveals high rates of phenotypic evolution early in the clade's history, a signal forecasted by the fossil record but not shown in extant-only analyses. We conclude that a posteriori insertion of fossils, even with designated rather than inferred branch lengths, can strengthen phenotypic inferences relative to analyses including only living taxa.

## Size Control of hnRNPK-based Nucleolar Condensates by RNA-Regulated Fusion Dynamics
- Source: bioRxiv (preprints)
- Date: 2026-09-17
- Categories: Proteins & structural biology
- Authors: Tejedor, A. R., Luengo, J., Llombart, P., Pedraza, E., Otero-Sobrino, A., Velasco-Estevez, M., Gallardo, M., Ocana, A., Collepardo-Guevara, R., Espinosa, J. R.
- DOI: 10.64898/2026.09.14.751346
- Source URL: <https://doi.org/10.64898/2026.09.14.751346>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.14.751346>

Abstract: Nucleoli are liquid-like condensates whose size is actively conserved--they remain small, numerous, and resistant to coalescence--yet the molecular mechanisms that constrain their fusion remain poorly understood. Heterogeneous nuclear ribonucleoprotein K (hnRNPK), an RNA-binding protein implicated in nucleolar organization and cancer, interacts directly with the scaffold protein Nucleolin. Combining residue-resolution coarse-grained simulations with biochemical experiments, we find that hnRNPK and Nucleolin condense through distinct interaction networks--a localized cation-\{pi\}/electrostatic hotspot in hnRNPK versus broadly distributed electrostatic contacts in Nucleolin, reorganized upon co-assembly. To probe how RNAs reshape these condensates, we develop and validate, against re-entrant phase-separation experiments and AlphaLISA binding data, a nucleotide-resolution coarse-grained model for single-stranded RNA. Using this framework, we show that RNA is asymmetrically and preferentially recruited by hnRNPK over Nucleolin, an asymmetry that grows stronger when the two proteins compete for the same RNAs. This selective recruitment sustains a dynamic fission-fusion equilibrium: hnRNPK-containing condensates repeatedly fuse and split rather than coalescing into a single condensate, whereas Nucleolin-containing and ternary condensates fuse into one dominant cluster. These results reveal a molecular mechanism, grounded in sequence-encoded, RNA-controlled fusion dynamics, by which nucleolar condensates conserve a controlled, non-coalescing size despite their liquid-like character.

## SmartHisto: Bayesian active learning for histology images
- Source: PLOS Computational Biology (journals)
- Date: 2026-09-17T00:00:00+00:00
- Categories: Biological imaging
- Authors: Sriram Vijendran, Bailey Arruda, Tavis K. Anderson, Oliver Eulenstein
- Journal: PLOS Computational Biology
- DOI: 10.1371/journal.pcbi.1013611
- Source URL: <https://doi.org/10.1371/journal.pcbi.1013611>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1371%2Fjournal.pcbi.1013611>

Abstract: Accurate and efficient characterization of biological images is crucial for advancing systems biology and medical research. Recent advancements in deep learning and image processing have enabled neural network models to rapidly accelerate image analysis by utilizing large expert-annotated datasets. However, in histopathology, the size of whole-slide images makes expert annotation expensive, limiting the acquisition of sufficiently large annotated datasets and posing a major challenge for developing automated, AI-driven image analysis pipelines. To address this limitation, we propose a novel active learning-based framework to train image segmentation models interactively. Our approach employs a Bayesian neural network to identify informative regions in unlabeled images rather than entire images, making expert labeling more cost-effective. We validate our framework on multiple benchmark datasets with variable staining at fixed magnifications, demonstrating substantial reductions in annotation requirements. Notably, our method achieves a mean IoU of 0.75, significantly outperforming competing approaches, which averaged 0.60.

## Spatiotemporal single-cell profiling reveals T cell clonal dynamics and phenotypic plasticity in human graft-versus-host disease.
- Source: Nature immunology (journals)
- Date: 2026-09-17
- Categories: Genomics & sequence analysis, Single-cell & spatial
- Authors: Lingting Shi, Ajna Uzuni, Ximi K Wang, Michael Pressler, David W Harle, Shami Chakrabarti, Rodney Macedo, Kirubel Belay, Christian A Gordillo, Thomas McMahon-Skates, Erik Raps, Jia Yi Ady Zhang, Achille Nazaret, Joy L Fan, Yinuo Jin, Xumin Shen, Joshua S Fuller, Tamjeed Azad, Jessie Huang, Pranik Chainani, Jose Pomarino Nima, Julian A Abrams, Armando Del Portillo, Markus Y Mapara, Mohamed Alhamar, Megan Sykes, José L McFaline-Figueroa, Elham Azizi, Ran Reshef
- Journal: Nature immunology
- DOI: 10.1038/s41590-026-02631-2
- External ID: 42754741
- Source URL: <https://doi.org/10.1038/s41590-026-02631-2>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1038%2Fs41590-026-02631-2>

Abstract: Allogeneic hematopoietic cell transplantation cures hematologic diseases but is limited by acute graft‑versus‑host disease. How human T cell clones drive epithelial injury remains poorly mapped. We studied 31 transplant recipients, integrating longitudinal T cell antigen receptor (TCR) profiling with single-cell RNA sequencing/TCR sequencing and spatial transcriptomics to track T cell clonal dynamics. We developed DecompTCR to resolve temporal dynamics and adapted computational tools to map clone phenotypes and niches in tissue. Our analyses revealed that cyclophosphamide selectively depletes alloreactive clones, although insufficient early expansion leads to incomplete depletion and severe disease. Severe graft‑versus‑host disease is marked by persistent expansion of alloreactive clones, rewiring of homeostatic cell types and diversification of donor-derived CD8+ clonotypes that acquire Hobit (ZNF683)+ tissue‑resident memory T (TRM) cell programs during migration to epithelium. Spatial deconvolution identified CD8+ effector/Hobit+ TRM hubs near intestinal stem‑cell-rich crypt bases and crypt‑loss regions. This clonotype‑resolved framework links tissue‑instructed TRM cell remodeling to localized epithelial injury, nominating early-repertoire dynamics and spatial hub burden as biomarkers.

## Spinal Recurrent Inhibition Shapes the Dynamics of TMS-induced Motor-Evoked Potentials: A Computational Modeling Study
- Source: bioRxiv (preprints)
- Date: 2026-09-17
- Categories: Systems & networks, Computational neuroscience
- Authors: Chien, V. S. C., Bernasconi, E., Müller, E., Wang, P., Lowery, M., Liegey, J., Wendt, K., O'Shea, J., Denison, T., Hlinka, J., Knösche, T. R., Weise, K., Schmidt, H.
- DOI: 10.64898/2026.09.11.750929
- Source URL: <https://doi.org/10.64898/2026.09.11.750929>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.11.750929>

Abstract: Motor-evoked potentials (MEPs) recorded via surface electromyography (EMG) from peripheral muscles following transcranial magnetic stimulation (TMS) of the motor cortex reflect the integrity of the entire corticospinal pathway and are widely used in both basic neuroscience and clinical practice. However, the relative contributions of spinal and peripheral mechanisms to the observed MEP waveform remain poorly understood, partly because computational models that capture individual MEP characteristics are lacking. Here, we present a biologically plausible and computationally efficient model of the descending motor pathway, spanning the spinal cord and hand muscles, that can be fitted to individual MEP waveforms across a range of TMS intensities. The model successfully reproduces individual MEP waveforms, accounting for approximately 90% of the observed variance in waveforms across 10 healthy participants. Crucially, we demonstrate that recurrent inhibition of Renshaw cells in the spinal cord is indispensable for reproducing the fine temporal structure of MEP waveforms, even when input-output curve fitting appears adequate without it. Beyond waveform reproduction, the fitted model provides interpretable estimates of latent neural dynamics and subject-specific pathway parameters, including motor neuron size distribution, synaptic receptor balance, axonal conduction delay, and hand muscle refractoriness, that are consistent with known biological ranges. These results suggest that individual MEP waveforms, when analyzed using a biologically grounded model, carry substantially more information about spinal and peripheral motor pathway integrity than conventional amplitude-based measures alone.

## Standing genetic variation drives polygenic adaptation to different environmental shifts
- Source: bioRxiv (preprints)
- Date: 2026-09-17
- Categories: Evolution & metagenomics
- Authors: Zhang, Y., Stetter, M. G.
- DOI: 10.64898/2026.09.14.751342
- Source URL: <https://doi.org/10.64898/2026.09.14.751342>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.14.751342>

Abstract: Most organisms are well adapted to the environment they have been exposed to for many generations. However, changing environments pose existential threats to populations, unless they are able to adapt. Hence, a central challenge in evolutionary genetics is to understand the adaptation to complex environmental changes. As most traits are controlled by a large number of loci with varying effects on a trait, it is challenging to understand the explicit role of genetic changes during such polygenic adaptation. We use forward-in-time simulations to investigate the evolutionary dynamics of phenotypic and genetic adaptation in populations facing different environmental shifts. Specifically, we simulated sudden, gradual, and fluctuating environmental shifts for multiple trait architectures. Our results show distinct evolutionary paths across different types of environmental shifts and a higher extinction risk during sudden and fluctuating environmental shifts than during gradual shifts. Comparing the contribution of mutations from different sources highlights the critical role of standing genetic variation in driving phenotypic adaptation. We summarize allele frequency trajectories by clustering them by their temporal pattern and reveal how large- and small-effect mutations jointly shape the successful adaptation to changing environments. Additionally, we trained a convolutional neural network (CNN) on "genetic architecture matrices" of populations to jointly infer the type and magnitude of environmental changes and the mutational effect size distribution. The CNN was able to predict all three input parameters with very high accuracy, even on unseen parameter combinations. Our results demonstrate the impact of ecological change on the evolutionary outcome and the assorted mutational changes that enable successful adaptation.

## STCGCar: Graph Contrastive Learning with Reliable Augmentation for Spatial Transcriptomics Clustering.
- Source: Genomics, proteomics & bioinformatics (journals)
- Date: 2026-09-17T00:00:00Z
- Categories: Genomics & sequence analysis, Single-cell & spatial, Tools & resources
- Authors: Li-Hong Peng, Long Yang, Min Chen, Geng Tian, Xin Liu, Zongzheng Bai, Jialiang Yang
- Journal: Genomics, proteomics & bioinformatics
- DOI: 10.1093/gpbjnl/qzag098
- External ID: e1e93566642e1943c92b1f7ee9f80d23ac2e1da5
- Source URL: <https://doi.org/10.1093/gpbjnl/qzag098>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1093%2Fgpbjnl%2Fqzag098>
- Code: <https://github.com/plhhnu/STCGCar>

Abstract: Accurately identifying spatial domains based on spatial transcriptomics (ST) data can greatly promote our understanding of cellular composition and tissue organization. While graph neural networks (GNNs) have shown significant advancements in spatial clustering, they tend to be insensitive to noisy edges, leading to intersections among identified spatial domains. Here, we introduce a GNN-based ST Clustering framework, called STCGCar, utilizing a Graph Contrastive learning model with reliable augmentation and redundancy reduction strategies. The framework begins by creating an enhanced view through a reversible network after data preprocessing. Subsequently, low-dimensional embeddings of spots are learned using a multi-head attention mechanism. Moreover, a redundancy reduction strategy is employed to reduce information redundancy in potential feature space. Finally, spatial domains are delineated through K-means clustering, followed by downstream analysis. STCGCar was benchmarked against six state-of-the-art clustering methods (i.e., Seurat, conST, CCST, STAGATE, DeepST, and GraphST) using five 10x Visium datasets, a STARmap dataset, and two Stereo-seq mouse embryo datasets. Through evaluation with adjusted rand index (ARI), normalized mutual information (NMI), and four internal indicators, it demonstrated outstanding clustering performance compared to other methods on four labeled and four unlabeled datasets. Additionally, STCGCar accurately identified spatial domains and discovered three potential differentially expressed genes (AZGP1, CD24, and CCND1) in human breast cancer tissues. Furthermore, it effectively delineated layer structures in human DLPFC and adult mouse brain tissues. STCGCar is a powerful tool for spatial domain identification, showcasing its effectiveness and scalability on diverse datasets. It is freely available at https://github.com/plhhnu/STCGCar.

## Stochastic delay derivatives of Newcastle disease application in epidemic model: Stability analysis and approximation
- Source: PLOS One (journals)
- Date: 2026-09-17T00:00:00+00:00
- Categories: Mathematical biology & statistics
- Authors: Naveed Shahid, Ali Raza, Marek Lampart, Sana Iqbal, Nauman Ahmed, Eman Ghareeb Rezk
- Journal: PLOS One
- DOI: 10.1371/journal.pone.0357918
- Source URL: <https://doi.org/10.1371/journal.pone.0357918>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1371%2Fjournal.pone.0357918>

Abstract: The illicit trade in wildlife for pet purposes poses a direct risk to animal populations through overharvesting, but it also serves as an indirect conduit for the spread of contagious diseases. The present study assessed the effects of the hypothetical release of captured, infected individuals of Newcastle disease into the natural population of the white-winged parakeet, the most trafficked psittacine species in Peru. This study analysed the computational dynamical analysis of the stochastic susceptible-exposed-infected-recovered model of Newcastle disease. We take two approaches to stochastic modelling: transition probabilities and the perturbation method. To investigate dynamical features such as positivity, boundedness, consistency, and stability, we introduced a stochastic non-standard finite-difference (SNSFD) scheme. Conventional numerical approaches such as Euler-Maruyama, stochastic Euler, and stochastic Runge–Kutta of order four were applied, but they failed to preserve the system’s essential dynamical properties. Consequently, we formulated the SNSFD method to address this limitation. To support the proposed method, we presented several theorems that demonstrate it satisfies all dynamical properties of the model. In the end, we presented several simulations to compare the proposed method’s efficiency to that of an existing method.

## TAPPR PCR Assay Design – Targeted, Automated, Primer and Probe Retrieval for Scalable Molecular Assay Design
- Source: Bioinformatics Advances (journals)
- Date: 2026-09-17T00:00:00+00:00
- Categories: Genomics & sequence analysis, Tools & resources
- Authors: Phillip E Davis, Colin W Price, Taylor Otwell, Anshika Kapoor, Vita Domnenko, Joseph A Russell
- Journal: Bioinformatics Advances
- DOI: 10.1093/bioadv/vbag274
- Source URL: <https://doi.org/10.1093/bioadv/vbag274>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1093%2Fbioadv%2Fvbag274>
- Code: <https://github.com/mriglobal/tappr>

Abstract: Motivation The rapid and reliable detection of infectious disease agents is critical for effective biosurveillance, diagnostics, and outbreak response. However, existing molecular assay design methods face significant limitations in scalability, speed, and adaptability to rapidly evolving pathogens, often leading to outdated or suboptimal assays. Here, we present Targeted, Automated Primer and Probe Retrieval (TAPPR), a novel, automated pipeline for scalable molecular assay design. TAPPR employs alignment-free methods to identify conserved regions and marker sequences across large-scale genomic datasets, supporting customizable design parameters for inclusivity, exclusivity, and assay specificity. To evaluate TAPPR's performance, assays were designed for diverse microbial targets, including Mpox, Mycobacterium tuberculosis, SARS-CoV-2, and Candida albicans, representing viral, bacterial, and fungal pathogens. The pipeline incorporates k-mer set operations and clustering strategies to address sequence diversity and streamline conserved region identification. Designed assays underwent in silico PCR simulations and laboratory testing to assess specificity, sensitivity, and exclusivity. Results demonstrated high accuracy across targets, with superior sensitivity to existing fielded assays where available. Additionally, we compare TAPPR to alternative available molecular assay design tools to demonstrate its advantages. This work highlights TAPPR’s capability to accelerate the development of molecular diagnostics by efficiently leveraging vast genomic datasets and addressing computational bottlenecks. TAPPR represents a scalable, adaptable tool for rapidly designing high-quality molecular assays, positioning itself as a critical asset for biosurveillance and public health response to emerging and re-emerging infectious disease threats. Results TAPPR demonstrates rapid and scalable data-driven molecular assay design through alignment-free estimations of conserved regions and marker regions. Through both in silico and lab bench evaluation, TAPPR assays are demonstrated to perform equivalently or better than previously utilized publicly available qPCR assays for emergent disease diagnostics. TAPPR is also shown to produce results where other available automated molecular assays design solutions fail to do so on the order of hours. Availability and implementation TAPPR is available at https://github.com/mriglobal/tappr under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International Public License.

## Target-driven optimization of feature representation and model selection for microbiome sequencing data with ritme
- Source: bioRxiv (preprints)
- Date: 2026-09-17
- Categories: Tools & resources
- Authors: Adamov, A., Mueller, C. L., Bokulich, N.
- DOI: 10.64898/2025.12.08.693045
- Source URL: <https://doi.org/10.64898/2025.12.08.693045>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2025.12.08.693045>

Abstract: Microbiome sequencing datasets are sparse, high-dimensional, compositional, and hierarchically structured, and predictive modeling from them typically relies on ad hoc feature representation choices that obscure their impact on performance and interpretation. We present ritme, an open-source Python package that jointly optimizes microbiome-specific feature representation and model selection - combined algorithm selection and hyperparameter optimization - tailored to these data. ritme systematically searches taxonomic aggregation, sparsity-aware selection, compositional transforms, and metadata enrichment together with model class and hyperparameters, using state-of-the-art optimizers that scale from a laptop to a compute cluster. Across three real-world use cases, ritme outperformed the original study pipelines by 7-29% on the primary task metric and surpassed three AutoML baselines in six of seven comparisons, while selecting substantially fewer features and exposing how feature and model choices drive performance. Open-source and modular, ritme supports reproducible, parsimonious predictive modeling, downstream biological investigation, and extension to other multi-omics modalities.

## The automated eukaryotic pangenome pipeline EukPan reveals accessory genome differentiation beyond core-gene phylogeny in Aspergillus oryzae
- Source: bioRxiv (preprints)
- Date: 2026-09-17
- Categories: Genomics & sequence analysis, Tools & resources
- Authors: Seki, K., Goto, M., Futagami, T., Nagano, Y.
- DOI: 10.64898/2026.09.13.751290
- Source URL: <https://doi.org/10.64898/2026.09.13.751290>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.13.751290>

Abstract: Pangenome analysis reveals recurrent gene-content variation beyond a single reference genome, but its application to eukaryotes is constrained by inconsistent gene annotation. ANNEVO predicts gene models from genome FASTA assemblies without RNA-seq data. We developed EukPan, an automated post-annotation pipeline that standardizes GFF/GTF files, selects representative isoforms, constructs proteomes, infers orthogroups, builds a concatenated single-copy core-protein alignment, and summarizes shared accessory orthogroups while excluding orthogroups detected in only one genome. Applied with ANNEVO to 123 Aspergillus oryzae genomes, EukPan identified 11,245 core and 4,407 shared accessory orthogroups. The core-protein phylogeny broadly recovered the reported A-H classification, whereas accessory-genome analyses clearly separated the 33 group-A strains from the other 90 strains. Directional analysis identified 62 group-A-associated and 158 group-A-depleted orthogroups, with major facilitator superfamily (MFS) transporter and fungal Zn2Cys6 transcription-factor domains prominent in the depleted set. Among 93 orthogroups present in all non-A strains and absent from all group-A strains, 59 mapped to 10 segments of RIB40, the standard A. oryzae reference genome and a non-A (group-F) strain. EukPan therefore enables reproducible, coordinated core- and accessory-pangenome analysis from eukaryotic genome assemblies.

## The CAHRA Challenge: A Community-Wide Assessment of Cryo-EM Heterogeneous Reconstruction Algorithms
- Source: bioRxiv (preprints)
- Date: 2026-09-17
- Categories: Proteins & structural biology, Biological imaging, Tools & resources
- Authors: Feathers, J. R., Heeter, R., Woollard, G., Hanson, S. M., Cossio, P., Greer, J., Burnley, T., Zhong, E.
- DOI: 10.64898/2026.09.15.751515
- Source URL: <https://doi.org/10.64898/2026.09.15.751515>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.15.751515>

Abstract: The ability of cryo-electron microscopy (cryo-EM) to interrogate the atomic structure and motion of biomolecules has motivated the development of a wide range of algorithms for heterogeneity analysis. However, evaluating and comparing these methods remains challenging because ground-truth structures are generally unknown for experimental samples. Here, we introduce the 2026 Community-Wide Assessment of Heterogeneous Reconstruction Algorithms (CAHRA), a community-wide challenge centered on three benchmark datasets that probe distinct tasks in heterogeneity analysis. These include (1) compositional heterogeneity arising from mixtures of distinct protein complexes, (2) continuous conformational variability and atomic modeling, and (3) entanglement between molecular conformation and particle pose. We describe the design, construction, and validation of these datasets, as well as their use in the recently completed CAHRA Challenge.

## The ModelSEED Biochemistry Database, 2026 update: grading multi-source thermodynamics
- Source: bioRxiv (preprints)
- Date: 2026-09-17
- Categories: Tools & resources
- Authors: Freiburger, A., Faria, J. P., Edirisinghe, J. N., Liu, F., Taylor, C., Setlur, V., Giessmann, R. T., Beber, M. E., Noor, E., Upadhyay, V., Anand, M., Maranas, C. D., Huss, S., Nikoloski, Z., Arkin, A. P., Cottingham, R. W., Wood-Charlson, E. M., Henry, C. S., Seaver, S. M. D.
- DOI: 10.64898/2026.09.15.751859
- Source URL: <https://doi.org/10.64898/2026.09.15.751859>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.15.751859>
- Code: <https://github.com/ModelSEED/ModelSEEDDatabase>

Abstract: The ModelSEED Biochemistry Database (https://modelseed.org) supplies foundational mass- and charge-balanced reaction networks for metabolic reconstructions. Here we present an update on the biochemistry where the database has been expanded to \[~\] 46,000 compounds, and \[~\] 56,000 reactions, featuring \[~\] 37,000 metabolic structures. We have expanded our approach for handling thermodynamic data, enabling multiple sources of data to be derived, integrated, and presented to the wider research community. We now publish predictions of pKa, reaction energy (and respective uncertainties), and estimates of reaction direction from multiple independent sources. Each reaction is graded gold, silver or bronze according to the strength of the evidence behind it, so that users can weigh its reliability directly. We also release reaction directions predicted by an ensemble of large language models. This multi-source approach exposes agreements and discrepancies between sources for \[~\] 33,000 reactions. Finally, to ensure data integrity, a new conflict-resolution pipeline reconciles structures across sources, documenting input from curators. Our work is publicly available at https://github.com/ModelSEED/ModelSEEDDatabase.

## The public health value of wastewater surveillance for viruses with pandemic potential: a modelling study
- Source: medRxiv (preprints)
- Date: 2026-09-17
- Authors: Dighe, A., Grassly, N. C., Whittaker, C.
- DOI: 10.64898/2026.09.16.26363050
- Source URL: <https://doi.org/10.64898/2026.09.16.26363050>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.16.26363050>

Abstract: Zoonotic spillover and early human-to-human transmission of emerging viruses are often missed by clinical surveillance. Wastewater surveillance (WS) offers low-cost complementary pathogen detection, but its value for emerging viruses remains uncertain. We reviewed evidence of viral shedding in human waste and combined shedding data, detection models and simulated transmission dynamics into a quantitative framework to identify for which types of emerging viruses WS could add most value. Simulations showed that WS improved probability or speed of outbreak detection for viruses shedding >1/100 to >10 times as much as SARS-CoV-2, depending on probability of clinical symptoms and diagnosis. Value added was highest for transmission scenarios with temporally concentrated infections resulting in spikes in daily shedders. Characteristics of SARS-CoV-2, Mpox, Influenza A, Lassa, and Zaire Ebola viruses appear more favourable for WS than others e.g. chikungunya virus. We anticipate this framework can support more targeted, evidence-based application of WS to emerging threats.

## The reasonable effectiveness of domain adaptation for inference of introgression
- Source: bioRxiv (preprints)
- Date: 2026-09-17
- Categories: Genomics & sequence analysis, Evolution & metagenomics
- Authors: Cobb, K., Smith, M. L.
- DOI: 10.1101/2025.01.17.633659
- Source URL: <https://doi.org/10.1101/2025.01.17.633659>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1101%2F2025.01.17.633659>

Abstract: Supervised machine learning approaches have proven powerful in population genetics. To use such approaches, training data with known inputs and outputs are required. Since such data are generally unavailable in population genetics, researchers typically rely on simulations under the models of interest to train machine learning algorithms. While powerful, this approach depends heavily on the models used to generate training data. Because of the variety and complexity of processes shaping genetic variation, it is inevitable that not all processes important in an empirical system will be included when generating training data. This leads to a mismatch between the data used to train a machine learning algorithm and the data to which the trained model is ultimately applied--i.e., a domain shift-- and can negatively impact inference. Here, we train a Convolutional Neural Network (CNN) to detect introgression between sister populations and demonstrate that it has near perfect accuracy when applied to data generated under the models used for training. To evaluate the impacts of domain shifts on inference, we generated new data with introgression from a third, unsampled population into one of the two focal populations (i.e., ghost introgression), and accuracy was substantially reduced on these data. Finally, we used domain adaptation, which aims to train a network that performs well in the presence of a domain shift. Notably, this requires no knowledge of the target or empirical domain. Our domain adaptation network was able to accurately detect introgression, even in the presence of unmodelled ghost introgression. We also applied this approach to empirical data to detect introgression between ABC Island brown bears and other populations of brown bears. Previous work has suggested that introgression between ABC Island bears and polar bears can mislead tests of introgression between populations of brown bears. We found that using domain adaptation reduced support for introgression between geographically isolated populations of brown bears, suggesting that our approach reduces false inferences of introgression due to ghost introgression.

## The Sequential Threshold Model: A Unified Framework for Microbiome-Driven Periodontitis Progression
- Source: bioRxiv (preprints)
- Date: 2026-09-17
- Authors: Duran-Pinedo, A., Reguera-Gomez, M., Rus, M. J., Teles, F., Frias-Lopez, J.
- DOI: 10.64898/2026.09.15.746783
- Source URL: <https://doi.org/10.64898/2026.09.15.746783>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.15.746783>

Abstract: Periodontitis affects nearly one billion people, yet its episodic, site-specific, age-dependent progression is not explained by linear pathogen-burden models. We propose the Sequential Threshold Model (STM), a bistable framework in which progression at a previously diseased but currently stable site requires two sequential events. First, a systemic host gate opens: butyrate-driven histone deacetylase (HDAC) inhibition and NF-\{kappa\}B blockade reduce the senescence-associated secretory phenotype (SASP) surveillance program below the level needed to contain a dysbiotic biofilm. Second, a local microbial gate is crossed when a critical hemin threshold initiates gingipain-dependent positive feedback. Two longitudinal paired-site cohorts, subgingival metatranscriptomic and gingival crevicular fluid, support a six-month transcriptomic breakpoint, a predicted cytokine hierarchy, and primarily cell-state divergence. Formalized as an age-dependent ordinary differential equation, the STM explains five clinical phenomena as consequences of bistability and yields six testable predictions, with implications for epigenetic biomarkers and host-targeted therapy.

## tidyGenR: tidy multilocus amplicon genotypes in R
- Source: PeerJ (journals)
- Date: 2026-09-17T00:00:00Z
- Categories: Genomics & sequence analysis, Tools & resources
- Authors: M. Camacho-Sanchez, Jennifer A. Leonard
- Journal: PeerJ
- DOI: 10.7717/peerj.21726
- External ID: 24074e9a8c3c74b54b303c2218ed2088925a79de
- Source URL: <https://doi.org/10.7717/peerj.21726>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.7717%2Fpeerj.21726>

Abstract: Multiplexed amplicon sequencing has become an important tool in phylogenetics and conservation genetics. Amplicon sequencing reads need to be processed to get final haplotypes. The bioinformatics involved is often limiting for embarking on these kind of projects and there are few tools designed to handle this type of data. tidyGenR is an R package for reproducible multilocus amplicon genotyping workflows from sequencing reads. It provides a modular workflow that starts by demultiplexing loci, variant determination with DADA2 , and ends with genotyping. Input data can be raw single-end or paired-end FASTQ reads and the main outputs are haplotypes in tidy tables. Results can also be exported as FASTA files. We successfully tested tidyGenR on amplicon libraries of 27 loci from a population genetics study in a rodent. The results from tidyGenR were reliable and robust across a wide range of read depths. In addition, tidyGenR offers greater flexibility and interoperability.

## Tractography from Serial Optical Coherence Tomography: How and Why?
- Source: bioRxiv (preprints)
- Date: 2026-09-17
- Categories: Biological imaging, Computational neuroscience
- Authors: Poirier, C., Petit, L., Lefebvre, J., Descoteaux, M.
- DOI: 10.64898/2026.08.14.744847
- Source URL: <https://doi.org/10.64898/2026.08.14.744847>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.08.14.744847>

Abstract: To disentangle complex fiber configurations that remain challenging for diffusion MRI tractography, insights might be gained from microscopy tractography. Indeed, by precisely following small white matter (WM) fascicles, invisible at the resolution of diffusion MRI, microscopy tractography can help explain how fiber populations are organized at the finest scales. Due to its high resolution and its 3D nature, serial optical coherence tomography (S-OCT) offers promise for studying WM at the microscale. However, whether the reflectivity contrast from S-OCT supports tractography at the microscale remains unknown. Furthermore, there is a gap in the literature regarding how an ideal microscopy tractography algorithm should behave with respect to the choice of tractography algorithm, tracking maps definition and microscale orientation distribution functions (ODF) estimation. In this work, we describe a tailored approach to reconstruct WM fascicles at the microscale from S-OCT acquisitions. We validate our approach on a simulated microscopy-like FiberCup dataset, and show that multiscale Frangi filters outperforms structure tensor analysis for estimating ODF. We also show that anatomically-constrained particle filtering tractography enables targetted, region-to-region tractography, and outperforms standard deterministic or probabilistic tracking approaches. We further demonstrate our method on a whole mouse brain S-OCT reconstruction at 10 m by reconstructing thalamocortical WM projections. Overall, our results show that S-OCT tractography recovers fine white matter fascicles that are supported by viral tracing experiments from the Allen Mouse Brain Connectivity Atlas. Moreover, this work shows the first ODF estimation and fully-3D probabilistic particle filtering tractography of the mouse brain from S-OCT reconstructions at 10 m isotropic resolution.

## Unicellular and Multicellular Modes of Selection Impose Distinct Constraints on Cellular Phenotype Evolution
- Source: bioRxiv (preprints)
- Date: 2026-09-17
- Authors: Kim, M., Pennell, M.
- DOI: 10.64898/2026.09.16.751889
- Source URL: <https://doi.org/10.64898/2026.09.16.751889>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.16.751889>

Abstract: Single-cell sequencing data have revealed that cellular phenotypes, such as gene expression states, are often low-dimensional, suggesting that cellular variation may arise from combinations of a smaller set of gene expression programs. A genome therefore defines a repertoire of cellular phenotypes that can be configured through different combinations of programs. However, organisms vary in how much of this repertoire is exposed to selection. In unicellular organisms, different phenotypes are often expressed across environments or life-cycle stages, so selection in a given context acts primarily through the phenotype expressed there. In multicellular organisms, multiple phenotypes can coexist within an individual and contribute jointly to fitness. Here, we use a geometric model to ask how selection acting through cellular phenotypes separately or jointly constrains the ability of a shared genome to evolve and maintain differentiated phenotypes across multiple functional demands. We vary the number of functional demands and how many corresponding phenotypes contribute jointly to fitness. We find similar evolutionary outcomes when demands are weakly divergent. Under strongly divergent demands, however, selection on one phenotype at a time leads to reduced differentiation as demands accumulate, even when sufficient programs are available. As more phenotypes contribute jointly to fitness, differentiation and performance improve. When all phenotypes contribute jointly, differentiation is maintained until demands outnumber programs. Our results suggest that how cellular phenotypes are organized in time and space can impose distinct constraints on the evolution of differentiation from a shared genome.

## UNIVERSAL EPIDEMIC SCALING: INFLUENZA AND COVID-19
- Source: medRxiv (preprints)
- Date: 2026-09-17
- Categories: Mathematical biology & statistics
- Authors: Below, D., Mairanowski, F.
- DOI: 10.64898/2026.09.16.26363217
- Source URL: <https://doi.org/10.64898/2026.09.16.26363217>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.16.26363217>

Abstract: Conventional compartmental models may become difficult to parameterize over multi-wave epidemic horizons when susceptibility and transmission conditions change between successive epidemic regimes. This study introduces a reduced macroscopic framework that changes the scale of epidemic description from individual-level transmission structure to effective population-level dynamics. Epidemic waves are formulated as transport processes governed by deterministic boundaries, mass balance, and a small set of effective macroscopic parameters. The susceptible population is represented by a dynamic Effective Susceptible Pool that can be re-initialized at transitions between biologically distinct epidemic regimes, while transmission resistance and external control measures are incorporated at the macroscopic level. The resulting equations admit a dimensionless similarity representation and closed-form analytical solutions, enabling analytical estimation of epidemic trajectories and peak healthcare demand without computationally intensive numerical simulation. The framework is evaluated using comparative time-series data for SARS-CoV-2 and seasonal influenza A within the geographically and demographically consistent setting of Rhode Island. Despite their different biological and immunological regimes, the analyzed trajectories exhibit a common reduced asymptotic scaling form. The results support the use of a macroscopic, scale-reduced representation for cross-calibration of heterogeneous surveillance signals and analytical assessment of healthcare-system demand. Further validation across pathogens, populations, and open-system settings is required.

## Unlocking Sensitive Data with SPHERE in the Age of AI
- Source: bioRxiv (preprints)
- Date: 2026-09-17
- Categories: Tools & resources
- Authors: He, Z., Park, J., Pulgrossi, R. C., Lee, J., Butler, R. R., Weber, A., Tian, L., Zhang, X., Wang, J., Sha, S., Mormino, E. C., Wyss-Coray, T., Henderson, V. W., Longo, F. M., Zou, J., Desai, M., Altman, R.
- DOI: 10.64898/2026.09.01.748580
- Source URL: <https://doi.org/10.64898/2026.09.01.748580>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.01.748580>

Abstract: Sensitive human data underpin discoveries across medicine, biology and the social sciences, yet privacy regulation often prevents sharing them with collaborators or artificial intelligence (AI) systems. We introduce SPHERE, a model-free method that makes sensitive datasets directly usable by AI and shareable for open science as a synthetic twin, while the original records never leave the local environment. Across 33 datasets spanning five scientific domains, SPHERE protects individual privacy against adversarial re-identification attacks while preserving the data's statistical structure: means, variances and correlations are reproduced exactly, effect size and P value in linear statistical analysis is numerically identical, nonlinear machine-learning utility is retained, and each twin is generated in seconds on a laptop. Frontier AI agents running on the twin reach the same scientific conclusions as on the original records. Analyses of the twin reproduce genome- and proteome-wide results at UK Biobank scale and recover the findings of landmark studies across three independent cohorts and consortia. The approach also extends to deep-learning embeddings across language, vision and time-series, with minimal utility loss. We make the Stanford Alzheimer's Disease Research Center cohort openly available for the first time, as a SPHERE twin spanning nine modalities that any registered researcher can analyze without an approval process. We release SPHERE with certification of each twin's privacy and fidelity, and an AI agent that autonomously executes research tasks on sensitive data without ever accessing it. Sensitive datasets that are currently closed to research could thus become routine inputs to open science and AI to enable key discoveries.

## A 1970s Lipid-Lowering Compound May Offer a New Approach to Obesity
- Source: Bio-IT World (feeds)
- Date: 2026-09-16T21:51:07+00:00
- Categories: Blog
- Source URL: <https://www.bio-itworld.com/news/2026/09/16/a-1970s-lipid-lowering-compound-may-offer-a-new-approach-to-obesity>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fwww.bio-itworld.com%2Fnews%2F2026%2F09%2F16%2Fa-1970s-lipid-lowering-compound-may-offer-a-new-approach-to-obesity>
- Abstract: not stored for this record.

## A Combined ODE Model of Carbohydrate Fermentation and Colorectal Cancer
- Source: arXiv (preprints)
- Date: 2026-09-16T19:48:46Z
- Categories: Systems & networks, Mathematical biology & statistics
- Authors: Alexandra Lawryshyn, Hermann J. Eberl, Thomas Hillen
- External ID: 2609.19371v1
- Source URL: <https://arxiv.org/abs/2609.19371v1>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Farxiv.org%2Fabs%2F2609.19371v1>
- PDF: <https://arxiv.org/pdf/2609.19371v1>

Abstract: We formulate and analyze a system of non-linear ordinary differential equations that describe key metabolic and immunological interactions between butyrate produced by fiber-fermenting gut microbiota, colorectal cancer cells and host cell populations. The model is studied both independently and in conjunction with a pre-existing carbohydrate fermentation model. The parameter space is explored through sensitivity analyses. Simulation experiments are conducted to illustrate the emergence of varying dynamical behaviour driven by butyrate availability. Our model predicts that butyrate production is driven by fiber consumption and further supported by probiotics in the case of microbial dysbiosis. It also suggests that butyrate may help in suppressing tumour growth. We also show that by adding noise with sufficiently high intensity, cancer elimination occurs almost surely in infinite time and that this threshold level of noise intensity decreases with increasing butyrate concentrations.

## dbGaP Modernization Update
- Source: NCBI Insights (feeds)
- Date: 2026-09-16T18:50:43+00:00
- Categories: Blog
- Source URL: <https://ncbiinsights.ncbi.nlm.nih.gov/2026/09/16/dbgap-modernization-update/>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fncbiinsights.ncbi.nlm.nih.gov%2F2026%2F09%2F16%2Fdbgap-modernization-update%2F>
- Abstract: not stored for this record.

## STUART: Sequence Triage and qUAntification of Read Transcripts for Rapid Ionizing Radiation Exposure Assessment
- Source: arXiv (preprints)
- Date: 2026-09-16T17:58:48Z
- Categories: Genomics & sequence analysis, Tools & resources
- Authors: Tomasz Strzoda, Lourdes Cruz-Garcia, Mustafa Najim, Christophe Badie, Joanna Polanska
- External ID: 2609.19139v1
- Source URL: <https://arxiv.org/abs/2609.19139v1>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Farxiv.org%2Fabs%2F2609.19139v1>
- PDF: <https://arxiv.org/pdf/2609.19139v1>

Abstract: Rapid medical triage following ionizing radiation exposure is critical for emergency management, yet traditional alignment-based bioinformatics are too computationally intensive for mass-casualty scenarios. To address this, we developed STUART (Sequence Triage and qUAntification of Read Transcripts), a mapping-free machine learning framework optimized for mobile biological dosimetry. Inspired by Natural Language Processing (NLP), the system converts raw sequencing reads into k-mer-based numerical profiles, completely bypassing standard alignment. While the architecture is universally applicable to any transcriptomic biomarker, this study focused on radiation exposure using the FDXR gene model. Evaluating Logistic Regression, Random Forest, and XGBoost, advanced signature selection strategies drastically reduced the initial 1024-dimensional feature space by over 98%. Highly robust performance - characterized by near-perfect balanced accuracy and F1-scores within the 95-100% range - was consistently achieved while retaining as few as 17 transcriptomic signatures. Crucially, learning curve analysis demonstrated that complete signal stabilization requires aggregating merely 1000 potentially related reads. Furthermore, external validation on an independent dataset yielded over 99% specificity, confirming the tissue-agnostic nature of the extracted signatures despite different cellular origins. The framework's exceptionally low data threshold enables a real-time, analyze-as-you-sequence diagnostic paradigm compatible with portable sequencers. By minimizing time-to-decision, this decentralized tool bridges the gap between advanced biomarkers and practical on-site biomonitoring, offering a scalable foundation for rapid epidemiological response and routine occupational radiation monitoring.

## Interpretable Multi-Instance Learning Enables Early Prediction of Key Molecular Alterations from Routine Flow Cytometry in Acute Myeloid Leukemia
- Source: arXiv (preprints)
- Date: 2026-09-16T15:30:56Z
- Authors: Jonathan Legrand, Aguirre Mimoun, Baudouin Denis de Senneville, Audrey Bidet, Pierre-Yves Dumas, Christèle Etchegaray
- External ID: 2609.18825v1
- Source URL: <https://arxiv.org/abs/2609.18825v1>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Farxiv.org%2Fabs%2F2609.18825v1>
- PDF: <https://arxiv.org/pdf/2609.18825v1>

Abstract: Background: Molecular testing for NPM1 and FLT3-ITD mutations guides critical early treatment decisions in acute myeloid leukemia (AML), but results can take weeks, long after these decisions must be made. Flow cytometry, already performed within hours of admission as part of routine care, may carry enough signal to predict these mutations directly, without added cost or delay. Methods: We developed an interpretable multi-instance learning classifier based on a decision tree, in which each patient sample is modeled as a collection of individual cells and mutation status is inferred from cell-level predictions. The model was benchmarked against a random forest trained on clinical variables and a deep convolutional neural network adapted for multitube flow cytometry data. Performance was assessed by cross-validation on a discovery cohort of 197 patients and tested on an independent cohort of 161 patients, using the area under the receiver operating characteristic curve (AUROC) and positive predictive value. Results: In cross-validation on the discovery cohort, the MIL model achieved mean AUROCs of 0.96 (SD=0.05) for NPM1 and 0.86 (SD=0.10) for FLT3-ITD, outperforming the clinical baseline and matching deep learning approaches. The model then successfully generalized to the independent test cohort of 161 patients, reaching AUROCs of 0.90 (NPM1) and 0.82 (FLT3-ITD), with positive predictive values of 0.87 and 0.68, respectively. Cell-level interpretation recovered established immunophenotypic signatures (CD33$\{\}^\{+\}$ /CD34\_\_\_ for NPM1-mutated cases, CD33$\{\}^\{+\}$ /low side-scatter for FLT3-ITD), directly linking model predictions to known biology. Conclusions: These results show that an interpretable model applied to data already collected in routine care can predict AML molecular status within hours, offering a practical route to earlier, biology-informed treatment decisions.

## When Edit Flows are Edit Jumps: replicating Edit Flows and EvoFlows
- Source: arXiv (preprints)
- Date: 2026-09-16T14:39:26Z
- Categories: Genomics & sequence analysis, Proteins & structural biology, Tools & resources
- Authors: Gabriel Bénédict, Melanie Buechler, Gerard Riera-Solà, Chloé de Ancos, Yves Gaetan Nana Teukam, Moritz Freidank
- External ID: 2609.18745v1
- Source URL: <https://arxiv.org/abs/2609.18745v1>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Farxiv.org%2Fabs%2F2609.18745v1>
- PDF: <https://arxiv.org/pdf/2609.18745v1>
- Code: <https://github.com/VisiumCH/editjumps>

Abstract: Antibody lead optimization calls for a small, bounded set of edits to an existing candidate: substitutions, but also insertions and deletions. Edit-based generative models are the only ones that allocate such an edit budget without fixing the edit positions, the edit count, or the output length in advance. However, the existing approaches Edit Flows and EvoFlows did not release code or complete training specifications. Here, we show that both methods follow the same underlying process -- edits firing one at a time, at learned rates, in continuous time -- the pure-jump case of generator matching over finite sequences. With EditJumps we introduce the first open implementation of this framework, with a single generalist antibody editor trained on 1.66M Observed Antibody Space homolog pairs to propose homolog-like variants of a seed sequence, editing unseen leads zero-shot, without the per-family retraining original approaches require. Replicating this system from scratch exposes why open code is essential for generative biology: reconciling published edit distributions required reverse-engineering an undocumented rate-scaling hyperparameter that dictates realized mutation counts. Moreover, we show that published evaluation metrics are highly sensitive to reference sample size, frequently flipping method rankings. We release our full codebase, automated test suite, and configurations at: https://github.com/VisiumCH/editjumps

## Automatic denoising and differentiation based on Savitzky-Golay filtering and Homogeneous Differentiators for attractor reconstruction via differential embedding
- Source: arXiv (preprints)
- Date: 2026-09-16T13:19:11Z
- Categories: Computational neuroscience
- Authors: Uros Sutulovic, Daniele Proverbio, Rami Katz, Giulia Giordano
- External ID: 2609.18631v1
- Keywords: computational neuroscience
- Source URL: <https://arxiv.org/abs/2609.18631v1>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Farxiv.org%2Fabs%2F2609.18631v1>
- PDF: <https://arxiv.org/pdf/2609.18631v1>

Abstract: Differential embedding methods aim to reconstruct attractors of dynamical systems from noisy measured time series, but require accurate estimates of signal derivatives. We introduce SHADED (Savitzky-Golay and Homogeneous-differentiator based Automatic DEnoising and Differentiation), a novel methodology for denoising and estimation of derivatives up to an arbitrary order, which enables attractor reconstruction via differential embedding from noisy time series data. Homogeneous Differentiators (HD) guarantee finite-time derivative estimates in the presence of noise, while subsequent Savitzky-Golay (SG) filtering attenuates chattering. Crucially, SHADED extracts all parameters required for application of both HD and SG automatically from the data, without requiring manual tuning that may lead to inaccurate reconstruction, and can also incorporate prior knowledge, if available, thereby yielding a flexible tool for data-driven numerical differentiation of noisy signals. The obtained differential embeddings can reveal features of the underlying dynamics that are useful, e.g., for system identification, pattern recognition and discrimination between dynamic regimes; the latter application is particularly important in biomedical settings, to help distinguish between different physiological and pathological states. We demonstrate the efficacy of SHADED by testing it on computational neuroscience models, LTspice-simulated chaotic electronic circuits, and photoplethysmography and arterial blood pressure experimental recordings: across all these case studies, SHADED produces accurate derivative estimates and accurate attractor reconstructions via differential embedding (whenever a ground truth is available) or geometrically coherent and reproducible reconstructions consistent with the expected dynamics (in the absence of a ground truth), without the need for manual parameter tuning.

## Optimum foraging area in a three-trophic food chain
- Source: arXiv (preprints)
- Date: 2026-09-16T12:59:30Z
- Categories: Evolution & metagenomics, Mathematical biology & statistics
- Authors: Lucas Massoni, Rafael Menezes, Marcus A. M. de Aguiar, Sabrina B. L. Araujo
- External ID: 2609.18609v1
- Source URL: <https://arxiv.org/abs/2609.18609v1>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Farxiv.org%2Fabs%2F2609.18609v1>
- PDF: <https://arxiv.org/pdf/2609.18609v1>

Abstract: Organisms' foraging strategies are shaped by a trade-off between search area and local capture efficiency. This trade-off leads individuals to adapt their foraging area to an optimal value, impacting population dynamics. Here we study the effect of multiple foraging areas in a predator-prey model composed of three trophic levels. The interactions between predators and prey occur only within a limited neighborhood of the predators, where adaptation can occur over generations. We assume a trade-off where local predation efficiency is inversely proportional to the foraging area. These dynamics were implemented computationally via cellular automata and analytically via Master Equations with mean-field and pair approximations. Unlike the mean-field approximation, the pair approximation reproduced the dependence of population density on foraging area observed in the simulations. However, the simulations showed that the optimal foraging area does not maximize population density. Moreover, we found that a polymorphic population emerged, where not a single optimal strategy but a range of optimal strategies can coexist. Using the framework of Adaptive Dynamics, we confirm that the range of optimal areas is not the one that maximizes population size, but the Evolutionary Stable Strategy that can invade a population and not be invaded.

## Learning Where to Focus: Self-Supervised Multi-Scale ViTs for Histopathology
- Source: arXiv (preprints)
- Date: 2026-09-16T12:38:41Z
- Categories: Biological imaging
- Authors: Anabel Stammer, Valay Bundele, Mehran Hosseinzadeh, Hendrik P. A. Lensch
- External ID: 2609.18578v1
- Source URL: <https://arxiv.org/abs/2609.18578v1>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Farxiv.org%2Fabs%2F2609.18578v1>
- PDF: <https://arxiv.org/pdf/2609.18578v1>

Abstract: Pathologists diagnose diseases by first locating suspicious tissue and then examining it at higher magnification, whereas self-supervised vision transformers (ViTs) allocate the same spatial resolution to every image region despite diagnostic evidence being sparse and spanning multiple biological scales. Recent pathology foundation models have substantially improved representation quality by scaling training data and model capacity, but largely retain uniform tokenization. We instead investigate whether pathology representations can be improved by learning where to allocate spatial resolution during self-supervised learning. To this end, we propose CRAFT (Coarse-to-fine Region-Adaptive Feature Tokenization), a DINO-based framework that learns image-dependent mixed-scale representations by using self-supervised attention to selectively refine informative regions while preserving coarse context, together with a symmetric cross-scale regularization objective that encourages complementary coarse and fine representations. Across CAMELYON16, TCGA-Lung subtype classification, and TCGA-LUAD survival prediction, CRAFT consistently outperforms comparable-scale self-supervised methods while requiring lower inference computation. Despite using only a compact 22M parameter backbone trained on comparatively small pathology datasets, CRAFT remains competitive with, and often surpasses, substantially larger pathology foundation models.

## The evolution of sex for artificial intelligence: a population-genetic framework for multigenerational model populations
- Source: arXiv (preprints)
- Date: 2026-09-16T12:20:30Z
- Categories: Evolution & metagenomics
- Authors: Giorgio F. Gilestro
- External ID: 2609.18560v1
- Keywords: population genetics, framework
- Source URL: <https://arxiv.org/abs/2609.18560v1>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Farxiv.org%2Fabs%2F2609.18560v1>
- PDF: <https://arxiv.org/pdf/2609.18560v1>

Abstract: Some aspects of AI development resemble a population process in which models are specialised, retrained on the output of peers, or combined by averaging weights. These practices lead to generations of models, in the biological sense studied by population genetics. Here, I develop this parallelism and interpret multigenerational model populations in terms of sexual and asexual reproduction, formally recombining the two fields. I test these analogies in an exact inheritance model, in trained networks (recurrent, feedforward and variational autoencoder generators) and in large language models, and show that they hold generally, with some measurable architecture-specific biases. Training recursively on model output is known to lead to model collapse, a process previously described as akin to genetic drift; I develop all that follows. A minimal model of a learner retrained on its parent's output reproduces the Wright-Fisher process exactly; verified real data added to each generation play the role of immigration, with the surprising finding that the absolute number of real data samples matters, not their share, exactly as in population genetics. Training a child on the average of its parents' outputs cancels the benefit of having several parents, matching blending inheritance (and reviving Jenkin's objection to Darwin), whereas combining parents so that each keeps its strongest contribution preserves it; merged language-model specialists exceeded every parent across seeds (the Fisher-Muller effect); and lineages become reproductively isolated, losing the ability to merge at all, when they have learned conflicting conventions and not when they have merely drifted apart. As AI societies become societies in time as well as in space, a mathematical framework for their inheritance acquires predictive power. Remarkably, that framework can be adapted almost wholesale from biology.

## Avoiding a sticky situation: how cells stop messenger RNAs from clumping together
- Source: RNA-Seq Blog (feeds)
- Date: 2026-09-16T11:14:25+00:00
- Categories: Blog
- Source URL: <https://www.rna-seqblog.com/messenger-rna-or-mrna-is-best-known-for-carrying-genetic-instructions-from-dna-to-the-cellular-machinery-that-makes-proteins-but-all-rna-molecules-have-another-less-appreciated-property-they-are/>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fwww.rna-seqblog.com%2Fmessenger-rna-or-mrna-is-best-known-for-carrying-genetic-instructions-from-dna-to-the-cellular-machinery-that-makes-proteins-but-all-rna-molecules-have-another-less-appreciated-property-they-are%2F>
- Abstract: not stored for this record.

## New AI approaches to help understand complex biological data
- Source: RNA-Seq Blog (feeds)
- Date: 2026-09-16T11:14:14+00:00
- Categories: Blog
- Source URL: <https://www.rna-seqblog.com/new-ai-approaches-to-help-understand-complex-biological-data/>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fwww.rna-seqblog.com%2Fnew-ai-approaches-to-help-understand-complex-biological-data%2F>
- Abstract: not stored for this record.

## Hyperbolic Graph Representation Learning for Differential Diagnosis on Biomedical Knowledge Graphs
- Source: arXiv (preprints)
- Date: 2026-09-16T11:11:46Z
- Categories: Proteins & structural biology
- Authors: Pietro Miotto, Lucia Mellini, Tommaso Marzi, Cesare Alippi, Elena Casiraghi, Alberto Paccanaro, Giorgio Valentini, Mauricio Soto-Gomez
- External ID: 2609.18481v1
- Keywords: representation learning
- Source URL: <https://arxiv.org/abs/2609.18481v1>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Farxiv.org%2Fabs%2F2609.18481v1>
- PDF: <https://arxiv.org/pdf/2609.18481v1>

Abstract: Biomedical knowledge graphs combine ontology-derived hierarchies with transversal associations among heterogeneous entities such as phenotypes, diseases, genes, proteins, and patients. This hybrid structure raises the question of whether hyperbolic embeddings, which naturally capture tree-like organization, remain useful beyond purely hierarchical graphs. We present a preliminary study of hyperbolic graph representation learning for Mendelian-disease differential diagnosis on a patient-integrated biomedical graph. Experiments on isolated ontology subgraphs show that hyperbolic models achieve strong performance in substantially lower dimensions than Euclidean baselines. We then evaluate the models on a link-prediction task that ranks candidate diseases for each patient. Results suggest that hyperbolic embeddings can exploit biomedical hierarchical structure while supporting diagnostic reasoning over heterogeneous patient-level graphs.

## HPOQuest: A Rare-Disease Diagnostic Agent Using Active Phenotype Acquisition
- Source: arXiv (preprints)
- Date: 2026-09-16T10:23:34Z
- Authors: Kamilia Zaripova, Nassir Navab, Azade Farshad, Annalisa Marsico
- External ID: 2609.18431v1
- Source URL: <https://arxiv.org/abs/2609.18431v1>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Farxiv.org%2Fabs%2F2609.18431v1>
- PDF: <https://arxiv.org/pdf/2609.18431v1>

Abstract: More than 300 million people worldwide are affected by one of over 7,000 known rare diseases, yet diagnosis remains difficult because patients initially present with incomplete and heterogeneous phenotypes. We present HPOQuest, a training-free framework for sequential phenotype acquisition in rare-disease diagnosis. Starting from a small set of observed patient phenotypes, HPOQuest maintains a probabilistic disease ranking and iteratively selects informative follow-up questions to support clinicians during patient assessment. Confirmed phenotypes update the disease ranking, while all responses update the candidate question set. Across four benchmark cohorts, HPOQuest substantially improves diagnosis from sparse initial phenotypes, with gains of up to 30% points at Recall@1 and 45% points at Recall@5. These results demonstrate that sequential phenotype acquisition can substantially improve rare-disease diagnosis from limited initial clinical evidence.

## NP-Hardness and a Fixed-Parameter Algorithm for Translocation Distance
- Source: arXiv (preprints)
- Date: 2026-09-16T09:53:48Z
- Categories: Genomics & sequence analysis
- Authors: Maria Constantin, Adrian Miclăuş, Alexandru Popa
- External ID: 2609.18397v1
- Keywords: genome, dna, algorithm
- Source URL: <https://arxiv.org/abs/2609.18397v1>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Farxiv.org%2Fabs%2F2609.18397v1>
- PDF: <https://arxiv.org/pdf/2609.18397v1>

Abstract: In this paper we study the genome rearrangements done by translocation events. Genome rearrangements were used to measure evolutionary distance between organisms since 1936 (Dobzhansky and Sturtevant). The chromosomes are represented as strings of DNA and the \\emph\{translocation operation\} is defined as the exchange of prefixes between two strings. This operation results in the creation of two new strings (chromosomes) that can then be utilized in subsequent translocations. A translocation is referred to as \\emph\{contiguous\} if the new strings are produced in a single copy, so each of them can be used in only one subsequent operation. When the words produced by a translocation operation are considered to have an infinite number of copies, the translocation is referred to as \\emph\{non-contiguous\}. If the exchanged prefixes are of equal length, the translocation is called \\emph\{uniform\}. Otherwise, the translocation is termed \\emph\{non-uniform\}. The \\emph\{translocation distance\} between two sets of strings, termed the input set and the target set, represents the minimum number of translocations necessary to obtain all the strings in the target set via translocation operations. We prove that both the non-uniform contiguous and the non-uniform non-contiguous translocation distance problems are NP-hard over arbitrary finite alphabets, where the alphabet is part of the input. For the case in which the target set consists of a single string, we give a fixed-parameter tractable algorithm parameterized by the length of the target string.

## PlainMap: a lightweight, restartable mapping pipeline for ancient and modern DNA
- Source: arXiv (preprints)
- Date: 2026-09-16T09:26:24Z
- Categories: Genomics & sequence analysis, Tools & resources
- Authors: Michael V. Westbury
- External ID: 2609.18372v1
- Source URL: <https://arxiv.org/abs/2609.18372v1>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Farxiv.org%2Fabs%2F2609.18372v1>
- PDF: <https://arxiv.org/pdf/2609.18372v1>
- Code: <https://github.com/BiodiversityExtinction/PlainMap>

Abstract: Mapping sequencing reads to a reference genome requires preprocessing and alignment choices that can vary with library type, fragment length, sequencing platform, and reference genome. These considerations are particularly important for ancient and historical DNA, where short and damaged fragments can make the most appropriate mapping strategy difficult to determine a priori. We present PlainMap, a lightweight and restartable mapping pipeline for modern and degraded DNA sequencing data. PlainMap accepts a simple manifest of FASTQ files, automatically identifies single-end and paired-end data from read headers, supports mixed sequencing platforms, and provides alternative mapping strategies for modern and degraded DNA. Deterministic chunking and checkpoint-based execution allow large analyses to resume after interruption, while optional pilot subsampling enables empirical comparison of mapping strategies using identical subsets of raw fragments. PlainMap produces duplicate-filtered BAM files together with fragment-aware mapping and coverage statistics. Evaluation using heterogeneous sequencing data confirmed the expected behaviour of the three mapping modes, while adaptive chunking reduced peak memory use by approximately 25% and allowed an interrupted analysis to resume from completed mapping chunks. PlainMap is implemented as a single Bash script and is freely available at https://github.com/BiodiversityExtinction/PlainMap.

## Four decades of discovery at EMBL
- Source: EMBL (feeds)
- Date: 2026-09-16T07:48:47+00:00
- Categories: Blog
- Source URL: <https://www.embl.org/news/people-perspectives/four-decades-of-discovery-at-embl/>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fwww.embl.org%2Fnews%2Fpeople-perspectives%2Ffour-decades-of-discovery-at-embl%2F>
- Abstract: not stored for this record.

## Building a clear picture of UK health research funding
- Source: EMBL (feeds)
- Date: 2026-09-16T07:41:32+00:00
- Categories: Blog
- Source URL: <https://www.embl.org/news/updates-from-data-resources/uk-healthcare-funding-dashboard-announcement/>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fwww.embl.org%2Fnews%2Fupdates-from-data-resources%2Fuk-healthcare-funding-dashboard-announcement%2F>
- Abstract: not stored for this record.

## Neural noise enables accurate internal simulation of rare events
- Source: arXiv (preprints)
- Date: 2026-09-16T02:31:27Z
- Authors: Heng Zhang, Pawel Herman, Zenas C. Chao
- External ID: 2609.18033v1
- Source URL: <https://arxiv.org/abs/2609.18033v1>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Farxiv.org%2Fabs%2F2609.18033v1>
- PDF: <https://arxiv.org/pdf/2609.18033v1>

Abstract: The brain needs an accurate internal model of the world to generate predictions and guide behavior. However, it must estimate the statistical structure of the environment from limited experience. This is particularly difficult for rare events, whose observed frequencies in a limited sample may substantially under- or overestimate their true frequencies. How the brain constructs an accurate internal model despite this sampling problem remains unclear. We address this problem using a Bayesian Confidence Propagation Neural Network (BCPNN) trained on event sequences from a Markov-chain random walk with controlled event frequencies. Treating the underlying Markov structure as the ground truth, we train the network on limited sample of event sequences and then allow it to generate autonomous replay based on the learned structure. We evaluate replay fidelity at the levels of both marginal event frequencies and conditional transition structure. We find that moderate neural noise, modeled as temporally correlated random fluctuations in unit activity during replay, is critical for faithful internal simulation. Without this variability, deterministic replay systematically under- or overrepresents rare events, whereas moderate noise restores both their marginal and conditional occurrence. Moderate noise also broadens the range of parameter values that produce accurate replay, making the model more robust to parameter variation. Together, these results support noise-assisted internal simulation as a potential mechanism for compensating for sampling errors arising from limited experience. Our model also provides a testable framework for investigating how altered neural variability may impair internal-model fidelity in disorders such as Parkinson's disease.

## 3D chromatin remodeling during domestication defines novel targets for crop improvement.
- Source: Cell (journals)
- Date: 2026-09-16
- Categories: Genomics & sequence analysis, Proteins & structural biology, Systems & networks
- Authors: Xianhui Huang, Yabin Peng, Xiubao Hu, Yuejin Wang, Zeyu Zhang, Xianzhe Huang, Xuanxuan Luo, Sainan Zhang, Zengyuan Zhao, Erin Farmer, Sheng-Kai Hsu, Corrinne E Grover, Zhengyang Qi, Lu Li, Jinglei Yang, Yinfang He, Zhiwei Chen, Yuanhang Zhang, Ye Mei, Pengcheng Deng, Yang Meng, Yufei Wang, Mengyuan Ji, Junyuan Lv, Liuling Pei, Fang Liu, Xinhui Nie, Lili Tu, Keith Lindsey, Adnane Boualem, Abdelhafid Bendahmane, Jonathan F Wendel, Michael A Gore, Xianlong Zhang, Maojun Wang
- Journal: Cell
- DOI: 10.1016/j.cell.2026.08.038
- External ID: 42748920
- Keywords: chromatin, genome, interactome
- Source URL: <https://doi.org/10.1016/j.cell.2026.08.038>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1016%2Fj.cell.2026.08.038>

Abstract: Three-dimensional (3D) genome folding shapes gene regulation, yet the genetic underpinnings linking 3D genome evolution to phenotypic innovation during domestication remain elusive. Using population-scale Hi-C profiling of 34 semi-wild and 267 cultivated allotetraploid cottons, we generated a pan-3D genome atlas capturing extensive diversity in topologically associating domains (TADs) and chromatin loops. Chromatin interactome-wide association studies identified 105 TAD reconfigurations and 58 loop rewirings that were established as the 3D chromatin basis of fiber quality, boosting heritability estimates for fiber strength by 16% and fiber length by 20%. We reveal that domestication selection within sequence-defined sweeps fixed 57% of 3D conformation signatures, thereby decoupling sequence-level from chromatin-level selection and shifting the subgenome expression balance of 39 homoeologs in cultivated cotton. Sequence-based modeling and mutational analyses identified the C2H2 zinc-finger protein YY1 as a conserved mediator of 3D genome organization. This study provides a resource for redefining precision-breeding paradigms by harnessing cryptic 3D chromatin targets.

## A comprehensive map of the bovine mobilome and their epigenetic regulation of mastitis
- Source: Functional & Integrative Genomics (journals)
- Date: 2026-09-16T00:00:00Z
- Categories: Genomics & sequence analysis
- Authors: Nai-Su Yang, Meng-Qi Wang, Sarah E. Abanda Mbili, Antony T. Vincent, Cheng-Yi Song, E. Ibeagha-Awemu
- Journal: Functional & Integrative Genomics
- DOI: 10.1007/s10142-026-01970-5
- External ID: 524f9382915682340200c6ef0715825cffb0a13c
- Source URL: <https://doi.org/10.1007/s10142-026-01970-5>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1007%2Fs10142-026-01970-5>

Abstract: Retrotransposons are major components of mammalian genomes, yet their genome-wide annotation and epigenetic regulation in cattle remain incompletely characterized. Here, we present a comprehensive annotation and integrative epigenomic analysis of major retrotransposon classes in the bovine genome, with emphasis on DNA methylation patterns and their potential roles in subclinical mastitis. Using a multi-step de novo pipeline, we identified and classified long interspersed nuclear elements (LINEs), short interspersed nuclear elements (SINEs), and endogenous retroviruses (ERVs) in the bovine reference genome, including three LINE families, seven SINE families, and 20 ERV families. Retrotransposons accounted for ~ 40% of the genome, with LINEs representing the largest fraction. Evolutionary analysis suggested that BosL1A1, BosSINEL, BosERV1, and BosERV16 are the most recently active families. DNA methylation profiles in milk somatic cells showed consistently high levels across retrotransposons, with widespread hypermethylation in cows with Staphylococcus aureus–induced subclinical mastitis. We identified 20,839 differentially methylated retrotransposons, 3,306 of which overlapped differentially expressed genes. Promoter- and exon-overlapping elements showed inverse correlations between methylation and gene expression, and enriched genes are involved in immune and transport pathways. These results provide a comprehensive bovine mobilome resource and indicate that retrotransposon methylation is associated with gene regulation and host responses to mastitis.

## A Deep Model Framework for Morphological Trait Imputation Across Taxonomic Groups.
- Source: Integrative zoology (journals)
- Date: 2026-09-16T00:00:00Z
- Categories: Evolution & metagenomics
- Authors: Yuang Wang, Xin-Ying Shi, Yu Bai, Hong-Jie Zhu, Xue-Mei Yin, Peng-Fei Song, Daji Ergu, Ta-Xing Zhang, Sheng-Kai Pan, Zhong-Ru Gu, Fang-Yao Liu, Xiangjiang Zhan
- Journal: Integrative zoology
- DOI: 10.1111/1749-4877.70178
- External ID: c5594cc6f13b33eefcd5d6064012041463b891a0
- Keywords: phylogenetics, framework
- Source URL: <https://doi.org/10.1111/1749-4877.70178>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1111%2F1749-4877.70178>

Abstract: Incomplete morphological trait data pose major hurdles for trait-based analyses, particularly when missing values, multicollinearity, and sparse sampling constrain inference. These issues limit our ability to quantify trait variation and explore broad patterns of functional differentiation across taxa. Here, we introduce FS-DeepRBFNet, which overcomes these pitfalls through integrating correlation-based feature selection with a dual-layer adaptive radial basis function (RBF) network. This end-to-end approach effectively reduces noise and captures both linear allometric trends and nonlinear morphological relationships. We tested the framework on a large species-level morphological trait dataset of Chinese birds and further validated its cross-taxon transferability using the Amphibian Database (Caudata). FS-DeepRBFNet consistently outperformed conventional methods such as KNN, Random Forest, and XGBoost, demonstrating superior predictive accuracy across multiple traits. Beyond improvements, the model revealed biologically interpretable trait associations and stable cross-taxon generalization. These results demonstrate that FS-DeepRBFNet provides a robust and biologically grounded solution for morphological trait prediction, enabling reliable imputation for comparative phylogenetics, functional ecology, and biodiversity forecasting in data-limited situations.

## A Fokker-Planck framework for control of epidemics
- Source: Journal of Mathematical Biology (journals)
- Date: 2026-09-16T00:00:00+00:00
- Categories: Mathematical biology & statistics
- Authors: Christian Parkinson, Souvik Roy
- Journal: Journal of Mathematical Biology
- DOI: 10.1007/s00285-026-02462-7
- Source URL: <https://doi.org/10.1007/s00285-026-02462-7>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1007%2Fs00285-026-02462-7>

Abstract: We present a control framework for stochastic compartmental models in epidemiology. In this framework, rather than directly controlling the stochastic system, we perform optimal control of an associated Fokker-Planck equation, with the goal of steering the distribution of possible solutions of the stochastic system to some desirable state. In particular, this allows for robust control mechanism with uncertainty not only in the dynamics, but also in the initial data. We formulate and fully analyze a partial differential equation constrained optimization problem, including a proof of existence of optimal controls via analysis of the control-to-state map, and a characterization of optimal controls via the Pontryagin minimum principle. We describe the application of the sequential quadratic Hamiltonian method to our problem, which provides numerical approximations of optimal control maps. We demonstrate our method using a minimal stochastic susceptible-infected-recovered model with different choices of cost functionals that represent different policy-maker concerns.

## A framework of Microbial Genomic Database for clinical metagenomic pathogen diagnosis: development and multi-cohort evaluation
- Source: Frontiers in Cellular and Infection Microbiology (journals)
- Date: 2026-09-16T00:00:00Z
- Categories: Genomics & sequence analysis, Evolution & metagenomics, Tools & resources
- Authors: Han Xia, Yan-Hua Wen, Xu-Ming Li, Long Hu, Ya-Qi Yuan, Juan-Juan Tian, Song Li, Yao Zhan, Xiao-Fei Dang, Yu-Ting Lin, Li-Li Li, Ying-Jie Chen, Ye Zhang, Yuan-Lin Guan, Jun Wang
- Journal: Frontiers in Cellular and Infection Microbiology
- DOI: 10.3389/fcimb.2026.1938149
- External ID: ca1eb91f6e6bc6379369d693cc6b4bba2f03a59d
- Source URL: <https://doi.org/10.3389/fcimb.2026.1938149>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.3389%2Ffcimb.2026.1938149>

Abstract: Clinical metagenomic next-generation sequencing (mNGS) enables broad, untargeted pathogen detection, but its analytical performance depends on host depletion strategy, reference database composition, and alignment methodology. We developed the Clinical Microbial Genomic Database (CMGD), a clinically focused reference resource prioritizing medically relevant taxa. CMGD was manually curated, clinically stratified, and included more than 18,000 microbial species. We evaluated host-depletion references, alignment and classification strategies, six published clinical cohorts, and 30 retrospective mNGS-positive clinical samples. The combined GRCh38-T2T reference achieved the highest human-read depletion rate while minimizing microbial-read loss. CMGD provided broader target-species coverage than the standard Kraken2 database, and BWA-CMGD showed lower erroneous assignment rates overall, although Kraken2 yielded higher unique species-level assignment rates for many shared taxa. Across six published clinical cohorts, CMGD achieved 91.0% detection concordance with BLAST-NT and a strong read-count correlation (R 2 = 0.97). In 30 retrospective samples, CMGD and NT showed strong correlations for total mapped reads (R 2 = 0.99) and uniquely mapped reads (R 2 = 0.89), with concordance correlation coefficients of 0.99 and 0.92, respectively. High sequence-mapping accuracy did not ensure reliable species-level discrimination for highly homologous taxa such as Escherichia coli and Shigella flexneri . Clinically stratified database curation improves the analytical performance, computational efficiency, and interpretability of mNGS-based pathogen detection. Species-complex-level reporting may be more appropriate when species-level discriminatory evidence is insufficient. Prospective multicenter validation is required to establish clinical diagnostic utility.

## A mathematical model for fitness effects on viral persistence
- Source: bioRxiv (preprints)
- Date: 2026-09-16
- Categories: Evolution & metagenomics, Mathematical biology & statistics
- Authors: Llopis-Almela, O., Lazaro, J. T., Duran, A., Perales, C., Domingo, E., Sardanyes, J.
- DOI: 10.64898/2026.09.14.751427
- Source URL: <https://doi.org/10.64898/2026.09.14.751427>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.14.751427>

Abstract: Persistent viral infections arise from complex interactions between viral replication, host cell responses, and ongoing viral evolution. A general framework linking viral fitness to persistence dynamics is lacking. Here, we develop for the first time a mathematical model of viral persistence that takes into consideration viral fitness variations. Essential to the model is the partition of the classical fitness parameter into three components: replicative, infective, and dispersal fitness. The model was initially inspired by a new experiment on hepatitis C virus (HCV) persistence, established in human hepatoma cells, also reported in this work. This experiment documents two strikingly different viral trajectories depending on the initial replicative fitness of the viral population used to establish persistence. The dynamical model describes the interactions among uninfected cells, infected cells, and infectious virions, and it incorporates, through a continuum, two alternative mechanisms of viral release from cells: budding and lysis. Analysis of the model reveals that viral fitness parameters organise infection outcomes into distinct dynamical regimes. Low replicative and dispersal fitness values lead to viral extinction, whereas high values enable persistence through either stable coexistence or recurrent infection waves. The space of fitness discloses a hierarchy among these parameters, with replicative and dispersal fitness being able to trigger important shifts in the outcome of the infection as opposed to infective fitness. The model identifies trade-offs between replication and dispersal that shape viral production and predicts slow dynamical regimes in which infection may persist despite low detectable viral loads. These dynamical transitions provide candidate mechanisms capable of generating the persistence patterns observed experimentally. Our results establish a computational framework linking multidimensional viral fitness to persistence dynamics and suggest general principles by which evolving RNA viruses transition between extinction (cell curing) and sustained persistence.

## A mechanistic digital twin model for epigenetic therapy optimization in triple-negative breast cancer
- Source: bioRxiv (preprints)
- Date: 2026-09-16
- Categories: Systems & networks
- Authors: Bruno, S., Indeglia, A., Lichterfeld, S., Schade, A. E., Cichowski, K., Michor, F.
- DOI: 10.64898/2026.09.11.750942
- Source URL: <https://doi.org/10.64898/2026.09.11.750942>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.11.750942>

Abstract: Epigenetic therapies offer a promising approach to cancer treatment by modulating chromatin states that govern tumor cell identity, plasticity and therapeutic response. However, predicting and optimizing the effects of such interventions remains challenging. Here, we developed a digital twin framework that integrates mechanistic models of chromatin regulation, in vitro cell-state and treatment response data, and pharmacokinetics to simulate tumor progression and therapeutic response. We applied this framework to triple-negative breast cancer (TNBC), an aggressive disease in which chromatin dysregulation contributes to tumor progression, and investigated combination treatment with an EZH2 inhibitor promiting chromatin opening and an AKT inhibitor, which together enhance expression of GATA3 and BMF. Parameterized and validated using in vitro treatment response data, the model enables in silico clinical trials of alternative combination regimens and treatment schedules. These simulations identify regimens that achieve comparable therapeutic effects to reference schedules while substantially reducing cumulative drug exposure. We further demonstrated the digitan twin's ability of identifying personalized therapeutic strategies by incorporating patient-specific treatment-response data. Our work establishes a mechanistic digital twin framework for predicting tumor responses to chromatin-modifying therapies and provides a quantitative approach for optimizing treatment combinations and schedules across diverse cancer contexts.

## A single nuclei expression resource for exploring dog brain cell transcriptomic diversity
- Source: bioRxiv (preprints)
- Date: 2026-09-16
- Categories: Genomics & sequence analysis, Single-cell & spatial
- Authors: Christmas, M. J., Pederson, E., Wallerman, O., Pyl, P. T., Reinsbach, S., Sundstrom, E., Wang, C., Karlsson, A., Arendt, M., Meadows, J. R. S., Lindblad-Toh, K.
- DOI: 10.64898/2026.09.14.751376
- Keywords: transcriptomic, rna, transcriptomes, gene expression, single nuclei, single nucleus, resource
- Source URL: <https://doi.org/10.64898/2026.09.14.751376>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.14.751376>

Abstract: Domesticated dogs present a unique case in nature where mutualism with and selective breeding by humans has led to profound changes in their environment, physiology, and behaviour compared to their wolf-like ancestors. Dog tameness, sociability, and trainability have likely evolved due to significant alterations to brain function. Exploring the effects of domestication on the dog brain will be facilitated by the characterisation of dog brain cell diversity, a task which is not yet complete. To fill this gap, we use single-nucleus RNA sequencing and survey cell types across the dog brain. We present transcriptomes from almost 60,000 cells sampled from the cerebellum, thalamus, and three regions of the cerebrum. Our analysis identified 24 major clusters representing 21 broad cell types, and 131 subclusters revealing regional variation. Comparisons with human and mouse datasets revealed a high level of conservation in neurons across mammalian brains, and greater divergence in glial cells, particularly oligodendrocytes. We demonstrate the utility of the dataset for interrogating specific gene expression patterns across the brain, including genes implicated in dog domestication, behavioural traits, and disease. We discover distinct differences in the expression of opioid receptors in the brains of dogs, compared to humans and mice, providing a potential explanation for dogs' higher tolerance and lower risk of severe adverse effects of opioid drugs. The data from this project are released via a web application for use by the wider community.

## AlphaBridge: Tools for the analysis of predicted biomolecular complexes.
- Source: Structure (London, England : 1993) (journals)
- Date: 2026-09-16
- Categories: Proteins & structural biology, Tools & resources
- Authors: Daniel Álvarez-Salmoral, Razvan Borza, Carmen Maiella, Bjørn P Y Kwee, Ren Xie, Robbie P Joosten, Maarten L Hekkelman, Anastassis Perrakis
- Journal: Structure (London, England : 1993)
- DOI: 10.1016/j.str.2026.08.011
- External ID: 42748924
- Source URL: <https://doi.org/10.1016/j.str.2026.08.011>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1016%2Fj.str.2026.08.011>

Abstract: Artificial intelligence (AI)-powered protein structure prediction transformed how scientists explore macromolecular function. AI-based prediction of macromolecular complexes is increasingly used for evaluating the likelihood of proteins forming complexes with other proteins, nucleic acids, lipids, sugars, or small-molecule ligands. Efficient tools are needed to evaluate such predicted models. Here, we combine confidence metrics of AlphaFold3 to enable clustering of sequence motifs participating in binary interactions and subsequently in 3D interfaces of complexes. Interaction interfaces within confidence limits are visualized via chord diagrams, network graphs, and summary tables of predicted interfaces and intermolecular interactions, and linked to interactive graphics. We validate AlphaBridge for scoring binary and multi-component protein complexes and discuss real-life examples of its use. AlphaBridge is a reproducible, objective, and automated toolkit available also as a web server, providing novice and experienced users with validation for assessing structure prediction of biomolecular complexes.

## An analytical stochastic stage-structured model for pest population dynamics: a case study on Cydia pomonella in Italy
- Source: Journal of Mathematical Biology (journals)
- Date: 2026-09-16T00:00:00+00:00
- Categories: Evolution & metagenomics, Mathematical biology & statistics
- Authors: Berk Tan Perçin, Serena Baiocco, Federico Cavina, Gianfranco Pradolesi, Sara Pasquali
- Journal: Journal of Mathematical Biology
- DOI: 10.1007/s00285-026-02461-8
- Source URL: <https://doi.org/10.1007/s00285-026-02461-8>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1007%2Fs00285-026-02461-8>

Abstract: In this work, we propose a novel phenological model to describe population dynamics based on a compound Poisson process driven development. The model offers an alternative to other well-established approaches founded on systems of ordinary or partial differential equations, in which, the stochasticity is driven by the Brownian motion. A key advantage of the proposed framework is that it prevents the age regression of individuals and allows for the analytical derivation of the number of individuals in each developmental stage into which the population is structured. The model is applied to Cydia pomonella , a major pest of pome fruit crops. The results obtained are highly promising and suggest that this approach constitutes a valuable alternative to traditional models based on partial differential equations for simulating the population dynamics.

## An Ensemble-Dense X-Net model for detection of complex regions in Parkinson’s disease using high-resolution MRI scans
- Source: Scientific Reports (journals)
- Date: 2026-09-16T00:00:00+00:00
- Authors: Madhavi Garimella, Ponnam Vidya Sagar
- Journal: Scientific Reports
- DOI: 10.1038/s41598-026-69832-5
- Source URL: <https://doi.org/10.1038/s41598-026-69832-5>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1038%2Fs41598-026-69832-5>

Abstract: Parkinson’s disease (PD) is a chronic neurological disorder that mainly affects daily life. The aim of this research was primarily to detect PD in its early stages based on abnormal behavior such as cognitive impairment, rapid changes in emotions, and self-control disorders. In this research, a fine-tuned pre-trained DenseNet-based deep learning (DL) model that reliably retrained on PD MRI images. The preprocessing techniques, such as affine transformations and Patch Extraction, are used to enhance the input images. Advanced tissue segmentation is another process that segments the brain output images into specific regions. Finally, the Ensemble-Dense X-Net (EDX-Net) model is used to detect PD based on significant brain regions like substantia-nigra and classifies the samples. The proposed model is developed as a Cross-modal system and was effectively evaluated on two neuroimaging datasets: the Parkinson’s disease functional magnetic resonance imaging (fMRI) Images dataset (D1) and the Parkinson’s disease Dementia (PDD) MRI dataset (D2), both collected from Kaggle. This research also focused on identifying affected regions using both fMRI and MRI images. These two datasets are two different imaging modalities such as fMRI and MRI. Experimental results show that the proposed approach achieves performance of Sn of 97.78, Sp of 98.34, P of 97.89, Acc of 98.99, F1S of 96.23. For D1, and Sn-98.31, Sp-96.99, P-97.78, Acc-98.45, and F1S-97.88 for D2 with significantly less processing time. Thus, we can say that the proposed approach works effectively on fMRI and MRI images.

## Atlas of Proteomic Technologies: an evidence-based framework for selecting and combining commercial proteomics platforms
- Source: bioRxiv (preprints)
- Date: 2026-09-16
- Categories: Proteins & structural biology, Tools & resources
- Authors: Whelan, C. D., Smith-Byrne, K.
- DOI: 10.64898/2026.09.10.750723
- Source URL: <https://doi.org/10.64898/2026.09.10.750723>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.10.750723>

Abstract: High-throughput proteomics now spans several technologies that differ in biological breadth, precision, specificity, sensitivity, cost, and method(s) of quantification; however, investigators currently lack an evidence-based framework to select or combine platforms. Here, we present the Atlas of Proteomic Technologies (APT) - a curated resource that evaluates six commercial platforms across ten analytical dimensions using published head-to-head evidence, expert evaluation, and iterative peer review. APT hosts two decision tools: the Help Me Choose tool maps a study's primary aim(s), scale, sample matrix, and technical constraints to a calibrated, weight-adjustable platform ranking, and the Help Me Combine tool ranks complementary platform pairs by their net-new protein coverage and user-specified technical priorities. Both tools were tested exhaustively across all possible combinations and behave in a balanced, merit-based manner. APT is openly accessible at https://aptatlas.org, providing downloadable tool specifications, seeded analysis scripts, and complete protein coverage lists, ensuring every score and recommendation is inspectable and reproducible.

## Automatic Quality Control and Error Correction in MRI linear registration via a Residual Parameter Prediction Network for T1w MRI
- Source: bioRxiv (preprints)
- Date: 2026-09-16
- Categories: Biological imaging, Tools & resources
- Authors: Chen, Z., Moqadam, R., Metz, A., Adame Gonzalez, W., Zeighami, Y., Dadar, M.
- DOI: 10.64898/2026.09.10.750693
- Source URL: <https://doi.org/10.64898/2026.09.10.750693>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.10.750693>
- Code: <https://github.com/ZhaojinChen/RACOON>

Abstract: Errors in linear registration can propagate to downstream nonlinear registration and bias volumetric estimations, deformation-based morphometry (DBM) and voxel-based morphometry (VBM) analyses. Subtle linear registration errors are particularly challenging as they are difficult to detect and may not result in obvious failures in nonlinear registration but still affect downstream results. Therefore, accurate identification and correction of these errors are critical. In this study, we present the Residual Affine COefficient Optimization Network (RACOON), a framework designed to identify and correct linear registration errors in T1w MRI scans registered to the MNI-ICBM152 space. RACOON's correction module achieved a residual misalignment RMSE of 0.778 mm on synthetic dataset, comparable to the variability observed among repeated QC-passed registrations using the same pipeline. For the classification module, RACOON achieved a balanced accuracy of 76.8% and a precision of 74.4%, outperforming existing state-of-the-art methods. RACOON is open source and publicly available at https://github.com/ZhaojinChen/RACOON.

## AutoRNA: RNA tertiary structure prediction using variational autoencoder.
- Source: bioRxiv (preprints)
- Date: 2026-09-16
- Categories: Proteins & structural biology
- Authors: Kazanskii, M. A., Uroshlev, L., Zatylkin, F., Pospelova, I., Kantidze, O., Gankin, Y.
- DOI: 10.1101/2024.06.18.599511
- Source URL: <https://doi.org/10.1101/2024.06.18.599511>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1101%2F2024.06.18.599511>

Abstract: Understanding the three-dimensional (3D) organization of RNA is essential for advancing therapeutic development and vaccine design. However, the limited availability of experimentally resolved RNA structures restricts the applicability of data-intensive machine learning approaches for tertiary structure prediction. This study aims to develop a data-efficient method for learning coarse-grained RNA structural organization directly from sequence information. We propose AutoRNA, a variational autoencoder-based model that learns sequence-conditioned structural priors in the form of inter-nucleotide distance matrices. The model was trained on RNA structures obtained from the Protein Data Bank and restricted to sequences up to 64 nucleotides. Predicted distance matrices were converted into 3D coordinates using multidimensional scaling, followed by template-based assembly and molecular dynamics refinement. Model performance was evaluated using mean absolute error (MAE), root mean square error (RMSE), global distance test (GDT), and template modeling (TM) scores. On the test dataset, AutoRNA achieved an RMSE of 4.49~\\AA\{\} and an MAE of 3.13~\\AA\{\} for predicted inter-nucleotide centroid distances. However, the reconstructed three-dimensional structures showed limited global fold recovery, as indicated by moderate GDT scores and low TM-scores. Performance decreased for longer sequences, indicating limitations associated with data scarcity and increased structural complexity. Molecular dynamics refinement provided modest improvements, particularly for initially low-quality predictions. AutoRNA demonstrates that variational autoencoders can learn meaningful coarse-grained structural representations of RNA from limited data. While not suitable for near-native tertiary structure prediction, the method generates candidate coarse-grained inter-nucleotide distance restraints that could potentially be incorporated into downstream structure-reconstruction or physics-based refinement workflows. This work highlights the potential of generative models for RNA structure prediction in low-data regimes.

## BioTester: an AI-driven automated testing framework for identifying potential quality risks in bioinformatics software
- Source: Briefings in Bioinformatics (journals)
- Date: 2026-09-16T00:00:00+00:00
- Categories: Tools & resources
- Authors: Xin Lian, Jiayin Wang, Xiaoyan Zhu, Sizhe Dang, Tianxiang Xu
- Journal: Briefings in Bioinformatics
- DOI: 10.1093/bib/bbag452
- Keywords: framework
- Source URL: <https://doi.org/10.1093/bib/bbag452>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1093%2Fbib%2Fbbag452>

Abstract: Bioinformatics software plays a critical role in clinical applications such as cancer screening and genetic disease diagnosis, where comprehensive quality management is essential for ensuring the accuracy and reliability of downstream analysis. However, current validation practices rely heavily on manually designed simulation experiments, which are labor-intensive and limited in their ability to systematically identify potential quality risks under certain scenarios. In this study, we first construct a benchmark by simulating subtle implementation-level defects in bioinformatics programs. We then propose BioTester, an oracle-based automated testing framework that integrates software testing techniques to support more comprehensive quality assessment of bioinformatics software. BioTester integrates retrieval-augmented LLMs with a differential testing strategy to address the long-standing oracle problem in bioinformatics software testing, demonstrating superior defect-detection performance over existing methods on the constructed benchmark. Finally, applying BioTester to real-world bioinformatics software demonstrates its practical effectiveness and highlights the value of automated testing as a generalizable complement to existing validation practices for improving software reliability in biomedical applications.

## Bridging data scarcity and explainability in EEG based Alzheimer’s prediction using cGAN augmented GATv2-LSTM networks
- Source: Scientific Reports (journals)
- Date: 2026-09-16T00:00:00+00:00
- Authors: Neha Prerna Tigga, Nandini Kumari, Fady Alnajjar
- Journal: Scientific Reports
- DOI: 10.1038/s41598-026-70929-0
- Source URL: <https://doi.org/10.1038/s41598-026-70929-0>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1038%2Fs41598-026-70929-0>

Abstract: Accurate and early diagnosis of Alzheimer’s Disease (AD) is still a key factor in the treatment of the disease and patient health. The proposed model for the current work is an EEG-based multi-classification model which is tested on two separate datasets with different diagnostic classes (Dataset I: Healthy Control (HC), Mild Cognitive Impairment (MCI), Alzheimer Disease (AD); Dataset II: Healthy Control (HC), Fronto-temporal Dementia (FTD), Alzheimer Disease (AD)). A Conditional GAN (cGAN) data augmentation technique was used to tackle the issue of class imbalance, especially MCI, FTD and HC were underrepresented. Analysis of the fidelity of the generated synthetic EEG signals was conducted using the Wasserstein distance and Kernel Density Estimation (KDE). The results reveal that there is a high similarity between real and synthetic signals for Dataset I in frontal channels (Fp1, F3) and with larger channel-specific deviations for the MCI class. Whereas, for Dataset II, both FTD and HC data showed good overall similarity between the real and synthetic EEG signals, with some larger channel-specific deviations observed for the FTD class. The proposed Hybrid GATv2-LSTM model integrates both spatial and temporal information from EEG signals for Alzheimer’s disease classification. Specifically, Graph Attention Network v2 (GATv2) learns the spatial relationships among EEG channels by modeling their connectivity, while Long Short-Term Memory (LSTM) captures the temporal dynamics of brain activity. Furthermore, adaptive attention fusion and residual connections are incorporated to effectively combine spatial and temporal features, enhance feature learning, and improve classification performance. Under a sample-level 80/20 split, it achieved the classification accuracies of 94.60% on Dataset I and 89.19% on Dataset II, better than the model without data-augmentation and standalone model. To improve the interpretation of the model, Integrated Gradients (IG) was employed to evaluate the contribution of individual EEG channels to the classification outcomes. The results demonstrated that the channels over the frontal, parietal, central, and occipital areas were the most relevant for the model, while the channels over other areas had a relatively small impact. This finding is consistent with previous studies on EEG-based Alzheimer’s detection. This proof-of-concept study is a promising exploration of the feasibility of using validated synthetic EEG data and interpretable deep learning models to achieve data efficient EEG-based AD classification.

## Calibrated structural homology transfer yields putative molecular functions for domains of unknown function in four model proteomes
- Source: bioRxiv (preprints)
- Date: 2026-09-16
- Categories: Proteins & structural biology, Tools & resources
- Authors: Vedanayagam, J.
- DOI: 10.64898/2026.09.11.750891
- Source URL: <https://doi.org/10.64898/2026.09.11.750891>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.11.750891>

Abstract: Improvements in computational protein structure prediction have enabled searches for remote homologs of proteins whose molecular function remain unknown. However, the reliability of functional annotations from such searches has not been systematically quantified. In this study, we assembled a time-split benchmark from Pfam families annotated as domains of unknown function (DUFs) in Pfam 28.0 and were subsequently assigned a function in the current release (Pfam 38.0). These retrospective-DUFs provide ground truth for assessing functional annotation from remote homology searches. In our benchmark analysis, we paired retrospective-DUFs with a difficulty-matched arm from known domains to distinguish query difficulty from the method's performance. Across four model proteomes (yeast, C. elegans, Drosophila, and mice), Foldseek searches against AlphaFold/Swiss-Prot, PDB100, and CATH50 recovered the later-assigned function for 15.9% of retrospective DUF queries, compared with 30.1% of matched known-domain queries, after masking for self-family and self-clan level hits to remove circularity. At a fixed confidence cut-off (qTM \[≥\] 0.5), 55.4% of informative calls on retrospective DUFs were incorrect, establishing an error model for prospective use. Furthermore, in our comparison of tools for homology searches, Foldseek outperformed MMseq2-based sequence search but was not statistically separable from the ESM-2 protein language model embedding baseline. Applying the calibrated structural homology search pipeline to 296 currently unannotated DUF queries in Pfam 38.0 yielded 50 confident, specific functional assignments. We provide the benchmark and ranked candidates as a resource to facilitate functional studies.

## CatRange enables robust prediction of enzyme variant kinetic regimes
- Source: PNAS Nexus (journals)
- Date: 2026-09-16T00:00:00Z
- Categories: Proteins & structural biology, Tools & resources
- Authors: Karuna Anna Sajeevan, A. Osinuga, B. Arunraj, Sakib Ferdous, Nabia Shahreen, M. S. Noor, Shashank Koneru, Laura Mariana Santos-Correa, Rahil Salehi, N. B. Chowdhury, R. Aryee, Brisa Calderon-Lopez, Supantha Dey, A. Mali, Rajib Saha, Ratul Chowdhury
- Journal: PNAS Nexus
- DOI: 10.1093/pnasnexus/pgag309
- External ID: c4d4877f81153deaf60d025177076ae5e1dcb28b
- Source URL: <https://doi.org/10.1093/pnasnexus/pgag309>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1093%2Fpnasnexus%2Fpgag309>

Abstract: Predicting enzyme kinetics directly from sequence remains a central challenge in computational biology, particularly in resolving the effects of mutations at catalytically essential residues. Existing models frequently overlook the functional consequences of such perturbations, defaulting to wild-type predictions even in cases of substantial activity loss, thereby limiting their reliability for enzyme design and mechanistic inference. Here, we introduce CatRange, a machine learning framework trained on CatLog-27k, a human-in-the-loop, AI-agent trustworthy dataset of 27,176 in vitro enzyme–substrate kinetic records created by a systematic audit and correction of BRENDA and SABIO-RK. All mutant entries are manually reconciled against 2,158 source articles. CatRange reframes kinetic prediction from exact numerical regression into classification over log10-spaced bins for catalytic turnover (kcat) and substrate affinity (KM), matching the order-of-magnitude scale at which experimental enzyme kinetic measurements are commonly interpreted. This biologically grounded formulation mitigates assay-level variability while preserving distinctions among functional catalytic and binding states. Using joint enzyme–substrate representations and gradient-boosted classifiers, CatRange predicts kinetic ranges for wild-type and mutant enzymes across standard held-out, out-of-distribution, and few-shot mutation settings. The model shows robust order-of-magnitude recovery with class-balanced discrimination and captures mutation-induced movement across kinetic regimes, including losses associated with perturbation of annotated catalytic residues. CatRange detects non-enzyme sequence inputs and emphasizes rigorous data curation, transparent training data dissemination (CatLog), biochemically informed task formulation, and balanced evaluation metrics. These position CatRange as an interpretable, mutation-sensitive framework with utility in enzyme engineering and kinetic metabolic modeling.

## Circle of Willis-Guided Localization for Simultaneous Detection and Classification of Large Vessel Occlusions in Brain CTA
- Source: Neuroinformatics (journals)
- Date: 2026-09-16T00:00:00+00:00
- Categories: Biological imaging
- Authors: Valeriia Abramova, Arnau Oliver, Uma M. Lal-Trehan Estrada, Rachika E. Hamadache, Paola Martínez Arias, Jordi Freixenet, Mikel Terceño, Yolanda Silva, Xavier Lladó
- Journal: Neuroinformatics
- DOI: 10.1007/s12021-026-09817-x
- Source URL: <https://doi.org/10.1007/s12021-026-09817-x>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1007%2Fs12021-026-09817-x>

Abstract: Large vessel occlusions (LVOs) are blockages in the brain’s major arteries that can cause severe neurological damage. Rapid and accurate detection using computed tomography angiography (CTA) is critical for timely stroke treatment. Here, we present a fully automated approach that detects LVOs and classifies the affected vessel simultaneously. Our method incorporates a spatial prior by using Circle of Willis (CoW) segmentation as additional input, guiding the model to anatomically relevant regions. We evaluated the two strategies, the global approach using the full CTA volume, and the local one focused on CoW regions. Both achieved high performance. Detection sensitivity was 0.97 at 0.20 false positives per image for the global approach, and 0.97 at 0.13 false positives for the local approach. Classification accuracy reached 94% and 91% for global and local strategies, respectively. Importantly, the local approach was 3.3 $$\\times $$ faster, offering a computationally efficient solution, a critical advantage in acute stroke care, where every minute impacts patient outcomes.

## Comparison of cell-cycle gene expression dynamics and mRNA kinetics across mouse and human pluripotent systems
- Source: bioRxiv (preprints)
- Date: 2026-09-16
- Categories: Genomics & sequence analysis, Single-cell & spatial, Tools & resources
- Authors: Nariya, M. K., Santiago-Algarra, D., Zanardelli, G., Boudjelthia, I. K., Ye, T., Thibault-Carpentier, C., Jarriault, S., Molina, N.
- DOI: 10.64898/2026.09.14.751360
- Source URL: <https://doi.org/10.64898/2026.09.14.751360>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.14.751360>

Abstract: Cell-cycle remodeling is fundamental to pluripotency and lineage commitment, yet whether its transcriptional and post-transcriptional architecture is conserved across species and developmental states has remained unresolved. Here we introduce Ciclopes, a biology-informed deep-learning framework that resolves continuous cell-cycle phase and phase-dependent mRNA transcription and degradation directly from single-cell RNA sequencing. Applying Ciclopes across six mouse and human pluripotent stem-cell systems spanning naive and primed states, we uncover striking divergence in transcriptional complexity and oscillatory control: mouse systems sustain elevated baseline expression of core cell-cycle regulators, while human systems trade higher baseline expression for larger oscillatory amplitude. Strikingly, mRNA degradation timing remain far more conserved across systems than transcription timing, exposing post-transcriptional regulation as a stable evolutionary backbone. As human iPSCs differentiate into definitive endoderm, cells progressively exit the cell cycle, cell-cycle-coupled gene networks contract, and surviving regulators oscillate with larger amplitude. Ciclopes establishes a general framework for dissecting how pluripotent cells tune proliferation across evolutionary and developmental transitions.

## Computational investigation of circuit mechanisms underlying short-latency responses to cortical stimulation
- Source: bioRxiv (preprints)
- Date: 2026-09-16
- Categories: Computational neuroscience
- Authors: Kumaravelu, K., Yu, G. J., Aberra, A. S., Sommer, M. A., Peterchev, A. V., Grill, W. M.
- DOI: 10.64898/2026.09.11.750854
- Source URL: <https://doi.org/10.64898/2026.09.11.750854>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.11.750854>

Abstract: Transcranial magnetic stimulation (TMS) over the primary motor cortex (M1) elicits a series of high frequency volleys termed D- and I-waves measured epidurally in the corticospinal tract of awake humans. Further, intracortical microstimulation (ICMS) in M1 of non-human primates evokes D- and I-wave responses similar to those observed in TMS. The cortical circuits and mechanisms involved in the generation of D- and I-waves by stimulation of M1 remain unclear. Here, we implemented computational models of cortical columns with laminarly-organized biophysically-based neurons, following existing models published in the literature: (1) M1 - single compartment (SC), (2) M1 - multi-compartment (MC), (3) primary auditory cortex (A1) - MC, and (4) primary somatosensory cortex (S1) - MC. The network connectivity of each model represented wiring found in the respective cortical regions. The direct effects of stimulation-induced electric fields were modeled as activation of different proportions of pyramidal neurons (PNs) across layers, and dose response curves were constructed for layer 5 (L5) PNs. Both the M1 and A1 models reproduced D- and I-waves, with the magnitude of I-waves increasing with higher recruitment of layer 2/3 and layer 5 PNs. The S1-MC model evoked rhythmic firing activity but with timings mismatched to experimental I-waves. The models replicated the experimentally observed effects of pharmacological agents on I-waves, and virtual lesions of specific neural populations across models revealed plausible microcircuit explanations for the first and later I-waves. This comprehensive comparison of models across multiple cortical regions identified consistent mechanisms underlying the cortical response to TMS and contributes to the refinement of computational strategies for optimizing stimulation paradigms.

## Connecting CKAN and Galaxy: from data repositories to analysis and back
- Source: Galaxy (feeds)
- Date: 2026-09-16T00:00:00+00:00
- Categories: Blog
- Source URL: <https://galaxyproject.org/news/2026-09-16-ckan-integration/>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fgalaxyproject.org%2Fnews%2F2026-09-16-ckan-integration%2F>
- Abstract: not stored for this record.

## Corpus-wide causality: Algorithm design & application for aggregating gene-disease causal evidence
- Source: bioRxiv (preprints)
- Date: 2026-09-16
- Categories: Systems & networks
- Authors: Bansal, N., Parsodkar, A. P., Pathak, A., Narayanan, M.
- DOI: 10.64898/2026.05.08.723796
- Keywords: gene regulatory, algorithm
- Source URL: <https://doi.org/10.64898/2026.05.08.723796>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.05.08.723796>

Abstract: Identifying causal relationships and distinguishing them from associations is a central scientific endeavor with many applications; knowing causal links between genes and diseases, for instance, can focus drug discovery on curing diseases beyond just symptom management. Despite several studies on automatically extracting relations between entities from large biomedical literature corpora like PubMed, only a few studies extract causal relations from abstracts and even fewer summarize corpus-level evidence for causal links. Recently, Large Language Models (LLMs) have been increasingly deployed to summarize biomedical information and extract relations; however, there is a distinct lack of explicit benchmarking comparing these generalized LLM-based methods against specialized, domain-aware frameworks for corpus-wide causal inference. In this work, we develop a method to infer Corpus-Wide Causal Score (CWCS) of a gene-disease (G-D) pair by integrating two pieces of evidence: (i) network-based causal signals in a prior gene regulatory network, quantified as a CWCS-Net score using an existing multilayer network centrality algorithm; and (ii) corpus-wide literature evidence, quantified as a CWCS-TD (TD for Truth Discovery) score using a newly-developed TD algorithm. Our CWCS-TD (scoring) algorithm jointly and iteratively estimates causal scores for multiple G-D pairs while modeling the reliability of PubMed abstracts co-mentioning them; and represents an advance in the field of TD algorithms due to its incorporation of bibliometric features of publications to address the challenge of sparsity of abstracts that assert a G-D causal relation. Using OMIM as an external expert-curated reference to evaluate classifications of G-D pairs as causal or not, our CWCS method achieved a causal class F1 score of 0.600 across ten diseases, outperforming both LLMs, GPT-4o and MMed-Llama 3 (this performance trend also persists when using area under the precision-recall curve as the evaluation metric). Both LLMs exhibit high recall accompanied by comparatively low precision, resulting in lower causal class F1 scores (0.505 for GPT-4o and 0.522 for MMed-Llama 3) due to large number of false positive predictions. Taken together, these evaluations and other ablation studies show the promise of our carefully designed algorithm in collating and integrating evidence of biomedical causal relations from both network- and literature-based sources, thereby supporting its broader applicability.

## Cyclic parthenogenesis helps populations cross fitness valleys
- Source: bioRxiv (preprints)
- Date: 2026-09-16
- Categories: Evolution & metagenomics
- Authors: Yang, Z., Hardy, N. B.
- DOI: 10.64898/2026.01.12.699030
- Source URL: <https://doi.org/10.64898/2026.01.12.699030>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.01.12.699030>

Abstract: Cyclic parthenogenesis is a form of occasional sexual reproduction. Here, we compare the evolvability of cyclic parthenogens and obligate sexuals in adaptive scenarios that entail the crossing of a fitness valley. We divide the valley crossing process into two phases. In the escape phase, a population produces genotypes that lie on the other side of the fitness valley. In the establishment phase, such genotypes are refined and promoted to high frequency. With individual-based models, we find that although cyclic parthenogens are slower to produce escape genotypes, they more efficiently establish such genotypes, and therefore more rapidly complete fitness-valley crossings. In obligate sexuals, genetic and mutational variance and covariances, G and M, are readily shaped by selection. In cyclic parthenogens, the shapes of G and M are less responsive to selection but have relatively little effect on the rate of fitness-valley crossing. In our model, the more adaptive evolution of G and M in obligate sexuals is not enough to overcome the establishment advantage of cyclic parthenogens. Nevertheless, the relative recalcitrance against selection on G and M in cyclic parthenogens could be an evolutionary cost that helps explain its paradoxical rarity.

## Deterministic and ANN-based modelling of measles transmission using real epidemiological data.
- Source: Scientific reports (journals)
- Date: 2026-09-16
- Authors: Kamil Shah, Changqing Du, Ali Akgül, Farad Sameer Alshammari, Sanaa Ahmed Bajri, Hamiden Abd El-Wahed Khalifa
- Journal: Scientific reports
- DOI: 10.1038/s41598-026-53188-x
- External ID: 42744821
- Source URL: <https://doi.org/10.1038/s41598-026-53188-x>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1038%2Fs41598-026-53188-x>

Abstract: This study aims to develop a nonlinear SVLIQR epidemic model to investigate measles transmission dynamics by incorporating vaccination, latency, quarantine, and treatment effects. The qualitative properties of the model, including positivity and boundedness of solutions, are established. The basic reproduction number is derived using the next-generation matrix method, and the existence of disease-free and endemic equilibrium points is examined. The local stability of the disease-free equilibrium is analysed using the Routh-Hurwitz criterion, while global stability is studied through the Castillo-Chávez approach. In addition, model parameters are estimated using the least squares method based on reported measles cases in China from 2005 to 2017. To improve the numerical approximation and capture the nonlinear dynamics of the system, an ANN-based solver coupled with the ode45 scheme is implemented. The results show that the model fits the reported data well, and the ANN framework provides highly accurate and convergent approximations under different epidemiological parameter settings. Numerical simulations further indicate that higher vaccination and quarantine rates significantly reduce measles transmission. Overall, the proposed framework provides a useful and reliable tool for understanding measles dynamics and assessing effective intervention strategies.

## Development and validation of a DNA methylation-based classifier for CNS tumors using a large Chinese cohort (n = 1,581)
- Source: Scientific Reports (journals)
- Date: 2026-09-16T00:00:00+00:00
- Categories: Genomics & sequence analysis, Tools & resources
- Authors: Xueping Xiang, Dexiang Huang, Hui Zhang, Lihua Guo, Xiaojing Ma, Linlin Ying, Jinghong Xu, Jimin Shao
- Journal: Scientific Reports
- DOI: 10.1038/s41598-026-70119-y
- Source URL: <https://doi.org/10.1038/s41598-026-70119-y>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1038%2Fs41598-026-70119-y>

Abstract: The clinical implementation of DNA methylation profiling for central nervous system (CNS) tumors faces significant challenges in China, particularly regarding the development and validation of locally applicable classifiers. We developed MethAI-CNS, a locally executable DNA methylation-based classifier for CNS tumors that strictly follows the established DKFZ classification framework. The classifier was trained on global databases augmented with 1,368 local Chinese samples, covers 122 subclasses. The model was independently validated on 213 Chinese samples, which included challenging pathological consultation cases and medulloblastoma molecular subtyping cases, and further validated on an independent public cohort (GSE289137, n = 687), with comparisons against DKFZ v12.8. MethAI-CNS achieved an overall accuracy of 0.988 and an AUC of 0.993 in 5 × 5 cross-validation. At a threshold of 0.9, the sensitivity and specificity were 0.976 and 0.964, respectively. In the Chinese validation cohort, high-confidence predictions from MethAI-CNS showed 99% concordance with the DKFZ classifier v12.8, demonstrating high fidelity to the DKFZ reference framework. On the independent GSE289137 cohort, MethAI-CNS achieved a concordance rate of 99% (543/546) with DKFZ v12.8 at the 0.9 threshold, with an accuracy of 0.956 and an AUC of 0.949. Among diagnostically challenging consultation cases and medulloblastoma cases, methylation-based classification led to revision rates of 32% and 5% of the original histopathological diagnoses, respectively. MethAI-CNS is a locally validated methylation classifier for CNS tumors in a Chinese cohort that demonstrates robust performance in diagnosing complex neuroepithelial tumors and medulloblastoma. It provides a reliable tool for supporting pathological practice, particularly in regions with restricted access to reference models.

## Development of a novel aggregated deep learning framework for small biological datasets using overlapping subsequences
- Source: Scientific Reports (journals)
- Date: 2026-09-16T00:00:00+00:00
- Categories: Genomics & sequence analysis
- Authors: Mohammad Ali Abbasi-Vineh, Naser Farrokhi, Pär K. Ingvarsson
- Journal: Scientific Reports
- DOI: 10.1038/s41598-026-69140-y
- Source URL: <https://doi.org/10.1038/s41598-026-69140-y>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1038%2Fs41598-026-69140-y>

Abstract: The development of deep learning models and techniques to harness the potential of data mining in biological sequences with limited data availability is a vital need. This necessity arises from unique biological characteristics and technical constraints that limit access to sufficient quantities of high-quality data across many biological and genetic fields. Building on a data augmentation strategy that generates overlapping augmented subsequences, an aggregated CNN-LSTM-Attention-Residual architecture was developed for analysis at the original full-length sequence level. The model was independently tested and validated on three sets of regulatory sequences from evolutionarily diverse organisms, chloroplasts sharing a common ancestor, and prokaryotic sequences, comprising 100, 50, and 100 sequences per group, respectively, within each dataset. Trained on augmented subsequences, the model effectively aggregated and transferred high-level features back to the original full-length sequences. It achieved high performance with approximately 96% accuracy, recall, precision, F1-score, and AUC at the subsequence level. The model correctly distinguished the sequences of the distinct groups at the full-length sequence level across all datasets. The strong performance confirmed that an ensemble of features learned from short, overlapping subsequences provide an informative signature for classifying full-length regulatory regions. The combination of first-level training on augmented subsequences with dual-level validation—evaluating both subsequences and their corresponding full-length sequences— provides a framework for advance model development in limited-data biological sequence analysis.

## dicast: a machine learning method for accurate structural variant detection from short-read sequencing data
- Source: Genome Biology (journals)
- Date: 2026-09-16T00:00:00+00:00
- Categories: Genomics & sequence analysis, Tools & resources
- Authors: Nico Alavi, M-Hossein Moeinzadeh, Jakob Hertzberg, Uirá Souto Melo, Lion Ward Al Raei, Paolo Infantino, Maryam Ghareghani, Marco Savarese, Stefan Mundlos, Martin Vingron
- Journal: Genome Biology
- DOI: 10.1186/s13059-026-04280-y
- Source URL: <https://doi.org/10.1186/s13059-026-04280-y>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1186%2Fs13059-026-04280-y>

Abstract: Structural variants are a common cause of human diseases, but their detection from short-read sequencing remains challenging, despite being the technology underlying most clinical workflows. We present dicast , a machine-learning method that scores SV calls from short-read data using alignment and genomic-context features. dicast is trained on a new multi-technology ground truth built from nine samples, with extensive manual curation. It outperforms existing short-read callers and consensus approaches, recovering substantially more true positives at high precision. We also demonstrate dicast’s applicability for diagnostics, identifying all pathogenic variants in multiple disease cohorts, and 20% more candidate pathogenic deletions than consensus approaches.

## Division-Specific Organization of a Shared Functional Scaffold in the Early-Life Human Brain
- Source: bioRxiv (preprints)
- Date: 2026-09-16
- Categories: Computational neuroscience, Tools & resources
- Authors: Hu, D., Cheng, J., Han, K., Wu, Z., Yin, W., Sun, Y., Liu, J., Hung, S.-C., Wang, L., Cohen, J. R., Lin, W., Li, G.
- DOI: 10.64898/2026.09.14.751555
- Source URL: <https://doi.org/10.64898/2026.09.14.751555>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.14.751555>

Abstract: Early cognition and behavior emerge from coordinated maturation of functional systems spanning the brain, making cross-division integration essential for defining whole-brain architecture and understanding its role in neurodevelopment. Yet early human functional development is still largely investigated through cortical networks or through isolated, coarsely resolved noncortical structures, leaving unclear how this integrated architecture is organized across the cerebral cortex, subcortex and cerebellum. Moving beyond structure-isolated and adult-derived maps, we create a reproducible, fine-grained Early-Life Whole-Brain Functional Parcellation (UNC-ELF) spanning birth to six years. Using this framework, we identify a shared functional scaffold expressed through distinct division-specific modes: differentiated and dual-polarity cortical organization, spatially graded cerebellar organization with blended functional transition zones, and nucleus-constrained subcortical mosaics of multi-system affiliation. Across early childhood, the scaffold is selectively refined through heterogeneous, nonlinear connectivity trajectories, with major inflections concentrated in infancy. Age-, sex- and prospective cognition-related variation is differentially encoded across distinct connectomic components. Together, these findings reframe early functional brain development as coordinated refinement of a shared but nonuniform whole-brain scaffold and establish UNC-ELF as a unified reference for resolving typical and atypical functional organization across the developing brain.

## Dual-contrastive learning for spatial domain identification in spatial transcriptomics with STAMGC
- Source: Genome Research (journals)
- Date: 2026-09-16T00:00:00+00:00
- Categories: Genomics & sequence analysis
- Authors: Zhuoyue Zhang, Qianmao Wen, Junlin Xu, Yajie Meng, Feifei Cui, Zilong Zhang
- Journal: Genome Research
- DOI: 10.1101/gr.281999.126
- Source URL: <https://doi.org/10.1101/gr.281999.126>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1101%2Fgr.281999.126>

Abstract: Spatial transcriptomics (STs) have become a valuable approach for understanding the growth and development of organisms. Despite the recent emergence of numerous ST models, accurately identifying spatial domains remains challenging owing to the trade-off between preserving local details and reducing noise. Here, we introduce STAMGC, which is a dual-contrastive learning framework built upon graph convolutional networks. This model leverages regional and topological contrastive learning to jointly optimize the model, effectively reducing the noise in spatial domain identification and enhancing the extraction of detailed features. In this study, Gaussian smoothing, originally developed in the image processing field, is introduced to process ST data, providing a foundation for region-level contrastive learning by mitigating spatial discontinuities of gene expression signals. Experimental results indicate that STAMGC outperforms existing methods across multiple data sets according to comprehensive evaluations. Furthermore, STAMGC not only identifies finer structures in the mouse brain but also brings new discoveries for human breast cancer research.

## Dynamic medical knowledge graph updating method based on LLM decision control
- Source: Scientific Reports (journals)
- Date: 2026-09-16T00:00:00+00:00
- Categories: Systems & networks
- Authors: Keyong Hu, Jiabin Hu, Menghuan Yue, Yan Sun, Ben Wang
- Journal: Scientific Reports
- DOI: 10.1038/s41598-026-71802-w
- Keywords: pathway
- Source URL: <https://doi.org/10.1038/s41598-026-71802-w>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1038%2Fs41598-026-71802-w>

Abstract: Medical knowledge exhibits significant timeliness and evidence-based dependency characteristics, and traditional static Medical Knowledge Graphs are difficult to adapt to the continuous evolution of clinical knowledge. Existing updating methods mostly rely on offline reconstruction or generative models, which suffer from update lag and potential medical safety risks, especially in primary healthcare scenarios with limited data and lack of continuous expert validation. To address this issue, this paper proposes a Dynamic Medical Knowledge Graph updating framework based on Large Language Model (LLM) decision control, modeling the knowledge updating process as a sequential decision problem driven by multi-source feedback and stability constraints. Different from existing methods, this paper innovatively transforms the LLM from a “knowledge generator” into an “update strategy controller,” which, at each time step, performs operations such as Add, Revise, Decay, or Reject on candidate knowledge based on task effectiveness, knowledge consistency, and temporal evolution feedback, thereby achieving a safe and controllable updating mechanism at the mechanism level. In the respiratory disease scenario, small-sample temporal evolution experiments (2,000 records) constructed based on cMedQA v2.0 data and real outpatient cases demonstrate that the proposed method improves Micro-F1 to 78.95% in medical question answering tasks, reduces the error knowledge introduction rate to 3.68%, and achieves a graph utility value of 0.8174, reaching a good balance between performance improvement and structural stability. The results indicate that, within the tested respiratory disease scenario using DeepSeek-R1 as the decision controller, the performance improvement of Dynamic Knowledge Graphs does not depend on knowledge generation capability, but rather on the decision control capability of the updating process. This method provides a new technical pathway for constructing safe and sustainable clinical knowledge systems, especially suitable for resource-constrained primary healthcare environments.

## Dynamics of mutators of arbitrary dominance in humans
- Source: bioRxiv (preprints)
- Date: 2026-09-16
- Categories: Evolution & metagenomics, Mathematical biology & statistics
- Authors: Saeidi, M., Sella, G., Przeworski, M., Milligan, W. R.
- DOI: 10.64898/2026.09.10.750760
- Source URL: <https://doi.org/10.64898/2026.09.10.750760>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.10.750760>

Abstract: Recent findings in humans and other species have revealed the presence of "mutator" alleles that increase germline mutation rate across the genome. Such mutators are expected to be selected against because of the additional deleterious alleles that they generate, to a degree that will depend on how much they increase the mutation rate in heterozygotes and homozygotes. To describe their dynamics, we develop a population genetic model of mutation rate modifiers with arbitrary dominance coefficients, in which fitness effects stem from additional germline mutations. We then use it to interpret findings for the seven human mutators identified to date. For six of the seven known mutators, the observed frequencies are well fit by the model and thus consistent with purifying selection arising solely due to their effects on germline mutation rates, although also consistent with a wide range of parameters; in two, the observed frequencies are readily explained by purely recessive fitness effects. The exception is a variant in MUTYH, which is more common than predicted under plausible parameters, for reasons that remain unclear. We also use the model to explore what types of mutators are most likely to be discovered in parents, based on identifying offspring with unexpectedly high numbers of de novo mutations. Although the two first mutators identified by this approach seem to be recessive, our modeling suggests that, for the same effect size, semi-dominant mutators are much more likely to be detected. These findings therefore imply that there are many more modifier sites with recessive effects than semi-dominant ones. More generally, our model provides a framework for interpreting properties of mutators in humans and other species and for learning about the genetic architecture of germline mutation rate variation.

## eCOMET: an R package for evaluating metabolic diversity and enrichment from LC-MS/MS data to test ecological hypotheses from individuals to ecosystems.
- Source: The New phytologist (journals)
- Date: 2026-09-16T00:00:00Z
- Categories: Tools & resources
- Authors: Min-Soo Choi, Dale L. Forrister, Guillaume J. Dury, Kyo Bin Kang, Brian E. Sedio, Youngsung Joo
- Journal: The New phytologist
- DOI: 10.1111/nph.71589
- External ID: 63a4f281731e07dd589c789f94e12c6e76064d78
- Source URL: <https://doi.org/10.1111/nph.71589>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1111%2Fnph.71589>

Abstract: Methods in metabolomics have grown exponentially in recent years, providing new insight into the ecological function and evolutionary impact of diverse plant metabolites. Metabolomics requires a command of numerous tools, the outputs of which are typically integrated through in-house custom code that presents a workflow bottleneck and a barrier to entry for researchers in ecology, evolution, and behavior who may benefit from adding a metabolomics perspective to their research. We introduce eCOMET, an R package for integrating and harmonizing the outputs of common metabolomics bioinformatics tools and conducting statistical analyses and data visualization methods useful for ecological metabolomics. Our package combines metabolome feature metadata with quantification tables (e.g. mzmine), feature dissimilarity matrices (e.g. modified cosine and DreaMS), and feature annotations (e.g. SIRIUS) into a cohesive R data object to facilitate downstream analyses, including the calculation of diversity and disparity metrics and differential accumulation analysis. We provide two tutorials, each explores herbivore-induced Arabidopsis thaliana metabolome and 10 co-occurring tropical trees species metabolomes. Our goal is to make metabolomics accessible to a wider range of researchers in ecology, evolution, and behavior to unlock the potential of ecological metabolomics to generate novel insight into these fields.

## Empirical Estimation of Ambient Contamination in Combinatorial Single-Cell Methods Using Multi-Reference Mapping
- Source: bioRxiv (preprints)
- Date: 2026-09-16
- Categories: Genomics & sequence analysis, Single-cell & spatial, Tools & resources
- Authors: Gomez-Cano, F., Jiang, L., Welch, J. D., Marand, A. P.
- DOI: 10.64898/2026.09.11.750809
- Source URL: <https://doi.org/10.64898/2026.09.11.750809>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.11.750809>

Abstract: Droplet-based microfluidics and combinatorial indexing (scifi-ATAC and scifi-RNA) have made single-cell experiments massively scalable. However, higher-order multiplexing complicates data quality, introduces noise, and affects the potential for biological discovery. Here, we show that ambient chromatin accumulates through the experimental workflow and distorts chromatin profiles, most drastically in low-depth nuclei and in minority populations. Standard cell calling relies heavily on read count thresholds, while existing decontamination methods generally operate on aggregated count matrices rather than the underlying reads. We introduce scifi-demux, for preprocessing scifi-ATAC libraries, and AmbientMapper, a generative model that maps reads competitively against multiple references, learns the ambient profile from empty and low-complexity barcodes, and separates nuclei from background and singlets from doublets by Bayesian Information Criterion. Using interspecies ground truth experiments, published multi-genotype libraries, and simulated read-level synthetic barcodes in which every contaminating read is traceable, we show that calls are robust to parameter variation and stable across designs, achieving a wrong-genome rate of 0.19% on a 26-genome reference panel. Finally, we evaluated the impact of removing contaminants, showcasing how AmbientMapper rescues low-depth nuclei discarded by standard pipelines and restores biological structure obscured by contamination.

## Enhancing interpretability in metabolomics: ranking metabolites by their impact on graph neural network predictions
- Source: BMC Bioinformatics (journals)
- Date: 2026-09-16T00:00:00+00:00
- Categories: Systems & networks
- Authors: Francisco Traquete, Carlos Cordeiro, Marta Sousa Silva, António E. N. Ferreira
- Journal: BMC Bioinformatics
- DOI: 10.1186/s12859-026-06633-7
- Keywords: metabolomics, pathways, interpretability
- Source URL: <https://doi.org/10.1186/s12859-026-06633-7>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1186%2Fs12859-026-06633-7>

Abstract: A key step in the biological interpretation of untargeted metabolomics data is the ranking of compounds by importance, usually using statistical significance or impact on classification as the basis for importance assignment. However, current approaches treat metabolites as independent variables and rarely incorporate the network structure underlying biochemical relationships. This disconnect limits the ability of existing ranking methods to highlight groups of interconnected metabolites that jointly contribute to biological differences. Here we developed a new method where Formula Difference Networks are used as inputs to predictive Graph Neural Network models. After fitting, metabolites were ranked by their impact on sample class prediction probabilities. When applied to three benchmark datasets, this ranking highlighted subnetworks of metabolites, favouring connectivity as a driving factor for importance assignment. This led to an enrichment of the number of edges between the top ranked compounds. Furthermore, using datasets containing metabolites with simulated significance, we found that there was a clear bias for assigning higher importance to nodes in connected subgraphs. This new strategy for Graph Neural Network interpretability, is an alternative to common approaches based on mapping of important metabolite onto biological pathways supported by enrichment analysis.

## Estimating effects on cumulative incidence probabilities by direct polytomous regression and polytomous log-odds product
- Source: Biometrics (journals)
- Date: 2026-09-16T00:00:00+00:00
- Categories: Mathematical biology & statistics
- Authors: Shiro Tanaka, Thomas H Scheike
- Journal: Biometrics
- DOI: 10.1093/biomtc/ujag160
- Source URL: <https://doi.org/10.1093/biomtc/ujag160>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1093%2Fbiomtc%2Fujag160>

Abstract: In competing-risks analysis, modeling each cause separately may yield specification of cumulative incidence functions that are not jointly coherent, because the resulting cause-specific probabilities need not satisfy the natural sum-to-one constraint. We address this problem by introducing a direct polytomous regression approach that models all causes jointly and enforces coherence through a reparameterization based on polytomous log-odds products. Our approach is applicable to semiparametric models with common multiplicative effects over time as well as models focused on a specific time point. For estimation under right-censoring, we develop stratified inverse probability of censoring weighted (IPCW) estimators for the effect parameters. Within a specified class of augmented IPCW estimators, the proposed estimators attain the minimum asymptotic variance under the stated regularity conditions, without the computational burden of deriving augmentation terms for each cause. The utility of our coherent modeling is demonstrated through simulation studies and its applications to a cohort study of type 2 diabetes and a randomized trial of prostate cancer.

## Estimating home range size under spatial constraints: a comparative approach using the semi-aquatic European mink (Mustela lutreola)
- Source: bioRxiv (preprints)
- Date: 2026-09-16
- Authors: Bodinier, R., Aulagnier, S., Bressan, Y., Beaubert, R., Fournier-Chambrillon, C., Devillard, S., Fournier, P.
- DOI: 10.64898/2026.04.13.718143
- Source URL: <https://doi.org/10.64898/2026.04.13.718143>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.04.13.718143>

Abstract: Accurate home range knowledge is essential for conserving species that are highly constrained by spatial features. The critically endangered European mink (Mustela lutreola) is a wetland specialist whose movements are constrained along rivers or in wetlands. In dendritic landscapes, conventional home range estimators such as Minimum Convex Polygons tend to include unsuitable areas in estimated home ranges. Using VHF telemetry data from 16 individual-years tracked in France between 1996 - 1999 and 2020 - 2022, we compared four methods: Kernel Density Estimator (KDE), an adaptative sphere-of-influence local convex hull (a-LoCoH), a newly developed Ecological Home Range method (EHR), and a Generalized Additive Model (GAM) approach integrating hydrographic covariates. Our objective is to determine which method best accounts for the European mink's specialization in wetlands, considering the spatial distribution of locations. Evaluation with a wetland-specific metric showed KDE consistently overestimated range extent and included unsuitable areas, and a-LoCoH yielded mixed results, but these indicated that the method was not effective in excluding unused areas. It was EHR and GAM methods that aligned more closely with ecological constraints. We therefore recommend GAM because it matches our objective and has the capacity to integrate additional environmental variables. Using the GAM, male home ranges averaged 3,074 ha - 26 times larger than female ranges (116 ha) - and were significantly larger in river than marsh landscapes. These are the largest ranges reported for the species. Large spatial requirements heighten vulnerability to road fatality and predation, both significant threats for remaining French populations. Our findings highlight the need for conservation strategies that integrate precise, spatial-constraint-based range estimates. The GAM method offers a robust, adaptable framework for managing European mink and other semi-aquatic species in complex landscapes.

## Estimating Mutual Information and Pearson Correlation on Neural Evoked Responses
- Source: Neuroinformatics (journals)
- Date: 2026-09-16T00:00:00+00:00
- Categories: Computational neuroscience
- Authors: Anni Hukari, Silvia Federica Cotroneo, Riitta Salmelin
- Journal: Neuroinformatics
- DOI: 10.1007/s12021-026-09784-3
- Source URL: <https://doi.org/10.1007/s12021-026-09784-3>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1007%2Fs12021-026-09784-3>

Abstract: In neural evoked responses, small variations in the timing or duration of responses can be observed when the same functional response is recorded in different trials, different experimental conditions or by different sensors. These variations limit the ability of correlation-based methods to detect similarities between signals. Mutual information (MI) provides an alternative similarity measure, capable of capturing both linear and non-linear dependencies, yet its practical use is hindered by lack of consensus on estimators for continuous data and the limited understanding of the behavior of the estimators on realistic signals. In this work, we investigate how to estimate the similarity of neural evoked responses by systematically comparing sample Pearson correlation with three of the most common MI estimators. We describe their behavior using both simulated signals and real magnetoencephalographic data. In the simulations, the estimators are tested against a set of transformations that depict realistic changes in neural evoked responses. Subsequently, we propose guidelines for defining adaptive lower bounds on the similarity estimates and analyzing the similarity rankings induced by the different estimators. Our findings reveal trade-offs between measures sensitivity and different signal properties. We confirm that Pearson correlation is reliable in describing linear relationships for low-noise signals, and we identify parameter settings that stabilize MI estimators, enabling them to capture complex signal dependencies. Together, these results introduce practical parameter choices and thresholding strategies for mutual information and provide guidance for selecting and interpreting similarity measures in the analysis of neural evoked time series.

## Estimating protein isoform abundances with \[Formula: see text\].
- Source: Proceedings of the National Academy of Sciences of the United States of America (journals)
- Date: 2026-09-16
- Categories: Genomics & sequence analysis, Single-cell & spatial, Proteins & structural biology, Mathematical biology & statistics, Tools & resources
- Authors: Lorenzo Testa, Lambertus Klei, Alesia Rengle, Anastasia K Yocum, David A Lewis, Bernie Devlin, Kathryn Roeder, Matthew L MacDonald
- Journal: Proceedings of the National Academy of Sciences of the United States of America
- DOI: 10.1073/pnas.2614319123
- External ID: 42748139
- Source URL: <https://doi.org/10.1073/pnas.2614319123>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1073%2Fpnas.2614319123>

Abstract: A single gene can encode multiple versions of a protein, dubbed isoforms, with varying functionality. Cellular control of isoform abundances is critical for multiple aspects of biology and is only partially regulated by transcript levels. While long-read sequencing facilitates transcript quantification, quantifying the resulting protein isoforms on a large scale is a major challenge, complicating biological interpretation of transcript alterations. Standard "bottom up" mass spectrometry can assess only short portions of isoforms called peptides, and these peptides often map onto more than one isoform. We introduce \[Formula: see text\] (Protein isoform Abundance Quantification), a Bayesian method that leverages multiomic information from the peptidome and transcriptome to provide accurate estimates of isoform abundance even when peptide mapping is ambiguous. \[Formula: see text\] offers several advantages over existing methods in a unified framework. It provides uncertainty quantification, integrates multiomic information for improved accuracy, and provides a rigorous framework for hypothesis testing. Extensive simulations show that \[Formula: see text\] consistently outperforms competing methods in detecting differentially abundant protein isoforms and estimating their abundances. We use \[Formula: see text\] to investigate differences in isoform abundance levels between people with schizophrenia and control subjects, confirming a long-held hypothesis that levels of the C4A isoform of Complement Component 4 are increased in schizophrenia while C4B is not. These results demonstrate that \[Formula: see text\] can identify significant variations in isoform abundance levels not previously possible.

## Evaluation of methods for AlphaFold-based integrative modeling
- Source: bioRxiv (preprints)
- Date: 2026-09-16
- Categories: Proteins & structural biology
- Authors: Majila, K., Viswanath, S.
- DOI: 10.64898/2026.09.13.751323
- Source URL: <https://doi.org/10.64898/2026.09.13.751323>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.13.751323>
- Code: <https://github.com/isblab/af_im>

Abstract: Motivation Recent methods enable incorporation of experimental data into AlphaFold for predicting structures consistent with the data. However, the applicability of these methods for integrative modeling remains to be determined. It is unclear how these methods balance the input experimental data with the learned structural priors. Results We assess the performance of state-of-the-art AlphaFold-based integrative modeling methods, including AlphaLink2, Boltz2, and GRASP, on a dataset of 37 complexes based on crosslinking data. We evaluate these methods based on their ability to predict structures that satisfy the input crosslinks. We further assess the robustness of these methods to noise in the crosslinking data. Finally, we probe their ability to predict distinct states using crosslinks from multiple states. Overall, our study highlights the limitations of the AF-based IM methods and points to directions for future improvements. Availability and implementation All scripts used to obtain the predictions and perform the analysis in this study are available at https://github.com/isblab/af\_im.

## Exploring Nile Red and machine learning for microplastics detection in Tridacna maxima
- Source: PLOS One (journals)
- Date: 2026-09-16T00:00:00+00:00
- Categories: Proteins & structural biology
- Authors: Irène Godéré, Taiamiti Edmunds, Nabila Gaertner-Mazouni, Fiona Gimenez, Pascal Wong-Wah-Chung, Stéphanie Lebarillier, Magalie Baudrimont, Chloé Pupier, Nicolas Maihota, Jean-Claude Gaertner
- Journal: PLOS One
- DOI: 10.1371/journal.pone.0357014
- Source URL: <https://doi.org/10.1371/journal.pone.0357014>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1371%2Fjournal.pone.0357014>

Abstract: Small Island Developing States (SIDS) face unique challenges for microplastics (MPs) monitoring due to limited infrastructure and resources. In this context, we propose and test innovative approaches toward a standardized, low-cost methodology for quantifying MPs in SIDS. We evaluate the giant clam T. maxima as a bio-integrator, combining Nile red (NR) fluorescence staining with automated machine-learning detection. We optimized a digestion protocol using KOH and HNO 3 for T. maxima viscera, and developed a DAPI-guided multi-spectra composite imaging approach based on triband fluorescence (DAPI, FITC, TRITC), to enhance polymer detection while reducing blooming artifacts. A semi-automated annotation pipeline using Labkit interactive segmentation with CLIP/UMAP clustering efficiently generated training data from 6711 fluorescence images. A U-Net model was trained on composite images to segment fluorescent particles. The workflow was applied to giant clams from three French Polynesian islands (Makemo, Hao, Tubuai), and NR-based estimates were validated against µFTIR spectroscopy. The model achieved F1-scores of 0.741 for giant clam samples and 0.657 for controls, comparable to human annotation (F1 = 0.680). MPs were detected across all islands, with highest concentrations in gills (16.9–52.7 particles·g −1 wet weight) compared to viscera (2.5–11.0 particles·g −1 ww). µFTIR validation revealed that NR overestimates MP counts (µFTIR: 0.80 ± 0.16 particles·g −1 ww at Tubuai), primarily due to false positives from proteins, cellulose, and stearates. In Tubuai, polyamide (28.9%), PVC (12.6%), and polystyrene (10.7%) were the dominant polymers, suggesting contributions from fishing gear, agriculture, and household waste. While NR-based quantification overestimates absolute MP counts, the automated pipeline demonstrates potential for high-throughput image processing, reproducible sample analysis, and methodological standardization. This workflow represents a first step toward scalable, low-cost approaches for MPs monitoring in insular systems, highlighting areas for further calibration and optimization. Future work should refine fluorescence thresholds and expand validation across species and locations.

## FetchPA: a guided, end-to-end solution for local ATAC-Seq data processing and analyses
- Source: bioRxiv (preprints)
- Date: 2026-09-16
- Categories: Tools & resources
- Authors: Fetch, D. R., Soshnev, A. A.
- DOI: 10.64898/2026.09.10.750716
- Source URL: <https://doi.org/10.64898/2026.09.10.750716>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.10.750716>

Abstract: Local genome accessibility strongly correlates with activity of cis-regulatory elements, and Assay for Transposase-Accessible Chromatin coupled with next-generation sequencing (ATAC-Seq) has emerged as method of choice to profile chromatin accessibility in both healthy and pathogenic conditions. The introduction of streamlined protocols and manufacturer kits has made this technique accessible to labs of a variety of disciplines, background, and research interests. Many bioinformatics tools have been created for the quality control, mapping, and visualization of ATAC-seq data, however these tools require familiarity with shell scripting, version control, UNIX directory structure, Python and/or R. Several pipelines for the processing of ATAC-seq data have been developed, yet even with these tools, bioinformatic analyses represents a bottleneck between wet-lab protocol execution and graphical representation of differentially accessible regions. To address this problem, we assembled FetchPA, an intuitive pipeline which allows users with virtually no scripting and version control experience to install and manage all software for end-to-end analyses of ATAC-Seq data. FetchPA handles both local and public repository sources of sequencing data, executes standard QC benchmarks, and handles genome assembly and alignment using industry-standard PEPATAC pipeline. Further, it guides the user through the identification of differentially accessible regions and allows basic exploratory analyses via a dialogue interface. FetchPA operates in Windows Subsystem for Linux (WSL) and is installed via a single script that handles all individual tools, as well as their dependencies and updates, reference genome annotations and system resource allocation.

## FlavoTyper: a genome-based in-silico serotyping tool for the fish pathogen Flavobacterium psychrophilum
- Source: bioRxiv (preprints)
- Date: 2026-09-16
- Categories: Genomics & sequence analysis, Tools & resources
- Authors: Mbarki, S., Debeljak, P., Carpentier, M., Jolley, K. A., Haddad, N., Rochat, T., Duchaud, E.
- DOI: 10.64898/2026.09.14.751350
- Source URL: <https://doi.org/10.64898/2026.09.14.751350>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.14.751350>

Abstract: Flavobacterium psychrophilum is a devastating pathogen of fish reared in freshwater worldwide. Serotyping is a relevant method for epidemiological surveillance and outbreak detection, as well as for a better understanding of host-pathogen interactions. Serological diversity may also have important consequences for the selection of appropriate strains for vaccine development and for selective breeding for increased disease resistance. F. psychrophilum serotyping relies on structural variations in the O-polysaccharide (O-PS) moiety of the cell surface lipopolysaccharide (LPS). However, conventional serotyping is costly, labor-intensive and requires significant technical expertise. Moreover, divergent scheme proposals highlighted the absence of harmonization among laboratories. In this context, the development of an mPCR-based serotyping scheme targeting wzy genes greatly improved the reliability and standardization of serotyping. Nevertheless, the proposed mPCR scheme did not capture the entire diversity of genomic variability. The aim of this study was to establish a robust and publicly available tool for F. psychrophilum genome-based serotyping. Extensive genome analysis of the O-antigen biosynthesis locus allowed the identification of biomarkers enabling the development of FlavoTyper, an in-silico-based serotyping tool. The FlavoTyper tool was evaluated on all F. psychrophilum genome assemblies publicly available, providing sound and sensitive predictions and easily interpretable results. When applied to a curated collection of publicly available genomes, the in-silico O-types assigned by the tool were statistically significantly associated with host fish species, confirming previous studies (coho salmon with O:0, rainbow trout with O:1 and O:2, and ayu with O:3) and their distribution across MLST clonal complexes revealed that the O-antigen locus is frequently rearranged independently of the core-genome lineage, consistent with the extensive recombination that shapes the evolution and genomic diversity of this species.

## Flow Orchestrated Regulatory Genomics Engine (FORGE): A Configurable Nextflow Pipeline for End-to-End snMultiome Analysis
- Source: bioRxiv (preprints)
- Date: 2026-09-16
- Categories: Genomics & sequence analysis, Single-cell & spatial, Systems & networks, Tools & resources
- Authors: Solano, L. E., Rahimzadeh, N., Shi, Z., Swarup, V.
- DOI: 10.64898/2026.09.10.750690
- Source URL: <https://doi.org/10.64898/2026.09.10.750690>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.10.750690>

Abstract: MOTIVATION Single-nucleus resolution multiome (snMultiome) assays concurrently profile gene expression and chromatin accessibility in the same nucleus. Yet regulatory inference from such analyses are difficult to scale, audit, and reproduce; moreover, as a field, snMultiomics and its toolset remains far from standardized. To address these challenges we developed FORGE, a configureable Nextflow workflow that carries paired data from raw counts and fragments through regulatory network inference with a comprehensive differential testing suite. Execution is containerized, tracks provenance, robust to interruption, optimized for cluster-based compute environments, and allows for nuanced customization of specific processes. SUMMARY Single-nucleus multiome assays jointly profile gene expression and chromatin accessibility, yet their analysis typically requires bespoke chaining of modality-specific tools, creating barriers to reproducibility, scalability, and regulatory interpretation. We present FORGE (Flow Orchestrated Regulatory Genomics Engine), a configurable workflow that automates standalone snRNA-seq and snATAC-seq analysis, integrates the pair through complementary linear and nonlinear latent-variable models, and carries them through regulatory-network inference and differential testing. We evaluated FORGE on four human and mouse datasets spanning blood, brain, and kidney and two multiome chemistries, including a twelve-sample CRND8 Alzheimers disease cohort. We report cross-modal agreement alongside missing-modality reconstruction and an accounting of computational cost. In the Alzheimer's cohort, FORGE nominated a glial Mef2c-associated program defensible across expression, co-accessibility, footprinting, and eRegulon evidence. In human PBMC, FORGE layered evidence models also provide nuanced interpretations that largely corroborate previously published regulatory links while also proposing an additional CD83 myeloid module.

## From fluke to fragment: A multifaceted method for molecular sex identification and mitochondrial haplotyping from environmental DNA samples
- Source: Methods in Ecology and Evolution (journals)
- Date: 2026-09-16T00:00:00Z
- Categories: Genomics & sequence analysis, Evolution & metagenomics
- Authors: L. K. Rodriguez, Sandra Schallhart, Philipp Hobmeier, T. Curran, S. Pérez-Jorge, Rui Prieto, Cláudia Oliveira, Mónica A. Silva, B. Thalinger
- Journal: Methods in Ecology and Evolution
- DOI: 10.1111/2041-210x.70400
- External ID: d79e7513f790275f601548aeb0e481675b203c84
- Source URL: <https://doi.org/10.1111/2041-210x.70400>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1111%2F2041-210x.70400>

Abstract: Environmental DNA (eDNA) analyses have become a powerful tool for non‐invasive biodiversity monitoring, yet the applicability of population‐genetic approaches to environmental samples remains largely unexplored. Even when genetic traces originate from a single individual, low target DNA concentrations and amplification or sequencing artefacts can compromise downstream genetic inferences. Here, we present a novel approach for obtaining demographic insights and lineage‐level mitogenomic information from aquatic eDNA samples collected near vertebrate individuals, while assessing the utility of paired tissue sampling for benchmarking eDNA‐based population genetic analyses. Paired eDNA and tissue samples were collected during sperm whale ( Physeter macrocephalus ) encounters in the Azores. Samples were screened for the presence of vertebrate eDNA and analysed with a novel molecular sex identification assay. Additionally, long‐range PCR was used to amplify up to five mitochondrial DNA fragments (~3–4 k bp) before subsequent sequencing on an Oxford Nanopore Technologies platform. A stringent three‐tier filtering framework capable of identifying true mitogenomic variation across eDNA samples was developed for maximum recovery of genetic diversity at the haplogroup level. By validating eDNA samples via their paired tissues, parameter values were optimized to maximize concordance and minimize spurious variant calls. Sexing was successful for 50% of eDNA samples, with 96% concordance to paired tissues and marine vertebrate DNA concentration significantly predicted sexing success. Further, Medaka polishing produced high identity mitochondrial consensus sequences (>16 kb) from eDNA samples. Across filtering regimes in the framework, curated SNP panels comprising up to 453 high‐confidence mitochondrial SNPs resolved 19 haplogroups, with 93% concordance between eDNA and tissue samples. An intermediate bioinformatics filtering strategy maximized biologically accurate haplogroup recovery while minimizing sequencing artefacts, providing the most reliable lineage‐level inferences. This integrative approach demonstrates that targeted nuclear assays combined with long‐range mitochondrial sequencing can recover individual‐level genetic information from aquatic eDNA. By defining analytical thresholds governing success and demonstrating how paired tissue benchmarking can calibrate eDNA‐based population genetic insights for future applications, the framework advances non‐invasive genetic monitoring of populations via eDNA and enables population‐level monitoring and conservation of endangered and genetically‐vulnerable species.

## Genetic drift decouples Fisherian trait-preference coevolution from genetic correlations in finite populations
- Source: bioRxiv (preprints)
- Date: 2026-09-16
- Categories: Evolution & metagenomics, Mathematical biology & statistics
- Authors: Xu, K.
- DOI: 10.64898/2026.09.10.750770
- Source URL: <https://doi.org/10.64898/2026.09.10.750770>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.10.750770>

Abstract: A central mechanism of sexual selection theory is Fisher's process, which refers to the coevolution of male traits and female preferences, in which preference is indirectly selected through genetic association with male trait alleles built up via mate choice. However, empirical studies often fail to detect strong trait-preference genetic correlations, raising doubts about the importance of Fisherian selection in nature. Notably, the theoretical expectation that genetic correlations are essential for trait-preference coevolution through Fisherian selection derives largely from models assuming infinitely large populations, whereas real populations are finite. Using population genetic models, I show that interactions between Fisherian selection and genetic drift can fundamentally decouple trait-preference genetic correlations from the evolution of male traits and female preferences. Genetic drift generally reduces expected trait-preference correlations but simultaneously promotes the expected increase in female preference frequency beyond deterministic predictions. Consequently, in populations of realistic sizes, trait-preference correlations may often be weak or even negative, particularly when recombination among trait and preference loci is infrequent, but substantial trait-preference coevolution can still occur. Therefore, the strength of genetic correlations may be an unreliable indicator of the extent of trait and preference coevolution, offering a potential resolution to a longstanding dilemma in sexual selection theory.

## Genolator enables protein function interpretation using a multimodal large language model fusing genomic and structural interpretation with natural language interaction
- Source: Genome Biology (journals)
- Date: 2026-09-16T00:00:00+00:00
- Categories: Proteins & structural biology
- Authors: Martin Danner, Tanhim Islam, Matthias Begemann, Florian Kraft, Miriam Elbracht, Ingo Kurth, Jeremias Krause
- Journal: Genome Biology
- DOI: 10.1186/s13059-026-04274-w
- Source URL: <https://doi.org/10.1186/s13059-026-04274-w>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1186%2Fs13059-026-04274-w>

Abstract: Background Decoding the genetic code to unveil its genome functionality is a monumental task which would greatly advance the understanding of disease mechanisms and development of targeted treatments. Although large language models (LLMs) have transformed natural language processing across diverse domains, translating the complex language of DNA into human-readable form remains challenging due to genomic data complexity and unexplored regions of the human genome. Current language models either are capable of processing natural language or the genomic code. Models fusing both aspects are largely lacking. Results Here we present Genolator, a multimodal large language model that integrates embeddings from DNA sequences, amino acid sequences, and protein structures with natural language queries. Fine-tuned on over 365,000 question–answer pairs generated using abstracted Gene-Ontology (GO) terms, Genolator effectively answers queries regarding protein subcellular localization, molecular function, and biological processes. Evaluation demonstrates high accuracy in confirming or denying protein function associations, outperforming baseline models such as openly available allrounder LLMs like GPT 4.1 as well as smaller domain-specific models integrating knowledge from a protein structure transformer. Explorations of Genolator’s hidden states unveil a biologically and linguistically plausible organization of its learned representations. Analysis of the attention heads of the underlying language model and an ablation study provide evidence for a benefit of the multi-modal approach. Conclusion Genolator enhances accessibility to genomic information by enabling natural language interaction with protein data, facilitating biological discovery, and clinical research. It represents a step towards bridging genomic code and human language through the integration of a multimodal LLM.

## Genome-scale perturbation signatures from primary human CD4+ T cells improve genetics-based prioritization of immune drug targets
- Source: bioRxiv (preprints)
- Date: 2026-09-16
- Categories: Tools & resources
- Authors: Han, A. L., Slotnik, M., Fox, D. A., Dhindsa, R. S., Gudjonsson, J. E., Kahlenberg, J. M., Welch, J.
- DOI: 10.64898/2026.09.12.748195
- Source URL: <https://doi.org/10.64898/2026.09.12.748195>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.12.748195>

Abstract: While human genetic evidence improves drug program success, target selection largely relies on associational features. Even functional features are mainly observational, capturing disease associations rather than the consequence of perturbing genes in the disease-relevant human cell type. To move beyond association, we present IGNITE (Immune Genomics and fuNctional Integration for Target Enrichment), a framework that prioritizes immune drug targets by integrating human genetic priors with features from genome-scale perturb-seq in 22 million primary human CD4+ T cells and polarized T-helper subset differential expression. IGNITE is a semi-supervised machine learning model that applies positive-unlabeled learning to approved immune targets, then ranks 19,502 protein-coding genes to prioritize new candidates. In 727 in-trial genes held out from training, functional genomics increased target enrichment among the top 50 nominations from 2.7- to 4.8-fold for immune trial targets at any phase, and from 4.5- to 5.9-fold for targets in Phase III. These performance gains were immune-specific, as IGNITE outperformed genetic comparators on predicting immune-exclusive targets but not on cardiac-exclusive targets. Using temporal validation with labels frozen in 2014, IGNITE outperformed the genetics-only model in ranking 134 genes that subsequently entered immune trials. The functional genomics layer also surfaced two novel, pharmacologically tractable candidates, ELOVL6 and RUVBL1, elevating their rankings from outside the top 1,000 into the top 50. These findings highlight that perturbational signatures in the disease-relevant human cell type harbor target-relevant signals beyond human genetics, offering a generalizable route to indication-specific prioritization as perturbation atlases expand. Full IGNITE scores are publicly available at https://ignite.eecs.umich.edu/

## Genomic characterization of antifungal resistance patterns in Candida auris clade I: A large-scale analysis of 647 global genomes
- Source: Northern Clinics of Istanbul (journals)
- Date: 2026-09-16T00:00:00Z
- Categories: Genomics & sequence analysis, Single-cell & spatial, Systems & networks, Evolution & metagenomics
- Authors: Ayhan Tosunoglu, Ozleyis Konyali, Mehmet Demirci
- Journal: Northern Clinics of Istanbul
- DOI: 10.14744/nci.2026.40040
- External ID: 35baf984a40064844ac5ced4d0fa32a7da2321db
- Keywords: genomic, genomes, genome, single nucleotide, pathways, phylogenetic
- Source URL: <https://doi.org/10.14744/nci.2026.40040>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.14744%2Fnci.2026.40040>

Abstract: OBJECTIVE: Candidozyma auris has emerged as a global “urgent threat” characterized by multidrug resistance and high mortality. While six distinct lineages have been identified, Clade I (South Asian) is the primary driver of global nosocomial outbreaks and exhibits the most profound resistance profiles. This study aims to provide a high-resolution genomic characterization of Clade I, focusing on the consolidation of resistance mechanisms and virulence factors that support its global dominance. METHODS: We performed a comprehensive genomic analysis focusing on a cohort of 647 Clade I isolates, selected from a total of 662 global C. auris genomes available in public repositories. Using a standardized bioinformatic pipeline, we conducted core-genome single-nucleotide polymorphisms-based phylogenetic reconstruction, non-synonymous mutation profiling of key resistance genes (TAC1B, FKS1, ERG11/6/3), and functional mapping of virulence-related pathways (ALS4, secretable aspartyl proteases \[SAP5\], LIP1). The remaining isolates from Clades II to VI were utilized as comparative reference groups to identify clade-specific signatures. RESULTS: Analysis of the Clade I cohort (n=647; 97.73% of the total dataset) revealed a significant consolidation of resistance markers. The Y132F and K143R substitutions in ERG11 were near-ubiquitous, often co-occurring with specific TAC1B variants (A640V, V742A). Notably, 24.32% of the Clade I isolates demonstrated a highly synchronized “genomic armor,” characterized by the simultaneous presence of TAC1B (A:YTDQ/A:GSVG), FKS1 (S:SL), and deletions in ERG6/ERG3. Viru-lence profiling showed high-frequency conservation of biofilm-associated (ALS4) and proteolytic (SAP5) genes, suggesting a synergistic evolution of resilience and pathogenicity. CONCLUSION: This study delineates the genomic landscape of C. auris Clade I, highlighting how the consolidation of multiple resistance and virulence markers contributes to its clinical success. The high frequency of multidrug-resistant genotypes within this lineage mandates a transition toward genome-led surveillance and personalized antifungal stewardship.

## Genomic foundation model-derived disruption profiling links somatic mutations to cancer biology and clinical outcomes
- Source: medRxiv (preprints)
- Date: 2026-09-16
- Categories: Genomics & sequence analysis, Proteins & structural biology
- Authors: Nayak, A., Lee, T.-R., Agarwal, V., Georgakopoulos-Soares, I.
- DOI: 10.64898/2026.09.15.26363174
- Keywords: genomic, genomics, dna, genome, chromatin, splicing, foundation model
- Source URL: <https://doi.org/10.64898/2026.09.15.26363174>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.15.26363174>

Abstract: Cancer genomics has concentrated on individual mutations, overlooking whether somatic mutations can accumulate to produce partial, gene-level disruption with biological and clinical consequences. Sequence-to-function models can quantify these effects directly from DNA sequence. Here, we use AlphaGenome and AlphaMissense to quantify the disruption imposed by somatic mutations across 8,800 patients and 33 cancer types from The Cancer Genome Atlas. At the individual-variant level, recurrent hotspot mutations showed substantially larger predicted protein-level effects, whereas non-hotspot mutations exhibited larger regulatory effects across most cancer types. We then aggregated the variant-level predictions to construct patient-gene disruption profiles capturing transcriptional activity, chromatin accessibility, transcription factor binding, and splicing. These profiles were gene- and modality-specific, and reflected tissue of origin, cancer type, and microsatellite-instability status while retaining information beyond tumor mutational burden. Among patients lacking recurrent hotspot mutations in a given cancer gene, higher predicted disruption was associated with overall survival, with the strongest and most consistent signal observed for chromatin accessibility. In an independent treatment-annotated cohort, gene-level disruption was also associated with survival within treatment-defined subgroups. Together, these findings show that recurrent hotspots are enriched for strong predicted protein-level effects, whereas regulatory consequences are distributed more broadly across other variants, supporting a continuous, multidimensional view of cancer-gene perturbation beyond discrete drivers.

## Global transfer learning pipeline for protein disease association in Alzheimer's disease.
- Source: Computational biology and chemistry (journals)
- Date: 2026-09-16
- Categories: Proteins & structural biology
- Authors: Hansa J Thattil, Arunkumar M N
- Journal: Computational biology and chemistry
- DOI: 10.1016/j.compbiolchem.2026.109408
- External ID: 42763975
- Source URL: <https://doi.org/10.1016/j.compbiolchem.2026.109408>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1016%2Fj.compbiolchem.2026.109408>

Abstract: Researchers have prioritized the study of protein-disease associations to decode triggers of clinical pathology and isolate high-value targets for drug development. Comprehensive modeling of genetic network dynamics is equally vital for advancing our functional understanding of these disorders. In this study, we applied a hierarchical, transfer-learning framework to address the challenge of protein-disease association prediction, specifically targeting scenarios with limited labeled data for specific diseases. We used Alzheimer's disease as a case study to demonstrate the efficacy of our proposed Global Transfer Learning Pipeline model. We combined the embeddings generated from protein-protein interactions along with protein-cluster association and protein sequences to train the proposed model. We addressed the scarcity of reliable negatives by employing PU learning strategies with deep fusion architecture to ensure robustness of the model. The ordinal regression was integrated to the learning pipeline to learn granular confidence levels for protein-disease associations which was later fed into the stacked meta model. The model also uses techniques of hyperparameter optimization to enhance the prediction performance. Our model achieved a weighted F1 score of 96% with average AUC 0.852 and AUPRC 0.967 which outperforms the individual gradient boosting, tree models and deep neural network models. These results demonstrate the effectiveness of the proposed model in predicting protein-disease associations for Alzheimer's disease.

## GSCA-UNet: a gated spatial-channel attention U-net for accurate skin lesion segmentation
- Source: Scientific Reports (journals)
- Date: 2026-09-16T00:00:00+00:00
- Categories: Biological imaging
- Authors: Lazhen Zhou, Wenjie Ou, Xiuhua Chen, Lingyan Zhang
- Journal: Scientific Reports
- DOI: 10.1038/s41598-026-67896-x
- Source URL: <https://doi.org/10.1038/s41598-026-67896-x>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1038%2Fs41598-026-67896-x>

Abstract: Medical image segmentation is a fundamental component of computer-aided diagnosis, where automatic skin lesion segmentation serves as a critical upstream task by providing pixel-wise delineations for subsequent analysis. Deep encoder-decoder architectures, such as U-Net and its variants, have advanced skin lesion segmentation. However, the task remains challenging. Key difficulties include low contrast between lesions and skin, ambiguous or irregular boundaries, acquisition artifacts. Furthermore, lesions exhibit large intra-class variations in scale, shape and texture. In this work, we propose GSCA-UNet (Gated Spatial-Channel Attention UNet), a novel segmentation architecture for skin lesions. At its core, a gated spatial attention block adaptively models horizontal and vertical spatial dependencies by multi-scale 1D convolutions with learnable gating to strengthen lesion boundaries while suppressing background clutter. A cross-dimensional attention interaction block establishes bidirectional guidance between spatial and channel attention through multi-head self-attention and gating fusion. Extensive experiments on public benchmark datasets demonstrate that GSCA-UNet consistently outperforms competitive baselines across multiple metrics and exhibits superior robustness on challenging cases with blurry borders, irregular shapes, and severe artifacts.

## HaloUMI: Physics-informed analysis of inhibition halo assays
- Source: bioRxiv (preprints)
- Date: 2026-09-16
- Categories: Biological imaging, Tools & resources
- Authors: Pembery, A., Nadir, H. H., MacDonald, C., Leake, M. C.
- DOI: 10.64898/2026.08.08.743694
- Source URL: <https://doi.org/10.64898/2026.08.08.743694>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.08.08.743694>

Abstract: Quantification of microbial growth inhibition is central to assays from antibiotic susceptibility to antifungal sensitivity, yet existing approaches struggle with irregular inhibition zones and variation in microbial lawn density. Here, we present Halo Unbiased Measurement of growth Inhibition (HaloUMI), an open-source Python graphical interface for automated, high-throughput analysis of lawn-based microbial assays. HaloUMI integrates robust image processing with physics-informed modelling to quantify inhibition zones without assuming circular geometry, enabling analysis of uniform and irregular halo phenotypes. Using diffusion- and growth-based physical modelling, HaloUMI experimentally validates a correction for variation in microbial lawn density, a major source of assay variability that can confound quantitative comparison of inhibition phenotypes. Validation using simulations and yeast killer-toxin assays demonstrates precise, reproducible measurement across diverse conditions. HaloUMI is applicable to multiple assay formats, including microbial interaction, mating, and conventional disc-diffusion assays, providing an accessible and generalisable framework for quantitative analysis of microbial growth inhibition.

## Hierarchical feature binding in a spiking neural network model of the primate ventral visual pathway
- Source: PLOS Computational Biology (journals)
- Date: 2026-09-16T00:00:00+00:00
- Categories: Computational neuroscience
- Authors: Brian Gardner, Patrick T. McCarthy, Joseph Chrol-Cannon, Dan F. M. Goodman, Simon R. Schultz, Giovanni Lo Iacono, Simon M. Stringer
- Journal: PLOS Computational Biology
- DOI: 10.1371/journal.pcbi.1014752
- Source URL: <https://doi.org/10.1371/journal.pcbi.1014752>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1371%2Fjournal.pcbi.1014752>

Abstract: Feature binding - how the brain encodes which features are part of other features to form representations of the coherent objects we perceive - remains an unsolved problem in neuroscience. Despite progress towards a solution, major theories either lack detailed explanations at the neuronal level or rely on biologically unrealistic simplifications, and none adequately account for the representation of hierarchical information, which is crucial to our perception of the world. To address this, a solution termed binding by polychrony has been proposed to explain how hierarchical feature relationships may be encoded at the neuronal level in a biologically realistic system. This theory relies on a phenomenon known as polychronization, where groups of neurons fire in precisely coordinated, time-locked sequences, leading to the emergence of regularly repeating spatiotemporal patterns that might encode these relationships. In this study, we explore binding by polychrony through simulations of a spiking neural network that closely aligns with the structural organisation of the primate ventral visual pathway, incorporating bottom-up, top-down, and lateral synaptic connections. By exposing the network to collections of related 2D object shapes from ecologically realistic datasets and applying spike-timing-dependent plasticity, the network self-organises such that individual neurons respond selectively to specific shape features. Furthermore, the network exhibits polychronization, giving rise to repeating spatiotemporal patterns, some of which form circuits that encode hierarchical feature relationships. Notably, these circuits are robust, even with the randomised, Poisson-distributed spike timings that represent the visual stimuli in the input layer. These results provide evidence for binding by polychrony as a feasible solution to the feature binding problem, and characterise the mechanism by which it may function. This mechanism can guide experimentalists in identifying such circuits in vivo , and could also be utilised in computer vision systems to capture more information and improve robustness to adversarial inputs.

## HyLnc: a hybrid deep learning and feature-based approach for long non-coding RNA prediction.
- Source: RNA biology (journals)
- Date: 2026-09-16
- Categories: Genomics & sequence analysis, Proteins & structural biology
- Authors: Amrit Venkatesan, Prashasti Sinha, Jolly Basak, Ranjit Prasad Bahadur
- Journal: RNA biology
- DOI: 10.1080/15476286.2026.2731913
- External ID: 42716909
- Source URL: <https://doi.org/10.1080/15476286.2026.2731913>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1080%2F15476286.2026.2731913>

Abstract: Long non-coding RNAs (lncRNAs) play important roles in gene regulation, development and disease, yet accurate identification of lncRNAs from transcriptomic data remains a major computational challenge. Existing methods often rely either on handcrafted sequence features or deep learning approaches, each with their inherent limitations in capturing the full complexity of RNA sequences. In this study, we proposed HyLnc, a computational framework that integrates transformer-based contextual embeddings with biologically meaningful sequence features for improved lncRNA prediction. A custom BERT-based model was first pre-trained on a large corpus of metazoan RNA sequences using a masked language modelling strategy to learn contextual nucleotide dependencies. The model was subsequently fine-tuned on curated datasets of lncRNAs and protein-coding transcripts and 256-dimensional deep sequence embeddings were extracted. Parallelly, 348 handcrafted features, including ORF characteristics, untranslated region (UTR) properties, nucleotide composition and Fickett scores, were computed. A multi-stage feature selection strategy was applied to identify the most informative features, resulting in optimized hybrid feature sets. Multiple machine learning classifiers were evaluated, with the RF model achieving the best performance. The proposed framework attained an accuracy of 91.30%, F1-score of 91.23% and MCC of 82.60 on an independent validation dataset, outperforming several existing lncRNA prediction tools. Thus, HyLnc demonstrates that integrating deep contextual representations with biologically interpretable features enhances lncRNA prediction. This approach provides a robust and scalable solution for large-scale transcriptome annotation and can be extended to other sequence-based prediction.

## Identification and characterization of bacterial repeat-in-toxin adhesins using long-read genome analysis
- Source: Bioinformatics Advances (journals)
- Date: 2026-09-16T00:00:00+00:00
- Categories: Genomics & sequence analysis
- Authors: Thomas Hansen, Laurie A Graham, Blake P Soares, Daniel Lee, Justin R Gagnon, Trina Dykstra-MacPherson, Shuaiqi Guo, Peter L Davies
- Journal: Bioinformatics Advances
- DOI: 10.1093/bioadv/vbag272
- Source URL: <https://doi.org/10.1093/bioadv/vbag272>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1093%2Fbioadv%2Fvbag272>

Abstract: Gram-negative bacteria attach to host surfaces using ligand-binding domains at the distal tips of fibrillar Repeats-in-ToXin adhesins. Blocking these initial interactions could prevent colonization, biofilm formation, and infection. To achieve this, the adhesins must be identified and, in species encoding multiple adhesins, the predominant type determined. These adhesins are often the largest proteins encoded by a genome (1,500-15,000 aa) and are frequently misannotated as incomplete or pseudogene products because their repetitive sequences complicate short-read genome assemblies. Our bioinformatic pipeline collects predicted proteins from long-read assemblies and clusters them according to similarity in their C-terminal regions, where ligand-binding domains are typically located. Adhesins are identified by their size and domain architecture and modelled using AlphaFold3. Analysis of multiple strains from seven species identified 35 adhesin isoforms distributed across 16 loci, exhibiting diverse combinations of putative binding domains such as carbohydrate-binding modules and von Willebrand factor A-like domains. Similar adhesins were sometimes shared among species through common ancestry or horizontal gene transfer. Three species encoded an adhesin of unknown function that lacked an obvious ligand-binding domain.

## Identification of broadly tumour-reactive γδ TCRs from multiple myeloma.
- Source: Nature (journals)
- Date: 2026-09-16T00:00:00Z
- Categories: Single-cell & spatial, Tools & resources
- Authors: Michael St. Paul, Liam D. Hendrikse, Fan Ying, Bryan E. Snow, P. Luo, Simone Helke, Logan K. Smith, Arwa Hilal, D. Abelman, Nisha Ramamurthy, Hayley Nault, E. Wei, Matthew J. Gold, Chantal Tobin, S. Lien, Yi Liu, Wen-Jing Zhou, Xin Zhang, Dat Nguyen, Oluwatobi Agbede, S. Pedersen, Jenna Eagles, M. Saunders, Thorsten Berger, A. Wakeham, D. Scott, Dalam Ly, C. B. de Campos, Gu W. Liang, Chun-Xing Zheng, Wesley V. Wilson, E. Masih-Khan, Darrell White, A. McCurdy, M. Louzada, R. Kotb, Michael P. Chu, Stephen Parkin, Donna Reece, E. Gul, R. Tiedemann, Trevor J. Pugh, N. Hirano, Pamela S. Ohashi, A. Stewart, S. Trudel, Tak W. Mak
- Journal: Nature
- DOI: 10.1038/s41586-026-11055-9
- External ID: 37909ff098fb8839f6ddab5f61f987f9d7a3cc02
- Source URL: <https://doi.org/10.1038/s41586-026-11055-9>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1038%2Fs41586-026-11055-9>

Abstract: γδ T cells are becoming increasingly appreciated for their antitumour capacity and role in mediating responses to immune checkpoint blockade1-3. Unlike classical αβ T cells, the degree to which γδ T cells rely on their T cell receptors (TCRs) to induce antitumour responses remains unclear. The challenge of distinguishing γδ T cells with tumour-reactive TCRs from bystander γδ T cells limits our understanding of tumour-reactive γδ T cell biology and the translation of their TCRs into immunotherapeutics. Here we present PreGame, a machine-learning algorithm capable of identifying tumour-reactive γδ T cells from single-cell CITE sequencing data. We use PreGame to identify tumour-reactive γδ T cells from patients with multiple myeloma or other solid cancers, and confirm the specificity of their TCRs to tumour cells. Clinically, we demonstrate that expansion of tumour-reactive γδ T cells is an early biomarker of response in patients with multiple myeloma receiving combination therapy with belantamab mafodotin. We also identify a γδ TCR epitope in the ubiquitously expressed HLA-C protein and a logic gate that enables tumour immunosurveillance. Thus, PreGame is a versatile tool that can accelerate our understanding of γδ T cell biology and facilitate the translation of γδ TCRs into universal therapeutics.

## Improved detection and spatiotemporal spectral analysis of neural traveling waves
- Source: bioRxiv (preprints)
- Date: 2026-09-16
- Categories: Computational neuroscience
- Authors: Vinck, M., Gasco-Galvez, C., Rodrigues, M., Schwenk, J. C. B., Komatsu, M., Chavane, F., Canales-Johnson, A., Alamia, A.
- DOI: 10.64898/2026.09.11.750831
- Source URL: <https://doi.org/10.64898/2026.09.11.750831>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.11.750831>

Abstract: Traveling waves (TWs) are a fundamental mode of neural dynamics, yet existing detection methods are limited by sensor geometry, spatial-frequency resolution, signal amplitude, and ambiguity between propagating and standing-wave patterns. Here we introduce the Traveling Wave Index (TWINDEX), a framework for three-dimensional spatiotemporal spectral analysis of TWs across temporal frequency, spatial frequency, and propagation direction. TWINDEX generalizes to irregular sensor layouts, and quantifies wave strength as the reduction in circular phase variance produced by a candidate planar wave. This normalization yields robust behavior at both low and high spatial frequencies and suppresses coherent in-phase activity. Directional moments further separate planar from standing waves. We derive analytical links between TWINDEX, parametric planar-wave fit, and distance-phase correlation, and introduce projected distance-phase correlation (ProDPC) for sensitive single-trial planar-wave detection. Applying these methods to large-scale marmoset ECoG and human EEG, we identify alpha/low-beta TWs localized in spatial and temporal frequency, with physiologically plausible propagation speeds. Across trials and epochs, waves occur in oppositely directed propagation modes, demonstrating that alpha/beta activity is associated with both feedforward- and feedback-directed large-scale dynamics.

## In silico characterization of SAP55: insights into a predicted phytoplasmal M41-like metallopeptidase effector with potential eukaryotic host dual lipidation motifs
- Source: Scientific Reports (journals)
- Date: 2026-09-16T00:00:00+00:00
- Categories: Genomics & sequence analysis, Proteins & structural biology, Systems & networks, Evolution & metagenomics
- Authors: Kayhan Derecik, Gul Oz, Isil Tulum
- Journal: Scientific Reports
- DOI: 10.1038/s41598-026-70902-x
- Keywords: genomic, peptide, pathway, phylogenetic
- Source URL: <https://doi.org/10.1038/s41598-026-70902-x>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1038%2Fs41598-026-70902-x>

Abstract: Phytoplasmas are cell wall-less, phloem-limited plant pathogenic bacteria that cause devastating agricultural losses globally. Although phytoplasma pathogenicity is driven by secreted effector proteins translocated via the Sec pathway, their identification and functional characterization remain severely hindered by the fastidious nature of these pathogens. Here, we present a comprehensive structural, evolutionary, and functional in silico characterization of SAP55, an uncharacterized candidate effector from the Aster yellows witches’-broom strain. AlphaFold 3 modeling predicted an N-terminal signal peptide with a cleavage-compatible structural architecture that is predicted to interact with phytoplasmal signal peptidase I. Genomic and phylogenetic analyses revealed that SAP55 is linked to potential mobile units and virulence islands, suggesting potential evolutionary mobility across lineages. Structural and sequence-based annotation identified a core domain with similarities to the M41 zinc-dependent metallopeptidase family with a conserved HEXXH motif. Notably, SAP55 is predicted to represent an atypical protease variant; structural comparisons and HSYMDOCK/PDBePISA thermodynamic simulations suggest that it lacks the AAA+ ATPase domain, the central loop, and hexameric subunit affinity, operating instead as a putative monomeric form with an elongated antiparallel β4-strand that may facilitate substrate interaction. Our analyses suggest that the conserved N-terminal methionine may represent a potential stabilization feature under N-end rule principle. Its C-terminal hypervariable region contains a conserved CXCAAL motif and polybasic cluster predicted to be compatible with host geranylgeranyltransferase type I and S-palmitoylation machinery, potentially supporting association with the cytoplasmic leaflet of the plasma membrane By providing a comprehensive computational framework for a candidate membrane-associated candidate effector, this study proposes a working model for SAP55-mediated host manipulation, and establishes a foundation for future experimental investigation.

## Inferring the demographic history of Chinese and Indian rhesus macaque ( Macaca mulatta ) populations from PacBio HiFi long-read sequencing data
- Source: Molecular Biology and Evolution (journals)
- Date: 2026-09-16T00:00:00+00:00
- Categories: Genomics & sequence analysis
- Authors: Erangi J Heenkenda, Cyril J Versoza, John W Terbot II, Vivak Soni, Gabriella J Spatola, Susanne P Pfeifer, Jeffrey D Jensen
- Journal: Molecular Biology and Evolution
- DOI: 10.1093/molbev/msag237
- Keywords: genome
- Source URL: <https://doi.org/10.1093/molbev/msag237>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1093%2Fmolbev%2Fmsag237>

Abstract: The rhesus macaque (Macaca mulatta) is one of the most widely used animal models in biomedical research, both as it resembles humans in key biological aspects and as it is characterized by a broad geographic range. Most of the individuals housed in U.S. research colonies have been sampled from either China or India, though notably the source population of these animals has significantly shifted over time. Given the substantial genetic and immunological differences between these populations, a deeper understanding of the underlying population structure is critically important for biomedical interpretation. Despite this, the demographic histories of these two populations remain poorly resolved. Here, we present an analysis of whole-genome, PacBio HiFi long-read sequencing data from ten unrelated individuals of each population, applying four related model- and non-model based demographic inference approaches, in order to reconstruct their ancestral history. We evaluated the fit of the subsequently estimated models against the empirical data, and incorporated underlying uncertainty in the mutation rates used for scaling. We inferred a well-fitting population history characterized by substantial structure between Chinese and Indian populations, with a split time ∼140,000 generations ago from an ancestral population of ∼65,000 individuals. We additionally inferred the subsequent history of size change within, and gene flow between, these populations, reaching the current estimated sizes of ∼220,000 individuals in the Chinese population and ∼14,000 individuals in the Indian population. The robust baseline demographic model established in this study will serve as a valuable resource for future research on this species, including for improved fine-scale recombination mapping, selection inference, and association studies.

## Integrating NMR and contact-response analysis reveals the allosteric network driving domain closure in Enzyme I
- Source: Proceedings of the National Academy of Sciences (journals)
- Date: 2026-09-16T00:00:00+00:00
- Authors: Aayushi Singh, Daniel Burns, Sergey L. Sedinkin, Sayan Das, Davit A. Potoyan, Vincenzo Venditti
- Journal: Proceedings of the National Academy of Sciences
- DOI: 10.1073/pnas.2612191123
- Source URL: <https://doi.org/10.1073/pnas.2612191123>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1073%2Fpnas.2612191123>

Abstract: Understanding how phosphoenolpyruvate (PEP) binding induces the large open-to-closed conformational transition of bacterial Enzyme I (EI) has remained a long-standing problem in structural biology. In EI, PEP binds the C-terminal EIC domain, yet catalysis requires docking of the distant N-terminal EIN domain, which carries the active-site H189 residue, onto EIC. How this local binding event is coupled to global domain rearrangement has been unclear. Here, we combine experimental Chemical Shift Covariance Analysis (CHESCA) with computational Chemically Accurate Contact Response Analysis (ChACRA) to map the allosteric network underlying EI closure. Using a library of active-site mutants that systematically tune the open-to-closed equilibrium, CHESCA identifies a dominant cluster of residues whose chemical shifts correlate with the small-angle X-ray scattering-derived population of the closed state, revealing long-range energetic coupling between the PEP-binding site and distal structural elements. To obtain atomistic resolution, ChACRA analysis of Hamiltonian replica exchange molecular dynamics simulations identifies a spatially continuous network of coupled interactions spanning the PEP-binding pocket, interdomain linker, domain interfaces, and dimer contacts. Ligand binding reshapes this network, stabilizing interactions that promote domain docking and global rearrangement. Together, these results show that EI closure is governed by an extended allosteric network rather than a direct local contact. Crucially, CHESCA and ChACRA report on different physical observables; their convergence provides experimental validation of an atomistic interaction map and atomic-resolution interpretation of sparse NMR correlations, establishing a general framework for resolving allostery in complex biomolecular systems.

## Interpretable Machine Learning Identifies an Emergent Absence Seizure Mechanism
- Source: bioRxiv (preprints)
- Date: 2026-09-16
- Categories: Computational neuroscience
- Authors: Hull, J. M., Denomme, N., Ganguli, S., Huguenard, J. R.
- DOI: 10.1101/2025.09.23.678032
- Source URL: <https://doi.org/10.1101/2025.09.23.678032>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1101%2F2025.09.23.678032>

Abstract: Absence seizures are widespread spike-and-wave oscillations disrupting consciousness. Several consciousness frameworks emphasize intercortical/thalamocortical feedback dynamics, but we lack explicit dynamical mechanisms for how seizures disrupt these circuits. Using interpretable machine learning, we derived dynamical equations directly from seizure electrocorticogram data, reproducing seconds-long multi-regional local field potentials with precision matching inter-mouse variability. The model contained a low-dimensional chaotic seizure attractor emerging from between-region synchronization at preferred phase-lags. Unit recordings revealed corresponding spiking synchrony at single-neuron resolution, linking somatosensory and motor cortex with posterior thalamic nucleus (PO). Model coupling terms predicted tonic and burst firing spatiotemporal organization across regions and guided multisite-optogenetic stimulations. These stimulations showed PO controls corticocortical connectivity gain and that motor/somatosensory corticocortical functional connectivity varies at predicted phase-lags to drive seizures. Our results define absence seizures as dynamics confined to a chaotic attractor within a circuit implicated in anesthetic unconsciousness, linking distributed network chaos to loss of consciousness.

## KRstereo: Predicting β-Hydroxy Stereochemistry in Polyketides Using Protein Language Models
- Source: Journal of Chemical Information and Modeling (journals)
- Date: 2026-09-16T00:00:00Z
- Categories: Proteins & structural biology, Tools & resources
- Authors: Hsin-Ying Tsai, Wen-Qiang Xu, Wen-Jun Xie, You-Song Ding
- Journal: Journal of Chemical Information and Modeling
- DOI: 10.1021/acs.jcim.6c02706
- External ID: a09e79003a0ea2ac0bf30c149bf04e245e9dfb0b
- Source URL: <https://doi.org/10.1021/acs.jcim.6c02706>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1021%2Facs.jcim.6c02706>

Abstract: Polyketides are a major class of bioactive natural products whose activities are often determined by the stereochemistry of β-hydroxyl groups. In type I polyketide synthases (PKSs), ketoreductase (KR) domains establish these stereocenters, making accurate prediction of KR stereochemistry important for natural product discovery and PKS engineering. Existing rule-based methods rely on a limited set of sequence motifs and often perform poorly across phylogenetically diverse taxa. Here, we present KRstereo, a machine learning framework that predicts KR stereochemistry directly from sequence. Analysis of β-modular KR domains from MIBiG 3.1 identified informative sequence features beyond canonical motifs, motivating the use of protein language model embeddings. KRstereo achieved accuracies of up to 95.7% across taxa and 93.0% for non-Streptomyces KRs, consistently outperforming existing rule-based approaches. Validation using newly characterized KR domains from MIBiG 4.0 confirmed strong generalizability, including cases misclassified by current methods. Application of KRstereo to 20,840 β-modular KR domains from antiSMASH enabled large-scale stereochemical annotation of previously uncharacterized PKS systems. By linking sequence to stereochemical function, KRstereo improves reconstruction of polyketide structures from biosynthetic gene clusters and facilitates stereochemistry-aware genome mining and engineering of PKS assembly lines.

## Longer Is Not Always Better: Effects of Equilibration Length on Umbrella Sampling Estimations for RNA Hairpin Folding Stabilities
- Source: bioRxiv (preprints)
- Date: 2026-09-16
- Authors: Akinyemi, O., Kierzek, E., Kierzek, R., McSally, J. P., Puthenpeedikakkal, A. M. K., Mathews, D. H.
- DOI: 10.64898/2026.09.15.751614
- Source URL: <https://doi.org/10.64898/2026.09.15.751614>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.15.751614>

Abstract: Umbrella sampling is widely used to estimate biomolecular free energy landscapes and relative folding stabilities. Although equilibration is a critical component of umbrella sampling workflows, the impact of equilibration length on thermodynamic predictions remains poorly understood. Here, we investigate the effect of equilibration length on relative folding free energy predictions for four RNA hairpins with loop sequences GUGAAA, CUGGGA, GUAAUA, and UUAAUU with helical stems of three base pairs. Umbrella sampling simulations were performed using an end-to-end distance reaction coordinate spanning 15-45 \[A\], where equilibrium simulations (windows) were spaced at roughly 1 \[A\] intervals. In these calculations, the hairpin stem-loops were allowed to equilibrate in an end-to-end distance window and then the coordinates were transferred to the next larger end-to-end distance window to equilibrate. Two equilibration lengths, 2 ns and 100 ns per window, were followed by 600 ns of production sampling. Potential of mean force (PMF) profiles were reconstructed using the Weighted Histogram Analysis Method (WHAM) and used to calculate pairwise free energy differences with thermodynamic cycles. Increasing the equilibration length produced substantial, sequence-dependent changes in the reconstructed free energy landscapes. The 100 ns protocol generated markedly flatter PMFs for GUGAAA and UUAAUU and pronounced reshaping of the free energy landscape for GUAAUA. These changes were accompanied by reductions in hydrogen-bonding and stacking interactions, particularly within the intermediate regions of the reaction coordinate. The resulting thermodynamic predictions, as free energy change differences, were therefore highly sensitive to equilibration length. Across nearly all hairpin pairs, the 100 ns equilibration yielded substantially larger magnitude free energy change difference values than the corresponding 2 ns equilibration, with differences that greatly exceeded replica-to-replica variability. Comparison with optical melting measurements and nearest-neighbor thermodynamic predictions revealed that 2 ns equilibration times more closely agreed with experimental values than those obtained using 100 ns equilibration. These findings demonstrate that longer equilibration can systematically alter the structural ensembles sampled during umbrella sampling and amplify predicted stability differences without improving agreement with experiment. More widely, our results highlighted equilibration length as a critical and nontrivial parameter in RNA free energy calculations and demonstrate that increased equilibration does not necessarily lead to more accurate thermodynamic predictions.

## Machine learned potentials with electrostatic embedding accurately capture Kemp eliminase reactivity
- Source: bioRxiv (preprints)
- Date: 2026-09-16
- Categories: Proteins & structural biology
- Authors: Lear, A., Chan, E. W., Zinovjev, K., van der Kamp, M. W., Bunzel, H. A., Mulholland, A. J.
- DOI: 10.64898/2026.09.14.751485
- Source URL: <https://doi.org/10.64898/2026.09.14.751485>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.14.751485>

Abstract: Kemp elimination has become a benchmark for de novo enzyme design due to its simplicity and detailed mechanistic understanding but accurate barrier calculations are required to understand the difference in activity between designed Kemp eliminases. QM/MM MD simulations are capable of calculating reaction barriers, but are limited by a cost/accuracy tradeoff in which the most accurate methods are too computationally expensive for extensive screening as would be required in a prospective enzyme design campaign. Recent advances in embedded ML/MM simulations using the electrostatic machine learning embedding (EMLE) method have enabled transferable potentials for ML/MM MD simulations with low computational cost but QM-level accuracy. Here, we trained a MACE MLIP and EMLE embedding model for fast and accurate simulations of the enzymatic, base-catalysed Kemp elimination of 6-nitro benzisoxazole. Applying our EMLE ML/MM scheme to a designed Kemp eliminase and its evolved counterpart captured the \{approx\}4 kcal/mol difference in barrier between the two observed in experiment. Furthermore, training was done using structures generated from simulations of a single variant and was applied to the second variant, achieving near-experimental barriers, with no further finetuning or computational overhead. Thus, this work establishes EMLE-based ML/MM MD simulations as a potential route for fast and accurate assessment of barriers in computational screening for enzyme design.

## MACS3: A Peak-calling Platform for Bulk and Single-cell Regulatory Genomics.
- Source: Genomics, proteomics & bioinformatics (journals)
- Date: 2026-09-16T00:00:00Z
- Categories: Genomics & sequence analysis, Tools & resources
- Authors: Philippa Doherty, Qiang Hu, Zi-Han Zhuang, Hong Zhang, Sai C. Penikalapati, Song Liu, Tao Liu
- Journal: Genomics, proteomics & bioinformatics
- DOI: 10.1093/gpbjnl/qzag097
- External ID: c6da8e6eedd050980b26d1d55407b5e7ecc85d30
- Source URL: <https://doi.org/10.1093/gpbjnl/qzag097>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1093%2Fgpbjnl%2Fqzag097>
- Code: <https://github.com/macs3-project/MACS>

Abstract: Since the original publication of Model-based Analysis for ChIP-Seq (MACS), the software has been widely used to identify enriched genomic regions in ChIP-seq, ATAC-seq, CUT&RUN, DNase-seq, and related regulatory genomics assays. Over the years, MACS has evolved substantially, with MACS version 3 (MACS3) now serving as the actively maintained implementation. MACS3 preserves the core MACS framework for fragment pileup, dynamic local background noise, statistical enrichment testing, and peak refinement, while adding functionality needed for contemporary bulk and single-cell workflows. It supports conventional bulk peak calling, paired-end and fragment-based file formats, modular signal processing, direct analysis of single-cell ATAC-seq fragment files, barcode-restricted pseudobulk and cluster-level peak calling, specialized ATAC-seq and variant-calling modules, as well as command-line and programmatic interfaces. MACS3 is distributed through standard software channels and supported by continuous testing across operating systems, Python versions, and CPU architectures. Here we describe the architecture, current capabilities, and recommended use of MACS3, providing an updated reference for applying the MACS framework in contemporary bulk and single-cell regulatory genomics workflows. MACS3 is open-source software available at https://github.com/macs3-project/MACS.

## MAPA: A Semantic Network Framework for Functional Module Discovery and Interpretation in Multi‐Omics Data
- Source: Advanced Science (journals)
- Date: 2026-09-16T00:00:00Z
- Categories: Single-cell & spatial, Tools & resources
- Authors: Yi-Fei Ge, Fei-Fan Zhang, Yi-Jiang Liu, Chao Jiang, Peng Gao, N. S. Tan, Sai Zhang, Yu-Chen Shen, Qian-Yi Zhou, Xin Zhou, Xiao Wang, Fang-Qing Zhao, Chu-Chu Wang, Xiao-Tao Shen
- Journal: Advanced Science
- DOI: 10.1002/advs.77774
- External ID: e9965efbeee696524d5d8a182fafa10333c6abe5
- Source URL: <https://doi.org/10.1002/advs.77774>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1002%2Fadvs.77774>

Abstract: Multi‐omics technologies generate high‐dimensional molecular signatures that provide unprecedented opportunities to uncover biological mechanisms. However, translating complex molecular alterations into coherent and interpretable functional insights remains a major challenge. Existing module discovery methods can identify groups of related features, but often lack direct biological interpretability, whereas pathway‐based approaches frequently yield redundant results that complicate interpretation. Here, we present MAPA (Modular Analysis and Phenotype‐informed Annotation using large language models \[LLMs\]), a semantic‐biological network framework for functional module discovery and interpretation in multi‐omics data. MAPA integrates molecular interactions and pathway‐level functional context into a unified semantic‐biological network, and applies random walk with restart to quantify global functional relatedness among molecules and pathways for coherent module discovery across omics layers. MAPA further incorporates LLM‐assisted interpretation with retrieval‐augmented generation (RAG) to produce structured, literature‐informed module interpretation. Benchmarking against existing approaches shows that MAPA achieves superior module reconstruction and expert‐aligned functional interpretation. Applied to aging‐related multi‐omics datasets, MAPA reveals biologically coherent modules and biological insights that are difficult to obtain from conventional pathway analyses alone. MAPA provides a generalizable framework for organizing fragmented and heterogeneous molecular features into functional modules and comprehensive interpretations.

## Martini 3 Coarse-Grained Model of DNA for Heterogeneous Molecular Systems
- Source: bioRxiv (preprints)
- Date: 2026-09-16
- Categories: Genomics & sequence analysis, Proteins & structural biology
- Authors: Dargis, R., Arya, G.
- DOI: 10.64898/2026.09.14.751576
- Keywords: dna, molecular dynamics
- Source URL: <https://doi.org/10.64898/2026.09.14.751576>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.14.751576>

Abstract: DNA often functions in heterogeneous molecular systems containing proteins, lipids, polymers, and other materials. All-atom molecular dynamics simulations can be used to study DNA in these multicomponent systems, but computational cost limits the accessible system sizes and time scales. Coarse-grained models extend these scales, but existing DNA models are generally not designed for interactions with a broad range of other molecular species. To fill this gap, we develop a coarse-grained model of DNA designed for use with the Martini 3 force field. The model was parameterized through an iterative Bayesian optimization workflow, which used a scaled Wasserstein metric to compare distributions of local geometrical features and global structure from coarse-grained simulations against all-atom reference simulations. The optimized model captures key structural and mechanical properties of single- and double-stranded DNA across varying strand lengths and ionic conditions, while retaining compatibility with the broader Martini 3 ecosystem. This compatibility enables DNA to be integrated with a broad range of molecular systems, as we illustrate through simulations of double-stranded DNA bound to a transcription factor, cholesterol-tagged DNA duplex interacting with a lipid bilayer, a crossover-containing DNA nanostructure, and single-stranded DNA adsorbing onto graphene. Together, these results establish a transferable coarse-grained model of DNA for simulations of heterogeneous biomolecular and engineered systems.

## Menger\_Curvature : a MDAKit implementation to decipher the dynamics, curvatures and flexibilities of polymeric backbones at the residue level
- Source: bioRxiv (preprints)
- Date: 2026-09-16
- Categories: Proteins & structural biology, Tools & resources
- Authors: Reboul, E., Marien, J., Prevost, C., Taly, A., Sacquin-Mora, S.
- DOI: 10.1101/2025.04.04.647214
- Source URL: <https://doi.org/10.1101/2025.04.04.647214>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1101%2F2025.04.04.647214>
- Code: <https://github.com/EtienneReboul/menger_curvature>

Abstract: Characterizing the dynamics of the backbone of flexible polymers such as Intrinsically Disordered Regions and Proteins (IDRs and IDPs) has proven to be a significant challenge in molecular dynamics (MD) simulations due to their high conformational variability. The widely-used mobility metric Root-Mean-Squared Fluctuations (RMSF) is powerless to provide information for highly flexible systems, as defining a relevant reference structure is often not possible. We previously introduced a new flexibility metric to remedy this gap : the Local Flexibilities (LFs), derived (alongside the Local Curvatures (LCs)) from the Proteic Menger Curvatures (PMCs). Here we present a numba accelerated implementation for any polymer of the calculation of Menger Curvatures as a MDAKit from the widely-used MDAnalysis package. We perform a benchmark with the RMSF and another flexibility metric derived from Proteic Blocks (PBs), the Equivalent Number of PBs (Neq), and show that the PMCs are an order of magnitude faster to compute on a modern CPU chip. We applied all 3 flexibility metrics to a \{beta\}III-tubulin monomer as an example, as tubulins are known to possess the entire range of proteic elements, from -helix and \{beta\}-sheets to a flexible loop and a disordered C-terminal tail (CTT). RMSF, LFs and Neq all succeed in identifying the flexible loops and the CTT, although the RMSF requires a system-specific alignment to do so. Finally, we expose different applications of PMCs, LCs and LFs ranging from mechanism characterization to NMR T2 predictions. We believe that Menger curvatures will prove to be a valuable metric to study protein dynamics and polymers in general. The MDAKit package Menger\_Curvature is readily available at https://github.com/EtienneReboul/menger\_curvature

## MetaproDB: A Flexible and Reproducible Framework for Biome-Informed Protein Sequence Database Construction for Metaproteomics
- Source: Journal of Proteome Research (journals)
- Date: 2026-09-16T00:00:00+00:00
- Categories: Genomics & sequence analysis, Proteins & structural biology, Tools & resources
- Authors: Muzaffer Arıkan
- Journal: Journal of Proteome Research
- DOI: 10.1021/acs.jproteome.6c00326
- Source URL: <https://doi.org/10.1021/acs.jproteome.6c00326>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1021%2Facs.jproteome.6c00326>
- Code: <https://github.com/arikanlab/MetaproDB>

Abstract: Metaproteomic analyses commonly rely on protein sequence databases, yet database construction remains one of the most variable steps in metaproteomic workflows. Here, I present MetaproDB, a flexible and reproducible framework for biome-informed protein-sequence-database construction in metaproteomics. MetaproDB integrates ecological taxon selection, build-plan generation, genome resource linkage, protein sequence assembly, exact-sequence deduplication, completeness assessment, and provenance tracking within a unified workflow. It supports both database generation from a curated reference panel of representative biomes and cohort-specific construction from user-provided microbiome profiles. I demonstrate the functionality of MetaproDB through three case studies that compare different database construction strategies. MetaproDB provides a practical framework for explicit and reproducible biome-informed database design in metaproteomics and is available at \[https://github.com/arikanlab/MetaproDB\].

## MIA-Jet: Multi-scale Identification Algorithm of Chromatin Jets
- Source: bioRxiv (preprints)
- Date: 2026-09-16
- Categories: Genomics & sequence analysis, Tools & resources
- Authors: Kim, S., Kim, M.
- DOI: 10.1101/2025.08.27.672730
- Source URL: <https://doi.org/10.1101/2025.08.27.672730>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1101%2F2025.08.27.672730>

Abstract: The mammalian genome is organized into large-scale chromosome territories, compartments, domains, and at the smallest scale, chromatin loops and stripes. The newest element is a chromatin jet, a diffused line perpendicular to the main diagonal in the Hi-C contact map, which was reported in quiescent mammalian lymphocytes supporting a two-sided symmetric cohesin loop extrusion model. A similar structure is observed in Repli-HiC and related data, where relatively thin and straight chromatin fountains indicate coupling of DNA replication forks. However, the precise biological implications of these jet-like structures are unknown due to the limitations in computational methods. We developed MIA-Jet, a multi-scale ridge detection algorithm that can accurately detect jets of variable lengths, widths, and angles. When tested on Hi-C, Repli-HiC, ChIA-PET, ChIA-Drop, and Micro-C data in mouse, human, roundworm, and zebrafish cells, MIA-Jet outperformed existing methods. In human cells, jets were enriched in cohesin loading sites and early replication initiation zones. Applying MIA-Jet to Hi-C data generated from protein-degraded cells revealed that jets are dependent on cohesin and that depleting CTCF results in longer and less angled jets than wild-type. We envision MIA-Jet to be broadly applicable to any 3D genome mapping data, thereby providing new insights into the functional roles of chromatin jets.

## ML4SD: Leveraging Machine Learning and High-Throughput Search Algorithms for an Iterative Growth-Coupled Design Innovation
- Source: bioRxiv (preprints)
- Date: 2026-09-16
- Categories: Systems & networks
- Authors: Gargantilla Becerra, A., Nogales Enrique, J.
- DOI: 10.64898/2026.09.11.750664
- Source URL: <https://doi.org/10.64898/2026.09.11.750664>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.11.750664>

Abstract: Optimizing microbial biomanufacturing is required if renewable and waste carbon are to replace petrochemical routes at competitive titers, rates, and yields. Growth-coupled (GC) production supports that goal by linking target synthesis to biomass formation, so product formation is required for growth. Constructing knockout strains yielding GC production from a list of candidate genes is labor and time demanding. This results in few in vivo tested designs, which hampers standard machine-learning methods to learn GC patterns for specific bioprocesses. We therefore developed ML4SD, an active-learning Design-Build-Test-Learn (DBTL) cycle that trains ensembles on genome-scale metabolic model (GEM) scores of knockout designs, sampling the next designs from predicted model performance and error. That cycle generalizes only if the initial library is large and diverse, including suboptimal and non-viable designs; libraries restricted to minimal designs or Pareto-optimal knockouts were found to generate models overfitting. To meet those specific demands a novel strain design algorithm, gcSwarms, was developed and tested for a diverse set of bioprocesses within Pseudomonas putida iJN1462. ML4SD was tested with an in silico case study converting lignin-derived 4-hydroxybenzoate to 6-caprolactam, the nylon-6 monomer. ML4SD results showed improvements of up to 164% on carbon yield, recovering a shared SHAP motif that redirects TCA flux through acetyl-CoA. Importantly it reaches that result using 2.5- to 7.1-fold fewer designs than a gcSwarms-only search, demonstrating the data efficiency of this method.

## Moderated designs can balance between batch-effect mitigation and cell loss due to hashtag-assisted pooling in single-cell experiments.
- Source: Genome research (journals)
- Date: 2026-09-16
- Authors: Budha Chatterjee, Katrina Gorga, Carly Blair, Yuko Ohta, Michelle Radov, Elizabeth M Hill, Christopher T Boughter, Martin Meier-Schellersheim, Nevil J Singh
- Journal: Genome research
- DOI: 10.1101/gr.281624.125
- External ID: 42532835
- Source URL: <https://doi.org/10.1101/gr.281624.125>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1101%2Fgr.281624.125>

Abstract: Minimizing experimental noise is integral to robust data generation in single-cell omics. The current standard for avoiding batch effects during sample processing is barcode- or hashtag-assisted combining of different experimental treatments into one pool, allowing all samples to be subject to the technical protocols uniformly. The final data points for each treatment group are then computationally separated based on the original hashtag labels. Clearly, whereas hashtagging all groups and pooling them in a single well is expected to minimize batch effects, the procedure can also lead to a loss of cells that cannot be confidently decoded during the computational demultiplexing step. Here, we examine four alternate experimental designs, namely, compound, reference, chain, and confounded, that could be used instead of a single-pool approach and quantify the batch effects as well as cell loss in each case. We find a linear relationship-the percentage of cells lost is double the number of hashtags used in the experiment. We use these analyses to identify experimental designs that can successfully mitigate batch effects while minimizing multiplexing, hence the cell loss, in each well. Although a reference design offers the best overall performance, this study can help individual investigators choose particular approaches that are best suited for their biological questions.

## MRSIPrep: A Standardized Post-Quantification Framework for Preprocessing Whole-Brain Magnetic Resonance Spectroscopic Imaging
- Source: bioRxiv (preprints)
- Date: 2026-09-16
- Categories: Biological imaging, Tools & resources
- Authors: Lucchetti, F., Celereau, E., Jenni, R., Ledoux, J.-B., Eliez, S., Delavari, F., Hagmann, P., Aleman-Gomez, Y., Klauser, A., Klauser, P.
- DOI: 10.64898/2026.09.11.750896
- Source URL: <https://doi.org/10.64898/2026.09.11.750896>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.11.750896>

Abstract: Magnetic resonance spectroscopic imaging (MRSI) enables non-invasive mapping of neu-rometabolites levels across the human brain. Although spectral fitting and metabolite quan-tification are increasingly supported by mature software tools, the downstream processing of quantified metabolite maps remains heterogeneous across laboratories. Here, we introduce MRSIPrep, an open-source, modular, and reproducible post-quantification framework for whole-brain MRSI. MRSIPrep standardizes quality control, tissue correction, spatial nor- malization, atlas projection, and derivative generation from quantified metabolite maps and associated quality metrics. The framework produces voxelwise, regional, and connectomics-ready outputs together with automated quality-control reports. We describe the architecture of MRSIPrep and demonstrate its utility for reproducible MRSI analysis across datasets,acquisition protocols, and downstream applications.

## Multi-agentic system for primer design in qPCR and LAMP diagnostics tests
- Source: bioRxiv (preprints)
- Date: 2026-09-16
- Authors: Lau, K. J. X.
- DOI: 10.64898/2026.09.15.751771
- Source URL: <https://doi.org/10.64898/2026.09.15.751771>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.15.751771>

Abstract: Primer design is a fundamental component of molecular diagnostics in both quantitative polymerase chain reaction qPCR and loop-mediated isothermal amplification LAMP assays. However, assay design is often performed manually as nucleotide databases, sequence alignment tools and resources are found at different places on the Internet. In this study, an AI-orchestrated bioinformatics workflow was developed to automate the end-to-end qPCR and LAMP primers and probes. The workflow was implemented using LangGraph, LangChain and Biopython, where a series of specialized agents were coordinated to execute sequential bioinformatics tasks with minimal human intervention. Target sequences were then retrieved based on the user request from the National Center for Biotechnology Information nucleotide database and the requested sequence records were then subjected to multiple sequence alignment for the identification of conserved genomic regions. The multi-agentic primer design system can be used for assay development for applications in infectious disease diagnostics, outbreak surveillance and environmental monitoring. This study also demonstrates how multi-agentic systems can be combined with established bioinformatics methods to automate qPCR and LAMP assay design.

## Multiphoton tomographic fluorescence lifetime imaging microscopy -TomoFLIM
- Source: bioRxiv (preprints)
- Date: 2026-09-16
- Categories: Biological imaging
- Authors: Collard, L., Jose, A. A., Rahman, F., Garcia-Aguirre, R., Treacy, C., Pallett, T., Culley, S., Ameer-Beg, S. M., Poland, S. P.
- DOI: 10.64898/2026.09.11.750847
- Source URL: <https://doi.org/10.64898/2026.09.11.750847>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.11.750847>

Abstract: Fluorescence lifetime imaging microscopy provides quantitative, concentration-independent contrast for probing molecular interactions, biochemical environments and cellular physiology. However, the requirement to acquire sufficient time-resolved photon statistics makes FLIM inherently slow, limiting its application to dynamic biological processes. In live-cell applications, including calcium signalling and vesicular trafficking, acquisition times can exceed the timescale of the underlying biology, causing temporal averaging, motion artefacts and loss of transient information. Methods that increase FLIM acquisition speed while preserving quantitative lifetime accuracy are therefore required. We report on a high-speed, compressive, multiphoton fluorescence lifetime imaging technique (TomoFLIM). Two-photon fluorescence is excited using a line focus projected tomographically across the sample, while time-tagged fluorescence is acquired using time-correlated single-photon counting. Time-resolved fluorescence data are reconstructed using Lucy-Richardson deconvolution followed by a computationally efficient centre-of-mass method lifetime estimator. In addition, TomoFLIM Net, a physics-informed neural-network, directly reconstructs fluorescence intensity and lifetime from compressed time-resolved tomographic data. TomoFLIM was benchmarked against raster-scanned fluorescence lifetime measurements using calibrated fluorescence lifetime beads and biological specimens. We demonstrate imaging at compression ratios exceeding 90%, with Pearson correlation coefficients above 80% relative to reference images. A raster-scanned FLIM dataset acquired in 60 s was reproduced using TomoFLIM in 3.75 s, representing a 16-fold increase in frame rate and equivalent reduction in accumulated dark counts. TomoFLIM Net recovered distinct experimental bead lifetime populations, demonstrating a direct route from compressed measurements to quantitative lifetime maps. TomoFLIM therefore offers significant potential for rapid live-cell imaging of dynamic biological processes, including deep within turbid biological specimens.

## Multiscale computational modeling to quantify how spiral artery remodeling alters wall shear stress on placental villi
- Source: PLOS Computational Biology (journals)
- Date: 2026-09-16T00:00:00+00:00
- Authors: Armita Najmi, Noelia Grande Gutiérrez
- Journal: PLOS Computational Biology
- DOI: 10.1371/journal.pcbi.1014769
- Source URL: <https://doi.org/10.1371/journal.pcbi.1014769>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1371%2Fjournal.pcbi.1014769>

Abstract: Proper placental development is essential for a healthy pregnancy. It depends on the remodeling of the maternal uterine vasculature to meet fetal demands while maintaining physiological intervillous space (IVS) hemodynamics for biochemical exchange. Terminal villi, the primary sites of feto-maternal exchange, exhibit impaired development in pregnancies complicated by intrauterine growth restriction and preeclampsia, which are also associated with incomplete spiral artery (SA) remodeling. Despite this association, the mechanistic link between maternal blood flow and villous development remains unclear. Here, we investigate whether incomplete SA remodeling alters IVS hemodynamics and increases wall shear stress (WSS) on placental villi, potentially impairing terminal villi formation. Computing WSS throughout an entire placentone is challenging due to uncertainty in placental microstructure and the computational cost of resolving microscale hemodynamics. We propose a novel multiscale computational framework to quantify WSS on placental villi at the end of the second trimester, when WSS may affect terminal villi development. A macroscale placentone model is used to compute IVS velocities, which are coupled with microscale models of intermediate villi to estimate villous WSS across physiologically relevant flow conditions. We simulate IVS hemodynamics in healthy pregnancy and varying degrees of incomplete SA remodeling. Our results show that IVS velocity is the primary determinant of mean villous WSS, whereas villous type and orientation have comparatively weaker effects. In healthy placentones, most villi experience a mean WSS of 0.001–1 Pa, with higher stresses localized near the free-of-villi cavity. By correlating these estimates with regions naturally devoid of terminal villi, we identify a mean WSS range of approximately 0.71–1.44 Pa that may inhibit terminal villi formation. Incomplete SA remodeling significantly increases WSS, reaching levels consistent with villous tissue loss and placental lake formation. These findings suggest a mechanistic link between uteroplacental hemodynamics and villi development, establishing physiological shear-stress thresholds relevant to placental health.

## Mutational constraints on RSV F and its neutralization by antibodies
- Source: Nature (journals)
- Date: 2026-09-16T00:00:00+00:00
- Authors: Cassandra A. L. Simonich, Teagan E. McMahon, Gavin Juviler, Lucas Kampman, Helen Y. Chu, Jesse D. Bloom
- Journal: Nature
- DOI: 10.1038/s41586-026-11030-4
- Source URL: <https://doi.org/10.1038/s41586-026-11030-4>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1038%2Fs41586-026-11030-4>

Abstract: New antibodies targeting the F protein of respiratory syncytial virus (RSV) have substantially reduced infant hospitalizations 1 . However, viral resistance is a concern: one antibody failed clinical trials because of a resistant strain 2 , and sporadic resistance mutations to the most widely used antibody (nirsevimab) have been identified 3–6 . Here we define how RSV F mutations affect antibody neutralization. We first provide a biophysical model of how the buffering of bivalent IgG binding combines with the lower Fab potency of nirsevimab to subtype B to make resistance to this antibody more common in subtype B than A strains. We then perform pseudovirus deep mutational scanning to safely measure how nearly all mutations to F affect its cell entry function and neutralization by IgG and Fab forms of nirsevimab, clesrovimab and several other key antibodies. We use these measurements to enable real-time surveillance of RSV sequences for antibody resistance, and show that resistant strains have arisen sporadically but are at present rare. Overall, our work shows how Fab potency and epitope specificity combine to determine how viral mutations affect antibody neutralization, enables monitoring for natural RSV strains resistant to antibodies of public-health importance, and can help guide development of future antibodies with resilience to viral escape.

## nanorepertoire: an end-to-end Nextflow pipeline for nanobody repertoire analysis
- Source: bioRxiv (preprints)
- Date: 2026-09-16
- Categories: Genomics & sequence analysis, Tools & resources
- Authors: Bagordo, D., Martelossi, N., Lescai, F.
- DOI: 10.64898/2026.09.11.750933
- Source URL: <https://doi.org/10.64898/2026.09.11.750933>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.11.750933>
- Code: <https://github.com/lescailab/nanorepertoire>

Abstract: Camelid heavy-chain antibodies, and particularly their variable domains known as nanobodies or VHHs, combine full antigen-binding capacity with a compact and highly stable scaffold, which makes them attractive for both fundamental immunology and therapeutic development. High-throughput adaptive immune receptor repertoire sequencing (AIRR-seq) allows nanobody repertoires to be profiled at great depth, but the analyses applied to VHH data are typically assembled ad hoc from standalone scripts, which limits standardisation and reproducibility across laboratories. Here we present nanorepertoire, an end-to-end Nextflow DSL2 pipeline dedicated to camelid VHH repertoires. It takes paired-end AIRR-seq FASTQ files through quality control, adapter trimming, read merging, in-silico translation, CD-HIT clonotyping and deep-learning CDR3 annotation with nanoCDR-X (Bagordo et al., 2026), and returns an interactive HTML report describing clonal architecture, CDR3 length and amino-acid composition, intra-clonal homogeneity and repertoire diversity, together with the computational carbon footprint of the run. Applied to two publicly available SARS-CoV-2 RBD-selected llama libraries sampled before and after phage-display enrichment (4.8 million paired-end reads in total), the pipeline completed in 59 min on a 16-vCPU cloud instance and recovered 41,363 distinct CDR3 paratopes, reproducing the expected contraction of clonal diversity upon selection. nanorepertoire is open source under the MIT licence at https://github.com/lescailab/nanorepertoire, is archived on Zenodo, and runs unchanged on local, HPC and cloud infrastructures.

## Network propagation in bipartite metabolite-reaction graphs for metabolomic data exploration.
- Source: Metabolomics : Official journal of the Metabolomic Society (journals)
- Date: 2026-09-16
- Authors: Julia Kuligowski, Marta Moreno-Torres, David Pérez-Guaita, Francesc Albert Esteve-Turrillas, Guillermo Quintás
- Journal: Metabolomics : Official journal of the Metabolomic Society
- DOI: 10.1007/s11306-026-02529-y
- External ID: 42749860
- Source URL: <https://doi.org/10.1007/s11306-026-02529-y>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1007%2Fs11306-026-02529-y>

Abstract: INTRODUCTION: Interpretation of metabolomic data is frequently limited by incomplete metabolite coverage and the predefined representation of biochemical organization provided by pathway-based approaches. OBJECTIVES: This study describes and evaluates a graph-based network propagation method designed to faciitate metabolomic data exploration by diffusing node-level statistical relevance scores across a metabolic network. METHODS: An undirected bipartite metabolite-reaction network was reconstructed using the KEGG database, comprising 8236 nodes (1601 metabolites and 6635 reactions). Simulated datasets with 75% sparsity mimicking two metabolic perturbations and two no-effect control groups, and real data from a previous study were used to test the strategy. Node-level statistical signal scores were distributed across the network topology using a random walk-based diffusion operator combined with a supervised clamping procedure to preserve initially observed measurements. A topological distance mask was subsequently applied to restrict propagation to nodes located within a predefined network distance from experimentally measured metabolites. RESULTS: Network propagation redistributed statistical relevance across connected subgraphs, expanding the number of metabolic features carrying topology-informed statistical scores beyond experimentally observed metabolites in simulated and real data. In both oxidative stress and mitochondrial dysfunction simulations, significant nodes displayed non-random topological organization consistent with the simulated perturbations. Propagation of real data also identified additional unmeasured metabolites that were topologically connected to experimentally observed metabolites. In multivariate analyses, network propagation expanded the feature space and showed that it could improve clustering performance. CONCLUSIONS: The proposed bipartite metabolite-reaction network propagation provides an exploratory tool for metabolomics. By integrating topological context with statistical relevance, this approach complements established pathway analyses for hypothesis generation and candidate feature recovery in partially observed metabolic systems.

## Neuronal loss reshapes survivor dynamics and limits mechanism inference in excitatory inhibitory neural fields
- Source: bioRxiv (preprints)
- Date: 2026-09-16
- Categories: Computational neuroscience, Mathematical biology & statistics
- Authors: Reyes, R. G., Valdes-Sosa, P. A.
- DOI: 10.64898/2026.09.11.750823
- Source URL: <https://doi.org/10.64898/2026.09.11.750823>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.11.750823>

Abstract: Does neuronal loss simply reduce measured activity, or also change how the surviving network behaves? We separate these effects in a next-generation excitatory-inhibitory neural field by writing the viable population measure as q\_a = lambda\_a f\_a, where lambda\_a is viable population mass and f\_a is the normalized survivor distribution. Under state-independent thinning with fixed Cauchy heterogeneity, normalization commutes with the Ott-Antonsen/Montbrio-Pazo-Roxin reduction on the specified analytic invariant manifold. The mortality term disappears from conditional transport, but viable mass remains in recurrent coupling: loss can reshape survivor dynamics, not merely scale their contribution to tissue activity. Conversely, for otherwise identical constant homogeneous parameters, viability, pathway integrity and compensation give exactly conjugate conditional deterministic dynamics whenever c\_ab lambda\_b^(1-nu\_ab) is preserved. Identical conditional activity therefore need not imply an identical biological mechanism. Equilibrium and oscillatory bifurcations, finite-population escape, and delayed propagation reveal consequences of these two principles. In particular, matched field simulations show that localized loss can increase whole-sheet firing through recurrent reorganization, while coherent-wave continuation quantifies viability-dependent propagation and phase relaxation. The framework distinguishes neuronal abundance from survivor state and places an exact limit on mechanism inference. Attributing activity changes to neuronal loss therefore requires information beyond conditional neural dynamics, such as tissue-level measurements or independent structural constraints, interpreted through an appropriate observation model.

## New Publication "Bioimage management and analysis in Galaxy: Tools, workflows, training, and community practices"
- Source: Galaxy (feeds)
- Date: 2026-09-16T00:00:00+00:00
- Categories: Blog
- Source URL: <https://galaxyproject.org/news/2026-09-16-bioimagingpub-jm/>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fgalaxyproject.org%2Fnews%2F2026-09-16-bioimagingpub-jm%2F>
- Abstract: not stored for this record.

## Ourotide: decoding the hierarchical peptide recognition for generative design
- Source: bioRxiv (preprints)
- Date: 2026-09-16
- Categories: Proteins & structural biology
- Authors: Shen, Y., Zhang, J., Wu, Z., Xing, Z., Yuan, Q., Zhang, W., Zhou, Q., Han, F., Jiang, N., Chen, X.
- DOI: 10.64898/2026.09.10.750272
- Source URL: <https://doi.org/10.64898/2026.09.10.750272>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.10.750272>

Abstract: The historical dichotomy between small-molecule pocket and extended protein interfaces misrepresents the physical reality of peptide recognition.1,2 Here we show that peptide binding is not a simple structural intermediate but a distinctly multimodal landscape comprising small-molecule-like pockets, protein-like interfaces, and a previously unrecognized third mode. This third regime is governed by a hierarchical subpocket architecture where a flattened surface achieves near-complete peptide engagement through spatially partitioned hydrophobic components. To maintain stability, an incompletely enclosed dominant anchor cooperates with highly hydrated auxiliary subpockets and an asymmetric receptor coupling mechanism that concentrates energy in an adjacent continuous water network. Because this unique binding mode suffers from extreme data scarcity, standard deep learning models fail to capture its physics.3,4 To resolve this, we mapped these specific geometric signatures to mine structurally faithful training distributions from global protein interactomes. Based on these data, we trained Ourotide, a deep learning framework coupling conditional geometric flow matching with interface-aware affinity learning. Evaluated across peptides up to 65 residues, Ourotide outperforms generalist models in backbone accuracy, interface recovery, and affinity prediction. This approach suggests that overcoming data scarcity in the physical sciences requires physics-guided data augmentation rather than naive statistical scaling.

## Palaeoproteomic Deconvolution of Physical and Genetic Collagen Mixtures.
- Source: Advanced science (Weinheim, Baden-Wurttemberg, Germany) (journals)
- Date: 2026-09-16
- Categories: Proteins & structural biology
- Authors: Ian Engels, Tristan Dedrie, Synnøve K Mo, Simon Van de Vyver, Thijs R A Vandenbroucke, Kévin Di Modíca, Jan Decher, Alice Toso, Dieter Deforce, Simon Daled, Alexandra Burnett, Grégory Abrams, Maarten Dhaenens
- Journal: Advanced science (Weinheim, Baden-Wurttemberg, Germany)
- DOI: 10.1002/advs.77807
- External ID: 42750214
- Source URL: <https://doi.org/10.1002/advs.77807>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1002%2Fadvs.77807>

Abstract: Species identification in palaeoproteomics relies on genome-derived protein sequences which are often poor-quality, and lacks tools to cope with multi-species samples. Here, we address both challenges through the analysis of "physical and genetic mixtures". Species that are absent from our database are considered a "genetic mixture", i.e. a patchwork of peptides from closely related species. Inversely, various overlapping peptide stretches allow us to resolve complex "physical mixtures". This is benchmarked by analysing physical mixtures of modern bone fragments, including genetic mixtures. We illustrate the impact of our approach via a rapid and high-throughput analysis of >2500 bone fragments, revealing the Eemian-era faunal environment around Scladina Cave, including the first Palaeoloxodon antiquus identified at this site.

## PERADS.net: Automated PE-RADS Grading with Named Anatomic Localization and Right-to-Left Ventricular Ratio Measurement on CT Pulmonary Angiography
- Source: medRxiv (preprints)
- Date: 2026-09-16
- Categories: Biological imaging
- Authors: Lanza, E., Catapano, F., Lisi, C., D'Orazio, F., Levi, R., Laghi, A.
- DOI: 10.64898/2026.09.15.26363142
- Source URL: <https://doi.org/10.64898/2026.09.15.26363142>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.15.26363142>

Abstract: PurposeTo develop an automated pipeline (PERADS.net) that segments acute pulmonary embolism on CT pulmonary angiography, assigns a Pulmonary Embolism Reporting and Data System (PE-RADS) grade with named anatomic localization, and measures the right-to-left ventricular (RV/LV) diameter ratio, and to assess radiologist agreement. Materials and MethodsIn this retrospective study, 120 CT pulmonary angiograms (70 peripheral, 35 central, 15 negative) from the public RSNA Pulmonary Embolism CT dataset were analyzed. A two-channel five-fold nnU-Net ensemble segmented the embolus; the pulmonary arterial tree was reduced to a branching graph, and the grade was set by the most proximal level with at least 1% of embolic volume. Three radiologists, blinded to the algorithm-assigned grade, independently assigned grades and rated RV/LV plausibility. Agreement was assessed with percent agreement and Cohen or Fleiss kappa on five-grade and grouped scales (grade 0 versus 1-2 versus 3-4). The automated ratio was compared with the datasets binary RV/LV label in 2119 examinations. ResultsAgreement between the algorithm and three-radiologist consensus (n = 119) was 60.5% (kappa, 0.38) for individual grades and 90.8% (kappa, 0.74; 95% CI: 0.59, 0.87) for the grouped scale; 34 of 47 discordant examinations fell within grades 3-4. As a fourth reader, the algorithm matched grouped-scale interobserver agreement (mean kappa, 0.62 versus 0.67). Radiologists showed no agreement rating RV/LV plausibility as favorable or incorrect (Fleiss kappa, -0.00), whereas the automated ratio agreed with the external label (kappa, 0.49), overestimating strain 3.1:1. ConclusionAutomated PE-RADS grading agreed with radiologist consensus comparably to a fourth reader on the grouped scale. Summary StatementAn automated pipeline assigned PE-RADS grades agreeing with three-radiologist consensus at a level approaching interobserver agreement on the clinically grouped scale, while automated RV/LV measurement showed no reader agreement on plausibility. Key PointsO\_LIIn 120 CT pulmonary angiograms, agreement between automated and consensus PE-RADS grading was 90.8% (kappa, 0.74; 95% CI: 0.59, 0.87) on the clinically grouped scale (grade 0 versus 1-2 versus 3-4) and 60.5% (kappa, 0.38; 95% CI: 0.24, 0.51) on the five-grade scale. C\_LIO\_LITreated as a fourth reader, the algorithm reached grouped-scale agreement with individual radiologists (mean kappa, 0.62) comparable to that observed among radiologists themselves (mean kappa, 0.67). C\_LIO\_LIAutomated right-to-left ventricular ratio agreed moderately with an independent binary label in 2119 examinations (kappa, 0.49; 95% CI: 0.45, 0.52), whereas three radiologists showed no agreement when rating the same measurement as favorable or incorrect (Fleiss kappa, -0.00; 95% CI: -0.09, 0.09). C\_LI

## Perseus: Lineage-Aware Refinement of Kraken2 Taxonomic Classification for Long Read Metagenomes
- Source: Bioinformatics (journals)
- Date: 2026-09-16T00:00:00+00:00
- Categories: Genomics & sequence analysis, Evolution & metagenomics, Tools & resources
- Authors: Matthew H Nguyen, Michael C Schatz
- Journal: Bioinformatics
- DOI: 10.1093/bioinformatics/btag687
- Source URL: <https://doi.org/10.1093/bioinformatics/btag687>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1093%2Fbioinformatics%2Fbtag687>
- Code: <https://github.com/matnguyen/perseus>

Abstract: Motivation Long-read metagenomic sequencing improves assembly contiguity and enables genome-resolved analysis of complex microbial communities, but accurate taxonomic classification of long reads and assembled contigs remains challenging. Highly scalable k-mer-based classifiers such as Kraken2 frequently over-assign fine-rank taxonomic labels when applied to long-read data, producing high false positive classification rates driven by sparse or localized k-mer matches, particularly in microbiomes with extensive taxonomic novelty. Results We present Perseus, a lineage-aware confidence estimation framework for taxonomic classification that models the spatial distribution and hierarchical consistency of k-mer evidence along sequences. This formulation reframes taxonomic classification as a hierarchical confidence estimation problem rather than a single-rank prediction task. Perseus refines k-mer-level taxonomic signals from Kraken2 using a multi-headed convolutional neural network that estimates calibrated confidence scores for taxonomic correctness at each canonical rank. Using these estimates, Perseus confirms assignments, backs off to higher taxonomic ranks, or abstains when evidence is insufficient, prioritizing correctness and lineage consistency over overly specific assignments. Across simulations of taxonomic novelty and real-world metagenomic datasets, Perseus consistently and substantially reduces the false assignment rate while improving precision and lineage-consistent accuracy. These improvements are most pronounced for long reads and assembled contigs, where spatial context enables reliable discrimination between consistent taxonomic signal and spurious matches. Availability and implementation Perseus integrates with existing Kraken2 workflows and is available at https://github.com/matnguyen/perseus.

## Physics-aware resolution enhancement of soft X-ray tomography with measurement-supervised learning
- Source: bioRxiv (preprints)
- Date: 2026-09-16
- Categories: Biological imaging
- Authors: Chueh, S., Gallagher, E., de Ceuninck van Capelle, C., Luo, L., Ishikawa, T., Evans, C., Fletcher, N., Lopez-Perez, M., Rogers, D., O'Connor, S., McIntyre, C., Donnellan, M., Sheridan, P., Simpson, J. C., Kapishnikov, S.
- DOI: 10.64898/2026.06.21.730079
- Source URL: <https://doi.org/10.64898/2026.06.21.730079>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.06.21.730079>

Abstract: Soft X-ray tomography (SXT) is an emerging modality for whole-cell 3D imaging in near-native states. However, the effective spatial resolution is limited by optical artifacts characterized by the point spread function (PSF). Standard reconstruction methods force a compromise between structural sharpness and noise, failing to fully resolve these depth-dependent artifacts. By embedding experimentally measured, depth-variant PSFs into a differentiable forward model, we demonstrate a physics-aware computational optimization that bypasses these limitations to recover high-frequency cellular ultrastructure. The structural fidelity was validated using split-tilt Fourier ring correlation (FRC), alongside an experimental bead phantom tomogram, providing supporting evidence that the recovered high-frequency features reflect genuine specimen structure rather than fabricated artifacts. Our method effectively increases FRC spatial resolution and recovers cellular ultrastructure. Furthermore, under sparse-angular subsampling, the framework maintained spatial resolution using half the projection angles, a computational proxy pointing toward the potential for reduced radiation exposure in future acquisitions. This hardware-free, computational approach offers a route toward mitigating the optical and dosimetric constraints that currently limit nanoscale soft X-ray tomography.

## Predicting non-specific binding of VHHs using machine learning models with cluster-aware validation.
- Source: mAbs (journals)
- Date: 2026-09-16T00:00:00Z
- Categories: Proteins & structural biology
- Authors: V. Stanev, Federico Devalle, Mehdi Boroumand, Maryam Pouryahya, Isabelle Sermadiras, Jenna G. Caldwell, Kuan-Lin Chen, Jay Hyun Jo, Rohan Jain, Bismark Amofah, Tony Pham, Mark Hutchinson, Sharfa Farzandh, Jennifer DiChiara, Chacko S. Chakiath, Tom Diethe, A. Dippel, Gilad Kaplan, Rebecca Croasdale-Wood
- Journal: mAbs
- DOI: 10.1080/19420862.2026.2732787
- External ID: 228e75ffe7f929f9b160e7ff4f6556c0713a1e0c
- Source URL: <https://doi.org/10.1080/19420862.2026.2732787>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1080%2F19420862.2026.2732787>

Abstract: Propensity for nonspecific binding-also known as polyreactivity-is a serious developability risk factor for biotherapeutic candidates. To minimize this risk, drug companies are increasingly relying on in silico tools utilizing machine learning methods, but developing these tools is challenging. For example, the available data often contains many closely related sequences originating from drug pipeline projects, which can introduce significant biases in the in silico models training and benchmarking, leading to poor generalizability on new data. We present here a workflow designed to diagnose and mitigate some of the problems associated with using pipeline data. The workflow is based on a custom cross-validation procedure that can evaluate model performance on unseen data in different contexts. As a demonstration of the workflow, we use it to train a model to predict variable heavy-chain only fragment antibodies (VHH) binding to baculovirus particles (BVP)-a widely used assay for nonspecific binding. Using descriptors based on computed protein structures, the workflow identifies several risk factors that correlate with higher polyreactivity levels.

## Probabilistic mapping of sub-genic intolerance reveals functional and disease-critical protein regions
- Source: bioRxiv (preprints)
- Date: 2026-09-16
- Categories: Genomics & sequence analysis, Proteins & structural biology, Evolution & metagenomics, Mathematical biology & statistics, Tools & resources
- Authors: Stavrianidis, C., Duan, Y., Rhodes, G. E., Hayeck, T. J., Majoros, W. H., Allen, A. S.
- DOI: 10.64898/2026.09.13.745535
- Source URL: <https://doi.org/10.64898/2026.09.13.745535>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.13.745535>

Abstract: Different regions of genes perform distinct functions and vary in their importance to human health. Evolutionary intolerance provides a powerful means of identifying regions where disruptive mutations are under strong purifying selection, informing genetic disease discovery and variant interpretation. However, estimating intolerance in small sub-genic regions from population variation alone is underpowered and unstable. We present PRIME, a Bayesian model that stabilizes estimates of regional missense intolerance by sharing information hierarchically across regions. Importantly, PRIME produces a full joint posterior across all genes, allowing complex inferential questions that are difficult or impossible to address with existing approaches to be answered. We utilize this to identify regions enriched for pathogenic and experimentally deleterious missense variants, improve prioritization of Mendelian disease genes by focusing on their most intolerant regions, and uncover conserved patterns of purifying selection across protein families. Integrating PRIME with existing computational variant predictors improves pathogenicity prediction, demonstrating that regional missense intolerance provides complementary information for clinical variant interpretation.

## Programmable plant nutrition through synthetic transportome engineering.
- Source: Journal of plant physiology (journals)
- Date: 2026-09-16
- Categories: Single-cell & spatial, Systems & networks
- Authors: Yao Lu, Jia Li, Keke Yi, Xianqing Jia
- Journal: Journal of plant physiology
- DOI: 10.1016/j.jplph.2026.154874
- External ID: 42763965
- Keywords: multi omics, synthetic biology
- Source URL: <https://doi.org/10.1016/j.jplph.2026.154874>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1016%2Fj.jplph.2026.154874>

Abstract: Plant nutrient homeostasis emerges from the coordinated transport, partitioning, and storage of nutrients across cellular compartments, tissues, and developmental stages. These processes form an interconnected transport system in which local changes in nutrient uptake, redistribution, or sequestration can propagate across biological scales and ultimately shape whole-plant nutrient status and growth. This system-level organization suggests that nutrient homeostasis should be viewed not simply as the output of individual transporters, but as an emergent property of a coordinated transport network. Here, we propose the synthetic transportome as a transport-centered, systems-level framework for rationally designing native and/or engineered transport components, regulatory circuits, and spatial architectures to achieve predefined nutrient flux and allocation states. Unlike descriptive systems-level analyses, this framework treats nutrient engineering as an inverse-design problem, in which desired nutrient fluxes and physiological outputs guide the selection and coordination of transport modules. We further propose design principles and enabling technologies, integrating multi-omics, machine learning, structural biology, and synthetic biology. Synthetic transportome engineering provides a conceptual framework for quantitatively predictable and environmentally robust nutrient engineering, paving the way toward programmable nutrient utilization in crops.

## PyEuk: a tool suite for catalogue-free multilocus typing
- Source: bioRxiv (preprints)
- Date: 2026-09-16
- Categories: Genomics & sequence analysis, Tools & resources
- Authors: Kosakovsky Pond, S. L., Callan, D., Nekrutenko, A.
- DOI: 10.64898/2026.09.10.750732
- Source URL: <https://doi.org/10.64898/2026.09.10.750732>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.10.750732>

Abstract: Multilocus sequence typing anchors molecular epidemiology, but traditional frameworks require centrally curated allele catalogues. For emerging and uncultivable eukaryotic parasites, maintaining these databases is impractical, leaving surveillance reliant on fragmented, assay-specific scripts. PyEuk eliminates this bottleneck by providing an open, catalogue-free suite that calls microhaplotypes directly from sequence differences relative to a reference within data-defined genomic windows. When amplicon coordinates are uncharacterized or unpublished, PyEuk reconstructs target panels de novo from raw read coverage peaks mapped to a draft assembly. Across benchmark cohorts spanning Cyclospora cayetanensis and Plasmodium vivax, PyEuk recovers epidemiological structure established by tracebacks, geography, and clinical recurrence without organism-specific tuning. In foodborne outbreaks, it resolves independent transmission chains using either curated or de novo panels and scales to national surveillance archives exceeding 8,000 isolates. In P. vivax malaria, its weighted identity-by-state distance separates continental lineages, discriminates liver-stage relapses from reinfections, and delineates transmission clusters. Rather than forcing an arbitrary partition on continuous variation, PyEuk evaluates bootstrap stability, reporting supported cluster count ranges alongside reproducible transmission cores. PyEuk provides a portable, reproducible foundation for eukaryotic pathogen surveillance.

## RamiGlyph Captures Microglial Morphological Diversity and Predicts Functional States in Ischemia Reperfusion and Amyloid Pathology
- Source: bioRxiv (preprints)
- Date: 2026-09-16
- Categories: Biological imaging
- Authors: Qu, Y., Xiao, Y., Lan, T., Xu, J., Liu, J., Qian, Q., Liu, J., Chi, Y.
- DOI: 10.64898/2026.09.10.750561
- Source URL: <https://doi.org/10.64898/2026.09.10.750561>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.10.750561>

Abstract: Microglia are resident immune cells of the central nervous system, whose ramified processes rapidly remodel in response to injury. However, how to capture subtle morphological changes and whether functional state can be predicted from morphology remain open questions. Here, we present RamiGlyph, a contrastive learning framework integrating topological and structural features, trained on more than 20,000 reconstructed microglia. RamiGlyph not only distinguishes physiological and pathological states of microglia but also generalizes to neuronal cell type classification. Projection of microglia morphological embeddings revealed a continuum rather than discrete classes, from which a morphology score was derived to quantify dynamic process remodeling. To link morphology with function, Gromov Wasserstein optimal transport was used to align unpaired morphological and functional data across stages of ischemia reperfusion injury and amyloid pathology. These alignment results enable prediction of microglial functional states using RamiGlyph embeddings alone, with prediction reliability increasing upon cell aggregation. In summary, RamiGlyph provides a robust framework for resolving continuous microglial morphological variation and linking morphology to functional states across acute and chronic neuropathological contexts.

## Rapid Assessment of Size, Shape, and Chemical Complementarity of Ligands for Computational Protein Design
- Source: bioRxiv (preprints)
- Date: 2026-09-16
- Categories: Proteins & structural biology
- Authors: Petrenas, R., Ozga, K., Chubb, J. J., Romanyuk, A. V., Alibhai, D., McManus, J. J., Leggett, G. J., Scrutton, N. S., Oliver, T. A. A., Woolfson, D. N.
- DOI: 10.1101/2025.06.30.662286
- Source URL: <https://doi.org/10.1101/2025.06.30.662286>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1101%2F2025.06.30.662286>

Abstract: Driven by deep-learning approaches, computational protein design is advancing rapidly, and it is now possible to generate many de novo protein structures quickly and robustly. This sets new frontiers for the field, including designing proteins that bind small molecules tightly and specifically, and understanding the non-covalent interactions that underpin such designs to make binding predictable and tunable. Here we address these challenges with a rapid physics-based computational method to generate isosteric and chemically complementary binding pockets for small-molecule targets in de novo designed proteins. We test this experimentally by constructing and characterizing binding proteins for several synthetic and natural chromophores. By evaluating only single-digit numbers of designs, the pipeline delivers stable proteins with pre-organized binding sites confirmed by X-ray crystallography, which bind the targets selectively with micromolar affinities or better. To illustrate the scope and applications of this approach, we incorporate distinct and coupled chromophore-binding sites in a two-domain de novo protein enabling controlled energy transfer between the two sites, and we develop a small de novo binding protein that can be used in live mammalian cells to visualize sub-cellular structures.

## Rapid patient-specific neural networks for X-ray to volume registration.
- Source: Nature (journals)
- Date: 2026-09-16
- Categories: Biological imaging
- Authors: Vivek Gopalakrishnan, David-Dimitris Chlorogiannis, Andrew Abumoussa, Anna M Larson, Nazim Haouchine, Darren B Orbach, Sarah Frisken, Neel Dey, Polina Golland
- Journal: Nature
- DOI: 10.1038/s41586-026-11045-x
- External ID: 42749809
- Source URL: <https://doi.org/10.1038/s41586-026-11045-x>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1038%2Fs41586-026-11045-x>

Abstract: Advanced navigation techniques in image-guided interventions and surgical robotics require the rapid and precise alignment of three-dimensional (3D) preoperative volumes (such as computed tomography and magnetic resonance imaging) to two-dimensional (2D) intraoperative images (such as X-ray fluoroscopy)1,2. However, existing 2D/3D registration methods fail to generalize across the broad spectrum of fluoroscopy-guided procedures: intensity-based optimizers require per-individual hyperparameter tuning3,4, while deep-learning approaches demand extensive manually labelled datasets and remain constrained to the specific anatomy on which they were trained5,6. Here, to address these limitations, we present xvr-a self-supervised framework that combines patient-specific neural networks with gradient-based optimization for automatic 2D/3D registration. xvr uses physics-based simulation to generate training data from a patient's own preoperative scan, eliminating the need for manual annotation. We present a foundation model pretrained on thousands of whole-body scans, achieving patient-specific adaptation to any anatomical region with only 5 min of fine-tuning. In to our knowledge the largest evaluation of 2D/3D registration on real fluoroscopy to date, xvr achieves high accuracy in seconds across diverse anatomical structures, volumetric imaging modalities and hospitals, improving on the accuracy of existing methods by an order of magnitude. xvr makes pan-anatomical 2D/3D rigid registration accessible to broad clinical and research communities through open-source software available online.

## Reimagining research papers as interactive and reliable AI agents.
- Source: Nature (journals)
- Date: 2026-09-16T00:00:00Z
- Categories: Genomics & sequence analysis, Single-cell & spatial
- Authors: Jia-Cheng Miao, Joe R. Davis, Yaohui Zhang, Jonathan K. Pritchard, James Zou
- Journal: Nature
- DOI: 10.1038/s41586-026-11044-y
- External ID: bfa74c3aa2d1223a816f30865954bb9e250a877f
- Keywords: genomic, transcriptomics, single cell, spatial transcriptomics
- Source URL: <https://doi.org/10.1038/s41586-026-11044-y>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1038%2Fs41586-026-11044-y>

Abstract: Here we introduce Paper2Agent, an automated framework that converts research papers into artificial intelligence (AI) agents. Paper2Agent transforms research output from passive artefacts into active systems that accelerate use and discovery. Conventional research papers require readers to understand and adapt the paper's code, data and methods to their work, creating barriers to dissemination and reuse. Paper2Agent addresses this challenge by converting a paper into an AI agent that functions as a virtual corresponding author, exposing its manuscript, supplementary materials, datasets, code and workflows as active, agent-native knowledge rather than static text. It analyses the paper and codebase using multiple agents to construct a model context protocol (MCP) server, then generates and runs tests to refine and increase robustness of the MCP. These paper MCPs can be connected to a chat agent (such as Claude Code) to carry out complex scientific queries through natural language while invoking tools and workflows from the paper. We demonstrate Paper2Agent's effectiveness through case studies. Paper2Agent created an agent that leveraged AlphaGenome1 to interpret genomic variants and agents based on Scanpy2 and TISSUE (transcript imputation with spatial single-cell uncertainty estimation)3 to conduct single-cell and spatial transcriptomics analyses. We validate that these agents reproduce the results of the original papers and carry out novel user queries. Paper2Agent created multiple agents that collaborate to prioritize a causal gene for psoriasis. By turning static papers into interactive AI agents, Paper2Agent introduces a paradigm for knowledge dissemination and a collaborative ecosystem of AI co-scientists.

## Rsearch: An R interface to VSEARCH supporting visualization and parameter tuning
- Source: bioRxiv (preprints)
- Date: 2026-09-16
- Categories: Tools & resources
- Authors: Stamsaas, C., Rognes, T., Rudi, K., Snipen, L., Vinje, H.
- DOI: 10.64898/2026.09.10.750626
- Source URL: <https://doi.org/10.64898/2026.09.10.750626>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.10.750626>
- Code: <https://github.com/CassandraHjo/Rsearch>

Abstract: Background: We present Rsearch, an R package that integrates the core functionality of VSEARCH into the R environment and extends it with visualization, parameter optimization, and conversion tools for compatibility with other R packages. By making VSEARCH directly accessible in R, Rsearch lowers the barrier for using VSEARCH and integrating it with downstream statistical and ecological analyses. Results: Comparative analysis with DADA2 using mock community data showed that both pipelines produced relative abundance profiles highly correlated with the expected composition. Compared to DADA2, Rsearch identified fewer OTUs, but these were more consistently prevalent across samples. In contrast, DADA2 appeared to overestimate diversity by splitting sequences into an excessive number of OTUs. In terms of computational performance, vs\_cluster\_unoise implemented in Rsearch was the fastest of all the clustering and denoising methods, while other Rsearch functions showed runtimes comparable to DADA2. In addition, Rsearch provides functions for systematic optimization of trimming and filtering parameters, an important feature for users who may not otherwise have a clear strategy for parameter selection. The package also includes functions to ensure compatibility with other R packages such as phyloseq. Conclusions: Rsearch offers a practical and accessible framework for analysing metabarcoding data within a single analytical environment and is freely available from The Comprehensive R Archive Network, with the development version hosted on GitHub (https://github.com/CassandraHjo/Rsearch).

## Sample-Specific Generalized Cross-Validation for Gene Network Analysis of Cytarabine Response in Cancer Cell Lines
- Source: International Journal of Molecular Sciences (journals)
- Date: 2026-09-16T00:00:00Z
- Categories: Genomics & sequence analysis, Systems & networks
- Authors: J. Oh, Heewon Park
- Journal: International Journal of Molecular Sciences
- DOI: 10.3390/ijms27188261
- External ID: 834919e629c9bf3e219a34a351d97566327330e9
- Source URL: <https://doi.org/10.3390/ijms27188261>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.3390%2Fijms27188261>

Abstract: Sample-specific gene regulatory network analysis can reveal molecular heterogeneity associated with individual characteristics, such as anticancer drug sensitivity. The varying coefficient model with kernel-based L1 regularization enables the estimation of such networks, but its performance depends strongly on hyperparameter selection. Conventional cross-validation is computationally intensive and provides only an averaged evaluation across samples, limiting its suitability for sample-specific analysis. To address these limitations, we propose doubleS-GCV, a sample-specific generalized cross-validation criterion for selecting hyperparameters in sample-specific gene network estimation. DoubleS-GCV provides a separate model evaluation for each sample while substantially reducing computational burden. Monte Carlo simulations demonstrated that doubleS-GCV achieved accurate gene selection and network estimation and outperformed conventional information criteria, including AIC, BIC, AICC, and HQC. Application to GDSC cancer cell lines identified Cytarabine sensitivity-specific gene networks and candidate biomarkers supported by previous studies. The estimated networks also exhibited nonlinear structural changes across Cytarabine sensitivity levels, indicating that molecular interactions vary with drug response. These results demonstrate that doubleS-GCV provides an efficient and reliable model selection framework for sample-specific gene network analysis.

## scACORN: Context-engineered agent orchestration of specialized small language models for single-cell transcriptomic interpretation
- Source: bioRxiv (preprints)
- Date: 2026-09-16
- Categories: Genomics & sequence analysis, Single-cell & spatial
- Authors: Rasti-Meymandi, A., Nahali, S., Paramithiotis, E., Cheung, A. M., Dolatabadi, E.
- DOI: 10.64898/2026.09.10.750801
- Source URL: <https://doi.org/10.64898/2026.09.10.750801>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.10.750801>

Abstract: Single-cell atlases now exceed 66 million cells, but turning a ranked expression profile and a free-form biological question into a reliable, evidence-grounded answer remains unsolved. Scaling a single model does not resolve this, because single-cell interpretation is a heterogeneous family of tasks whose correct answer depends on tissue, cohort, perturbation and annotation resolution. Here we present scACORN, an agentic alternative to monolithic single-cell language models that combines specialized small language models with context-engineered agent orchestration for their selection and composition at inference time. Each expert is built in two stages: domain-aligned contrastive adaptation fits a pretrained cell-to-text backbone to the transcriptomic geometry of a target dataset, and geometry-preserving specialization learns question-conditioned biological completions without eroding that geometry. A fixed orchestrating language model agent then selects and combines experts under a natural-language playbook that is itself optimized from textual feedback, with no gradient updates to the orchestrator. Across 10 Tabula Sapiens tissues, domain alignment raised transfer macro-F1 from 0.36 to 0.64 and Recall@5 from 0.87 to 0.97; specialized experts reached 0.89 mean exact-match annotation accuracy; and playbook optimization reduced unsupported gene citations from 14.5% to 3.5%. Our findings support specialization and orchestration as complementary responses to the heterogeneity and evidentiary demands of single-cell analysis.

## Scalable near-real-time Bayesian phylogenetics for outbreaks with Delphy.
- Source: Nature (journals)
- Date: 2026-09-16T00:00:00Z
- Categories: Genomics & sequence analysis, Evolution & metagenomics, Tools & resources
- Authors: P. Varilly, Mark Schifferli, Katherine Yang, P. Cronan, I. Specht, T. Burcham, O. Glennon, Olivia Jacks, E. Laning, L. Marrs, K. Oba, Shannon Yeung, Karlie Zhao, E. Parker, I. Omah, Jonathan E. Pekar, Laura Luebbert, Kristian G. Andersen, Daniel J. Park, Stephen F. Schaffner, B. MacInnis, C. Happi, Jacob E. Lemieux, A. Ozonoff, Michael D. Mitzenmacher, Ben Fry, P. Sabeti
- Journal: Nature
- DOI: 10.1038/s41586-026-11012-6
- External ID: 306ea311cfec54c6acdab779e63ca39e50649962
- Source URL: <https://doi.org/10.1038/s41586-026-11012-6>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1038%2Fs41586-026-11012-6>

Abstract: Pathogen genomic analysis is central to tracking, understanding and containing outbreaks1-13, but the complexity and cost of state-of-the-art phylogenetic tools limit global access and impact. Here we introduce Delphy, an exact reformulation of Bayesian phylogenetics14-17 designed to transform its speed, scalability and accessibility while retaining Bayesian state-of-the-art accuracy. Delphy's central data structure, an explicit mutation-annotated tree, takes advantage of the high sequence similarity of large-scale epidemic datasets18-20 for efficient tree exploration and convergence. By reproducing key analyses from recent major epidemics, including Ebola1,21, Zika2, SARS-CoV-2 (ref. 22), mpox3,4 and H5N1 (refs. 23,24), we demonstrate state-of-the-art accuracy with up to 2-3 orders of magnitude improvements in speed. Assessing Delphy's scalability, we show that a simulated dataset of 100,000 sequences can be analysed within a day. We distribute Delphy as a client-side web application that enables local, interactive analysis of raw data on the user's machine. Delphy automatically identifies key viral lineages and mutations, as well as their emergence and prevalence through time, with quantified uncertainties grounded in Bayesian theory. Delphy establishes Bayesian phylogenetics as a fast, accessible frontline tool for future outbreak response.

## Seamless dose optimization design accounting for unknown patient heterogeneity in cancer clinical trials
- Source: Biometrics (journals)
- Date: 2026-09-16T00:00:00+00:00
- Authors: Rebecca B Silva, Bin Cheng, Shing M Lee
- Journal: Biometrics
- DOI: 10.1093/biomtc/ujag161
- Source URL: <https://doi.org/10.1093/biomtc/ujag161>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1093%2Fbiomtc%2Fujag161>

Abstract: Project Optimus, an initiative by the FDA’s Oncology Center of Excellence, seeks to reform the dose-optimization and dose-selection paradigm in oncology. We propose a dose-optimization design that considers plateau efficacy profiles, integrates pharmacokinetic data to inform the exposure-toxicity curve, and accounts for patient characteristics that may contribute to heterogeneity in response. The dose-optimization design is carried out in two stages. First, a toxicity-driven stage estimates a safe set of doses. Then, a dose-ranging efficacy-driven stage explores the set using response and patient characteristic data, employing Bayesian Sparse Group Selection to understand patient heterogeneity. Between stages, the design integrates pharmacokinetic data and uses futility assessments to identify the target population among the general phase I patient population. An optimal dose is recommended for each identified subpopulation within the target population. The simulation study demonstrates that a model-based approach to identifying the target population can be effective; patient characteristics relating to heterogeneity were identified, and different optimal doses were recommended for each identified target subpopulation. Most designs that account for patient heterogeneity are intended for trials where heterogeneity is known, and pre-defined subpopulations are specified. However, given the limited information at such an early stage, subpopulations should be learned through the design.

## Self-buckling of undulating flagella: an elastohydrodynamic mechanism for double waves in spermatozoa
- Source: bioRxiv (preprints)
- Date: 2026-09-16
- Categories: Mathematical biology & statistics
- Authors: Htet, P. H., Ishimoto, K.
- DOI: 10.64898/2026.09.11.750800
- Source URL: <https://doi.org/10.64898/2026.09.11.750800>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.11.750800>

Abstract: The relatively long flagella of spermatozoa from insects, birds, and octopuses display double waves, characterized by two superimposed helical waves. The prevalance of these highly organized waveforms across diverse taxa and distinct flagellar architectures hints at shared underlying physics, motivating a model of the flagellum as an elastic filament immersed in a viscous fluid, actively driven by a single set of internal bending moment waves. Simulations of a clamped filament show that it can buckle under its own activity into whirling and flapping states. A multiple-scales analysis of the elastohydrodynamic equations reveals how nonlinear interactions between fast undulations generate an effective compression driving buckling, and connects wave-driven buckling to classical follower-force instabilities. Extending the model to a swimming spermatozoon, the same instability produces double waves. Parameter estimates across species show that most observed double waves lie within the regime where buckling is permitted, supporting self-buckling as a generic physical mechanism for double waves.

## Self-organized mechanochemical instabilities drive the emergence of digit tissue morphogenesis
- Source: bioRxiv (preprints)
- Date: 2026-09-16
- Categories: Mathematical biology & statistics
- Authors: Tsutsumi, R., Diez, A. N., Plunder, S., Kimura, R., Oki, S., Takizawa, K., Nakano, R., Akiyama, H., Takada, R., Takada, S., Musy, M., Sharpe, J., Eiraku, M.
- DOI: 10.1101/2025.08.31.673315
- Source URL: <https://doi.org/10.1101/2025.08.31.673315>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1101%2F2025.08.31.673315>

Abstract: The emergence of complex anatomical structures -such as the hands- from unstructured tissues remains a fundamental question in developmental biology. Turing-type reaction-diffusion models have provided a molecular explanation for the periodic pre-patterning of digits; however, the physical principles driving 3D morphogenesis remain incompletely understood. To identify the biophysical design principles leading to digit formation, we develop a limb-mesenchymal organoid system that spontaneously forms elongated, digit-like protrusions. Iterations between experiments and agent-based models at the cellular level identify sufficient microscopic mechanisms leading to morphogenesis of digit-like structures: symmetry-breaking and the elongation of digits result from a combination of differential cell adhesion and morphogen-induced chemotaxis and convergent-extension. Lastly, to describe tissue-scale deformations, we perform a coarse-graining analysis of the agent-based model and derive a continuum model that reveals a structural analogy to Cahn-Hilliard-type equations. These equations are typically used to describe fluid phase separation and so-called ''fingering instabilities'' in fluid physics. Here, we show that they also accurately describe organoid morphogenesis. These findings suggest that ''finger'' formation is driven by a mechanical fingering instability acting in concert with chemical patterning, shedding a new light on vertebrate limb morphogenesis.

## Self-Organized Pattern Formation of a Common Tropical Alga.
- Source: Journal of theoretical biology (journals)
- Date: 2026-09-16
- Authors: Dylan E McNamara, Conner W Lester, Clinton B Edwards, Jennifer E Smith, Stuart A Sandin
- Journal: Journal of theoretical biology
- DOI: 10.1016/j.jtbi.2026.112597
- External ID: 42749013
- Source URL: <https://doi.org/10.1016/j.jtbi.2026.112597>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1016%2Fj.jtbi.2026.112597>

Abstract: Recent in-situ observations within a tropical coral reef have revealed novel polygonal patterns of the calcifying green alga Halimeda. The observed patterns showed no evidence of a matching exogenous template in structural reef morphology, distribution of biological competitors for space, or other environmental factors, suggesting that pattern formation is consistent with endogenous, nonlinear dynamics. A simplified, spatially explicit numerical model is proposed that simulates a feedback whereby Halimeda preferentially grows in regions less conducive to the growth of corals (when corals are the dominant spatial competitor), and coral growth is inhibited in regions of dense Halimeda. Model results reveal self-organized emergent polygons of Halimeda cover that qualitatively match observations.

## Sex-specific biological aging clocks across organs and omics.
- Source: Nature medicine (journals)
- Date: 2026-09-16T00:00:00Z
- Authors: Zhi-Yuan Song, Derek Feng, Naowal Azraf Rahman, Michael R. Duggan, Qu Tian, Jian Zeng, Xia Zhou, Chun-Rui Zou, M. Rafii, Li Shen, Paul M. Thompson, E. Simonsick, Keenan A. Walker, A. Zalesky, C. Davatzikos, Paul Aisen, L. Ferrucci, S. Resnick, Jun-Hao Wen
- Journal: Nature medicine
- DOI: 10.1038/s41591-026-04662-6
- External ID: 6cdefed63c70bbfff598aaaf05db6e558422441b
- Source URL: <https://doi.org/10.1038/s41591-026-04662-6>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1038%2Fs41591-026-04662-6>

Abstract: Sex differentially shapes aging, neurodevelopment and neurodegenerative diseases such as Alzheimer's disease (AD). However, most biological aging clocks (artificial intelligence-predicted age minus chronological age) were trained on sex-pooled samples and implicitly assume sex invariance.Here we developed 38 sex-specific biological aging clocks across 15 organ systems. We first demonstrate the importance of sex-stratified training for constructing sex-specific healthy normative references and then reveal marked divergence between female and male clocks. Key genetic parameters and Mendelian randomization results indicate that organ-specific aging liability and its relationships to cardiometabolic, endocrine and mental traits are configured differently in females and males. Proteomic analyses identify distinct, organ-resolved synaptic, immune, vascular and metabolic networks that differentially track female and male biological aging. In longitudinal survival analyses, sex-specific clocks predict whole-body systemic diseases and all-cause mortality in a sex-dependent and organ-dependent manner. Further analyses reveal sex-dependent associations between the brain aging clock and cognitive decline trajectory during a preclinical AD clinical trial. Sex-stratified clocks may offer distinct value by defining biological age against sex-appropriate normative references and revealing sex-dependent genetic, molecular and clinical signatures that pooled models may obscure. Meanwhile, sex-pooled and sex-interaction approaches remain valuable, as human aging and disease also share fundamental biological similarities between females and males. Together, these findings reveal sex-specific biological aging signatures in aging, AD and systemic health, highlighting the need for explicitly sex-stratified modeling approaches.

## SixPack-AbScan: a web server to discover cross-reactivity of antibodies across species
- Source: bioRxiv (preprints)
- Date: 2026-09-16
- Categories: Tools & resources
- Authors: Grillo, M.
- DOI: 10.64898/2026.09.10.746162
- Source URL: <https://doi.org/10.64898/2026.09.10.746162>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.10.746162>

Abstract: Availability of commercial antibodies for immunochemistry is typically limited to major model species. Researchers working on non-model organisms therefore often need either to generate novel species-specific antibodies or venture into expensive empirical screening from available antibody catalogues, hoping to find cross-reactive reagents. SixPack-AbScan is a free web server aiding researchers in transferring antibodies across species: using the available epitope-mapping information, the software performs a simple computational pre-screening of potential cross-reactivity. The workflow is species-agnostic and designed to help non-model-species researchers prioritize antibodies for experimental validation. A hit indicates sequence-level conservation of the epitopes and a high probability of cross-reactivity; the server does not model substitutions, structure, accessibility, expression or binding affinity. The web server is available at https://sixpack-abscan.serve.scilifelab.se.

## SpaMOAL is a deep learning method that enables accurate spatial domain identification from multi-omics data
- Source: PLOS Biology (journals)
- Date: 2026-09-16T00:00:00+00:00
- Categories: Genomics & sequence analysis, Single-cell & spatial, Biological imaging
- Authors: Jinxia Wang, Yuying Huo, Rui Zhao, Yan Pan, Jianqiang Wu, Han Wang, Xiangyu Li
- Journal: PLOS Biology
- DOI: 10.1371/journal.pbio.3003690
- Source URL: <https://doi.org/10.1371/journal.pbio.3003690>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1371%2Fjournal.pbio.3003690>

Abstract: Recent advances in spatial multi-omics technologies have opened new avenues for characterizing tissue architecture and function in situ, by simultaneously providing multimodal and complementary information—such as spatially resolved transcriptomic, epigenomic, and proteomic features. Current computational approaches face substantial challenges, such as effective integration of multi-omics molecular information with spatial information and corresponding high-resolution histology images. To address this challenge, we proposed SpaMOAL ( Spa tially M ulti- O mics graph contr A stive L earning), a graph-based contrastive learning approach for spatial domain identification. SpaMOAL learns clustering-friendly representations from spatial multi-omics data by integrating spatial coordinates, histological image features, and molecular profiles, enabling accurate delineation of spatial tissue domains. Benchmarking across multiple recent paired spatial multi-omics datasets from mouse and human demonstrated that SpaMOAL consistently outperforms existing methods. By enabling accurate spatial domain delineation, SpaMOAL provides a powerful framework for interpreting tissue organization and cellular microenvironments.

## stably: error-controlled stability selection for biomarker panel discovery
- Source: bioRxiv (preprints)
- Date: 2026-09-16
- Categories: Proteins & structural biology, Tools & resources
- Authors: Byrne, D., McNamara, M., Unwin, R.
- DOI: 10.64898/2026.09.11.750833
- Source URL: <https://doi.org/10.64898/2026.09.11.750833>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.11.750833>
- Code: <https://github.com/byrnedaniel5-eng/stably>

Abstract: Motivation Data-independent acquisition mass spectrometry (DIA-MS) has become increasingly popular for clinical proteomics due to its sensitivity and reproducibility. Univariate statistical analysis tools, such as limma and MSstats, are widely used for identifying differentially abundant proteins but cannot capture multivariate relationships between proteins that may provide greater discriminatory power as a panel. Machine learning approaches can address this gap, but typically prioritise predictive performance over feature stability, producing biomarker panels for downstream validation that vary depending on data splitting and are poorly suited to clinical translation. Results We developed stably, a Python package that implements stability selection with formal false positive control for DIA proteomics data. Using synthetic data with known ground-truth biomarkers, we show that the Shah and Samworth complementary pairs stability selection framework recovers more true synthetic biomarkers than the Meinshausen and Buhlmann framework at moderate effect sizes typical of proteomics (d = 0.5 - 2.0), while both maintain false positive rates well below their theoretical guarantees. Applied to a publicly available serum proteomics dataset from patients with all stages of pancreatic ductal adenocarcinoma (n=176), stably identified a stable 17-protein biomarker panel in the discovery cohort (n=120), which achieved higher predictive power (AUC = 0.93) in the validation cohort (n=56) than the panel selected by Byeon et al. (2024)(AUC 0.82). stably represents a principled, error-controlled method for biomarker panel discovery for translation into second cohorts. Availability and implementation stably is available on GitHub (https://github.com/byrnedaniel5-eng/stably); the version used in this study is archived on PyPI (https://pypi.org/project/stably/).

## Stacked EEG spectrograms and an attention-augmented CNN-LSTM for subject-independent emotion recognition.
- Source: Journal of neuroscience methods (journals)
- Date: 2026-09-16
- Authors: Guiyoung Son
- Journal: Journal of neuroscience methods
- DOI: 10.1016/j.jneumeth.2026.110906
- External ID: 42748977
- Source URL: <https://doi.org/10.1016/j.jneumeth.2026.110906>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1016%2Fj.jneumeth.2026.110906>

Abstract: BACKGROUND: EEG provides direct neural measurements with high temporal resolution for emotion recognition. However, many spectrogram-based approaches process channels independently or integrate channel information only at later stages. NEW METHOD: We propose a stacked spectrogram representation that vertically concatenates channel-wise EEG spectrograms, preserving temporal, spectral, and channel-structured information at the input level. This representation is combined with an attention-augmented CNN-LSTM architecture for feature extraction, temporal modeling, and adaptive weighting of informative segments. RESULTS: Under LOSO cross-validation, the proposed framework achieved 82.34% accuracy and a macro-F1 score of 0.83 in four-class emotion classification. On SEED-IV, it achieved 84.51% accuracy and a macro-F1 score of 0.86. COMPARISON WITH EXISTING METHOD: The proposed method outperformed single-channel CNNs, stacked VGG16, and CNN-LSTM baselines, demonstrating competitive performance with lower computational complexity. CONCLUSIONS: Input-level channel integration effectively improves subject-independent EEG emotion recognition and provides a practical solution for consumer-grade EEG applications.

## Structure-informed theoretical modeling defines principles governing avidity in bivalent protein interactions.
- Source: The Journal of biological chemistry (journals)
- Date: 2026-09-16
- Categories: Proteins & structural biology
- Authors: Reagan Portelance, Anqi Wu, Alekhya Kandoor, Kristen M Naegle
- Journal: The Journal of biological chemistry
- DOI: 10.1016/j.jbc.2026.113559
- External ID: 42749259
- Source URL: <https://doi.org/10.1016/j.jbc.2026.113559>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1016%2Fj.jbc.2026.113559>

Abstract: In signaling cascades, signaling proteins often encode multiple domains or motifs, which presents the possibility for avidity -- where multivalent binding drastically increases interaction strength and duration. However, predicting and validating multivalent interactions that interact with avidity is a challenge. Here, we integrate mechanistic modeling, structure-based analysis, and experimental approaches as a framework for defining the conditions under which avidity plays a role. We explore the tandem SH2 domain family of interactions with bisphosphorylated partners as a multivalent archetype, which encompasses key secondary messengers in tyrosine kinase signaling networks. Theoretical modeling suggests that maximum avidity occurs with closely spaced tyrosine phosphorylation sites combined with moderate monovalent affinities - exactly around the innate range of SH2 domain affinity - or with phosphorylation sites separated by sufficiently flexible linkers. Surprisingly, despite sequence diversity, structure-based analysis showed relatively conserved three-dimensional spacing between SH2 domains across all tandem SH2 families, which we corroborate experimentally, suggesting evolutionary optimization for avidity interactions. The combination of structure-based analysis of domain spacing with available monovalent experimental data appears, along with iterative experimental refinement of biophysical parameters, can identify high affinity interactions of tandem SH2 domain recruitment to the EGFR C-terminal tail. Using these principles, we extended bivalent predictions into the full phosphoproteome space and structural parameterization of other partners of SH2 domain binding, providing resources and methods for more rapid expansion of bivalent analysis. These approaches lay the groundwork for larger utility in multivalent prediction and testing to help better understand protein interactions that drive cell signaling.

## Structured cross-omics interaction discovery with a triple-graph model
- Source: bioRxiv (preprints)
- Date: 2026-09-16
- Categories: Single-cell & spatial, Systems & networks
- Authors: YU, J., Lin, H., Chen, S.
- DOI: 10.64898/2026.09.10.749540
- Source URL: <https://doi.org/10.64898/2026.09.10.749540>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.10.749540>

Abstract: Multi-omics analyses often yield fragmented pairwise associations that obscure coordinated relationships among molecular features. We developed TriGer, a triple-graph framework that identifies many-to-many cross-omics modules by combining cross-layer associations with dependency structures within each layer. In simulations with sparse or nested signals, TriGer recovered planted modules while balancing sensitivity and specificity. In inflammatory bowel disease, it identified subtype-associated metabolite--transcript modules; in colorectal cancer, it identified genus--metabolite modules whose organization was attenuated in cancer. TriGer provides an interpretable approach for studying coordinated cross-omics structure in high-dimensional molecular data.

## SwinePan for pig graph-based pangenome and multiomics data mining
- Source: Genome Research (journals)
- Date: 2026-09-16T00:00:00+00:00
- Categories: Genomics & sequence analysis, Tools & resources
- Authors: Meng Lin, Langqing Liu, Gengyuan Cai, Sixiu Huang, Yibin Qiu, Zekai Yao, Shaoxiong Deng, Shiyuan Wang, Yiyi Liu, Donglin Ruan, Fuchen Zhou, Jiajin Wu, Zebin Zhang, Enqin Zheng, Jie Yang, Zhenfang Wu
- Journal: Genome Research
- DOI: 10.1101/gr.281750.125
- Source URL: <https://doi.org/10.1101/gr.281750.125>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1101%2Fgr.281750.125>

Abstract: Pigs are one of the most important livestock species worldwide. Although multiple high-quality reference genomes exist, reliance on a single linear reference limits the detection of structural variants (SVs) and the characterization of population-specific genetic diversity. To address this limitation, we developed SwinePan, a comprehensive and integrated multiomics database for pigs built on a graph-based pangenome framework. SwinePan incorporates a variome derived from the graph-based pangenome, covering 2,598 individuals across 35 breeds, including 185,759 SVs, 117 million SNPs, and 6.8 million indels. The database also integrates transcriptomic data from liver, loin muscle, abdominal fat, and backfat, along with over 150,000 phenotypic records. The online toolkit deployed in SwinePan enables genome-wide association studies (GWAS), expression quantitative trait locus (eQTL) mapping, and colocalization, while interactive modules visualize population structure and multiomics associations, streamlining candidate gene and variant exploration. Additionally, two proof-of-concept analyses demonstrate how SwinePan pinpoints trait-associated loci and deciphers their potential regulatory mechanisms.

## System biology analysis reveals circadian rhythm disorder associated with development and progression in colorectal cancer
- Source: npj Precision Oncology (journals)
- Date: 2026-09-16T00:00:00Z
- Categories: Genomics & sequence analysis
- Authors: Shi-Qian Zhang, Shan-Shan Cai, Nai-Jing Hou, Qian Guo, Pengpeng Zhang, Zhi-Jie Zhao, Song-Bin Guo, Xu-Feng Huang, Hua-Qing Wang, Hao-Nan Zhang, Chao-Yang Yu, Ru-Hao Wu, Chun-Ze Zhang, S. Tam, Ge Zhang
- Journal: npj Precision Oncology
- DOI: 10.1038/s41698-026-01699-1
- External ID: 600e2aa933f76cd2de198d0da17f58b26fed39f7
- Source URL: <https://doi.org/10.1038/s41698-026-01699-1>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1038%2Fs41698-026-01699-1>

Abstract: Circadian rhythm disorders represent an abstract concept lacking standardized quantitative metrics. Existing circadian indicators, including traditional rhythm parameters and a limited set of clock gene or physiological biomarkers, are insufficient to robustly capture steady-state endogenous circadian homeostasis in complex disease contexts, thereby constraining quantitative assessment of circadian disruption and limiting its translational applicability. Chronic circadian rhythm disruption is associated with various diseases, including metabolic disorders and malignancies. However, the mechanisms by which circadian disruption influences tumor microenvironment formation and colorectal cancer progression remain incompletely understood. This study employs systems biology analysis to decipher the molecular characteristics of circadian rhythm disruption in colorectal cancer progression. We analyzed single-cell RNA sequencing data from 13 CRC tissue samples and 12 normal mucosal samples, combined with 3733 samples from 34 public batch RNA, microarray, and single-cell RNA sequencing cohorts. We developed and validated the ClockProCRC system, which detects and quantifies intrinsic circadian misalignment in CRC. The ClockProCRC score elucidates how circadian misalignment drives CRC progression trajectories, shapes clinical phenotypes, regulates disease manifestations, and reshapes the tumor microenvironment. SYNE1 gene was identified as a key mediator of circadian misalignment, promoting tumorigenesis by driving epithelial-like phenotypic conversion and demonstrating therapeutic potential in colorectal cancer management. This study establishes a foundation for integrating rhythmic information into clinical practice and advances circadian biology research in the field of CRC.

## Targeted ortholog search in unannotated genome assemblies with fDOG-Assembly
- Source: bioRxiv (preprints)
- Date: 2026-09-16
- Categories: Genomics & sequence analysis, Tools & resources
- Authors: Muelbaier, H., Arthen, F., Tran, V., Schaefer, I., Balint, M., Ebersberger, I.
- DOI: 10.1101/2025.09.19.677253
- Source URL: <https://doi.org/10.1101/2025.09.19.677253>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1101%2F2025.09.19.677253>

Abstract: Whole genome shotgun sequencing and assembly is routine. However, identifying protein-coding genes in newly assembled genomes remains complex, time-consuming, and labour-intensive. Therefore, most eukaryotic genome assemblies in public databases lack gene annotations reducing their value for evolutionary and functional genomics. Here, we present fDOG-Assembly, a novel tool for targeted, feature architecture-aware ortholog searches directly in unannotated genome assemblies. Benchmarking shows that fDOG-Assembly performs similarly to BUSCO and Compleasm in ortholog identification while offering the advantage of not being restricted to universal single-copy genes. Applied to identify orthologs of 5,000 human genes in rat and Nematostella vectensis, fDOG-Assembly approaches the performance of traditional ortholog search tools that rely on pre-annotated proteomes. Importantly, it can recover orthologs missed by conventional methods because of incomplete gene annotations, helping to fill gaps in phylogenetic profiles. As a case study, we screened 176 soil invertebrate genome assemblies for genes involved in antibacterial compound production. We found that orthologs of \{beta\}-lactam biosynthesis genes are widespread in springtails, with individual species possessing nearly complete cephamycin biosynthetic gene sets, suggesting they may represent previously unrecognized natural producers of \{beta\}-lactam antibiotics. Overall, fDOG-Assembly is a powerful resource for orthology-based analyses of the rapidly growing collection of unannotated genome assemblies.

## Task-Specific Quality Gating for Retinal Optical Coherence Tomography B-Scans: Learned Representations Over Scalar Metrics in Choroid Segmentation
- Source: medRxiv (preprints)
- Date: 2026-09-16
- Categories: Biological imaging
- Authors: Hiras, A., Jayaraman, A., Gadari, A., Mankumare, A. S., Mynampati, A., Chhablani, J. K., Bollepalli, S. C., Vupparaboina, K. K.
- DOI: 10.64898/2026.09.14.26363079
- Source URL: <https://doi.org/10.64898/2026.09.14.26363079>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.14.26363079>

Abstract: Automated segmentation of Optical Coherence Tomography (OCT) images is a critical component of structural biomarker extraction for retinal diagnostics. Deep learning models achieve state-of-the-art performance on controlled datasets, yet exhibit unpredictable failures on real-world data. Current quality gates rely on device-reported scan quality scores, which have been shown to be unreliable predictors of segmentation performance. We define scan quality in a task-specific sense that is, whether a given B-scan will yield a reliable segmentation from a particular trained model. Under this definition, we perform a systematic evaluation of No-Reference Image Quality Assessment (NR-IQA) metrics, general-purpose and domain-specific pretrained representations as alternative quality gates. To this end, we use choroid segmentation as the prototype task, with a dataset of 6,076 OCT B-scans from 80 subjects. These quality gate candidates are evaluated at three levels: scalar metrics (BRISQUE, NIQE, PIQE, SNR, PSNR), supervised linear probing, and unsupervised partitioning (K-Means) of the feature vectors and the learned representations. All scalar NR-IQA metrics proved inadequate (|r| < 0.20). General-purpose ImageNet-based pretrained representations (EfficientNet-b0, ResNet-50, ViT-B/16) improve upon NR-IQA, achieving ROC-AUC up to 0.77, indicating that learned representations are better suited to task-specific quality gating than hand-crafted scalar statistics. Retinal foundation models (FMs) further close the gap: RETFound (OCT-specific FM) achieves ROC-AUC \{approx\} 0.81. Unsupervised K-Means partitioning indicates that general-purpose ImageNet-pretrained embeddings, despite carrying a linearly decodable quality signal, do not reliably organize scans by quality geometrically, whereas the retinal FMs produce quality-aligned clusters that exceed a patient-level permutation null, suggesting that domain-specific pretraining provides additional, complementary benefit on top of general-purpose learned representations.

## Test-Retest Reproducibility of Single- and Cross-Population White Matter Atlases in Diffusion MRI Tractography
- Source: bioRxiv (preprints)
- Date: 2026-09-16
- Categories: Biological imaging, Computational neuroscience
- Authors: Li, Y., Wang, X., Sun, J., Zhang, W., Wu, Y., Yin, L., Chen, Y., Rathi, Y., Makris, N., O'Donnell, L. J., Zhang, F.
- DOI: 10.64898/2026.09.11.750840
- Source URL: <https://doi.org/10.64898/2026.09.11.750840>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.11.750840>

Abstract: Diffusion MRI tractography enables noninvasive mapping of white matter fiber tracts. Atlas-based white matter parcellation supports automated tract identification by assigning individual streamlines to atlas-defined clusters and anatomical tract labels. Because clustering-based white matter atlases are constructed from cohort-specific tractography data, they capture common white matter organization represented in the atlas-construction population. Although major white matter anatomy is shared across populations, subtle population-related anatomical variability may influence atlas representation and test-retest correspondence. Therefore, the reproducibility and cross-population generalizability of tractography-based white matter atlases are important considerations for quantitative neuroimaging studies. In this study, we evaluated whether incorporating cross-population anatomical variability during atlas construction improves test-retest reproducibility. To do so, we compared a single-population ORG atlas constructed from a Western cohort with the cross-population East-West White Matter Atlas constructed from both Eastern and Western cohorts. Test-retest diffusion MRI scans from the Human Connectome Project Young Adult (HCP-YA) dataset and the Connectivity-based Brain Imaging Research Database (C-BIRD) were analyzed as independent Western and Eastern test-retest cohorts, respectively. Whole-brain tractography was reconstructed for each dMRI scan and parcellated using both atlases. Reproducibility was assessed using tract detection rate at both cluster level and anatomical tract level, weighted Dice coefficient, and the relative difference of mean fractional anisotropy (FA). Both atlases showed stable tract detection across test-retest scans in both cohorts. Compared with the ORG atlas, the East-West White Matter Atlas achieved higher overall spatial overlap and lower test-retest variability in mean FA, although atlas performance varied across individual tracts. These findings suggest that integrating cross-population information during atlas construction improves the reproducibility and generalizability of white matter atlas mapping across independent populations and imaging protocols.

## The construction and operation of type IV pili impose a variable energetic burden across phylogenetically distant bacteria
- Source: bioRxiv (preprints)
- Date: 2026-09-16
- Categories: Evolution & metagenomics
- Authors: Yusuf, A. O., Modi, Z. K., Koch, M. D.
- DOI: 10.64898/2026.09.14.751554
- Keywords: phylogenetically
- Source URL: <https://doi.org/10.64898/2026.09.14.751554>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.14.751554>

Abstract: Type IV pili (T4P) are dynamic surface appendages that mediate essential biological functions and virulence traits, yet their energetic burden on cellular budgets in light of fluctuating host environments remains unexplored. Here, we present a comprehensive economic analysis of T4P construction and operation in ATP equivalents, following established frameworks of flagella analyses. Using Pseudomonas aeruginosa as a model system, we quantify the total cellular burden to synthesize the T4P machinery, maintain the inner-membrane pool of major pilin (PilA), and drive repeated cycles of pilus extension and retraction over a generation. We estimate that the T4P system consumes ~0.7% of the total cellular energy budget, dominated by PilA monomer production. Conversely, the operational cost of dynamic T4P fibers is negligible due to their intermittent activity - contrasting sharply with the high continuous cost of rotating a polar flagellum. Extending this framework across five phylogenetically diverse species (P. aeruginosa, Vibrio cholerae, Caulobacter crescentus, Neisseria spp., and Myxococcus xanthus) reveals that T4P investment varies tenfold (0.2 - 1.8% of the cellular budget), driven by differences in pilin size, machine number, pilus extension rates, and cell volume. Neisseria is a distinct outlier whose high extension rate makes operational costs approach construction costs, while in all other species construction dominates. These findings indicate that changes in nutrient availability or surface association may modulate pilus number and length as a strategy to optimize energetic burdens during host-pathogen interaction.

## The Role of Smooth Muscle Cell Heterogeneity in Cerebral Autoregulation: A Multi-Scale Physics-Based Modeling Study
- Source: bioRxiv (preprints)
- Date: 2026-09-16
- Authors: Demeersseman, N., Maes, L., Depreitere, B., Famaey, N.
- DOI: 10.64898/2026.09.14.751526
- Source URL: <https://doi.org/10.64898/2026.09.14.751526>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.14.751526>

Abstract: Background: Cerebral autoregulation stabilizes cerebral blood flow over a range of cerebral perfusion pressures, but the precise shape of the pressure-flow relationship remains debated. The classical triphasic pressure-flow relationship was recently challenged by experiments demonstrating a quadriphasic response, hypothesized to arise from vessel-size-dependent pressure-diameter responses. We tested this hypothesis and investigated whether these size-dependent responses originate from heterogeneity in smooth muscle cell (SMC) abundance, SMC behavior, or neither. Methods: We developed a computational multi-scale physics-based model of cerebral autoregulation linking SMC activity to vessel-scale diameter regulation and organ-scale blood flow. Four scenarios were evaluated: passive vessels, homogeneous SMC abundance and behavior, heterogeneous SMC abundance, and heterogeneous SMC behavior. Predicted pressure-diameter responses and pressure-flow relationships were compared across scenarios and against experimental observations. Results: In contrast to passive vessels, homogeneous SMC activation produced partial flow stabilization, highlighting the key role of SMCs in autoregulation. However, only heterogeneous SMC behavior reproduced the experimentally observed vessel-size-dependent trends in pressure-diameter responses. This scenario also showed the best agreement with the experimental organ-scale pressure-flow relationship (R-squared = 0.93, nRMSE = 5.96%). Conclusion: The model suggests that vessel-size-dependent SMC behavior underlies vessel-size-dependent pressure-diameter responses and shapes the relationship between cerebral perfusion pressure and cerebral blood flow.

## TorchGWAS2: Cost-Effective Phenome- and Genome-Wide Association Testing in Related Samples
- Source: medRxiv (preprints)
- Date: 2026-09-16
- Categories: Tools & resources
- Authors: Zhang, M., Xie, Z., Salehi nasab, S., Wang, N., Zhao, X., Alkis, T., Barnard, J., Blackwell, T. W., Bowler, R. P., Chung, S., Cho, M., Clish, C. B., Drzymalla, E., Evans, A. M., Franceschini, N., Gerszten, R. E., Gillman, M. G., Grove, M. L., Heard-Costa, N., Hutton, S. R., Kelly, R. S., Kooperberg, C., Larson, M. G., Lasky-Su, J., Meyers, D. A., Ockerman, F. P., Raffield, L. M., Reiner, A. P., Rich, S. S., Rotter, J. I., Smith, A. V., Taylor, K. D., Vasan, R. S., Weiss, S. T., Wong, K. E., Wood, A. C., Woodruff, P. G., Wu, L., Yarden, R. I., Yu, J., Zhou, L. Y., Yu, B., Zhi, D., Chen, H.
- DOI: 10.64898/2026.09.10.26362744
- Source URL: <https://doi.org/10.64898/2026.09.10.26362744>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.10.26362744>

Abstract: Modern large-scale genetic association analyses of imaging and omics data reveal unprecedented details of the genetic architecture of complex traits. Such analyses involve scanning thousands of phenotypes using linear mixed model-based genome-wide association study tools to control for sample relatedness. However, current LMM tools are not designed for such scale, creating a computational burden that hinders discovery. We propose TorchGWAS2, a cost-effective solution that overcomes the bottleneck using a deterministic variance-correction algorithm for LMMs, making it well-suited for GPU acceleration. TorchGWAS2 is applicable to unrelated and related individuals, cross-sectional and longitudinal studies, with and without missing data, and its computational complexity scales linearly with the number of phenotypes, genetic variants, and individuals. TorchGWAS2 showed more powerful association testing across 128 retinal image-derived endophenotypes of pairs of eyes from 64,703 UK Biobank participants and achieved two orders of magnitude speed-up analyzing 1,023 circulating metabolites in 16,352 Trans-Omics for Precision Medicine participants.

## Tumour region identification guided scoring (TRIGS) and foundation model-based Tumour Infiltrating Lymphocyte scoring are prognostic for pathological complete response/event free survival in the triple negative patients in the PARTNER randomized controlled trial
- Source: medRxiv (preprints)
- Date: 2026-09-16
- Categories: Biological imaging
- Authors: Schouten, P. C., Irfan, M. O., Kinsella, Z., Sionakidis, A., Riddell, A., Worley, J. R., Lay, J., Casford, S., Pinilla, K., Kane, J., Whitehorn, D., Tarantino, S., Dayimu, A., Demiris, N., Earl, H. M., Simidjievski, N., Provenzano, E., Abraham, J. E.
- DOI: 10.64898/2026.09.15.26363116
- Source URL: <https://doi.org/10.64898/2026.09.15.26363116>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.15.26363116>

Abstract: Assessment of tumour infiltrating lymphocytes (TILs) is a robust prognostic biomarker for HER2-positive and triple negative breast cancer. We aimed to update a previously established pipeline for automated TIL assessment to align to clinical scoring guidelines (tumour region identification guided scoring), foundation model-based lymphocyte detection (SAM-TIL) and compare with alternative methods (HoVerNet, muTILs) and gold standard clinical assessment. TRIGS and SAM-TIL had an odds ratio of 1.95 (95% confidence interval(ci): 1.22-3.03, p=0.005) and 2.32 (95% ci:1.43-3.77, p=0.001) for predicting pathological complete response (pCR) rate in 166 neoadjuvantly treated patients in the TransNEO cohort. Hazard ratios for overall survival in 277 triple negative and HER2-positive patients in The Cancer Genome Atlas were 0.79 (95% ci: 0.63-1.00, p=0.05) and 0.80 (95% ci: 0.67-0.97, p=0.02) for TRIGS and SAM-TIL. Comparator methods showed similar results. Correlation with gold standard clinical assessment in 285 patients from the PARTNER randomized trial ranged from 0.59-0.69, which is substantially more than interobserver variability for the gold standard. Despite differences with gold standard assessment, no substantial difference in predicting pCR (AUC 0.60-0.64) or event free survival (Integrated Brier score approximately 0.10) were observed between the tested methods and gold standard assessment, suggesting AI tools that do not follow manual scoring guidelines could be validated and subsequently used. Although we reached prognostic performance similar to published literature and gold standard assessment, the lymphocyte detections produced by the models are not interchangeable with gold standard clinical assessment (correlation 0.59-0.69) and therefore cannot support pathologist assessment. Further studies to validate independent use and/or to improve prognostication are required.

## Uncertainty-Aware Model Selection with a Calibrated Probability-Generating-Function-Based Bayesian Information Criterion
- Source: bioRxiv (preprints)
- Date: 2026-09-16
- Categories: Genomics & sequence analysis, Mathematical biology & statistics
- Authors: Wang, Y., Shu, Z., Gao, F., Cao, Z.
- DOI: 10.64898/2026.09.11.750968
- Source URL: <https://doi.org/10.64898/2026.09.11.750968>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.11.750968>

Abstract: Selecting stochastic gene-expression models from single-cell counts requires balancing goodness of fit against unnecessary mechanistic complexity. The probability-generating-function-based Bayesian information criterion (PGF-BIC) combines covariance-weighted fitting in generating-function space with a complexity penalty, allowing candidate models to be compared without reconstructing their full count distributions. However, its conventional zero-threshold rule does not account for sampling uncertainty in the fitted score difference and may therefore favor overly complex models in finite samples. To address this limitation, we develop an uncertainty-aware PGF-BIC rule that selects the more complex model only when its score advantage exceeds a data-driven threshold. We use Cantelli's one-sided inequality to motivate a selection margin expressed in terms of a standard deviation. To determine this scale, we use influence functions to quantify sensitivity to small perturbations in the data distribution and obtain a first-order description of sampling fluctuations. The resulting variance estimate accounts for variability in both the empirical probability generating function and the estimated covariance weights, yielding a data-driven threshold for assessing the complex model's score advantage. A Poisson versus Bursty benchmark shows that the calibrated rule reduces incorrect selection of the more complex model. The calibration requires neither resampling nor additional optimization, incorporating sampling uncertainty into model selection while retaining the computational efficiency of PGF-BIC.

## Unifying multimodal single-cell data with a mixture-of-experts β-variational autoencoder framework
- Source: Genome Research (journals)
- Date: 2026-09-16T00:00:00+00:00
- Categories: Genomics & sequence analysis, Single-cell & spatial, Tools & resources
- Authors: Andrew J. Ashford, Trevor Enright, Julia Somers, Olga Nikolova, Emek Demir
- Journal: Genome Research
- DOI: 10.1101/gr.281431.125
- Source URL: <https://doi.org/10.1101/gr.281431.125>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1101%2Fgr.281431.125>

Abstract: Multimodal single-cell assays profile complementary layers of cell state, but integration is complicated by modality mismatch, sparsity, and uneven cohort coverage. Here, we present Unified Variational Inference (UniVI), a scalable mixture-of-experts β-variational autoencoder that learns a shared latent space while preserving modality-specific structure. UniVI couples modality-specific encoders/decoders with a shared latent prior and a symmetric cross-modal alignment objective, enabling consistent integration of paired measurements without curated feature-link graphs or preannotated reference atlases; optional supervised heads can be added when labels are available. Across paired RNA–protein (CITE-seq) and RNA–chromatin (10x Genomics Multiome, SHARE-seq) data spanning human PBMCs and mouse back skin—a nonhematopoietic tissue with continuous differentiation hierarchies—UniVI produces coherent embeddings, improves label transfer, and enables cross-modal reconstruction and denoising. Extending to trimodal measurements, UniVI maintains robust three-way alignment among RNA, chromatin accessibility, and surface proteins (TEA-seq), and accommodates DNA methylation in a paired scNMT-seq mouse gastrulation proof-of-concept under beta-binomial likelihoods. Performance degrades gracefully under severe cell type imbalance and in the presence of modality-exclusive populations. In an acute myeloid leukemia mosaic design, a paired RNA–protein bridge anchors independent RNA-only and protein+genotype cohorts, revealing genotype-associated neighborhoods that sharpen with mutation-aware fine-tuning. UniVI thus provides a flexible, interpretable framework for multimodal integration across paired, trimodal, and mosaic study designs and supports practical reference-to-query projection in partially observed studies.

## Unravelling the genetic basis of stuttering: GWAS meta-analysis highlights link with rare speech disorders
- Source: medRxiv (preprints)
- Date: 2026-09-16
- Categories: Genomics & sequence analysis
- Authors: Jackson, V. E., Shin, J. J., Horton, S., Boyce, J. O., Eising, E., van Reyk, O., Parker, R., Thompson-Lake, D. G. Y., Evans, M., Beilby, J., Below, J. E., Boomsma, D. I., Bridges, E., Corfield, E. C., Franken, M.-C. J., Gordon, S. D., Havdahl, A., Koenraads, S. P. C., Kraft, S. J., Luciano, M., Moen, G.-H., Mountford, H. S., Musial, A., Pennell, C. E., Polikowsky, H. G., Pool, R., Rebattu, V. A., Rimfeld, K., Scartozzi, A. C., St Pourcain, B., Szilagyi, I. A., Valand, S. B., Viljoen, K. Z., Wang, C. A., Whitehouse, A. J. O., Wren, Y. E., van Bergen, E., Gillespie, N. A., Vogel, A. P., Scheffer
- DOI: 10.64898/2026.09.15.26363183
- Keywords: genome, genomic, meta analysis
- Source URL: <https://doi.org/10.64898/2026.09.15.26363183>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.09.15.26363183>

Abstract: BackgroundDevelopmental stuttering affects up to 11% of children globally, with around one-fifth developing a persistent lifelong stutter. Twin and family studies indicate a strong genetic contribution and comorbidity with other heritable traits. Despite efforts to investigate the common genetic architecture of stuttering, much of variation contributing to clinically ascertained stuttering, persistence and recovery remains uncharacterised. MethodsWe performed a genome-wide association study (GWAS) meta-analysis of stuttering across 18 cohorts (6,096 cases, 81,629 controls) of European ancestries, with secondary analyses of stuttering persistence and sex-stratified GWAS. FindingsNo variant reached genome-wide significance in the primary meta-analysis, but 24 loci showed suggestive association (p<1x10-), with SNP-based heritability estimated at h\{superscript 2\}\{approx\}0\{middle dot\}26. FLAMES-prioritised genes at suggestive loci overlapped with those previously implicated in childhood apraxia of speech, including PTBP2, KIRREL3, CAMTA1, GRIN2A, and SETBP1, with significant enrichment for apraxia-associated genes overall (p=1x10-). Meta-analysis with an independent self-reported stuttering GWAS identified a genome-wide significant association at MPPED2 and gene-level convergence at CAMTA1 and PTBP2. A polygenic risk score derived from this independent GWAS was associated with stuttering susceptibility and severity within clinically ascertained cases. Partitioned heritability analysis pointed to enrichment in conserved regulatory regions, and integration with imaging data highlighted motor circuitry including decreased pallidum volume and cerebellar and white-matter microstructural differences. InterpretationOur findings support common variant associations in stuttering converging on genes implicated in speech and neurodevelopmental conditions, pointing to basal ganglia-cerebellar motor circuits as central to speech motor control. FundingAustralian National Health and Medical Research Council. Research in contextO\_ST\_ABSEvidence before this studyC\_ST\_ABSWe searched PubMed for genome-wide association studies (GWAS) of stuttering, using terms including "stuttering," "stammering," and "genome-wide association," for studies prior to July 2026, with no language restriction. Prior GWAS of stuttering are limited. The International Stuttering Project combined clinically ascertained and self-reported cases with population controls and identified one genome-wide significant locus near SSUH2 and 15 loci at suggestive significance. Another study investigating predicted stuttering within Vanderbilts Electronic Health Records, identified one locus surpassing genome-wide significance near CYRIA. A larger GWAS using self-reported stuttering status identified 57 genome-wide significant loci. Twin and family studies estimate stuttering heritability at 0\{middle dot\}42-0\{middle dot\}85, and rare variant studies have implicated genes including GNPTAB, GNPTG, NAGPA, AP4E1, PPID, and ZBTB20 in familial persistent stuttering, but it remains unclear whether these genes are also relevant to common genetic variation in the general population. Added value of this studyWe conducted the largest GWAS meta-analysis of stuttering to combine clinically ascertained cases with population-based cohorts, comprising 18 cohorts, 6,096 cases, and 81,629 controls. Unlike prior studies based solely on self-report, many of our ascertained cases had detailed phenotyping including measures of persistence and quantitative severity, allowing us to examine genetic overlap between stuttering onset, persistence, and severity. We identified suggestive genetic loci that converge with genes previously implicated in a rare, severe motor speech disorder (childhood apraxia of speech), and found that combining our data with the independent, previous GWAS of self-reported stuttering identified a genome-wide significant association. We further used imaging genetics approaches to link genetic risk for stuttering to specific brain regions and circuits involved in motor control, and used evolutionary genomic analyses to show that stuttering-associated regions are enriched in ancient, conserved parts of the genome. Implications of all the available evidenceOur findings suggest that common genetic variation contributing to stuttering converges on the same genes and brain circuits implicated in rare, severe speech disorders. This strengthens the case that stuttering, at least in part, shares a biological basis with other neurodevelopmental and speech-motor conditions, and points to the basal ganglia- cerebellar motor circuit as a promising target for future mechanistic research. For clinicians and people who stutter, these findings do not yet have direct treatment implications, but they lay groundwork for better understanding why stuttering persists in some individuals and not others, and highlight the value of collecting detailed speech and language phenotypes in future large-scale genetic studies.

## Vertex-wise biomechanical sensitivity mapping of subcortical structures under atrophy in Parkinson's disease.
- Source: Medical image analysis (journals)
- Date: 2026-09-16
- Categories: Computational neuroscience
- Authors: Arina Olentcevich, Shuli Guo, Lina Han
- Journal: Medical image analysis
- DOI: 10.1016/j.media.2026.104332
- External ID: 42762597
- Keywords: hippocampus, hippocampal
- Source URL: <https://doi.org/10.1016/j.media.2026.104332>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.1016%2Fj.media.2026.104332>

Abstract: Parkinson's disease (PD) is characterized by progressive neurodegeneration and pronounced subcortical asymmetry, yet the extent to which anatomical geometry modulates mechanical responses to atrophy remains insufficiently understood. We propose a vertex-wise biomechanical framework to quantify structure-specific vulnerability by linking MRI-derived morphology to finite-strain mechanical behavior. First, subject-specific 3D meshes of key subcortical regions are reconstructed from T1-weighted MRI data of 141 PD subjects from the PPMI dataset. Second, finite-strain simulations are implemented using a one-term Ogden hyperelastic model, which is then evaluated against experimental reference data. Third, region-specific atrophy is modeled via an isotropic expansion analogy, inducing mechanical deformation. Finally, we introduce three novel indices SALDI, MechSALDI, and DeformSALDI, which characterize surface-based asymmetry, strain-based sensitivity, and displacement-based sensitivity, respectively. These indices are employed to generate vertex-wise 3D sensitivity maps. The principal component analysis (PCA) reveals a dominant low-dimensional structure across the biomechanical indices, with the first component explaining 89.7% of the total variance, indicating strong coherence among geometry- and deformation-derived measures of mechanical sensitivity. The vertex-wise analysis demonstrates a reproducible hierarchy of regional vulnerability, with the hippocampus and amygdala exhibiting the highest mechanical sensitivity, while the thalamus and pallidum show relative resilience. Comparative analyses between PD and healthy controls reveal systematic region-specific differences in biomechanical sensitivity, particularly in striatal and limbic circuits. Hierarchical regression further demonstrates that several SALDI-derived biomechanical indices retain independent associations with cognitive performance and asymmetric motor manifestations, including hippocampal and amygdalar indices for cognition and thalamic, putaminal, pallidal, and accumbens indices for motor asymmetry measures, even after adjustment for conventional volumetric asymmetry. Overall, the observed subcortical sensitivity in PD follows a structured organization driven by local geometry and material-dependent mechanical behavior rather than uniform atrophy. The proposed framework provides a quantitative basis for characterizing geometry-dependent biomechanical sensitivity beyond conventional volumetric measurements and may facilitate future longitudinal studies of individualized degeneration trajectories, patient stratification, and disease progression in neurodegenerative disorders.

## Whole-body 3D kinematics of freely behaving Drosophila
- Source: bioRxiv (preprints)
- Date: 2026-09-16
- Categories: Biological imaging, Tools & resources
- Authors: Ispizua, J. I., Abe, E. T. T., Yan, J., Othayoth, R., Sawtelle, S., Atkins, F., Shiozaki, H., Meier, N. R., Wong, J., Tran, T. T., Mori, C. K., Chen, W., Voigts, J., Stern, D. L., Brunton, B. W., Tuthill, J. C., Johnson, R. E.
- DOI: 10.64898/2026.05.03.722293
- Source URL: <https://doi.org/10.64898/2026.05.03.722293>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.64898%2F2026.05.03.722293>

Abstract: Understanding how nervous systems generate coordinated movement requires precise measurement of body kinematics during natural behavior. The fruit fly, Drosophila, is a model organism with sophisticated behavior and well-studied neural circuits, but tracking fly movements in 3D remains challenging because of their teeny bodies, rapid movements, and frequent self-occlusions. Here we present a pipeline for markerless, full-body 3D pose estimation of fly terrestrial behavior, combining seven synchronized high-speed cameras to capture whole-body kinematics at 800 frames per second. We trained a hybrid 2D/3D deep learning model to track 50 keypoints, then refined them to produce anatomically feasible kinematic trajectories through a retargeting process that solved an inverse kinematics problem constrained by a biomechanical body model. Analysis of 3D kinematics revealed that flies perform grounded running across their full speed range, without transitioning between discrete gaits. Using multi-animal tracking, we found that courting males coordinate both wings during song and modulate body pitch to track the female's vertical position. Our open-source pipeline and large 3D kinematic dataset of fly behavior provide a foundation for neuromechanical modeling and mechanistic studies of motor control in a genetically tractable model organism.

## X-GCN: An Explainable and Uncertainty-Aware Graph Convolutional Framework for Multi-Omics Disease Risk Prediction
- Source: Natural Resources for Human Health (journals)
- Date: 2026-09-16T00:00:00Z
- Categories: Single-cell & spatial
- Authors: Suresh Kulandaivelu, Mohan Mani
- Journal: Natural Resources for Human Health
- DOI: 10.53365/nrfhh.336
- External ID: 594935688b9cd887faf17f036b6c01884e234d7a
- Source URL: <https://doi.org/10.53365/nrfhh.336>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Fdoi.org%2F10.53365%2Fnrfhh.336>

Abstract: To provide accurate and helpful insights, healthcare early illness risk prediction requires the integration of diverse multi-omics data. Unlike existing models (MOGONet, ExplainMix, CNNs), X-GCN presents an integrated explainable graph-based framework that combines graph convolution with hierarchical attention in a unified multi-omics supra-graph, enabling uncertainty-aware disease risk prediction and biologically grounded explanations. Multi-omics data such as transcriptomics, proteomics, genomics, and epigenomics are represented by X-GCN as graph data, where nodes replace biological entities and edges replace molecular interactions. X-GCN focuses on critical biomarkers and ensures prediction transparency through hierarchical attention techniques.X-GCN surpasses complex models such as MOGONet (88.7% accuracy, 0.912 AUC) and ExplainMix (90.2% accuracy, 0.920 AUC) with tests on The Cancer Genome Atlas (TCGA) and Gene Expression Omnibus (GEO) datasets, recording 92.4% accuracy, 0.935 AUC, and 0.91 F1-score. X-GCN also decreases model uncertainty by 18% and identifies experimentally validated biomarkers for cancer and cardiovascular disease prediction. X-GCN provides an interpretable and uncertainty-aware computational framework for multi-omics disease risk prediction, serving as a foundation for biomarker discovery and future clinical validation.

## Posture selection in active elastic filaments
- Source: arXiv (preprints)
- Date: 2026-09-15T23:32:40Z
- Authors: Adam Pearl, Ludwig A. Hoffmann, L. Mahadevan
- External ID: 2609.17924v1
- Source URL: <https://arxiv.org/abs/2609.17924v1>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Farxiv.org%2Fabs%2F2609.17924v1>
- PDF: <https://arxiv.org/pdf/2609.17924v1>

Abstract: Posture control in slender bodies such as snakes and eels arises from the interplay between passive deformation, active internal actuation, and task-level constraints. We formulate a general framework for the selection of stable postures in active elastic filaments subject to distributed forcing from gravity and fluid drag, by combining the constraints of mechanical equilibrium with optimal control theory. Our theory leads to a minimal description in terms of parameters governing the competition between hydrodynamic and gravitational loading, elasticity, and activity. We show that posture selection reflects a trade-off between control cost, function and dynamical stability, leading to the coexistence of distinct solution branches and abrupt transitions between them. Applying the theory to sessile eels in flow, we recover the experimentally observed transition from upright to reclining postures and predict scaling laws for body shape and exposed length. More generally, our results provide a unified perspective on how active filaments can regulate geometry to maintain function in external fields, with implications for biological and artificial systems.

## FlowLOT: Linearized Optimal Transport for Flow Cytometry Analysis
- Source: arXiv (preprints)
- Date: 2026-09-15T22:57:26Z
- Categories: Single-cell & spatial
- Authors: Naqib Sad Pathan, Mohammad Shifat-E-Rabbi, Kristofor E. Pas, Ivan Medri, Bartek Rajwa, Gustavo K. Rohde
- External ID: 2609.17906v1
- Keywords: single cell
- Source URL: <https://arxiv.org/abs/2609.17906v1>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Farxiv.org%2Fabs%2F2609.17906v1>
- PDF: <https://arxiv.org/pdf/2609.17906v1>

Abstract: Multiparameter flow cytometry generates high-dimensional, unordered single-cell mea- surement data for disease diagnosis and monitoring, yet analysis often remains dependent on manual gating, limiting scalability and reproducibility. Existing machine-learning ap- proaches can reduce annotation burden but frequently require large training cohorts and offer limited interpretability. To address these challenges, we introduce FlowLOT , an optimal-transport-based framework that models the single-cell measurement data of a pa- tient sample as an empirical cellular distribution and maps it directly into a fixed-length feature vector. Within a single transparent architecture, FlowLOT unifies high-dimensional classification, interpretable visualization, and continuous quantitative inference. In few-shot regimes, using as few as 16 patients per class on FlowCAP-II and 8 patients per class on BLAST110, it accurately distinguishes healthy from acute myeloid leukemia (AML) sam- ples, reaching 94.3% and 98.0% balanced accuracy, respectively. The underlying embedding exposes marker-level variation driving disease-associated population shifts and enables quantitative measurable residual disease (MRD) estimation, achieving a Pearson correlation of 0.82 on held-out samples and 0.79 under cross-dataset transfer. Furthermore, at the clinically relevant 0.1% threshold for leukemia-associated immunophenotype (LAIP) residual disease, FlowLOT detects positivity with 72% sensitivity at 100% specificity. By replacing subjective manual gating and black-box deep learning with a distribution-aware framework, FlowLOT offers a sample-efficient, scalable, and interpretable solution for high- dimensional cytometry under realistic clinical and experimental constraints.

## METALICA: METAdynamics and repLICA exchange for enhanced diffusion sampling
- Source: arXiv (preprints)
- Date: 2026-09-15T20:38:39Z
- Categories: Proteins & structural biology
- Authors: Alireza Omidi, Jiajun He, Jörg Gsponer, Saifuddin Syed
- External ID: 2609.17823v1
- Source URL: <https://arxiv.org/abs/2609.17823v1>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Farxiv.org%2Fabs%2F2609.17823v1>
- PDF: <https://arxiv.org/pdf/2609.17823v1>

Abstract: Many proteins function through transitions between conformational states, yet rare states are rarely sampled by diffusion models trained on an equilibrium ensemble, demanding better sampling methods. We introduce METALICA, which implements Metadynamics on a pretrained diffusion model via Replica Exchange. It accumulates a bias potential along a Collective Variable, repels new samples from previous ones through biased sampling, and reweights samples onto the unbiased distribution. METALICA holds one replica per diffusion level, forming a Markov Chain that evolves through inter-replica communication and is refined in place as the bias grows. METALICA is the dual of sequential control, in which Sequential Monte Carlo parallelizes the sampler over a batch of particles. Parallelism over the levels of the diffusion-time schedule instead allows METALICA to generate samples from long chains, essential for the discovery of rare events, with accuracy set by run length rather than by the memory available. We validate on a bimodal target with known free energies, then apply METALICA to the unfolding of a protein. At a budget for which sequential control yields no unfolded structure, METALICA populates the basin and resolves a second free energy minimum.

## GIA: Germline-Informed Aging with AlphaGenome Finds Genetically Regulated CpGs
- Source: arXiv (preprints)
- Date: 2026-09-15T20:10:59Z
- Categories: Genomics & sequence analysis
- Authors: Sean Lim
- External ID: 2609.17801v1
- Keywords: epigenetic, dna, methylation, chromatin, rna
- Source URL: <https://arxiv.org/abs/2609.17801v1>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Farxiv.org%2Fabs%2F2609.17801v1>
- PDF: <https://arxiv.org/pdf/2609.17801v1>

Abstract: Epigenetic clocks estimate age and aging-related phenotypes from DNA methylation at selected CpG sites, but the extent to which these inputs are influenced by germline genetic variation is unclear. Because methylation at many CpGs is genetically regulated, some between-person variation in clock estimates may reflect inherited genetic differences rather than aging-related change alone. Here we developed GIA (Germline-Informed Aging), a framework that maps CpGs selected from 13 published epigenetic clocks to blood methylation quantitative trait loci (meQTLs) and scores associated genetic variants with AlphaGenome. We show that clock CpGs were enriched for blood meQTLs relative to matched unused Illumina 450k probes (62.7% versus 39.6%; OR 2.57), across multiple clock families, suggesting that age-informative methylation sites are heavily influenced by germline genetic variation. Ranking by predicted chromatin effect isolated rs10190186, a cis-acting variant at FHL2 predicted to increase blood chromatin accessibility (ATAC +1.00; DNase +1.64) and FHL2 RNA (+0.30). This locus illustrates how inherited variation may shape methylation features repeatedly used by epigenetic clocks, motivating direct tests of whether such variants shift baseline clock estimates or longitudinal aging trajectories.

## Flash-Radiomics: A Scalable Hybrid CPU-CUDA Engine for Standardized Scalar Radiomics and Accelerated Spatial Mapping
- Source: arXiv (preprints)
- Date: 2026-09-15T19:54:28Z
- Categories: Biological imaging, Tools & resources
- Authors: Shanli Ding, Yiyi Hu, Ziyu Fu, Chia-Hsin Lin, Ruihan Luo, Jaehee Chun, Xinyue Zhang, Osama Mawlawi
- External ID: 2609.19190v1
- Source URL: <https://arxiv.org/abs/2609.19190v1>
- Dashboard article: <https://tagirshin.com/bioradar/article?u=https%3A%2F%2Farxiv.org%2Fabs%2F2609.19190v1>
- PDF: <https://arxiv.org/pdf/2609.19190v1>

Abstract: Background and Objectives: Spatial mapping retains the spatial distribution of radiomic features, but computational cost and fragmented software limit its use. We developed Flash-Radiomics with scalar extraction and spatial mapping, a central processing unit (CPU) backend, a hybrid Compute Unified Device Architecture (CUDA) backend, consistent feature names, and Hierarchical Data Format version 5 (HDF5) storage. Methods: We evaluated Image Biomarker Standardisation Initiative (IBSI) compliance, CPU-CUDA concordance, and end-to-end processing time. Compliance testing included 825 chapter 1 (IBSI-1) tests covering 165 high-consensus features and 323 chapter 2 (IBSI-2) tests with numerical references. Concordance testing included 1,148 scalar pairs and 93 spatial-map pairs. End-to-end processing time was measured five times per input volume of interest (VOI) size. Comparisons included the Medical Image Radiomics Processor (MIRP) and PyRadiomics for 102 shared scalar features and PyRadiomics for 93 shared spatial maps. Results: Both backends passed all 1,148 IBSI tests, and all paired results were concordant. At the largest scalar input, CPU required 76.343 s and hybrid CUDA 81.915 s; CPU was 4.7 times faster than MIRP and 190.7 times faster than PyRadiomics. At the largest spatial input completed by both backends, hybrid CUDA reduced processing time by 68.6% relative to CPU (79.280 versus 252.791 s). At PyRadiomics' largest completed spatial input, hybrid CUDA was 84.9 times faster. Conclusions: Flash-Radiomics unified standardized scalar extraction, spatial mapping, concordant CPU-CUDA results, and HDF5 storage. CPU processing time was similar or shorter for scalar extraction, whereas hybrid CUDA was faster for spatial mapping under the tested conditions.
