Quantitative Biology

Browse today’s new papers as interactive HTML on Academus.


FISHER: Gradient-Decoupled Hierarchical Multi-Task Learning for Fine-Grained Aquatic Species Recognition

Fine-grained recognition of aquatic species is challenging due to subtle morphological differences and long-tailed distributions, where ultra-rare species are underrepresented. A natural solution is to jointly model segmentation, morphological traits, and species classification within a multi-task learning (MTL) framework. However, existing MTL methods suffer from negative transfer caused by gradient conflicts between low-level dense tasks and high-level classification objectives, degrading fine-grained representations. To address this limitation, we identify gradient interference across hierarchical tasks as a fundamental bottleneck and propose FISHER, a gradient-decoupled hierarchical multi-task learning framework. FISHER aligns optimization with the biological hierarchy of aquatic species by enforcing a unidirectional information flow from segmentation to trait prediction and finally to species classification, while explicitly decoupling gradients across task boundaries. This design prevents high-level objectives from corrupting low-level morphological representations, effectively mitigating negative transfer while preserving the benefits of shared supervision. Furthermore, we introduce a prototype-based segmentation head with orthogonality regularization to encourage disentangled anatomical representations, and employ homoscedastic uncertainty weighting to dynamically balance task contributions during training. Our analysis shows that robust trait representations serve as a critical bridge for transferring knowledge to ultra-rare species. Extensive experiments on the Fish-Vista benchmark demonstrate that FISHER achieves 97.7% mAP for unseen trait identification and improves ultra-rare species classification accuracy by 13.4% over strong baselines, highlighting the effectiveness of gradient-decoupled hierarchical learning for long-tailed biodiversity recognition.


Auditing pretraining contamination in single-cell foundation model benchmarks

Single-cell foundation models (scFMs) such as Geneformer, scGPT, and Universal Cell Embeddings (UCE) are pretrained on tens of millions of cells drawn from public repositories. The same repositories underlie widely used integration benchmarks, creating an unmeasured risk that zero-shot benchmark performance reflects pretraining exposure rather than genuine generalization. We introduce \textbf{scContam}, a per-cell audit framework that combines a MinHash-based gene-set fingerprint signal against the explicit pretraining corpus with a loss-based membership inference attack (MIA-scFM). Applied to four scIB benchmarks and three scFMs, we find that two of the most-cited benchmarks, PBMC 3k and the CELLxGENE human pancreatic islet atlas, contain extensive pretraining-overlap evidence ($80.4\%$ and $77.0\%$ of cells with fingerprint $p < 0.05$ against Genecorpus-30M), whereas the post-cutoff datasets AIDA v2 and Tahoe-100M show no overlap evidence ($0\%$). A controlled re-pretraining experiment establishes that MIA-scFM AUROC scales monotonically with the model's capacity-to-data ratio (AUROC $0.494 \to 0.690 \to 0.881$ across properly-regularized, mildly-overfit, and aggressively-overfit regimes), demonstrating that production scFMs resist instance-level memorization but distributional contamination must be detected separately. A donor-matched, within-cell-type analysis with three architectures shows that contaminated cells embed measurably more tightly than donor-matched clean cells (permutation $p = 0.030, 0.014, < 0.002$, respectively), with a perfectly null AIDA negative control. Pretraining audits are tractable and should accompany scFM benchmark reporting.


Foundation-model-guided radiogenomic discovery linking cancer genomes to cancer scans

The function of many genes is still unknown, and conventional driver-discovery methods, which rely on how frequently a gene is mutated, cannot assess genes that are only rarely affected. Here we pair Evo~2-based genome analysis with routine clinical imaging to identify gene--phenotype associations at genome-wide scale. For every somatic mutation across three TCGA cohorts (cRCC=clear cell renal cell carcinoma, HCC=hepatocellular carcinoma, and BC=breast cancer; $n = 340$ total), Evo~2 predicts a severity score, with no task-specific training. Per-gene severity summaries are then correlated with radiomic features extracted from paired tumor segmentations, controlling for total mutation burden. In TCGA-cRCC ($n = 162$), this sweep recovers established renal-cancer drivers and identifies 46 additional genes reaching false discovery rate (FDR) significance absent from curated cancer-gene panels, several of which are Mendelian ciliopathy and cytoskeletal-disease genes. These results demonstrate that pairing a genomic language model with widely available clinical imaging can serve as a hypothesis-free discovery tool for gene--imaging associations invisible to conventional approaches.


Spectral theory for population density dynamics of spiking neurons with refractoriness

Incorporating an absolute refractory period into the population density approach for spiking neurons remains an open problem, despite evidence that refractoriness can strongly affect nonlinear transfer functions and network stability. We develop a rigorous operator-theoretic framework for neuronal population dynamics with a finite refractory time by augmenting the state space to include refractory history and formulating the problem as a non-self-adjoint boundary eigenvalue problem for the Fokker-Planck operator. This yields a complete spectral characterization of the generator, proves dissipativity and the existence of a contraction semigroup, and identifies defective eigenvalues as exceptional points where oscillatory modes emerge from coalescing relaxational modes. Within the framework of linear response theory, we also derive an exact transfer function that accounts for boundary conditions modulated by external input, correcting previous heuristic derivations and revealing additional threshold-noise contributions. Using this transfer function under a mean-field approximation, we further show that refractoriness in populations of interacting neurons can facilitate the onset of limit cycles, that is, stable oscillations in the firing rate. These results provide a rigorous foundation for spectral decomposition methods in computational neuroscience, opening the way to their further rigorous mathematical analysis.


Transition-Related Potentials as Markers of Narrative Comprehension in Continuous EEG

Harnessing the potential of electroencephalography (EEG) for brain research is fundamentally limited by intrinsic noise and the diffuse projection of brain-generated activity over the scalp. The standard event-related potential (ERP) paradigm addresses this limitation by relying on repeated independent trials, albeit at the cost of moving away from naturalistic experimental conditions. As a more naturalistic alternative, we collected continuous EEG while participants watched short films and extracted potentials aligned to sharp cinematic transitions (cuts). We demonstrate that such transition-related potentials (TRPs) exhibit canonical ERP-like temporal structure associated with significant information processing. By comparing coherent films with scene-scrambled versions containing matched post-cut sensory input, we find that these responses are systematically shaped by narrative context. We then show that the cut-related EEG signature can be recovered directly from group-averaged continuous recordings with a compact deep neural network (DNN). The detector generalized across films and subject groups, and the resulting TRPs reproduced the main context-dependent effects observed for manually annotated cuts. These results indicate that narrative context leaves a measurable signature in EEG responses, that this signature can be detected directly in continuous recordings, and that such detections provide a semi-automated framework for analyzing how viewers process and understand film narratives. We propose that the method outlined here can be adapted to parse EEG responses to other forms of continuous stimulation, providing a general tool for probing experimental conditions that are closer to natural human experience.


Analyzing \b{eta}-Lactamase Evolution from the Principle of Least Action Perspective

Protein sequences change over time due to the accumulation of mutations in the genes that encode them. Nonetheless, accurately identifying the most probable evolutionary pathways and more efficient trajectories between an ancestral and a derived sequence remains a challenge. This difficulty stems from a limited understanding of the critical factors that drive the evolutionary process. This study aims to address this issue by validating a newly proposed mechanistic model (grounded on the principle of least action) for analyzing protein evolution, using beta-lactamase as an empirically characterized evolutionary model. The initial findings indicate that, after accounting for how mutations and the context-dependent interactions between them (epistasis) affect protein stability and, hence, the kinetics of protein folding, a resort to the principle of least action allows us to simultaneously identify the most probable pathways and the most efficient trajectories swiftly and accurately. These findings also suggest that the most probable evolutionary protein pathways are the ones that exhibit the most efficient trajectories. All in all, our study addresses several unanswered questions regarding the main factors that govern how protein evolves at the molecular level (under constant selection pressure) and outlines directions for future research in key fields such as directed evolution and ancestral sequence reconstruction.


A Bayesian-optimization framework coupling a multiphase PDE tumor model to efficiently design combination therapy schedules

Designing combination cancer therapies requires choosing not only which agents to combine but also their relative doses and timing decisions that critically shape the trade-off between efficacy and toxicity. High-fidelity mechanistic models of tumor growth, formulated as systems of coupled PDEs, can in principle resolve how these scheduling choices interact with the tumor microenvironment, but each evaluation is computationally expensive, rendering brute-force exploration of the design space intractable. We present a Bayesian Optimization framework that treats a multiphase, vascularized, two-dimensional PDE tumor simulator as a black box and uses a Gaussian-process surrogate to find schedules that maximize therapeutic outcomes within a small budget of expensive simulations. We orchestrate the COMSOL Multiphysics solver from Python, producing a fully automated optimization loop in which a single simulation of ~650 days of tumor evolution requires roughly 80 hours of wall time. The framework is applied to three clinically relevant scenarios: (i) a two-agent regimen (docetaxel + bevacizumab), (ii) a three-agent regimen (docetaxel + bevacizumab + radiation) under reduced and full intensity, and (iii) a single-agent dose-fractionation problem in which efficacy is balanced against healthy-tissue toxicity through a weighted multi-objective formulation. The BO loop converges to clinically plausible optima with one to two orders of magnitude fewer simulations than an equivalent grid search, identifies docetaxel-induced radiosensitization as a decisive factor in the triple-therapy optimum, and recovers a fractionation regime consistent with clinical protocols when both efficacy and toxicity are considered. The framework is agnostic to the specifics of the underlying PDE model and provides a transferable methodology for design optimization of expensive engineered or biological simulators.


Evolutionary Le Chatelier's Principle: Phenotypic Plasticity and Genetic Assimilation via Timescale Separation in the Price Equation

Phenotypic plasticity and genetic assimilation play key roles in adaptive evolution, yet their underlying mechanism has lacked a unified physical description. A major theoretical difficulty lies in the fundamental difference in timescales, as phenotypic plasticity occurs rapidly within a generation whereas genetic changes accumulate slowly across generations. Here, we formalize these processes by bridging the continuous-time Price equation, a foundational equation of evolutionary dynamics, with the physical concept of timescale separation. A sudden environmental change induces a fast plastic displacement of the phenotype relative to the slow genotypic variable. Through genotype--phenotype coupling, this displacement generates an internal genetic stress. We demonstrate that genetic assimilation is a dynamical relaxation process in which the genotype evolves to resolve this self-generated stress. These evolutionary dynamics mathematically realize Le Chatelier's principle, where the slow genetic response naturally amplifies the initial plastic shift in the same direction. The theory predicts that a weaker restoring force, which can manifest as larger clonal phenotypic fluctuations, requires a longer evolutionary timescale for assimilation. In the ideal limit of cost-free, perfectly adaptive plasticity, the relaxation time diverges, so assimilation effectively stalls. This formulation provides a macroscopic physical mechanism for genetic assimilation, offering a universal response law for evolutionary systems in which rapid phenotypic responses precede slower genetic change.


From Berg-Purcell precision bounds to clock-limited information capacity

Physical limits to chemical sensing are traditionally expressed as Berg-Purcell bounds on estimation accuracy. Whether these bounds also limit the total amount of information a molecular receptor can transmit has remained unclear, despite the fact that cellular signaling performance is naturally quantified in bits rather than precision alone. Here we derive an explicit link between Berg-Purcell-type sensing limits and the information capacity of a single two-state receptor, yielding a compact expression that separates contributions from concentration range, receptor copy number, and averaging time. We show that, in the ideal fixed-time occupancy model, diffusion-limited sampling alone does not define a finite global information bound: with a perfectly specified integration window and unbounded input range, information capacity grows without bound with dynamic range, albeit slowly. A finite saturation arises when time integration is treated as an explicit physical resource. Finite timing precision yields a clock-limited bound on information capacity, and in the high-occupancy regime information transmission crosses over from diffusion-limited to clock-limited behavior. Together, these results establish a receptor-level information bound in bits that is finite once constraints on timing precision are taken into account.


Local intercellular coupling is sufficient for long-range calcium signaling

Long-range intercellular calcium (Ca2+) signaling coordinates biological processes ranging from fertilization to contraction and cell death. The classical model attributes this long-range propagation to rapid diffusion of inositol 1,4,5-trisphosphate (IP3) through gap junctions. However, recent evidence that IP3 diffuses far more slowly than previously believed, and that Ca2+ oscillations persist even when gap junctions are disassembled, indicates that an alternative mechanism must sustain long-range communication. Here we develop a computational model showing that local coupling between neighboring cells is sufficient to generate and propagate regenerative Ca2+ oscillations across a cell population without fast molecular diffusion. Each cell is treated as an oscillator whose intrinsic frequency is set by its local IP3 concentration through an IP3-dependent refractory period, and neighboring cells are coupled using a Kuramoto nearest-neighbor framework. In a dual-stiffness regime, cells on a stiff extracellular matrix entrain their soft-matrix neighbors, producing an offset traveling wave of Ca2+ release. This reproduces the finite spatial range of influence (~8 cell lengths) observed experimentally. Our findings propose a diffusion-independent paradigm for calcium signaling in which local intercellular coupling drives long-range communication, offering insight into how localized ECM stiffening in asthma and fibrosis may produce systemic effects.


Magnetosensitivity of amphibian morphological pigmentation is light- and eye-dependent and consistent with the radical pair mechanism

Weak magnetic fields influence a wide range of biological processes, yet the underlying mechanisms are poorly understood. The radical pair mechanism (RPM), which involves quantum spin dynamics, is a leading hypothesis. Here we show that weak magnetic fields modulate morphological pigmentation -- specifically, the number of perioptic melanophores -- in Xenopus laevis tadpoles in a field-strength-dependent manner. The response is light- and eye-dependent. The observed field-strength dependence is quantitatively consistent with a radical pair model. These properties are reminiscent of the light-dependent magnetoreception that is thought to operate in migratory birds, and establish amphibian pigmentation as a tractable vertebrate system for the study of radical-pair quantum biology.


SenCos-GEM: SENet-Calibrated and Law-of-Cosines-Constrained Geometry-Enhanced Molecular Representation for Property Prediction

Effective molecular representation learning is crucial for accurate molecular property prediction. Recently, numerous self-supervised learning (SSL) approaches leveraging 3D GNNs have been developed to capture comprehensive 3D structural information for drug discovery. However, existing methods lack explicit physical constraints and are highly susceptible to geometric noise induced by coarse empirical force fields during large-scale this http URL, they overlook dynamic feature modulation during downstream adaptation, often resulting in catastrophic forgetting and negative transfer. To address these limitations, we introduce SenCos-GEM, a novel explicitly decoupled geometry-enhanced molecular representation learning framework that incorporates SENet-calibrated and law-of-cosines-constrained enhancements. SenCos-GEM employs a physics-guided geometric consistency loss based on the law of cosines to derive high-fidelity and mathematically invariant 3D spatial priors. In addition, lightweight Squeeze-and-Excitation (SE) modules are integrated into the backbone as task-specific adapters, while a dual-modulation prediction head combines Feature-wise Linear Modulation (FiLM) and SENet mechanisms to enable dynamic feature recalibration. SenCos-GEM demonstrates highly competitive performance across diverse classification and regression tasks on MoleculeNet benchmark, establishing new state-of-the-art results specifically on 3D conformation-sensitive regression tasks, such as FreeSolv, Lipophilicity, and QM9, achieving relative error reductions of 12.9% (RMSE), 5.3% (RMSE), and 8.2% (MAE), respectively. Moreover, our model exhibits superior capability in distinguishing stereoisomers and discriminating conformational perturbations, underscoring its robust spatial modeling performance. Collectively, SenCos-GEM represents a significant breakthrough in accurate molecular property prediction.


Machine-Learned Compact Subspace Generation for Quantum Selected Configuration Interaction within Density Matrix Embedding Framework

Sample-based Quantum Diagonalization (SQD), an extension of Quantum Selected Configuration Interaction (QSCI), has emerged as a promising hybrid quantum-classical paradigm for computing molecular ground state energies. By leveraging quantum sampling instead of variational optimization, QSCI avoids barren plateaus and enables direct reconstruction of correlated electronic wavefunctions. However, existing configuration recovery techniques primarily enforce symmetry constraints without guaranteeing optimal selection of the most physically relevant configurations, often leading to unnecessarily large subspaces and increased classical diagonalization costs. In this work, we introduce a machine-learned compact subspace generation protocol based on Restricted Boltzmann Machines (RBMs), termed QSCI-RBM, and integrate it within the Density Matrix Embedding Theory (DMET) framework. The RBM is trained on quantum-sampled configurations to learn the underlying probability distribution of dominant determinants, enabling the targeted generation of high-probability configurations. We apply this framework to the simulation of a protein-ligand complex involving the inhibitor Carmofur bound to the SARS-CoV-2 main protease ($M^{\text{pro}}$). Our results demonstrate that DMET-QSCI-RBM achieves energies within the chemical accuracy threshold by accessing only approximately 4% of the configuration subspace. In contrast, standard DMET-SQD simulations failed to reach chemical accuracy while accessing up to 20% of the subspace, even as the chemical potential itself nearly converged. These findings highlight that RBM-assisted configuration generation produces significantly more compact subspaces while preserving physical accuracy, thereby reducing classical computational overhead and enabling the scalable quantum embedding simulation of complex biological systems.


Perspective Latents as an Architectural Condition for Causal Emergence in Active Inference Agents

A recent line of work measures causal emergence in reinforcement learning agents through Integrated Information Decomposition, reporting that $\Phi_r$ grows with training and tracks reward improvement. For active inference, this raises the question of how reward-free predictive organization relates to such information-theoretic signatures. I test this within an active inference agent whose architecture separates a fast perception latent $z$ from a slow global latent $g$, where $g$ is driven by prediction error and structurally decoupled from policy gradients. In a reward-free environmental regime-switching protocol, $\Phi_r$ concentrates in $g$; its aggregate magnitude is largely architectural and decreases with training. The substantive effect of learning becomes legible only at the atom-compositional level: decoupling flips sign from negative to positive and becomes regime-invariant under environmental change, while downward causation carries the regime-dependent adjustment. These results identify $g$ as the architectural locus of $\Phi_r$-relevant temporal organization in an active inference agent, and argue against reading scalar $\Phi_r$ as a direct index of learned integration.


Importance-Sampling Estimation of Gaussian Molecular Shape Overlap: Exact Union Volumes and Confidence-Bounded Virtual Screening

Gaussian descriptions of molecular shape underpin 3D shape-based virtual screening, but existing methods evaluate Gaussian overlap analytically. The widely used first-order approximation is fast but systematically overestimates overlap, whereas the exact molecular volume requires a combinatorial inclusion-exclusion expansion. We introduce the first stochastic estimator of Gaussian shape overlap: an unbiased Monte Carlo method that importance-samples directly from a molecule's Gaussian mixture. The estimator reproduces analytic overlap without bias and extends to the exact union volume of all inclusion-exclusion orders with O(N) cost per sample. On drug-like molecules, the union estimator matches high-resolution grid quadrature with a mean relative error of 0.07 percent, while the first-order approximation overestimates the true union volume by 3.4x on average. The estimator provides analytic standard errors, enabling confidence-bounded screening that reduces sampling by 94 percent while preserving ranking. Implemented in JAX, it is fully differentiable and supports gradient-based rigid alignment on CPU, GPU, and TPU. On the DUD-E and LIT-PCBA benchmarks, the method achieves shape-only enrichment comparable to existing single-conformer approaches while additionally providing unbiased absolute volumes and uncertainty estimates.


Fluctuation impossibility results for stochastic burst networks

Stochastic reaction networks often involve components at low copy number, where individual production and degradation events generate substantial fluctuations. Yan et al.\ conjectured in 2019 that for networks with linear degradation and arbitrary cross-regulatory production rates, feedback cannot suppress the stationary fluctuations for each component below the fluctuations of its constant-rate counterpart. Their formulation allows random burst sizes $K_i$: the unit-birth case has $K_i\equiv1$, the biologically important burst model takes $K_i$ to be geometrically distributed, but more general positive integer-valued burst laws are also permitted. The conjecture was recently proved for unit births, $K_i\equiv1$. We show here that the conjecture is \textit{false in general} by constructing a two-component network with bounded production rates and burst sizes in $\{1,2\}$ for which both stationary Fano factors lie below their common constant-rate baseline. We then prove the conjecture for positive geometric bursts, the canonical burst model in stochastic gene expression. For arbitrary regulatory architecture and nonlinear cross-regulatory production rates, we prove an exact weighted tradeoff that rules out simultaneous suppression of every component below its geometric-burst baseline. We also prove a complementary structural impossibility result for arbitrary positive integer-valued burst laws with finite second moments: if the activating and inhibiting interactions have a globally consistent sign structure, in the sense that every cycle of the regulatory interaction graph contains an even number of negative interactions, then every component individually satisfies $F_{X_i}\ge B_i$, where $F_{X_i}$ is its stationary Fano factor and $B_i$ is its constant-rate burst baseline.


HierarchicalDAEW: Domain-Aware Edge-Weighted Graph Convolution with Evidential Uncertainty for Multi-Section Spatial Gene Expression Prediction from H&E Histology

Spatial transcriptomics assays remain costly and technically demanding, restricting transcriptome-wide profiling to specialist settings and preventing routine clinical deployment. Predicting spatially resolved gene expression from H&E histology could close this gap, yet current methods largely ignore the underlying tissue architecture and rarely quantify how their predictions can be trusted. We introduce HierarchicalDAEW, a dual-graph architecture that addresses both gaps. On the spot graph, a Domain-Aware Edge-Weighted convolutional operator learns separate projections for inter-domain, intra-domain, and boundary edges derived from Leiden clustering, allowing the model to treat tissue heterogeneity as an explicit structural signal rather than an implicit one. A second gene-level graph then fuses protein-protein interaction priors from STRING-DB with tissue-specific co-expression through learned attention gating, propagating predictions from a landmark gene set to a broader gene panel. Reliability is handled through evidential uncertainty estimation, which produces far better calibrated confidence intervals than Monte Carlo dropout under identical conditions. Across six human Visium sections spanning breast, colorectal, prostate, and cerebellar tissue, and against thirteen published baselines, HierarchicalDAEW achieves the strongest correlation with ground-truth expression, with gains that hold up under multi-seed reproducibility checks and negative controls that rule out positional shortcuts. Ablations further confirm that both the domain-aware edge typing and the hierarchical depth are necessary to this improvement, and calibrated uncertainty estimates identify low-confidence predictions for pathologist review before clinical action.


A Kuramoto phase model to explore the synchronisation of a network of circadian clocks

We propose a model of the circadian clock in a population of cells based on a network of oscillators, derived from the Kuramoto model. The coupling between oscillators is described by a global interaction term but we introduce a phase-dependent coupling mechanism, such that oscillators interact only within a specific interval of the cycle corresponding to a specific stage of the circadian cycle. We analytically demonstrate that this modified system achieves complete asymptotic phase synchronisation, provided specific conditions on the initial phase distribution and the coupling window length are met. To bridge this theoretical framework with experimental observations, we introduce a signal processing procedure based on wavelet decomposition to extract quantitative oscillatory features from Per2::luciferase reporter traces. We then calibrate the model against datasets from both wild type and Cry2 knockout hepatocyte spheroids using a two-step quasi-Monte Carlo filtering algorithm. The comparative analysis reveals significant phenotypic divergence, showing that the Cry2KO condition alters the identifiability landscape of the model's parameters and introduces new compensatory mechanisms that confound the initial population heterogeneity with long-term macroscopic signal decay.


Graph Learning on Ensembles of Cyclic Peptides: An Investigation of Molecular Ensemble Modeling

Molecular property prediction from structure often uses a single representative conformation, even though many molecules exist as conformational ensembles in solution. We introduce EnsembleEGNN, a molecular ensemble foundation model that encodes an ensemble by first encoding each conformer with shared Equivariant Graph Neural Network (EGNN) layers, then pooling the resulting conformer representations with a Set Attention Block. We pretrain the model on CREMP, a cyclic peptide ensemble dataset, using a multi-task self-supervised objective combining masked token recovery, noisy-coordinate reconstruction, and pairwise distance reconstruction. On the CREMP-CycPeptMPDB dataset, training EnsembleEGNN from scratch fails entirely ($R^2=0.005$). However, the pretrained model reaches $R^2=0.477$ and Pearson $r=0.699$, outperforming the sequence-only BERT baseline ($R^2=0.439$, Pearson $r=0.667$). When EnsembleEGNN is co-trained end-to-end with the BERT sequence encoder, the hybrid model improves further to $R^2=0.538$ and Pearson $r=0.737$. These results demonstrate that encoding conformational ensembles into a single thermodynamically informed embedding improves cyclic-peptide property prediction.


Necessary and sufficient condition for hysteresis in the mathematical model of the cell type regulation of \textit{Bacillus subtilis}

The key to a robust life system is to ensure that each cell population is maintained in an appropriate state. In this work, a mathematical model was used to investigate the control of the switching between the migrating and non-migrating states of the Bacillus subtilis cell population. In this case, the motile cells and matrix producers were the predominant cell types in the migrating cell population and non-migrating state, respectively, and could be suitably controlled according to the environmental conditions and cell density information. A minimal smooth model consisting of four ordinary differential equations was used as the mathematical model to control the B. subtilis cell types. Furthermore, the necessary and sufficient conditions for the hysteresis, which pertains to the change in the pheromone concentration, were clarified. In general, the hysteretic control of the cell state enables stable switching between the migrating and growth states of the B. subtilis cell population, thereby facilitating the biofilm life cycle. The results of corresponding culture experiments were examined, and the obtained corollaries were used to develop a model to input environmental conditions, especially, the external pH. On this basis, the environmental conditions were incorporated in a simulation model for the cell type control. In combination with a mathematical model of the cell population dynamics, a prediction model for colony growth involving multiple cell states, including concentric circular colonies of B. subtilis, could be established.


Evaluation and Prognostic Validation of Deep Regression Models for WSI-Based Gene-Expression Prediction

Gene-expression profiling is widely used in research and central to many areas of precision oncology, but remains costly and not universally accessible. Recent advances in computational pathology enable prediction of transcriptomic profiles directly from hematoxylin and eosin (H&E)-stained whole-slide images (WSIs), although optimal modeling strategies and clinical relevance remain unclear. In this study, we systematically evaluate deep regression models for WSI-based gene-expression prediction across multiple regression formulations and pathology foundation models (PFMs), and assess whether the resulting predicted transcriptomic signals retain prognostic utility. Across four TCGA datasets, we find that direct regression using attention-based multiple instance learning together with PFM feature extractors provides a strong and computationally efficient baseline, with no consistent benefit from separately training multiple models on subsets of genes. We then externally validate the selected configuration on an independent cohort of 997 breast cancer patients, demonstrating robust generalization for clinically relevant gene sets such as PAM50. To assess clinical relevance, we further evaluate predicted gene-expression scores in two independent population-representative breast cancer cohorts comprising 4,172 patients with survival endpoints, where predicted scores retain prognostic value in both the full patient cohort and the ER+ & HER2- subgroup. Together, these results demonstrate that WSI-based gene-expression prediction can generalize across independent cohorts and recover biologically and clinically meaningful molecular structure, supporting its potential as a scalable approach for transcriptomic phenotyping and risk stratification.


Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks

Although computational phenotyping is a central informatics activity with resulting cohorts supporting a wide variety of applications, it is time-intensive because of manual data review. We previously assessed the ability of LLMs to perform computational phenotyping tasks using computable phenotypes for ARF respiratory support therapies. They successfully performed concept classification and classification of single-therapy phenotypes but underperformed on multi-therapy phenotypes. To better understand issues with these complex tasks, we expanded PHEONA, a generalizable framework for evaluation of LLMs, to include methods specifically for evaluating faulty reasoning. We assessed the responses of two lightweight non-reasoning LLMs (Mistral Small 24 billion and Phi-4 14 billion) and one lightweight reasoning LLM (Qwen-distilled DeepSeek-r1 32 billion) both with and without prompt modifications to identify explanation correctness errors and unfaithfulness errors during phenotyping. For experiments without prompt modifications, both errors were present in responses from all models. For experiments with prompt modifications, we measured the mean absolute change in accuracy relative to the unbiased prompt across biasing conditions. Adding specific few-shot examples aligned with an incorrect phenotype reduced accuracy by at least 5% and up to 10% depending on the model and CoT type. Since reasoning errors were ubiquitous across models, our enhancement of PHEONA to include a component for assessing faulty reasoning provides a practical framework for evaluating LLM reasoning and empirical evidence that reasoning errors occur during complex computational phenotyping.


Epidemic "momentum" and a conservation law for infectious disease dynamics

Infectious disease outbreaks have precipitated a profusion of mathematical models. Epidemic curves predicted by these models are typically qualitatively similar, despite distinct model assumptions, but there is no theoretical explanation for this similarity in terms of any recognised common structure. We introduce a unifying concept of "epidemic momentum"---prevalence weighted by potential to infect---which is more informative than prevalence, yet analytically tractable. Epidemic momentum reveals a common underlying geometry in which outbreak trajectories always follow contours of a conserved quantity. This previously unrecognised conservation law constrains how epidemics can unfold, enabling us to disentangle transmissibility from prior immunity and to infer each separately from the same time series. Epidemic momentum also exposes the true final size of an outbreak and a universal phase-plane description that links generic renewal models to the classical SIR system.


The Influence of Width Ratios on Structural Beauty in Male Faces

This study investigates the relationship between interocular distance relative to overall facial width (width ratio) and perceived subjective beauty in male faces. Building on the methodology of Pallett et al. (2010), who found that average proportions in female faces were rated as most attractive, the current study aimed to test this hypothesis in male faces. Faces from the Chicago Face Database (Ma et al., 2015) were morphed into average faces within three groups (with low, medium, and high width ratios), each composed of 96 or 97 individual images. These three average faces were then systematically manipulated in their width ratios across three levels in both directions, respectively, resulting in a total of 21 comparable faces. The use of multiple base faces served as a control for potential artifacts of image processing. Consequently, comparisons were restricted to within-group pairs to avoid confounding by co-varying facial features (e.g., skin tone), which precluded direct cross-condition comparisons but ensured internal validity. In a two-alternative forced-choice task, participants selected the more beautiful face from each pair. The data were analyzed using a Bayesian model which enables inference of the width ratio perceived as most beautiful. Results support the hypothesis that averageness in facial proportions correlates with higher perceived attractiveness. The study highlights the importance of controlling for image manipulation, including attempts at methodological implementation, and of considering ethnicity as a potential moderating variable. These findings offer a data-driven foundation for understanding facial aesthetics and cognitive processes of human perception, with applications in advertising, artificial face generation, and plastic surgery.


Omics Data Discovery Agents: Agent-Supported Retrieval, Reanalysis, and Synthesis of Published Omics Data

The biomedical literature contains a vast collection of omics studies, yet most published data remain functionally inaccessible for computational reuse. When raw data are deposited in public repositories, essential information for reproducing reported results is dispersed across main text, supplementary files, and code repositories, and in the rarer cases where intermediate data (e.g. protein abundance files) are shared, their location is irregular. Here we present an agentic framework for the agent-supported retrieval, reanalysis, and synthesis of published omics data. The system employs large language model (LLM) agents with access to tools for fetching omics studies, extracting article metadata, identifying and downloading published data, executing containerized quantification pipelines, and synthesizing results across studies. Applied at corpus scale, the pipeline cataloged dataset references across thousands of PubMed Central articles; we report these as descriptive system outputs rather than as a validated measure of extraction accuracy. Using model context protocol (MCP) servers to expose containerized analysis tools, the agents retrieved and re-quantified data in five end-to-end reanalyses spanning data-dependent and data-independent proteomics and bulk RNA-seq. All five reanalyses completed, each with documented human guidance and workflow accommodations, and reproduced the authors' deposited abundances with high per-sample correlation (0.85-0.997) and strongly concordant differentially expressed features (fold-change Spearman 0.88-0.91), with no direction reversals among features called differentially expressed in both analyses; residual differences in significant-feature lists were attributable to threshold placement, tool-version, and preprocessing differences rather than to the underlying quantities.


Neutralization titers reveal the structure of polyclonal antibody responses

The composition of a polyclonal antibody response is hard to measure experimentally but contains vital information about the robustness of immunity. Here, we argue that the statistics of neutralization titers alone can be used to make quantitative predictions about the composition of the response, circumventing challenges arising through sequencing and monoclonal antibody expression. We show that the response against influenza within a cohort can be either driven by a collective phenomenon where many antibodies contribute to neutralization, or dominated by just a few strong binders, leading to a broad distribution of titers across individuals described by a Gumbel distribution from extreme value theory. Comparing titers across cohorts, we find that Gumbel statistics {accurately describe} individuals prior to an immune challenge. We propose an equilibrium binding model that quantitatively captures titer data and illustrates the structure of the polyclonal response. Our approach extends generically to immune responses to other pathogens.


The Sensation Modulating Network:Haltability as the architectural ground for object-directed phenomenology

We propose the Sensation Modulating Network (SMN): the cognitive agent as the whole body, organized at every scale by opponent dynamics, built from Sensation Modulators -- tissue that senses and acts through one substrate -- paired into Coordinated Action Zones routed by a body-wide broadcast. It is an inclusive model of the body, in which gravity, elasticity, and the body's topology and geometry do constructive cognitive work. The paper is scoped to what such a body constructs at its foundation -- a self-model, a world-model in that self's frame, and object-directedness -- each built by the body's physics, not assumed as a primitive. The architecture is generative: one small kit of primitives whose morphological variations (chain, sheet, tube, layered, appendicular) construct experience by the same mechanism, an invariance shown for the self-model across body plans and scales. The central thesis: haltability -- the active holding of an opponent equilibrium -- is the architectural condition object-directed phenomenology requires; a second principle, that an object is a bundle of more than one property, carries it from felt resistance to a genuine object. A companion bench realizes each construction as a runnable, falsifiable experiment with a pre-registered order parameter and matched foil. We place the principal competing accounts -- sensorimotor enactivism, active inference, and ecological and affordance-based theories -- as limiting cases within a wider landscape, stating in each case the criterion that would tell them apart, and give systems and cognitive neuroscience its place: the nervous system as the integrating core that makes the body one, not a commander over it. On this account, the cognitivism-4E impasse reflects an incomplete architecture of the embodied agent: its resolution begins not with the brain alone but with the whole body.


Evaluating the Impact of Epidemic Control via State-Dependent Markovian Switching Modeling

We develop an exact finite-population stochastic framework for SIR epidemics evolving under Markovian switching between intervention regimes. The epidemic state is augmented by a finite phase component, allowing transmission, recovery, and direct immunity-acquisition rates to depend on the active regime. Phase-transition intensities may depend on the current epidemic state, so that policy escalation can react to the number of infectious individuals. Exploiting the monotonicity of the susceptible compartment, we derive level-wise recursions for the joint Laplace--Stieltjes transform and probability generating function of the extinction time and the number of infections generated before extinction. These recursions yield the infection-count distribution, conditional extinction-time transforms, and mixed moments linking epidemic duration and infection burden, while replacing a large global linear system with small phase-level solves. The framework is illustrated using weekly mpox incidence data from Luxembourg. A baseline one-phase SIR model is calibrated by maximum likelihood under a Poisson observation model. The calibrated baseline is then used for conditional comparisons of fixed control regimes, early versus delayed strict intervention, vaccination-supported control, and state-dependent escalation. The results show how switching mechanisms affect both the total number of infected individuals and the extinction time, including their dispersion. Since the switching mechanisms are specified rather than estimated from the intervention history, the results are conditional model-based comparisons rather than estimates of the historical effects of interventions in Luxembourg.


Gravity-Driven Eco-Epidemiological Dynamics in Tri-Trophic Food Chains

Ecological communities are shaped by the interplay between trophic interactions and infectious disease, yet how spatially mediated interactions influence disease-driven ecosystem dynamics remains poorly understood. Here, we develop a gravity-based eco-epidemiological framework for a tri-trophic food chain in which trophic interaction depends on species abundances and effective interaction distance. The disease-free food chain system supports a stable coexistence equilibrium, providing a baseline for investigating disease-induced ecological transitions. Introducing infection at the intermediate trophic level destabilizes this equilibrium through a Hopf bifurcation, leading to sustained oscillations, whereas infection at the top predator level results in a qualitatively different transition from persistence to extinction. By systematically varying the gravity coupling strength, we show that gravity-mediated trophic interactions regulate the thresholds separating these ecological regimes, while the trophic position of infection determines the nature of the transition. Together, these findings establish a unified framework for understanding how spatially mediated trophic interactions and infectious disease jointly govern ecosystem stability, providing new insights into disease-driven dynamics in ecological communities.


Evolutionarily Stable Stackelberg Equilibrium

We present a new solution concept called evolutionarily stable Stackelberg equilibrium (SESS). We study the Stackelberg evolutionary game setting in which there is a single leading player and a symmetric population of followers. The leader selects an optimal mixed strategy, anticipating that the follower population plays an evolutionarily stable strategy (ESS) in the induced subgame and may satisfy additional ecological conditions. We consider both leader-optimal and leader-pessimal selection among ESSs, which arise as special cases of our framework. Prior approaches to Stackelberg evolutionary games either define the follower response via evolutionary dynamics or assume rational best-response behavior, without explicitly enforcing stability against invasion by mutations. We present algorithms for computing SESS in discrete and continuous games, and validate the latter empirically. Our model applies naturally to biological settings; for example, in cancer treatment the leader represents the physician and the followers correspond to competing cancer cell phenotypes.


Multimodality Stacking with Blockwise missing values and application to the PIONeeR biomarkers study for prediction of resistance to immunotherapy

Integrating multimodal datasets in clinical oncology is frequently hindered by high dimensionality and blockwise missingness, where entire data sources are unavailable for specific patient subsets. Standard survival models often struggle with these gaps, leading to biased results or patient exclusion. We introduce Multimodality Stacking with Blockwise missing values (MSB), a late-fusion framework for survival analysis that independently models modality-specific features before aggregating predictions via a cross-validated stacking meta-learner. MSB was validated on the PIONeeR study (n=443 patients, 378 biomarkers across eight heterogeneous sources) to predict progression-free survival in advanced non-small cell lung cancer patients receiving immunotherapy. MSB yielded higher predictive performance (C-index) than baseline algorithms. Improvements varied by baseline strength: linear models showed a 15.9% increase (p<0.001 for the Wilcoxon signed-rank test), random survival forests gained 5.4% (p=0.002), and gradient boosting methods improved by 2.1% (p=0.030). Beyond discrimination, MSB reduced the generalization gap (train-test difference in 5 folds cross-validation repeated 3 times: 0.055 vs 0.380 for linear models). Permutation importance analysis identified routine laboratory markers, clinical features, and PD-L1 expression as primary predictive drivers. Missing block indicators showed negligible importance, suggesting the model learned from biomarker values rather than data availability patterns. MSB provides a statistically validated framework for multimodal survival prediction with blockwise missingness. By enabling systematic biomarker evaluation without requiring complete data, MSB offers a practical tool for predictive modeling in biomedical research, pending external validation. Implementation is available at this https URL under Inria license.


A multi-ensemble mean-field reduction method for networks of globally coupled phase oscillators with arbitrary parameter distributions

Understanding the dynamical properties of coupled phase oscillator systems with heterogeneous oscillator frequencies has been a long-standing challenge of complex systems theory. While the seminal work of Ott and Antonsen dramatically improved our theoretical understanding of coupled phase oscillators for a small family of oscillator frequency distributions, we here present a mean-field reduction method for arbitrary frequency distributions. Our method leverages the drastic dimensionality reduction obtained for Lorentzian frequency distributions, and combines it with a data-driven multi-ensemble approach. As such, the method renders the Ott-Antonsen equations directly applicable to empirical distributions of phase oscillator frequencies, often achieving a drastic dimensionality reduction and allowing to study real-world physical and biological systems by means of stability, sensitivity, and bifurcation analyses.


The Site Frequency Spectrum in an Exponentially-Growing Population with Selection

We consider a supercritical two-type continuous-time linear birth-death process with mutation and selection, in which wild-type individuals give rise to mutant offspring with a larger net growth rate. In this setting, we investigate the site frequency spectrum (SFS) of driver mutations, describing the number of driver mutations present at any given frequency in the population. First, we derive exact moments for the SFS and establish asymptotic power laws at large times and frequencies. Then, strong laws of large numbers for the driver SFS are proven by constructing suitable $L^2$-approximations. These results apply both to the case when all driver clones have the same selective advantage and when the selective advantage is random. Finally, we allow the frequency to vary with time to examine the number of "intermediate" and "large" driver clones, identifying a cutoff frequency at which there are order 1 number of mutant clones. Overall, our results provide quantitative insights into how selection shapes the site frequency spectrum both at small and large frequencies, which can in principle be leveraged to construct estimators of relevant evolutionary parameters, including the selective advantage of driver mutations.