Quantitative Biology

Browse today’s new papers as interactive HTML on Academus.


Emergent topological structure in spontaneous brain-organoid activity

Neural activity is widely held to organize on low-dimensional structure embedded in a high-dimensional state space. Persistent homology reads such structure directly from the pattern of pairwise correlations, without assuming in advance which variables are relevant. We apply persistent homology to microelectrode-array (MEA) recordings of spontaneous activity from human (Lancaster) and mouse (Paşca) cortical organoids, spanning $26$--$234$ simultaneously sorted units, and ask whether topological data analysis resolves structure at the node counts that neural recordings actually deliver. Building weighted networks in correlation space and characterizing them by Vietoris--Rips filtration, we find that the first homology ($H_1$, loops) rises significantly above a rate- and population-preserving null in $14$ of $18$ datasets. This loop structure occupies a non-redundant core: it is robust to random removal of units yet disrupted by targeted removal of the units that carry it. Topological richness grows with network size, and second homology ($H_2$) emerges significantly above the null only in the larger networks. These results show that persistent homology resolves structured topology in neural recordings at the scale experiments actually deliver.


The Origins of Transient Bimodality

Many dynamical systems exhibit diverse modes of behavior. In biology, such modes can represent individual or cell fates. While the emergence of multimodality is commonly studied, transient bimodality is much less well understood. Under transient bimodality, a system moving from a well-defined initial to a final state transiently undergoes a bifurcation into multiple probability modes. This noise-driven phenomenon can significantly impact processes such as cell differentiation and speciation in the presence of changing environmental conditions. We detail a theoretical approach for understanding transient bimodality connecting results from ecology, optics, chemical reaction networks and cell biology, propose a ``minimal model'' of transient bimodality and derive a general criterion for its presence. We show that fast-to-slow dynamics can lead to transient bimodality in addition to the well-known case of slow-to-fast dynamics. Finally, we discuss the role of transient bimodality across the scientific literature, with emphasis on biochemical kinetics and gene regulation.


Orientation Reading by Production Vision-Language Models on Optotype Charts: A Controlled Multi-Model Evaluation Across Reasoning Modes, Prompts, and Access Modalities

OBJECTIVES: Vision-language models are increasingly used to interpret medical and everyday images through consumer chat interfaces, yet their ability to read orientation - the single perceptual operation tested by the tumbling-E acuity optotype - is poorly characterized on the surfaces through which they are actually used. METHODS: We evaluated four production vision-language models (referred to as Claude, GPT, GROK, and Gemini) through their consumer chat interfaces on a locked set of seven optotype charts: four uniform tumbling-E charts (one per cardinal orientation), two mixed-orientation tumbling-E charts, and one Snellen letter chart as a specificity control. Each model was run in two reasoning modes (Fast and Thinking) under two prompt variants (with and without an explicit orientation-decoding rule) by up to three operators. The corpus comprised 920 scoreable trials and 50,420 glyph judgements. The primary outcome was glyph-level accuracy against the chart's designed orientation, summarized with Wilson 95% confidence intervals. RESULTS: Accuracy ranged from 43.0% to 97.0% across models on identical charts, and the strongest model depended on reasoning mode (GPT 97.0% in Fast mode; GROK 96.6% in Thinking mode). Errors were not random but collapsed onto a model-specific attractor direction. Models were 96-100% internally self-consistent yet ranged widely in accuracy, dissociating reliability from validity. An answer-key-free ensemble-consensus estimate tracked accuracy closely (r = 0.998). For one model, consumer-interface accuracy fell 25-27 points below programmatic access, almost entirely on a single orientation. CONCLUSIONS: A single accuracy figure conceals clinically relevant, orientation-specific failure modes; vision-language models should be evaluated along multiple axes and on the deployment surface before image-interpretation outputs are trusted.


The Positive Experience Principle: Forecasting Conscious Choices with AI Embeddings

A fundamental challenge in the science of consciousness is the lack of a universal, predictive framework for motivated behavior. While existing theories excel at describing specific mechanisms, from neural pathways to computational models, they do not provide a foundational principle that explains the consistent direction of conscious systems toward certain states and away from others. To address this gap, we propose the Positive Experience Principle (PEP), a unifying principle stating that conscious systems have an inherent tendency to move toward states of higher positive subjective experience. This tendency is quantified by a Positive Experience Value (PEV), a scalar metric derived from the physical configurations defined by our earlier Universal Consciousness Code (UCC) theory. The PEP bridges physics, neuroscience, and psychology by positing that diverse behaviors are manifestations of a single, fundamental drive to optimize PEV. The PEP generates testable predictions for the dynamics of conscious systems, offering a path toward a unified science of behavior.


Evolutionary dynamics in public goods games with general frequency-dependent returns

The public goods game serves as a significant paradigm for investigating the emergence and maintenance of cooperation in conflicting situations. In the traditional public goods game, the multiplication factor characterizing the synergy effect of common efforts is typically assumed to be constant. In real-world scenarios, however, investment returns are often dynamic and vary with the strategic composition of the interaction group. To date, the evolutionary dynamics of the public goods game with such frequency-dependent returns have remained not fully understood. In this work, we introduce a general frequency-dependent multiplication factor that depends on the strategy composition within the game group. Through theoretical analysis, we derive the mathematical conditions under which cooperation is favored. Our results show that whether cooperation has an evolutionary advantage over defection depends on the investment return rate in the full-contribution state of the game group, irrespective of the return rates in other states. An increase in this rate leads to a higher abundance of cooperators. Furthermore, we introduce a general frequency-dependent multiplication factor into the public goods game with peer punishment, and systematically explore its effects on the cooperation dilemma and the second-order free-rider problem. Our results highlight that the abundance of cooperators or punishers is governed solely by the investment return values in the full-cooperation and full-punishment compositions of the group. A higher return rate in the full-cooperation state facilitates the promotion of cooperation, whereas a higher return rate in the full-punishment state favors the emergence of punishment. Our theoretical findings are verified by individual-based simulations.


Laboratory Trajectories Improve Kidney Failure Risk Estimation

Accurate kidney failure risk assessment is critical to timely intervention in chronic kidney disease (CKD). Existing equations (e.g. Kidney Failure Risk Equation; KFRE) rely on single laboratory measurements to estimate short- and long-term kidney failure risk, leaving longitudinal laboratory patterns unused. Here we introduce Clalit Longitudinal Assessment of Risk of Kidney Failure (CLARK), an interpretable longitudinal extension of latest-value methods which incorporates routinely collected repeat laboratory measures. We develop CLARK using data from 5.4 million individuals, identifying 270,009 patients with CKD to create one of the largest longitudinal CKD cohorts to date, with 12,087 kidney replacement therapy initiation events and a median follow-up of 10.4 years. Across laboratory configurations and prediction horizons, CLARK demonstrated improved discrimination over static models (e.g., 2-year average precision 0.541 vs 0.516 in the eGFR-only setting). At intervention thresholds, trajectory-based models improved identification of high-risk patients, especially for longer-term prediction, suggesting that interpretable longitudinal laboratory features may enhance kidney failure risk assessment through improved identification of patients most likely to benefit from timely intervention.


Post-transcriptional Regulation of Stochastic Gene Expression Conditioned on Large Deviations

Gene expression is a stochastic process that gives rise to large fluctuations in protein levels leading to phenotypic heterogeneity in clonal cell populations; post-transcriptional regulation plays a crucial role in controlling the level of phenotypic variability within a population, which is directly tied to cell-fate decisions. As such, substantial efforts have been directed towards quantitatively modeling the effects of various post-transcriptional mechanisms on the strength of fluctuations in protein levels (noise). However, the corresponding effects of post-transcriptional regulation on the occurrence of rare events corresponding to large deviations are far less explored and have only been considered for a special model. Here, we take a general model of post-transcriptional regulation and apply the partitioning of Poisson arrivals (PPA) framework to map it onto a model that resembles promoter-based regulation of transcription, leading to a general framework to obtain objects of interest in large deviations (i.e. large deviation rate function for quantifying the likelihood of observing rare protein production rates and the corresponding driven process that characterizes the system dynamics conditional on the rare event) for models of post-transcriptional regulation directly from prior results for promoter-based models. The results derived create new avenues to analyze rare events in general models of post-transcriptional regulation pertaining to various different biological settings.


Network Characteristics of Individual Pigments in Cyanobacterial Photosystem II Core Complexes

Part of the excitation energy transfer (EET) characteristics of the photosystem II (PSII) comes from the interconnection between pigments. To understand the correlation between the EET and the pigments' interaction structure, we construct a network from the EET rates, which are related to both the distance between the pigments (chlorophylls and pheophytins) and their spatial orientations. Especially, we investigate how well the PS II core complex's EET functionality can be explained by using only the network topology in Thermosynechococcus vulcanus 1.9 Å. Starting from the Förster theory, we construct a network of EET pathways. For an analysis of the network structure, we calculate common network-structural measures like betweenness centrality, eigenvector centrality and weighted clustering. These measures can reflect the role of individual pigments in the EET network. In our work, we found that some well-known properties were reproduced by the network analysis of the simplified network, which means that the topology of the network encodes functionally relevant information. For example, from the network structural analysis, we can infer that most of the chlorophyll molecules (chlorophylls) in the pigment-protein complex CP47 have a heightened probability of transferring energy compared with other chlorophylls. We also see that the active branch chlorophylls in the reaction center are characterized by a high eigenvector centrality, a high betweenness centrality, and a low weighted clustering coefficient. This is indicative of functionally important vertices.


Harmonised benchmarking of foundation models for single-cell and spatial transcriptomics reveals context-dependent generalisation

Single-cell and spatial foundation models promise transferable biological representations, yet their generality remains largely untested across modalities, biological domains and analytical tasks. We benchmarked six representative models, Nicheformer, CellPLM, scGPT-spatial, GenePT, scELMo and Novae, using a harmonised framework spanning scRNA-seq, spatial transcriptomics and Perturb-seq. We evaluated zero-shot and continually pretrained clustering, supervised annotation, marker-gene concordance and perturbation prediction. Model performance was strongly conditional: expression-trained cell-level transformers best resolved many cell-identity tasks, spatial and graph-aware models better preserved tissue architecture, and language-derived gene embeddings were competitive for selected perturbation-response metrics. No model dominated across tasks, and rankings shifted with modality, preprocessing, tokenisation, biological prior, domain shift and metric choice. This benchmark provides practical guidance for model selection and argues that future models should be judged by biological generalisation, interpretability and perturbation-grounded validity, not by scale or leaderboard performance alone.


Evaluating Conformal Reliability of Pathway-Level Transcriptomic Signatures Under Cross-Cohort Shift in Sepsis Mortality Prediction

Blood transcriptomic profiling enables prognostic modeling by capturing the host immune response at the molecular level. Yet, the within-cohort evaluation strategies employed by many transcriptomic models inadequately reflect deployment across independent hospitals. Outside deployment scenarios introduce a cohort shift that can substantially degrade predictive performance and reliability of uncertainty estimates. We present a framework for evaluating transcriptomic sepsis mortality prediction under realistic cross-cohort deployment, systematically comparing gene-level, pathway-level and hybrid molecular representations. Four publicly available whole-blood transcriptomic cohorts consisting of 936 patients and 248 mortality events were harmonized into a shared 7,660-gene feature space and evaluated under leave-one-cohort-out validation using logistic regression, random forests, XGBoost and LightGBM. Beyond AUROC and AUPRC, model behavior was evaluated via conformal prediction, calibration analysis, selective prediction and the proposed Pathway Stability Index. Gene-level and hybrid representations were found to generally achieve the strongest discriminative performance, whereas pathway-level representations exhibited greater robustness across model families, more reliable uncertainty behavior under cross-cohort shift and stable molecular signatures enriched for immune and host-defense processes identified through Gene Ontology and KEGG enrichment analyses. These findings demonstrate that molecular representation influences not only predictive discrimination but also calibration, uncertainty reliability, biological coherence and transferability under external validation.


Exploring Brain Networks Using Noninvasive Electrophysiological Measurements: Methods and Applications

Electroencephalography (EEG) and magnetoencephalography (MEG) provide noninvasive measurements of brain activity with millisecond temporal resolution, enabling the investigation of functional and effective interactions within large-scale brain networks. This chapter presents a comprehensive overview of the methodological foundations and practical workflows for EEG/MEG-based brain network analysis. We first review the physical principles underlying EEG and MEG, emphasizing their complementary strengths and limitations. We then describe the forward and inverse problems, including subject-specific head modeling, source reconstruction techniques, and the importance of accurate anatomical modeling for reliable source localization. Strategies for mitigating volume conduction and signal leakage are discussed, together with best practices for source-space connectivity analysis. The chapter reviews widely used functional and effective connectivity measures, including coherence, phase synchronization metrics, amplitude envelope correlation, Granger causality, dynamic causal modeling, and transfer entropy, highlighting their assumptions, advantages, and limitations. Modern end-to-end analysis pipelines are presented, with particular emphasis on Brainstorm and complementary open-source software for reproducible EEG/MEG research. Finally, we discuss emerging approaches, including time-varying connectivity, cross-frequency interactions, and network-based analyses, illustrating how noninvasive electrophysiology contributes to understanding brain organization in health and disease. The chapter provides both conceptual foundations and practical guidance for researchers and advanced students seeking to map and interpret human brain networks using EEG and MEG.


How genome redundancy can promote evolutionary innovation

Polyploidy is defined as the existence of more than two complete sets of homologous chromosomes. Despite it being a widespread phenomenon across the tree of life, its role as either an evolutionary innovation or a dead end is still debated. Here, we investigate how under varying selective pressures the degree of ploidy interacts with two key biological factors: the mode of inheritance and the genotype-phenotype mapping. Through a minimal evolutionary model we find that polyploidy is especially advantageous during abrupt environmental changes, confirming that polyploidization is often associated with ecological upheavals. We observe that stochastic inheritance combined with a nonlinear (maximum-based) genotype-phenotype mapping maximizes both phenotypic exploitation and landscape exploration across all environments. By contrast, structured inheritance with an additive phenotype mapping systematically underperforms, yet displays a pronounced optimum at low-to-intermediate ploidy level that mirrors the distribution observed in natural plant and bacterial populations. When individuals are free to carry different chromosome numbers, selection drives the population toward values that reflect an interplay between exploitation, exploration, and convergence speed rather than any single evolutionary objective. The relative weight of these three factors depends on the fitness landscape, providing a unifying framework for understanding when and why polyploidy is favored by natural selection.


The energy landscape of DNA-binding proteins along the genome

Reconstructing the energy profile of DNA-binding proteins along the genome requires an algorithm that quantifies efficiently the binding free energy. We assembled a dataset of protein structures and DNA binding sites, together with their binding energies, and used it to train a machine-learning algorithm that learns a latent invariant representation of the protein interface and of the DNA sequence, combining them to predict the free energy. After validating the method, we used it to determine the energy profile of a single-domain transcription factor (PU.1) sliding on mammalian chromosomes, predicting its binding regions and quantifying the statistical properties that determine their stability and their kinetic accessibility.


Diffusion-induced instabilities promote cooperation in eco-evolutionary networks

Understanding how cooperation persists despite the advantage of selfish behavior remains a central challenge in evolutionary dynamics. Classical models of public goods dilemmas predict dominance of defectors, yet natural and social systems often sustain cooperation. We study an eco-evolutionary public goods game on complex networks where cooperators and defectors diffuse at different rates. When the isolated system is in a defector-dominated coexistence regime, faster dispersal of defectors than cooperators leads to a symmetry-breaking transition that produces localized clusters of cooperators. In heterogeneous networks, nodes with higher connectivity become significantly more likely to exhibit cooperative dominance. A degree-based mean-field reduction supports this result by showing that network connectivity controls an effective coupling strength proportional to node degree, thereby producing a bifurcation that separates defector-dominated and cooperative states. We also address why not all hubs become cooperative by means of a multistability analysis. These results reveal how asymmetric mobility and heterogeneous connectivity jointly promote cooperation in structured populations.


Overcoming the BCI Calibration Bottleneck: A Clinically-Grounded Architecture using Riemannian Alignment and Stochastic Weight Averaging

Brain-Computer Interfaces (BCIs) face a severe calibration bottleneck due to cross-subject spatial covariance shifts and physiological artifacts. To enable zero-calibration BCI, a deep learning pipeline was engineered combining Per-Session Independent Component Analysis, Riemannian Euclidean Alignment, and EEGNet stabilized by Stochastic Weight Averaging (SWA). Evaluated on the strict MOABB BNCI2014-001 benchmark, the proposed architecture successfully isolates true sensorimotor rhythms. For the primary case study (Subject 1), a clinically robust SWA stable accuracy of 90.97% (AUC: 0.976, Cohen's $\kappa$: 0.819) was achieved. Furthermore, expanded 9-fold Leave-One-Subject-Out (LOSO) cross-validation yielded a globally stable mean accuracy of 74.31%, proving hardware-agnostic zero-shot efficacy for binary motor imagery.


Preference-based Antibody Expression Ranking: Scaling with Large-scale Weak Supervision

Antibody expression ranking is a critical task in antibody design, yet its modelling is severely hindered by the scarcity of labeled experimental data. To address this, we propose a unified preference-based learning framework that integrates scarce quantitative expression data with large-scale weak positive supervision from immunization data. We adapt Direct Preference Optimization (DPO) to protein language models by introducing a union-masked log-likelihood approximation and IMGT-based alignment, enabling efficient training on variable-length sequences. Evaluating on a diverse internal dataset of 1254 labeled sequences and 4 million unlabeled camelid-derived antibodies, we show that our method consistently outperforms baselines on most metrics. Our results demonstrate that preference learning can effectively learn from weak supervision, providing a scalable solution for antibody expressibility optimization in data-constrained settings. Project page: this https URL.


The Site Frequency Spectrum in an Exponentially-Growing Population with Selection

We consider a supercritical two-type continuous-time linear birth-death process with mutation and selection, in which wild-type individuals give rise to mutant offspring with a larger net growth rate. In this setting, we investigate the ``driver'' site frequency spectrum (SFS), or the random measure that records mutant allelic frequencies in the population. We examine various regions of the SFS. First, we derive exact moments and prove results concerning the mean behavior of the driver SFS at large times and frequencies. Strong laws of large numbers for the driver SFS are proven by constructing suitable $L^2$-approximations. These results are extended to the setting in which the fitness increase associated with each clone is random. Next, we allow the frequencies to vary with time to examine the number of ``intermediate'' and ``large'' clones. Using this, we find a cutoff frequency at which there are order 1 number of clones. Our results allow for estimation of relevant evolutionary parameters, such as the fitness increase of mutant versus wild-type cells.


A multiverse-consensus pipeline for reproducible feature selection in untargeted LC-MS metabolomics

Background: Untargeted LC-MS metabolomics requires a long chain of preprocessing decisions, each with several equally defensible options. Analysts typically commit to one pipeline and report the resulting feature shortlist. How strongly that shortlist depends on choices that were never varied stays invisible. Results: We adapt multiverse analysis to untargeted metabolomics feature selection. We present an auditable, configuration-driven pipeline that (i) applies a ten-stage quality-control filter cascade in which every feature's fate is logged, and (ii) runs the downstream analysis as a multiverse over four contrasting preprocessing philosophies, each combined with four feature-ranking methods under bootstrap stability selection and label-permutation testing. Only features recurring across paths enter a tiered consensus. On a demonstration dataset of five breast-cancer cell lines (30,370 detected features), the four single pipelines individually returned shortlists of 4-20 features whose pairwise agreement was as low as Jaccard = 0.05. The multiverse consensus retained 15 features (>=2/4 paths), of which one recurred across all four, although two paths (sharing normalization and drift-correction methods) dominate the consensus. A pipeline-wide label-permutation test found no false discoveries in 50 null permutations. Conclusions: Reporting only preprocessing-robust features, with a complete kept/dropped audit trail, converts hidden analytical degrees of freedom into an explicit, inspectable output. We discuss scope and limitations, including single-batch design and the need for independent validation.


Direct Clinical Joint Angle Extraction from Parametric Body Model Rotation Matrices

Quantitative joint angles are rarely available in routine care because the tools are slow, costly, or confined to a laboratory. We show that clinical joint angles can be read directly from the per-segment rotation matrices a parametric body model already produces, with no inverse-kinematics or musculoskeletal-model fitting step. On the OpenCap LabValidation cohort, using the GEM-X body-model estimator on single-smartphone video, our pooled mean absolute error is 4.50 degrees over the fifteen joint angles that match the OpenCap Monocular reference set, the same accuracy range as OpenCap Monocular's 4.8 degrees on the same cohort and reference standard, from a much simpler pipeline. The step that connects a body model to clinical angles is a small calibration table rather than an optimisation, so the same procedure transfers unchanged to other body models: repeating it on SAM 3D Body, changing only the table, gives 4.66 degrees, statistically indistinguishable from GEM-X, and runs in real time from a live single-camera stream. The method needs no per-recording inputs beyond the video itself: no participant height, no camera-intrinsics database, no per-subject model scaling. This broadens where movement analysis is practical, from in-clinic and at-home recording to telerehabilitation and large-scale decentralised studies.


Graph-Induced Tensor Liftings for Networked SEIR Models: Dimensional Reduction and Residual Analysis

Networked SEIR models describe epidemic spread within and between interacting subpopulations through contact-supported nonlinear transmission. Standard polynomial liftings based on complete ordered Kronecker tensors yield linear higher-dimensional representations, but their dimensions grow rapidly because they retain interactions absent from the transmission graph. This paper develops a graph-induced tensor lifting whose observables are selected from the effective transmission support. An exact edge-based quadratic representation separates linear compartmental transitions from nonlinear infection terms. A homogeneous hierarchy is then constructed recursively. The quadratic transmission field generates the next degree. The linear compartmental field saturates the resulting dictionary within that degree. The first edge-closure dynamics are linear up to an explicit cubic truncation residual, and higher-order truncations contain only next-degree terms. The first lifted dimension scales with the numbers of subpopulations and effective transmission channels. At fixed order, graph-induced dictionaries grow linearly with network size under uniformly bounded local connectivity, whereas complete polynomial liftings retain order-dependent polynomial growth. Uniform first edge-closure residual bounds depend on the transmission rate and the maximum weighted incoming transmission intensity. Numerical illustrations compare equal intensity per active channel with equal total incoming intensity. They confirm that dictionary dimensions depend only on graph support, whereas residual trajectories also reflect weight accumulation, weight distribution, and nonlinear propagation. These results provide a structured basis for reduced modeling and subsequent model-specific analysis and control.


Feedback-mediated circulation and persistence of stochastic fluctuations in gene regulatory circuits

Feedback plays a significant role in biochemical networks that govern a multitude of cellular functions, including development, adaptation, and homeostasis. Yet, how feedback topology controls stochastic fluctuations remains incompletely understood. Here, we develop a theoretical framework for two-node feedback motifs composed of activating and repressive regulatory interactions between two transcription factors. Under the linear noise approximation, we identify a feedback-driven contribution to node-wise fluctuations, termed cyclic noise, that arises specifically from loop closure. Cyclic noise is the component of fluctuations that circulates through the regulatory circuit. Its sign and magnitude distinguish whether feedback amplifies or attenuates node-wise fluctuations. We further show that feedback-mediated noise circulation leaves a temporal signature in the decay of steady-state autocorrelation, revealing how loop closure modifies the persistence of fluctuations. We thus provide a minimal framework for understanding how feedback architecture regulates both the magnitude and the temporal persistence of noise in gene regulatory circuits.


An Early Warning of Emerging Biosecurity Risks in Frontier LLMs

Frontier large language models (LLMs) are increasingly integrated into scientific workflows, yet their growing biological capabilities may outpace current safeguards. To assess the biological risks of frontier models, we develop Intern-BioBreaker, a specialized bio-red-teaming model, together with an integrated computational-to-physical framework that couples model-level stress testing with wet-lab validation. Within this framework, Intern-BioBreaker generates targeted jailbreak prompts to test whether aligned models can be induced to provide operational guidance for safety-sensitive biological tasks or produce sequence-level outputs with potentially harmful properties. Selected sequence outputs are then carried forward for DNA synthesis, host expression, and orthogonal protein verification to assess whether model-generated designs can yield the intended biological products. Our evaluation reveals a concerning gap between text-level safeguards and the risks posed by capable scientific models: (i) Intern-BioBreaker outperforms baseline attack models and reveals widespread bio-risk jailbreak vulnerabilities across both open-weight and proprietary frontier LLMs, with several targets reaching near-saturated or 100% task-level attack success rate (ASR); (ii) in sequence-level case studies, GPT-5.5 can be induced to generate modified viral candidate sequences with pathogenic potential; the corresponding translated proteins may exhibit even stronger receptor-binding affinity and thus enhanced infection potential; and (iii) end-to-end verification shows that selected model-generated biological designs are not merely textual artifacts, but can be physically realized under controlled experimental settings. These findings underscore the need for stronger biological red-teaming, nucleic acid synthesis screening, and safety mechanisms that keep pace with model capabilities.


A Mathematical Model of Dengue Transmission Incorporating Hospital Capacity and Threshold-Based Fogging Interventions

Dengue remains a major public health challenge in tropical regions, and recurring outbreaks suggest that current intervention strategies are not yet fully effective. Existing mathematical models typically assume unlimited hospital capacity and continuously applied fogging, neglecting practical constraints that strongly influence disease control. We develop a non-smooth ordinary differential equation model of dengue transmission that incorporates finite hospital capacity and a threshold-triggered fogging strategy activated when reported infections exceed a prescribed fraction of the available capacity. The model exhibits three epidemiologically relevant operating regimes, reflecting changes in hospitalization and vector-control policies as the epidemic progresses. We establish the existence and local stability of the disease-free and endemic equilibria. Numerical continuation confirms the analytical results and reveals boundary-equilibrium bifurcations at the switching thresholds, a Hopf bifurcation after hospital capacity is exceeded leading to sustained oscillatory outbreaks, and a fold bifurcation near the epidemic threshold that generates additional unstable equilibria. We further investigate periodic solutions with respect to the fogging rate and activation threshold, identifying locally optimal intervention regimes that reduce epidemic peaks while avoiding unnecessarily intensive control efforts. The results demonstrate that hospital capacity, reactive fogging, and intervention thresholds fundamentally shape dengue dynamics and provide quantitative insights for designing effective state-dependent control strategies under limited healthcare resources.


Adaptive High-Level Tight Control of Prostate Cancer: A Path from From Terminal Disease to Chronic Condition

Metastatic prostate cancer is one of the leading causes of cancer-related morbidity and mortality worldwide. It is characterized by a high mortality rate and a poor prognosis. In this work, we explore how a clinical oncologist can apply a Stackelberg game-theoretic framework to prolong metastatic prostate cancer survival, or even make it chronic in duration. We utilize a Bayesian optimization approach to identify the optimal adaptive chemotherapeutic treatment policy for a single drug (Abiraterone) to maximize the time before the patient begins to show symptoms. We show that, with precise adaptive optimization of drug delivery, it is possible to significantly prolong the cancer suppression period, potentially converting metastatic prostate cancer from a terminal disease to a chronic disease for most patients, as supported by clinical and analytical evidence. We suggest that clinicians might explore the possibility of implementing a high-level tight control (HLTC) treatment, in which the trigger signals (i.e. biomarker levels) for drug administration and cessation are both high and close together, typically yield the best outcomes, as demonstrated through both computation and theoretical analysis. This simple insight could serve as a valuable guide for improving current adaptive chemotherapy treatments in other hormone-sensitive cancers.


The Illusion-Illusion: Vision Language Models See Illusions Where There Are None

Illusions are entertaining, but they are also a useful diagnostic tool in cognitive science, philosophy, and neuroscience. A typical illusion shows a gap between how something `really is' and how something `appears to be', and this gap helps us understand the mental processing that led to how something appears to be. Illusions are also useful for investigating artificial systems, and much research has examined whether computational models of perception fall prey to the same illusions as people. Here, I invert the standard use of perceptual illusions to examine basic processing errors in current vision language models. I present these models with illusory-illusions, neighbors of common illusions that should not elicit processing errors. These include such things as perfectly reasonable ducks, crooked lines that truly are crooked, circles that seem to have different sizes because they are, in fact, of different sizes, and so on. I show that many current vision language systems mistakenly see these illusion-illusions as illusions. I suggest that such failures are part of broader failures already discussed in the literature.


Reshaping Biomolecular Structure Prediction through Strategic Conformational Exploration with HelixFold-S1

Generating large ensembles of candidate conformations is standard for improving biomolecular structure prediction. Yet aimless sampling is inefficient and costly, producing many redundant conformations with limited diversity, particularly for complex multimeric assemblies. Here, we present HelixFold-S1, a guided planning approach specifically designed to enhance the structural prediction of biomolecular complexes by strategically targeting the most informative regions of conformational space to produce accurate conformations. For each complex, predicted inter-chain contact probabilities serve as a blueprint of the conformational space, guiding computational effort toward higher-probability, low-redundancy contacts that constrain structure generation. Across diverse biomolecular complex benchmarks, HelixFold-S1 achieves markedly higher structural accuracy than traditional unguided methods while reducing sampling requirements by an order of magnitude. Predicted contact probabilities also provide a rough indicator of prediction difficulty and sampling utility. These results demonstrate that guided planning reshapes conformational exploration and enables more efficient and accurate structural inference.


An Intelligent Infrastructure as a Foundation for Modern Science

Infrastructure shapes societies and scientific discovery. Traditional scientific infrastructure, often static and fragmented, leads to issues like data silos, lack of interoperability and reproducibility, and unsustainable short-lived solutions. Our current technical inability and social reticence to connect and coordinate scientific research and engineering lead to inefficiencies and impede progress. With AI technologies changing how we interact with the world around us, there is an opportunity to transform scientific processes. Neuroscience's exponential growth of multimodal and multiscale data, together with its urgent clinical relevance, demands an adaptive infrastructure that can expose computable states, coordinate across systems, and improve through use. Using neuroscience as a stress test, this perspective argues for a paradigm shift: infrastructure must evolve into a dynamic, AI-aligned ecosystem to accelerate science. Building on several existing principles for data, collective benefit, and digital repositories, I recommend operational guidelines for implementing these principles to create this dynamic ecosystem, aiming to foster a decentralized, self-learning, and self-correcting system where humans and AI can collaborate seamlessly. Addressing the chronic underfunding of scientific infrastructure, acknowledging diverse contributions beyond publications, and coordinating global efforts are critical for this transformation. A coordinating role, even more than analysis, is where AI becomes transformative rather than merely assistive. By prioritizing an intelligent infrastructure as a central scientific instrument for knowledge generation, we can overcome current limitations, accelerate discovery, ensure reproducibility and ethical practices, and ultimately translate neuroscientific understanding into tangible societal benefits, setting a blueprint for other scientific domains.


The proportional scaling of mRNA and ribosome concentrations controls eukaryotic cell growth

Cell growth underlies nearly all eukaryotic physiology, yet its quantitative principles remain unclear. Using single-molecule ribosome tracking, spike-in RNA sequencing, and quantitative proteomics across 15 nutrient-limited conditions in budding yeast, we define how growth is controlled in the budding yeast Saccharomyces cerevisiae. Ribosome concentration scales linearly with growth rate, while peptide elongation speed remains constant at ~9 amino acids/s. While elongation is not a regulatory lever, total mRNA concentration increases proportionally with ribosomes to accelerate growth. A simple kinetic model of mRNA-ribosome binding accurately predicts the fraction of active ribosomes, growth rate, and responses to transcriptional or size perturbations. Consistent with this model, transient inhibition of mRNA degradation boosts growth by elevating mRNA concentration. These results reveal that eukaryotic cells accelerate proliferation primarily by proportionally scaling mRNA and ribosome abundance, establishing a quantitative framework for understanding eukaryotic biosynthesis.


Revealing the building blocks of tree balance: fundamental units of the Sackin and Colless Indices

(Im)balance indices can be used to quantify the (im)balance of trees by assigning numerical scores to them. An easy way to generate a new index is to construct a compound index, e.g., a linear combination of established indices. Two of the most prominent and widely used imbalance indices are the Sackin index and the Colless index. In this study, we show that these classic indices are themselves compound in nature: they can be decomposed into more elementary components that independently satisfy the defining properties of a tree (im)balance index. We further show that the difference Colless minus Sackin results in another imbalance index that is minimized (amongst others) by all Colless minimal trees. Conversely, the difference Sackin minus Colless forms a balance index. Finally, we compare the building blocks of which the Sackin and the Colless indices consist to these indices as well as to the stairs2 index, which is another index from the literature. Our results suggest that the elementary building blocks we identify are not only foundational to established indices but also valuable tools for analyzing disagreement among indices when comparing the balance of different trees. Along the way, we investigate the so-called echelon tree, which plays an important role for several (im)balance indices, and present the first non-recursive algorithm to construct it.


CORE -- A Cell-Level Coarse-to-Fine Image Registration Engine for Multi-stain Image Alignment

Accurate and efficient registration of whole slide images (WSIs) is essential for high-resolution, nuclei-level analysis in multi-stained tissue slides. We propose a novel coarse-to-fine framework CORE for accurate nuclei-level registration across diverse multimodal whole-slide image (WSI) datasets. The coarse registration stage leverages prompt-based tissue mask extraction to effectively filter out artefacts and non-tissue regions, followed by global alignment using tissue morphology and accelerated dense feature matching with a pre-trained feature extractor. From the coarsely aligned slides, nuclei centroids are detected and subjected to fine-grained rigid registration using a custom, shape-aware point-set registration model. Finally, non-rigid alignment at the cellular level is achieved by estimating a non-linear displacement field using Coherent Point Drift (CPD). Our approach benefits from automatically generated nuclei that enhance the accuracy of deformable registration and ensure precise nuclei-level correspondence across modalities. The proposed model is evaluated on three publicly available WSI registration datasets, and two private datasets. We show that CORE outperforms current state-of-the-art methods in terms of generalisability, precision, and robustness in bright-field and immunofluorescence microscopy WSIs


ATP-Independent Entropy-Driven dsRNA Unwinding by DDX3X Revealed by Coarse-Grained Simulations and Deep Learning

DEAD-box RNA helicases (DDXs) are traditionally known as ATP-dependent motors that unwind double-stranded RNA (dsRNA). Recent experiments, however, show that some DDXs promote dsRNA unwinding even in the absence of ATP, raising a fundamental question about the physical mechanism underlying ATP-independent strand separation. Here, we develop a minimal, physics-based coarse-grained RNA model and incorporate weak, specific interactions between DDX3X and dsRNA, revealing the inherently stochastic nature of unwinding events. The unwinding process must overcome an energy barrier, but thermal fluctuations and entropy gain provide a driving force for RNA remodeling. We identify that dsRNA separation proceeds through rare yet obligatory strand-displacing intermediates facilitated by DDX3X. By combining deep learning-assisted analysis, we further rank the contributions of different entropic components, revealing hydrogen bonding as the dominate factor, followed by base stacking and then the backbone conformation. These findings reveal a previously unrecognized physical mechanism for RNA duplex unwinding and offer an effective framework for studying RNA remodeling kinetics.


The embodied brain: Bridging the brain, body, and behavior with biorealistic neuromechanical models

Animal behavior reflects interactions between the nervous system, body, and environment. Therefore, biomechanics and environmental context must be considered to understand algorithms for behavioral control. Computational models that embed artificial neural controllers within body models in simulated environments are a powerful tool for this purpose. Here, we review advances in biorealistic neuromechanical models while also highlighting emerging opportunities ahead. We first show how these models enable inference of biophysical variables that are difficult to measure experimentally. Through systematic perturbations, one can generate new experimentally testable hypotheses using these models. We then examine how neuromechanical models facilitate the exchange among neuroscience, robotics, and machine learning, and showcase their applications in healthcare. We envision that coupling experimental studies with active probing of their neuromechanical surrogates will significantly accelerate progress in neuroscience.


Quantifying Avian Morphological Evolution through Deep Representation Learning

The evolution of biological morphology is fundamentally linked to ecological adaptation and species survival, yet traditional morphological evolution relies on landmark-based geometric morphometrics, a process constrained by subjective manual annotation, strict requirements for anatomical homology, and an inability to easily quantify complex, non-rigid traits such as plumage and texture. To overcome these limitations, we propose a scalable, landmark-free morphometric framework driven by deep learning. By extracting high-dimensional feature vectors from a Convolutional Neural Network (ResNet34) trained on images of over 10,000 bird species, we project raw visual semantics into a high-dimensional morphospace. Even without a priori taxonomic knowledge, this visual morphospace naturally recovers classical hierarchical taxonomy and effectively captures both homology and convergence. Analyses reveal a highly significant phylogenetic signal within the network's embeddings, with principal components correlating strongly with established ecological and morphological traits. Furthermore, by implementing a novel spherical Ancestral State Reconstruction algorithm, we uncover a pronounced "early-burst" pattern of disparity following the K-Pg mass extinction, supporting the niche-filling hypothesis of adaptive radiation.


Dispersal diversity buffers species vulnerability to local extinction

Predicting species persistence within ecological communities is a fundamental challenge for both empirical and theoretical ecology. Existing methods span from mechanistic models, whose parameters are difficult to estimate from data, to statistical tools whose context-specific parameters are less interpretable. Here, we present a general framework, grounded in the statistical physics of complex systems, that integrates the key processes governing species survival into a single measurable quantity: the competitive balance. This metric quantifies a focal species' vulnerability to competitive exclusion beyond what is captured by its abundance alone by incorporating the diversity of dispersal strategies and the structure of interspecific interactions within the community. Crucially, it can be inferred from spatial abundance data, thus circumventing the need to estimate species traits or dispersal parameters. Our results reveal that greater heterogeneity in dispersal strategies reduces vulnerability to competitive exclusion for a given abundance. Although we validate the framework using tropical and temperate forest data, it can be applied to a range of different ecosystems, providing a systemic and interpretable tool for assessing a context-dependent species vulnerability that accounts for its interactions with the entire community.


Microsecond-precision sound localization emerges from slow equilibrium dynamics

Precise sound localization relies on microsecond sensitivity to interaural time differences (ITDs), yet binaural perception exhibits sluggish tracking of dynamic acoustic cues. How such extraordinary temporal precision arises despite comparatively slow neural responses remains unresolved. This study proposes that ITD is represented as a stable equilibrium of neural population dynamics rather than through the classical place-coding framework based on delay-line coincidence detection. In this framework, excitatory and inhibitory interactions across frequency channels drive the system toward an equilibrium corresponding to the estimated ITD. The resulting dynamics achieve microsecond-level precision and reproduce key physiological observations, including frequency-dependent best-delay distributions, without requiring explicit delay lines or precisely timed inhibition. These results challenge the classical place-coding framework and suggest a fundamentally different principle for binaural computation. More generally, the findings suggest that microsecond-level sensitivity and sluggish binaural perception are complementary consequences of the same equilibrium dynamics, offering a potential resolution to a long-standing paradox in auditory neuroscience.


A portable solution for simultaneous human movement and mobile EEG acquisition: readiness potential for basketball free-throw shooting

Advances in wireless electroencephalography (EEG) technology promise to record brain-electrical activity in everyday situations. To better understand the relationship between brain activity and natural behavior, it is necessary to monitor human movement patterns. Here, we present a pocketable setup consisting of two smartphones to simultaneously capture human posture and EEG signals. We asked 26 basketball players to shoot 120 free throws each. First, we investigated whether our setup allows us to capture the readiness potential (RP) that precedes voluntary actions. Second, we investigated whether the RP differs between successful and unsuccessful free-throw attempts. The results confirmed the presence of the RP over fronto-central channels, with significant negative deflection at channel Cz, from -400 to 0 ms before movement onset ($M$ $\pm$ $SE$: -6.54 $\pm$ 2.26 to -13.52 $\pm$ 2.42 $\mu$V; $z$ = -2.53 to -3.92; FDR-corrected $p$ = 0.049 to 0.003; $r$ = 0.50 to 0.77). However, the amplitude of the RP was not related to shooting success (all FDR-corrected $p$ > 0.05; maximum mean $R^2$ = 0.047, i.e., 4.7% explained variance). Preliminary exploratory pose analysis conducted offline indicated the presence of participant-specific variations in posture between successful and unsuccessful shots in 38.5% of participants (10/26), with 4.5% explained variance (maximum mean landmark $R^2$ = 0.045). We conclude that a highly portable, low-cost and lightweight acquisition setup, consisting of two smartphones and a head-mounted wireless EEG amplifier, is sufficient to monitor complex human movement patterns and associated brain dynamics outside the laboratory.


Geometric origin of adversarial vulnerability in deep learning

Balancing training accuracy and adversarial robustness has beeen a challenge since the birth of deep learning. Here, we introduce a geometry-aware deep learning framework that leverages layer-wise local training to sculpt the internal representations of deep neural networks. This framework promotes intra-class compactness and inter-class separation in feature space, leading to manifold smoothness and adversarial robustness against white or black box attacks. The performance can be explained by \blue{data-dependent statistical mechanics of integrating out the network parameters}, \blue{supplemented by a phenomenological model} with Hebbian coupling between elements of the hidden representation. Based on the current geometry-aware learning framework, the deep network can assimilate new information into existing knowledge structures while reducing representation interference.


Computing Evolutionarily Stable Strategies in Imperfect-Information Games

We present an algorithm for computing evolutionarily stable strategies (ESSs) in symmetric perfect-recall extensive-form games of imperfect information. Our main algorithm is for two-player games, and we describe how it can be extended to multiplayer games. The algorithm is sound and computes all ESSs in nondegenerate games and a subset of them in degenerate games which contain an infinite continuum of symmetric Nash equilibria. The algorithm is anytime and can be stopped early to find one or more ESSs. We experiment on an imperfect-information cancer signaling game as well as random games to demonstrate scalability.


Where Do We Poop? City-Wide Simulation of Defecation Behavior for Wastewater-Based Epidemiology

Wastewater surveillance, which regularly measures pathogen biomarkers in wastewater samples, is a valuable tool for monitoring infectious diseases circulating in communities. Yet, most wastewater-based epidemiology methods that use wastewater surveillance results to infer disease trends implicitly assume that individuals excrete only at their residential locations and that the populations contributing to wastewater samples are static. These simplifying assumptions ignore daily mobility, social interactions, and heterogeneous toilet-use patterns, which can bias the interpretation of wastewater results, especially at upstream sampling locations such as neighborhoods, institutions, or buildings. Here, we introduce an agent-based geospatial simulation framework. Building on an established Patterns of Life model, we simulate daily human activities within a realistic urban environment and extend the framework with a physiologically motivated defecation cycle and toilet-use patterns. We couple this behavioral model with an infectious disease model to simulate transmission through spatial and social interactions. When an infected agent defecates, a pathogen-shedding model determines the amount of pathogen released in the feces. By integrating population mobility, disease transmission, toilet-use behavior, and pathogen shedding, the framework can simulate the spatiotemporal dynamics of wastewater pathogen loads. Using a case study of 10,000 simulated agents in Fulton County, Georgia, we examine how varying infection rates alter epidemic trajectories, wastewater pathogen loads, and the spatial distribution of pathogen shedding over time. Our results show that mobility and toilet use can substantially decouple residential disease prevalence from wastewater pathogen loads and demonstrate how behaviorally grounded simulations can support interpretation, scenario analysis, and wastewater surveillance


Gene genealogies in haploid populations evolving according to sweepstakes reproduction

Sweepstakes reproduction may be generated by chance matching of reproduction with favorable environmental conditions. Gene genealogies generated by sweepstakes reproduction are in the domain of attraction of multiple-merger coalescents where a random number of lineages merges at such times. We consider population genetic models of sweepstakes reproduction for haploid panmictic populations of both constant ($N$), and varying population size, and evolving in a random environment. We construct our models so that we can recover the observed number of new mutations in a given sample without requiring strong assumptions regarding the population size or the mutation rate. Our main results are {\it (i)} continuous-time coalescents that are either the Kingman coalescent or specific families of Beta- or Poisson-Dirichlet coalescents; when combining the results the parameter $\alpha$ of the Beta-coalescent ranges from 0 to 2, and the Beta-coalescents may be incomplete due to an upper bound on the number of potential offspring an arbitrary individual may produce; {\it (ii)} in large populations we measure time in units proportional to either $ N/\log N$ or $N$ generations; {\it (iii)} incorporating fluctuations in population size leads to time-changed multiple-merger coalescents where the time-change does not depend on $\alpha$; {\it (iv)} using simulations we show that in some cases approximations of functionals of a given coalescent do not match the ones of the ancestral process in the domain of attraction of the given coalescent; {\it (v)} approximations of functionals obtained by conditioning on the population ancestry (the ancestral relations of all gene copies at all times) are broadly similar (for the models considered here) to the approximations obtained without conditioning on the population ancestry.


Universal Approximation Theorems for Dynamical Systems with Infinite-Time Horizon Guarantees

Universal approximation theorems establish the expressive capacity of neural network architectures. For dynamical systems, existing results are limited to finite time horizons or systems with a globally stable equilibrium, leaving multistability and limit cycles unaddressed. We prove that Neural ODEs achieve $\varepsilon$-$\delta$ closeness -- trajectories within error $\varepsilon$ except for initial conditions of measure $< \delta$ -- over the \emph{infinite} time horizon $[0,\infty)$ for three target classes: (1) Morse-Smale systems (a structurally stable class) with hyperbolic fixed points, (2) Morse-Smale systems with hyperbolic limit cycles via exact period matching, and (3) systems with normally hyperbolic continuous attractors via discretization. We further establish a temporal generalization bound: $\varepsilon$-$\delta$ closeness implies $L^p$ error $\leq \varepsilon^p + \delta \cdot D^p$ for all $t \geq 0$, bridging topological guarantees to training metrics. These results provide the first universal approximation framework for multistable infinite-horizon dynamics.


Emergence of generic first-passage time distributions for large Markovian networks

First-passage times are often the most relevant aspect of a complex Markovian network because they signify when information processing has resulted in a definite decision. Previous studies have shown that for kinetic proofreading networks in the limit of large network size the first-passage time distribution converges either to a delta or to an exponential distribution. Remarkably, these two forms correspond to the two extreme distributions of minimal and maximal entropy for a fixed mean, respectively. Here we build on the connection between first-passage times and graph theory to show that these two limits are not model-specific, but arise generically in Markovian networks from the distribution of the eigenvalues of the generator matrix. A deterministic peak emerges when infinitely many eigenvalues contribute, while the exponential limit arises from a single dominant eigenvalue. We also show that the exponential limit emerges robustly for reversible networks when the mean first-passage time from the initial state to the target state becomes much larger than the mean first-passage time in the reverse direction. In contrast, the deterministic limit is not obtained from a simple reversal of this condition, but follows from a non-vanishing conductance or a mean-residual lifetime of the process which becomes small compared to the mean first-passage time in the long-time limit. This reveals a fundamental asymmetry between the two regimes. Our theoretical analysis is illustrated and validated by computer simulations of one-step master equations and random networks.


Induction Meets Biology: Mechanisms of Repeat Detection in Protein Language Models

Protein sequences are abundant in repeating segments, both as exact copies and as approximate segments with mutations. These repeats are important for protein structure and function, motivating decades of algorithmic work on repeat identification. Recent work has shown that protein language models (PLMs) identify repeats, by examining their behavior in masked-token prediction. To elucidate their internal mechanisms, we investigate how PLMs detect both exact and approximate repeats. We find that the mechanism for approximate repeats functionally subsumes that of exact repeats. We then characterize this mechanism, revealing two main stages: PLMs first build feature representations using both general positional attention heads and biologically specialized components, such as neurons that encode amino-acid similarity. Then, induction heads attend to aligned tokens across repeated segments, promoting the correct answer. Our results reveal how PLMs solve this biological task by combining language-based pattern matching with specialized biological knowledge, thereby establishing a basis for studying more complex evolutionary processes in PLMs.


Life as Plasmas: Autonomy and Interactivism in-materio

When is a material system a candidate for life at all? We argue that this question is prior to behavior, functional architecture, or computational capacity, and that at root it is one of physical admissibility. We develop a framework in which minimal autonomy, taken in the interactivist sense of normativity grounded in self-maintaining far-from-equilibrium organization, corresponds to a distinct non-equilibrium phase of matter, and we take complex plasmas, a physical and non-biological system, as its in-materio exemplar. We formalize a diagnostic phase-space whose criteria (sustained free-energy throughput, organizational closure, active information maintenance, and regulated noise sensitivity) constitute necessary conditions for life-attribution. We instantiate the diagnostics across contrasting systems and fix the boundaries of the phase space via Bénard convection as a driven baseline lacking closure, and a digital self-replicating soup that carries measured informational heredity while its physical closure remains a structural zero. We demonstrate that plasmas satisfy every admissibility condition for minimal physical autonomy while carrying none of the informational heredity that open-ended evolution requires, sharpening the distinction between physical admissibility and biological sufficiency, and bounding downstream questions of machine sentience.