New articles on Electrical Engineering and Systems Science


[1] 2607.15297

Large Language Model-Enhanced Multi-hop Parallel Image Semantic Communication

This paper proposes a large language model-enhanced multi-hop parallel image semantic communication (LLM-MHPSC) framework to mitigate distortion accumulation in multi-hop wireless image transmission. Unlike conventional single-hop semantic communication schemes, LLM-MHPSC deploys an extra residual compensation link at each hop to counteract accumulated distortions. To minimize additional bandwidth overhead, a coarse-to-fine residual compression scheme is designed by integrating a deep learning-based compressor with adaptive arithmetic coding (AAC). Furthermore, a large language model-based residual transmission optimizer (LLM-RTO) is developed to accurately estimate residual distributions and enable channel state and hop-aware rate adjustment, thereby improving residual compression efficiency under varying channel and hop conditions. An adaptive hop selection strategy is also proposed to activate the residual link on demand, striking a balance between transmission performance and computational cost. Experimental results show that LLM-MHPSC outperforms state-of-the-art semantic communication and traditional schemes, realizing robust image transmission with a marginal increase in bandwidth. This framework provides a flexible and effective solution for extending semantic communication to practical multi-hop application scenarios.


[2] 2607.15298

Data-driven Video Codec with Implicit Neural Representations

A conventional codec stores a video as compressed pixel data. We instead store the video, together with its audio track, as the weights of a single sinusoidal representation network (SIREN) that maps space-time coordinates to RGB values and audio amplitudes. The network uses separate audio and video initialization layers, a stack of shared fully connected hidden layers, and three output branches: one for video and two Siamese audio branches whose disagreement is used to estimate and subtract residual noise. The overfitted teacher network is then compressed by response-based knowledge distillation into a smaller student, followed by 16-bit symmetric weight quantization and lossless LZMA2 (xz) encoding. On a 6.08 MiB test video, the quantized student reaches a video PSNR of 28.72 dB with SSIM of 0.75, and an audio PSNR of 24.18 dB with a log spectral distance of 10.69 dB, while the pipeline shrinks the representation from 9.05 MiB to 2.33 MiB, an overall compression ratio of 2.61. A bit-width sweep from 1-bit to 32-bit quantization shows that reconstruction quality saturates at 16 bits. We compare against H.264, HEVC, and MP3, report where the approach falls short of them, and describe a browser-based prototype that trains, transfers, and decodes these models over WebRTC.


[3] 2607.15322

A Distributed Cluster Economic Dispatch Scheme for Cross-regional Microgrids Induced by Well-designed Communication Weights

A large-scale microgrid typically consists of several cross-regional subgrids aggregated by a virtual power plant (VPP). However, current consensus based schemes can-not guarantee the feature of differential demand between subgrids. Thus, distributed cluster consensus control induced by communication weights is investigated in this paper to solve the ED problem of a large-scale microgrid, which can achieve the expected cluster via well-designed communication weights. A communication weight matrix design method for a directed and connected graph based on eigenvector centrality is designed, which enables the adjacency matrix of the communication network to have a given leading eigenvector and allows agents in each cluster to have the same eigenvector center value. Based on this, a distributed cluster ED scheme, namely a leader-follower cluster consensus controller, is designed to drive marginal cost (MC) to achieve multiconsensus, thus allocating power among DGs. In addition, the power deficit of each subgrid collected by a VPP can be allocated to utility grids according to predetermined ratios, thus maintaining power supply-demand balance of each subgrid. For this scheme, it should be emphasized that the weighted network used is directed and connected; meanwhile, leader information only can be accessed by a few clusters. Correspondingly, relevant simulations are attached to verify the effectiveness of the designed scheme.


[4] 2607.15323

Distributed Cooperative Control of BESSs in AC and DC Hybrid Microgrid and Its Energy Internet Paradigm

An AC and DC hybrid microgrid, which inherits advantages of AC and DC microgrids and discards some disadvantages, is considered to be the most promising power network structure and gradually applied in the community. Usually, the AC subgrid and the DC subgrid are interconnected by Bidirectional Interlink Power Converters (BILPCs). Besides, in view of the different droop characteristics of AC subgrid and the DC subgrid, it is necessary to design a suitable distributed secondary controller for an AC and DC hybrid microgrid. Accordingly, this paper proposes a flexible and scalable distributed control framework for an AC and DC hybrid battery energy storage system (ADHB) with BILPCs in an Energy Internet (EI) paradigm. An ADHB governed by multi agent systems via a cloud server can reach the State-of-Charge balance, proportional power sharing, frequency and voltage restoration. The proposed control framework provides the group play-and-plug by adding or removing an inter-MASs interaction link. For a single BILPC in an ADHB, active/reactive power, frequency and voltage are adjusted by an AC BESS. For the parallel BILPCs in EI, a decentralized secondary control scheme is proposed. Communication delay issues and stability are analysed. Then, the relevant simulation results verify the correctness of the proposed scheme.


[5] 2607.15329

Co-Design of Aeroelastic Systems with Deep Reinforcement Learning

Control co-design considers the physical system and its controller together, enabling the strong coupling between system design and control to be uncovered and exploited. This is especially relevant in aeroelastic flight systems, where structural, aerodynamic, and control design choices jointly determine manoeuvrability and efficiency. This paper presents a model-free nested co-design framework for aeroelastic systems using deep reinforcement learning, in which a design-conditioned control policy is trained with proximal policy optimisation while an outer loop updates a distribution over candidate design parameters. The approach is evaluated on three case studies of increasing complexity: a spring-mass-damper system, a pitch-plunge-flap aerofoil, and a highly flexible high-aspect-ratio glider performing a thermal-soaring mission in a stochastic environment. Across these case studies, the framework is shown to progressively concentrate the design search towards high-performing regions and to outperform policies trained on randomly sampled designs. The results also show that reward shaping plays an important role in enabling stable learning in partially observed and stochastic environments. In the final glider case, the method jointly addresses wing design, flight control, and mission-level behaviour in the presence of aeroelastic coupling and atmospheric uncertainty. These results highlight the potential of model-free co-design for complex aeroelastic systems in which design, control, and mission objectives are tightly coupled.


[6] 2607.15365

Induction-heated resonant reactors for electrified thermochemistry

We present induction-heated resonant reactors, a new concept in electrified thermochemistry in which the reactor itself serves as a volumetric electromagnetic resonator heated through resonant wireless power transfer. We use the Swiss roll resonator as a model system and show that it can be designed to support uniform volumetric heating profiles and enhanced heat transfer characteristics, creating opportunities for process intensification in scaled systems. Compared to conventional (i.e., non-resonant) induction heating systems, resonant reactors can achieve exceptionally high system efficiencies through the combination of near-unity power-to-heat efficiencies and low thermal losses, both enabled by the utilization of resonant energy transfer. These concepts demonstrate how the integration of electromagnetic power transduction with thermochemical reaction engineering enables new opportunities for utilizing green electricity in sustainable chemical conversion.


[7] 2607.15430

Six-sigma Quality Management of Additive Manufacturing

In this paper, we propose to design, develop, and implement the new DMAIC methodology for Six-Sigma quality management of AM. First, we define the specific quality challenges arising from AM layer-wise fabrication and mass customization (even one-of-a-kind production). Second, we present a review of AM metrology and sensing techniques, from materials through design, process, environment, to post-build inspection. Third, we contextualize a framework for realizing the full potential of data from AM systems, and emphasize the need for analytical methods and tools. We propose and delineate the utility of new data-driven analytical methods, including deep learning, machine learning, and network science, to characterize and model the interrelationships between engineering design, machine setting, process variability and final build quality. Fourth, we present the methodologies of ontology analytics, design of experiments (DOE) and simulation analysis for AM system improvements. In closing, new process control approaches are discussed to optimize the action plans, once an anomaly is detected, with specific consideration of lead time and energy consumption.


[8] 2607.15441

Low Complexity Neural Network Digital Predistortion of Wideband Power Amplifiers through Feature Selection

Due to the continuous increase in communication bandwidth and the use of highly efficient yet nonlinear power amplifiers, Digital Predistortion (DPD) algorithms are becoming increasingly complex. In particular, neural network (NN) based DPD approaches using Phase-Normalized NN architectures often incur substantially higher computational costs than widely deployed polynomial-based methods, such as the Memory Polynomial (MP) and Generalized Memory Polynomial (GMP) models. To bridge this gap between research performance and practical implementation, we propose a low-complexity Feature Selection NN DPD architecture. The proposed method employs an offline feature-engineering pipeline based on the Least Absolute Shrinkage and Selection Operator (LASSO) and the Minimum Redundancy Maximum Relevance (MRMR) algorithm to construct a compact and informative input representation. Using measured wideband FR3 power amplifier datasets that are publicly released with this work, we demonstrate up to 30% reduction in computational complexity while maintaining comparable linearization performance.


[9] 2607.15460

Stochastic Multi-Segment Scheduling of Variable-Speed Pumped Storage Hydropower for Energy and Ancillary Services Provision

Variable-speed pumped storage hydropower (VS-PSH) offers long-duration energy storage alongside ancillary services in competitive electricity markets. However, its operation and scheduling are challenged by head-dependent nonlinearities, discrete mode transitions, and energy-continuity constraints. This study proposes a stochastic framework for VS-PSH that employs a multi-segment bidding structure to generate market-consistent energy and synchronized reserve offers in compliance with market rules. The framework explicitly incorporates physical constraints, including head-dependent capability limits, discrete pumping and generating modes, as well as state-of-charge (SoC) and head dynamics, within a stochastic mixed-integer linear programming (MILP) formulation. Price uncertainty is represented through a scenario-based modeling approach that scales base-case prices and allows variations in (dis)charging incentives. The stochastic MILP produces optimal energy and mode schedules that maintain feasible SoC trajectories across scenarios and ensure physically feasible operating strategies. Case studies under different levels of price variability demonstrate the operational feasibility and market applicability of the proposed framework, showing effective coordination between energy arbitrage and reserve provision under uncertainty. These results highlight the operational and economic value of VS-PSH as a grid-scale energy storage resource.


[10] 2607.15541

StarCodex: Dynamic Coding Harness for Starlink Measurement Analysis and Experiment Automation

Starlink and other low Earth orbit (LEO) satellite broadband systems are producing increasingly diverse measurement data across regions, time periods, and access conditions. These measurements are valuable for throughput prediction, adaptive bitrate (ABR) evaluation, and network experimentation, but converting continuously arriving data into reusable experimental evidence still relies heavily on manually developed analysis code and expert-guided data inspection and failure-case organization. This paper proposes StarCodex, a dynamic coding harness for Starlink measurement analysis and experiment automation. StarCodex detects analysis gaps from the current measurement state, converts them into structured coding tasks, uses Codex to generate or repair executable analysis artifacts, and accepts artifacts through code, data-interface, measurement-semantics, and output validation. Experiments on real Starlink measurements show that StarCodex discovers 49 of 56 uncovered system-risk cases, attains higher average precision than the strongest predefined analysis baseline, and constructs a benchmark with denser and broader system-risk evidence. The generated prediction and replay artifacts further reveal prediction risks and quality-of-experience (QoE)--risk differences among ABR controllers. These results demonstrate the feasibility of using a dynamic coding harness to convert evolving Starlink measurements into validated analysis artifacts for automated experiment workflows.


[11] 2607.15547

Pinching Antenna-Assisted ISAC with Waveguide Mode Selection

Conventional pinching antenna (PA)-assisted integrated sensing and communication (ISAC) architectures typically assume static receiver locations or predetermined receive waveguides, thereby underutilizing the inherent spatial degrees of freedom. This paper proposes a novel mode-selectable PA-assisted ISAC framework to maximize the post-combining sensing signal-to-noise ratio while satisfying multi-user quality-of-service constraints by jointly optimizing the waveguide mode selection, transmit beamforming, and transmit/receive PA positions. To tackle the resulting mixed-integer nonconvex optimization problem, we develop a low-complexity block-coordinate descent algorithm that leverages a penalty-based majorization-minimization method to achieve high-quality suboptimal solutions. Numerical results demonstrate that the proposed design significantly outperforms both traditional PA and fixed-antenna benchmarks by synergistically harnessing spatial adaptability and modal reconfigurability. In particular, the mode-selectable design enables the coordinated optimization of transmit/receive operations and sensing-communication resource allocation, thereby maintaining sensing robustness under stringent communication requirements.


[12] 2607.15570

Price-Based Distributed Scheduling of Flexible Demands in Energy Communities

We study price-based distributed scheduling of flexible demand in an energy community, where a coordinator broadcasts electricity prices and individual households schedule their consumption. Household demand includes deferrable and non-deferrable loads, such as electric vehicle charging with completion deadlines and thermostatically controlled loads. The coordinator transacts with a distribution utility on behalf of community members under the regulated Net Energy Metering tariff. We formulate distributed demand scheduling as a bilevel stochastic dynamic program. The upper level optimizes the coordinator's pricing policy to minimize the community's energy costs subject to operating, revenue adequacy, and individual rationality constraints. The lower level involves stochastic dynamic programs that maximize households' consumption benefits subject to the availability of renewable generation. The computational cost of such a distributed stochastic dynamic program is prohibitive in general. By uncovering the structure of optimal centralized scheduling, we derive Threshold Pricing Rule (TPR) -- a simple community pricing policy with linear computational costs for the upper- and lower-level optimizations. Being independent of parameters of the underlying stochastic dynamic program, TPR is robust against modeling uncertainties and is shown to guarantee revenue adequacy for the community and individual rationality for community members. As the community size grows, TPR is shown to be asymptotically optimal.


[13] 2607.15572

A PI+R Control Scheme Based on Multi-agent Systems for Economic Dispatch in Isolated BESSs

Battery energy storage systems (BESSs) are widely used in smart grids. However, power consumed by inner impedance and the capacity degradation of each battery unit become particularly severe, which has resulted in an increase in operating costs. The general economic dispatch (ED) algorithm based on marginal cost (MC) consensus is usually a proportional (P) controller, which encounters the defects of slow convergence speed and low control accuracy. In order to solve the distributed ED problem of the isolated BESS network with excellent dynamic and steady-state performance, we attempt to design a proportional integral (PI) controller with a reset mechanism (PI+R) to asymptotically promote MC consensus and total power mismatch towards 0 in this paper. To be frank, the integral term in the PI controller is reset to 0 at an appropriate time when the proportional term undergoes a zero crossing, which accelerates convergence, improves control accuracy, and avoids overshoot. The eigenvalues of the system under a PI+R controller is well analyzed, ensuring the regularity of the system and enabling the reset mechanism. To ensure supply and demand balance within the isolated BESSs, a centralized reset mechanism is introduced, so that the controller is distributed in a flow set and centralized in a jump set. To cope with Zeno behavior and input delay, a dwell time that the system resides in a flow set is given. Based on this, the system with input delays can be reduced to a time-delay free system. Considering the capacity limitation of the battery, a modified MC scheme with PI+R controller is designed. The correctness of the designed scheme is verified through relevant simulations.


[14] 2607.15575

DFT-p-FDMA Based Chirp Transmission in CP-OFDM for Unified ISAC Waveform Design

We propose an integrated sensing and communications (ISAC) framework that supports chirp signal transmission in CP-OFDM-based multiple access communication systems, enabling efficient coexistence of communication and sensing capabilities. Our framework employs the discrete Fourier transform phase rotated and permuted frequency division multiple access (DFT-p-FDMA) waveform to transmit chirp signals using a portion of the frequency resources, while ensuring interference-free concurrent CP-OFDM data transmissions on other bands. We analyze the effective channel behavior under the DFT-p-FDMA waveform, characterizing how delays and Doppler shifts impact radar target echoes. We also show how processing multiple received symbols improves Doppler resolution in practical scenarios. Our framework allows flexible adjustment of range-Doppler resolution through optimized time-frequency resource allocation, offering a versatile solution for ISAC applications. Simulation results validate the framework's performance in delay and Doppler estimation, highlighting its potential to support ISAC in next-generation wireless networks.


[15] 2607.15583

Distributed Continuous Aerial Surveillance by UAS Swarms Under Formal Mission Specifications

Persistent aerial surveillance using multi-unmanned aerial systems (UASs) requires decentralized coordination, continuous team reconfiguration, and provable mission correctness despite limited onboard energy and communication constraints. This paper develops a distributed framework for continuous aerial surveillance under bounded Linear Temporal Logic (LTL) mission specifications. The proposed approach partitions the UAS team into stationary anchors and mobile workers operating under cyclic replacement modes, and constructs a deep neural network (DNN)-inspired communication topology that enables fully decentralized coordination through local interactions. A hierarchical bounded LTL specification formally captures mode-to-mode reference consistency, cyclic team rotation, finite-time reachability, trajectory tracking, and prescribed surveillance coverage. By proving the finite-time convergence of the worker-agent coordination dynamics, the paper guarantees the finite-time satisfaction of the mission specification. To maximize sensing effectiveness, an information-theoretic optimization framework synthesizes the reference configuration of newly deployed worker agents by minimizing the Kullback--Leibler divergence between the surveillance-node distribution and the induced coverage density. The resulting reference configuration uniquely determines a deterministic, mode-dependent communication topology, eliminating online communication-graph optimization while preserving the formal mission guarantees. Finally, a decentralized quadrotor controller realizes the distributed references using only local communication. Numerical simulations demonstrate cyclic team reconfiguration, decentralized communication-topology synthesis, finite-time formation convergence, and certified persistent surveillance coverage.


[16] 2607.15594

An Open-Source, Autonomous Platform for High-Resolution Energy Monitoring in Manufacturing

High-resolution energy data is increasingly central to Industry 4.0, where electrical signals such as three-phase voltage and current carry rich information about machine condition, tool wear, and process dynamics. Capturing this information in practice remains difficult: commercial power analysis are largely proprietary, offer limited or no access to high-sampling rate data for transient analysis, restrict access to raw waveform data, and offer no customization, while general-purpose open hardware lacks the front-end accuracy, isolation, and robustness required for industrial measurement. This paper presents Autonomous Energy Monitoring System (AEMS), an open-source, low-cost, and modular platform supported by a host, edge-gateway, and optional cloud software stack that enables autonomous, long-duration acquisition independent of a continuously connected host and thereby closes this gap by combining research-grade fidelity with industrial deployability. The system acquires three-phase voltage and current through an isolated front-end and a 24-bit, simultaneously sampling analog-to-digital converter, managed by a dual-core architecture that separates deterministic acquisition and on-board logging from host communication and control. Industrial interfaces (Ethernet, RS-485/Modbus, and BLE) together with hardware-level synchronization enable scalable, time-aligned acquisition across multiple machines, supported by a complete host, edge-gateway, and optional cloud software stack. We validate the platform on a three-axis CNC machining center, where it resolves spindle, feed-drive, rapid-traverse, and material-removal energy states and detects feed-rate changes as small as 50 mm/min. By releasing the full hardware and firmware openly, this work aims to democratize access to high-fidelity energy monitoring for both researchers and small and medium-sized manufacturers.


[17] 2607.15602

Optimal Sampling and Reconstruction of Graph Signals in the Fractional Fourier Domain

Graph signal sampling and reconstruction are commonly formulated in the graph Fourier transform (GFT) domain. However, the reconstruction performance may be limited when practical graph signals are not sufficiently concentrated in the GFT spectrum. To address this issue, this paper proposes a graph signal sampling and reconstruction framework based on the graph fractional Fourier transform (GFRFT) domain. The fractional order is introduced as an adjustable spectral domain parameter, and the optimal order is selected to provide a more suitable representation domain for a given graph signal and sampling model. Under a unified sampling reconstruction formulation, subspace, smoothness, and stochastic priors are incorporated, and both unconstrained and predefined reconstruction mechanisms are considered, leading to several fractional domain sampling and reconstruction methods. Furthermore, the theoretical analysis shows that the optimal GFRFT domain can provide a more suitable low-dimensional spectral representation by improving energy concentration and reducing projection residual. The effects of residual leakage and noise amplification are further considered to explain how this representation advantage is translated into reconstruction error reduction. Experimental results show that, GFRFT domain sampling and reconstruction generally achieve better recovery performance than GFT domain methods.


[18] 2607.15628

MSTF-Net: A UAV-Oriented Multi-Spectral Video Segmentation Method via Modality-Robust, Scale-Adaptive, and Consistent Fusion

Multi-spectral video segmentation is essential for robust scene understanding in unmanned aerial vehicle (UAV) applications such as city planning, land use monitoring, traffic monitoring, and crowd estimation. While the fusion of RGB and thermal modalities offers complementary information for perception under varying lighting and visibility conditions, two fundamental challenges remain: (1) the modal fusion dilemma, arising from significant discrepancies between RGB and thermal features that obscure complementary cues, and (2) temporal variation, induced by rapid motion and viewpoint changes on UAV platforms, which leads to appearance inconsistency and misalignment across frames. To address these issues, this study proposed MSTF-Net, a modality-robust scale-adaptive fusion framework for multi-spectral video segmentation that effectively models cross-modal fusion and temporal consistency. The Modality Spatial Complementary Suppression and Enhancement (MSCSE) module generates unified instance queries via cross-modal attention and suppresses modality-specific noise using residual-guided discrepancy filtering and consistency constraints. To model temporal dynamics, the Multi-scale Temporal Cross-modality Semantic Consistency (MTCSC) module adaptively adjusts the temporal receptive field based on frame distance, capturing both coarse global context and fine local structure across time. Extensive ablation experiments on public RGB-T datasets demonstrate that MSTF-Net achieves state-of-the-art segmentation performance, especially under challenging conditions such as small targets, occlusion, and modality degradation. Specifically, reached 56.42\% mIoU on the MVSeg dataset and 51.80\% mIoU on the CART dataset, respectively.


[19] 2607.15651

FSIM: Fluid Element Stacked Intelligent Metasurface for Multiuser Downlink Networks

A fluid element (FE)-aided stacked intelligent metasurface (FSIM) for multiple-input single-output (MISO) communication system is investigated, where a multi-antenna base station (BS) serves multiple single-antenna users through FSIM. Unlike conventional SIM with fixed meta-atom deployment, the architecture allows the meta-atoms in each layer to move within a predefined fluidic region to further increase the spatial diversity. By jointly optimizing the two-dimensional meta-atom positions, BS transmit beamforming, and FSIM phase-shifts, the cascaded BS-FSIM-user channels can be flexibly reconfigured to enhance the desired signals and suppress multiuser interference. The proposed sum-rate maximization problem is highly non-convex and nonlinear due to the coupled solutions of element position, beamforming, and phase-shift. To address this challenge, an alternating optimization (AO) algorithm is developed to iteratively update these variables. The beamforming and FSIM phase-shift subproblems are transformed into semi-definite programming problems and solved by using successive convex approximation (SCA), first-order Taylor approximation, and penalty-based rank-one relaxation, whilst the FE position subproblem is handled through a projected gradient-based update. Simulation results reveal that a compact FSIM with a small inter-layer thickness is preferable, as increasing the thickness weakens inter-layer coupling and degrades the achievable sum rate. Results also demonstrate that the proposed FSIM significantly outperforms conventional SIMs with fixed positions, patch-based structures, partial fluidity, restricted fluid regions, and existing flexible intelligent metasurfaces. Furthermore, the proposed AO-based algorithm achieves superior rate performance compared to sub-schemes, metaheuristic methods, and conventional beamforming benchmarks.


[20] 2607.15653

Energy-Efficient Resource Allocation for Six-Dimensional Movable Antenna Systems

This paper investigates the energy-efficiency (EE) maximization problem for a multiuser wireless network equipped with six-dimensional movable antennas (6DMAs), where the three-dimensional (3D) positions and orientations of the antennas are jointly optimized to fully exploit the additional spatial degrees of freedom offered by dynamic channel reconfiguration. However, the practical operation of 6DMAs incurs non-negligible mechanical energy consumption. Moreover, orientation-dependent phase variations, together with the strong coupling among antenna positions, rotation angles, transmit beamforming, and time allocation, render the resulting problem highly non-convex and analytically challenging. To address this issue, we develop a block coordinate descent (BCD) optimization framework that integrates Dinkelbach's transformation with the majorization-minimization (MM) approach to efficiently obtain a high quality suboptimal solution with guaranteed convergence. Simulation results unveil that the proposed design achieves significant EE improvements over conventional benchmarks, thereby highlighting the critical importance of accounting for practical mechanical energy costs in 6DMA enabled systems. Furthermore, our results reveal a fundamental trade-off between throughput enhancement and mechanical overhead: although larger antenna reconfigurations can improve channel conditions, their EE gains gradually diminish due to the increased mechanical energy consumption.


[21] 2607.15654

Energy Efficient Active Stacked Intelligent Metasurfaces

This paper investigates an energy-efficient active stacked intelligent metasurfaces (ASIM)-assisted downlink transmission framework, where a multi-antenna base station (BS) serves multiple users through a multi-layer metasurface architecture. Unlike conventional passive intelligent surfaces, the considered ASIM employs active amplification and multiple transmissive layers to enhance electromagnetic wave manipulation. We aim to maximize the system energy efficiency (EE) by jointly optimizing the BS beamforming and ASIM configurations under user quality-of-service and amplification constraints. The resulting problem is highly coupled and non-convex due to the cascaded near-field channel and multi-layer metasurface structure. To address this challenge, we first transform the original problem through epigraph, Lagrangian dual, and quadratic transformations. An alternative optimization framework is then developed, where the BS beamforming subproblem is solved via successive convex approximation (SCA), while the ASIM configuration is optimized using Bayesian optimization based on a Gaussian-process surrogate model. Numerical results demonstrate that the proposed scheme significantly improves the achievable EE compared to conventional passive SIM and heuristic benchmark methods. Furthermore, the impacts of amplification capability, number of metasurface layers, and inter-layer spacing on system performance are investigated, providing useful design insights for future active metasurface-assisted wireless networks.


[22] 2607.15663

Adaptive Model-Based Transfer Learning for Dynamic HVAC Control

In this paper, we aim to automate the adjustment of air handling unit (AHU) setpoints within heating, ventilation, and air conditioning (HVAC) systems to maintain indoor temperatures at user-specified levels. A key challenge lies in obtaining sufficient high-quality sensor data from real buildings. To address this, we explore transfer learning and leverage simulation software to generate training data. We propose an adaptive model-based transfer learning approach for dynamic HVAC control, where the agent directly controls the source domain under conditions identical to the target domain. This eliminates the need for extensive target-specific knowledge to define data generation schedules and reduces the risk of collecting irrelevant samples, while also providing greater flexibility during learning. At the control level, we enhance performance through physics rule embedding, which ensures physical consistency, and long-term-aware setpoint selection strategy, which mitigates abrupt setpoint changes. Finally, to accelerate and stabilize deployment in new buildings, we enable knowledge transfer directly between similar real-world buildings, reducing the need to construct virtual source domains repeatedly.


[23] 2607.15680

Global Survey of Technologies and Industrial Applications of Grid Forming Energy Storage Systems

Grid-forming (GFM) energy storage system (ESS) is a key enabler for stabilizing future power systems with high penetration of converter-based resources (CBRs). To get a better overview of the state-of-the-art and challenges for implementing and deploying GFM-ESS, a global survey has been initiated by Cigre Working Group B4.101 - industrial implementation and application of grid forming energy storage systems. Feedback was collected from universities, transmission system operators (TSOs), power plant developers, original equipment manufacturers (OEMs), research institutes, as well as consultants. It is interesting to note that while many common understandings have been established in practice, certain gaps persist among different stakeholders. This article intends to bridge this gap by presenting a summary of the survey, including the questionnaire, responses from various stakeholders, and in-depth analysis of the survey results. The key challenges faced by different stakeholders in deploying GFM-ESS are identified, shedding light on future research in this direction.


[24] 2607.15690

Obstacle-Aware Four-Dimensional Trajectory Design for Urban Air Mobility

Urban Air Mobility (UAM) with electric Vertical TakeOff and Landing (eVTOL) vehicles can help address ground traffic congestion. The design of an eVTOL trajectory that is safe and reduces travel time is key for UAM adoption. Existing works on trajectory design either may not adequately incorporate dense obstacles in urban environments, complex eVTOL flight dynamics, or one or more flight phases. Not considering these factors can result in low-quality, or worse infeasible, trajectories. We develop a hybrid framework that can integrate building obstacles data, wind data, eVTOL flight dynamics, and other real-world operational constraints to estimate a four-dimensional eVTOL flight trajectory in ascent, cruise, and descent that aims to minimize travel time. Our framework first fills the obstacle-free regions with intersecting convex polygons, then identifies potentially low-travel time candidate sequences of these polygons using a Graph of Convex Sets-based path planner, and then uses an Optimal Control Program to give the final trajectory that passes through the polygons in a sequence identified before. We evaluate our framework on routes within New York City. Our framework can design trajectories respecting the above constraints in the presence of as many as 250 building obstacles. We show that not including the above constraints can underestimate the flight time by as much as 20\%.


[25] 2607.15694

A Geometry-Limited Identification Floor and Its Consequences for Voice-Clone Attribution in Professional Voice Actors

A voice actor's voice is their asset, and AI cloning directly threatens it. The natural defense flags the enrolled actor whose embedding similarity to a suspect recording crosses a threshold. We show it fails where it is most needed: trained voices crowd the embedding space, and each actor performs many styles. On 1,168 Japanese voice actors (56,568 segments, ~63 h), a misidentification floor survives calibration, score normalization, and discriminative re-ranking (linear and nonlinear, including PLDA): the residual is a limit of the embedding geometry, not of the back-ends we evaluate. The best ensemble still leaves ~2.6% closed-set misidentification, several-fold above matched controls; session-disjoint, re-ranking lowers the floor only to 13.0%. The same crowding drives false attribution: on a generic English encoder, roughly half the clones of non-enrolled people falsely accuse an enrolled actor, while -- by a separate real-vs-synthetic shift -- 32% of Seed-VC clones of enrolled targets are missed at the same threshold; one operating point couples the two, and none escapes both. A domain-matched, voice-actor-trained encoder mitigates substantially (a four-fold gender gap vanishes; wrongful misattribution falls to 1.5-10%), but does not remove the floor. Controls (codec, channel, vocoder, content) support reading the miss rate as a real-versus-synthetic covariate shift, not missing speaker information. Fixed-threshold clone attribution is thus unreliable here, and on a generic encoder unfair. Robust attribution must extend spoofing-aware speaker verification to open-set 1:N (anti-spoofing gate, domain-matched encoder, per-speaker calibration, abstain option), and even then supports detection, not autonomous enforcement.


[26] 2607.15710

Dual-Security for Indoor OFDM-ISAC Systems via Temporal Artificial Noise

With the rapid development of integrated sensing and communication (ISAC) as a key enabler for future wireless networks, ensuring the security of both communication and sensing functions has become increasingly important. Current secure ISAC studies focus restrictively either on the communication or the sensing security, but not both. To bridge this gap, this paper investigates security for both, i.e., dual-security, in indoor orthogonal frequency division multiplexing (OFDM) based ISAC systems. Specifically, we consider a scenario in which a sensing user (SU) is authorised for sensing but may eavesdrop on communication data, while a communication user (CU) is authorised for communication but may perform unauthorised sensing. We chose this scenario as the pathological case where an authorised eavesdropper has more information and is more effective than an unauthorised one. To address this case, we propose the use of temporal artificial noise (AN) to prevent malicious CU sensing by enlarging its time-domain sensing error, and simultaneously degrade SU data eavesdropping by reducing its frequency-domain signal-to-noise-plus-interference ratio (SINR) with standard OFDM receiver processing. Meanwhile, our proposed scheme guarantees the sensing performance of the SU and the communication performance of the CU. We present numerical results that demonstrate AN can effectively provide dual protection for sensing and communication in OFDM-ISAC systems while guaranteeing the performance of legitimate users.


[27] 2607.15713

Map as a Prompt: Learning Multi-Modal Spatial-Signal Foundation Models for Cross-scenario Wireless Localization

Accurate and robust wireless localization is a critical enabler for emerging 5G/6G applications, including autonomous driving, extended reality, and smart manufacturing. Despite its importance, achieving precise localization across diverse environments remains challenging due to the complex nature of wireless signals and their sensitivity to environmental changes. Existing data-driven approaches often suffer from limited generalization capability, requiring extensive labeled data and struggling to adapt to new scenarios. To address these limitations, we propose SigMap, a multimodal foundation model that introduces two key innovations: (1) A cycle-adaptive masking strategy that dynamically adjusts masking patterns based on channel periodicity characteristics to learn robust wireless representations; (2) A novel "map-as-prompt" framework that integrates 3D geographic information through lightweight soft prompts for effective cross-scenario adaptation. Extensive experiments demonstrate that our model achieves state-of-the-art performance across multiple localization tasks while exhibiting strong zero-shot generalization in unseen environments, significantly outperforming both supervised and self-supervised baselines by considerable margins.


[28] 2607.15716

Energy-Efficient Target-Aware Hybrid Beamforming for THz Near-Field ISAC with Sparse Connectivity

Integrated sensing and communication (ISAC) at terahertz (THz) frequencies enables ultra-high-resolution perception while facing a key limitation: highly directional THz beams cannot illuminate extended targets within a single beam. Conventional solutions rely on sequential beam scanning, reducing sensing accuracy and increasing energy consumption. Moreover, in conventional sparse-array, grating lobes are generally treated as undesirable artifacts that should be suppressed to avoid ambiguity and interference. In contrast, this paper adopts a reverse design philosophy by intentionally engineering sparsity-induced grating lobes as controllable auxiliary illumination beams for extended-target sensing. This paper exploits grating lobes and proposes a sparse-connected hybrid beamforming architecture that intentionally engineers and exploits grating lobes to enable single-shot, full-aperture illumination of extended targets while supporting multi-user downlink communication. A switch-controlled sparse RF network preserves the array aperture and generates a dominant main lobe with structured secondary lobes covering the entire target extent. A covariance-driven alternating-minimization framework jointly optimizes digital precoders, quantized phase shifters, and antenna-RF switching. Simulations at 140 GHz demonstrate near fully-digital Cramer-Rao sensing accuracy, competitive communication performance in low-rank THz channels, rapid convergence, and significant hardware and energy savings, establishing structured sparse connectivity as a scalable and energy-efficient solution for extended-target THz ISAC.


[29] 2607.15809

From Similarity to Feasibility: Diffusion-Refined Retrieval-Augmented Generation for Distribution Network Optimization

Rapidly shifting operational scenarios driven by uncertain Distributed Energy Resource (DER) profiles render conventional distribution network optimization methods either computationally expensive or poorly generalizable. This paper introduces GridRAG, a pioneering retrieval-augmented framework that transforms optimization into a ``retrieve-and-refine'' paradigm. GridRAG first embeds scenario features and optimal solutions into a joint representation space to ensure semantic consistency. Based on the hybrid semantic information, the similar historical scenarios are then retrieved from a pre-constructed database. Then an SDEdit-style diffusion module is integrated to refine retrieved solutions by modeling the conditional distribution over near-feasible manifolds. This process effectively pulls retrieved solutions into near-optimal attraction basins, providing a high-quality warm-start for the final solver. Validated on three optimization tasks across four standard topologies, GridRAG demonstrates superior cross-scenario generalization and a multi-fold speedup in solution time compared to existing learning-based and model-based baselines. Our code is available at this https URL.


[30] 2607.15840

Converging Safety and Security: IO-Link Wireless and OPC UA over 5G under prEN 50742

The integration of wireless communication technologies in industrial automation offers greater flexibility, but also exposes safety systems to a broader threat vector. Emerging regulations, such as the draft standard prEN 50742, mandate the convergence of functional safety and cybersecurity by requiring cryptographic security mechanisms directly in safety-critical communication. This paper presents an empirical evaluation of this safety-security convergence across a complete control chain, spanning from an IO-Link Wireless Safety device to a PLC via an OPC UA backbone. We measure the latencies and jitter of different Safety-Related Security Levels under prEN 50742 over Ethernet, Wi-Fi 6, and private 5G. Our results reveal that while cryptographic execution time is negligible, the resulting frame payload expansion severely restricts wireless fieldbus capacity, reducing the maximum number of devices per IO-Link Wireless track from 8 to 2. Furthermore, we demonstrate that, despite higher average latency, a private 5G provides sufficiently deterministic latency characteristics to preserve functional safety watchdog margins, unlike unlicensed Wi-Fi 6.


[31] 2607.15867

Scalable Supervisory HVAC Control for Linear Objectives

Advanced control of heating and cooling systems can substantially reduce energy costs and pollution. However, real-world adoption of popular algorithms among researchers, such as model predictive control (MPC) and reinforcement learning (RL), remains limited due in part to their high deployment and commissioning costs. Here, we develop two nearly commissioning-free controllers tailored to objectives that depend linearly on the controlled thermal load, such as energy costs and pollution. The controllers require at most two thermal parameters. In representative heating simulations, controller performance is robust to large parameter specification errors, suggesting potential for deployment with no tuning. The controllers maintain good occupant comfort while achieving 43 to 98% (depending on the electricity pricing and controller variant) of the performance improvement achieved by an omniscient policy with perfect model information and forecasts. These results suggest that simple, structure-exploiting controllers may capture most of the attainable value of advanced control while avoiding the data, modeling, tuning, and computational burdens that can arise with conventional MPC or RL.


[32] 2607.15900

Day-Ahead Forecasting of Largest Single Infeed/Outfeed on the Irish Power Grid: A Generative Artificial Intelligence Approach

This paper presents a generative artificial intelligence (Gen AI) approach for forecasting, at a day-ahead stage, the largest single infeed (LSI) and largest single outfeed (LSO) on the Irish power system to assist in reserve dimensioning. Developed collaboratively between EirGrid, the electric transmission system operator (TSO) for Ireland, and this http URL using the this http URL platform, the system delivers accurate forecasts up to 38 hours ahead of real-time using limited data available before the day-ahead and intra-day energy market gate closure timings. Initial performance demonstrates an accuracy with a mean absolute percentage error (MAPE) that is only 1.1\% higher than the results possible using full market data (8-hours ahead). Thus, if this approach is integrated into operational systems and such high levels of accuracy are maintained, reserve procurement costs could be significantly reduced. The results also demonstrate the practicality and extensibility of AI-powered resource planning for TSOs.


[33] 2607.15949

Gaussian behaviors and stochastic data-driven control

We propose a stochastic behavioral modeling framework, termed Gaussian behaviors, which augments a deterministic linear time-invariant (LTI) behavior with a Gaussian noise component. We show that this notion is a tractable subclass of stochastic behaviors and encompasses classical parametric stochastic LTI state-space system models as special cases. Analogously to deterministic LTI behaviors, the framework enables simple and tractable stochastic data-driven control methods. To this end, we obtain a method for prediction by conditioning the Gaussian behavior on the known part of the trajectory, which is identified directly from the sample covariance of trajectory data. Building on this method, we develop predictive control formulations that optimize over feedforward or disturbance affine feedback policies. The resulting formulations are shown to be convex. We further derive a finite-sample confidence bound on the prediction accounting for both aleatoric and epistemic uncertainty, and incorporate it into a robust control method, for which a tractable convex upper bound is obtained. Within this framework, subspace predictive control is recovered when only the mean prediction is used, while data-enabled predictive control is shown to account for the prediction uncertainty in an optimistic fashion. Numerical case studies illustrate the benefits of the proposed methods.


[34] 2607.15959

Multibit Quantized Precoding for MU-mMIMO

We propose a novel multibit quantized precoding method for the downlink of multi-user massive MIMO systems with low-resolution digital-to-analog converters. The new method, termed multibit quantized precoding (MQP), enforces the finite-alphabet constraint through an l0-norm penalty, approximated by a smooth surrogate so as to yield a reformulated problem, which is then convexized via fractional programming, ultimately extending quantized precoding beyond 1-bit alphabets. The regularization parameter of the proposed method is selected via a discrepancy principle integrated with graduated non-convexity continuation, resulting in a principled and reproducible hyperparameter tuning method and an efficient iterative algorithm with a closed-form, least-squares-type update per iteration. In order to further reduce the computational complexity of the method, we include a Gaussian belief propagation (GaBP) step for turning the least-squares update in linear-time. Simulations performed for systems with different sizes demonstrate that both methods, namely the MQP with and without GaBP, achieve competitive or superior error-rate performance compared to state-of-the-art quantized precoding algorithms under various channel conditions.


[35] 2607.15961

Dynamic Constraint Reconstruction Based Control Barrier Functions for Safety-Critical Control of High-Dimensional Manipulators

Control barrier functions (CBFs) provide formal safety guarantees for constrained nonlinear systems, but their effectiveness relies on accurate system dynamics. In high-dimensional manipulators subject to unknown disturbances and model uncertainties, fixed safety constraints constructed from nominal dynamics may become inconsistent with the actual system behavior, leading to safety degradation or excessive conservatism. This paper proposes a dynamic constraint reconstruction based control barrier function (DCR-CBF) framework for safety-critical control of disturbed robotic manipulators. An extended state observer is employed to estimate lumped disturbances online, and the estimated disturbance is incorporated into high-order control barrier functions to reconstruct safety constraints according to the estimated true dynamics. To address estimation inaccuracies, a safety margin is introduced, and a sufficient condition is derived to guarantee forward invariance under bounded estimation errors. Simulation studies on a 4-DOF excavation manipulator demonstrate that the proposed DCR-CBF method achieves zero safety violation under strong unknown disturbances while significantly improving trajectory-tracking performance compared with standard and robust CBF methods.


[36] 2607.15969

Vessel Trajectory Prediction using COLREGs-aware Optimal Planning

This paper presents a trajectory prediction method for marine vessels based on optimal planning. Crude initial trajectories respecting static obstacles are first generated using A*-search to provide a feasible warm start. In the second step, a numerical optimizer is used to ensure COLREG compliance. The prediction problem is posed as sequential trajectory planning from the perspective of each surrounding vessel, requiring only their current positions, velocities, and intended destinations as input. As the latter is included in AIS messages, this enables faster predictions than learning-based methods that typically require longer data histories. The proposed method is validated using real-world scenarios constructed from AIS data.


[37] 2607.16000

Inertial Human Motion Capture: From Biomechanics to Recent Sensor Fusion Methods and Back

Inertial measurement units (IMUs) are a promising means to capture human motion, yet obtaining meaningful biomechanical quantities from IMU measurements remains non-trivial. This tutorial-style review focuses on kinematics and introduces four key aspects (inertial human motion capture objective, environmental conditions, subject & attributes, and motion characteristics) to determine how to translate biomechanical problems into adequate formulations for the fusion of inertial sensor measurements. We identify three fundamental challenges for kinematics estimation from IMUs: IMUs do not provide direct information about the joint angle, IMUs do not measure their own orientation, and real-world environments and dynamics compromise sensor reliability. Though there exist widely-used methods to overcome these challenges, they suffer from severe limitations in real-life applications, e.g., the need for sensor-to-segment calibration, and the fact that magnetic field disturbances degrade joint angle accuracy. The full potential for many use-cases hence remains untapped in terms of accuracy and reliability. We share insights into recently proposed methods, e.g. exploiting the human body's kinematic chain constraints, having the potential to overcome these limitations. We also present guiding questions related to the four key aspects and illustrate their use for navigating the methodological landscape for the use-case of lower-extremity joint angle estimation, for which we share open-access code and compare the traditional workflow with three alternatives. Our aim is to bridge the gap between the sensor fusion community developing methods for human motion capture and the biomechanics community in need of accurate, easy-to-use, and reliable methods to study human motion outside of the laboratory.


[38] 2607.16004

Robustness of Reinforcement Learning-Based Congestion Management in Low-Voltage Grids

Increases in photovoltaic generation, charging of electric vehicles and heat-pump demand challenge operating limits in low-voltage distribution grids. This requires curative curtailment methods that can operate under sparse observability, noisy measurements, and imperfect grid models. Unlike prior end-to-end reinforcement-learning approaches for partially observable curtailment, this work decouples congestion detection and control by combining a random-forest violation pre-classifier with an actor-critic controller, and evaluates its robustness to measurement noise and grid-parameter mismatch. The framework is tested on a real low-voltage grid using synthetic future operating scenarios with low observability and controllability. With accurate grid parameters, the controller reduces total violation magnitude by 98.9%, and this performance remains nearly unchanged under the tested measurement-noise settings. Grid-model mismatch proves to be more challenging, but the controller still mitigates most violations under the tested mismatch assumptions.


[39] 2607.16013

Robust Monitoring of Arc Welding Processes: A Generalizable Framework with DVAE and Particle Filter

Arc welding processes are essential for continuous fabrication but prone to disturbances that impair weld quality, making real-time monitoring critical yet difficult due to complex visual patterns and nonlinear, time-varying dynamics. Deep learning shows promise but faces scalability limits because of its dependence on large labeled datasets and application-specific tuning. We explore whether a unified approach can characterize major arc welding processes across applications and improve scalability through consistent state monitoring. This paper introduces a robust and generalizable monitoring framework for arc welding. It combines unsupervised deep latent representation learning, which extracts compact features from weld pool images, with Bayesian filtering to handle persistent and fluctuating disturbances such as arc radiation and specular reflections. Specifically, a Dynamic Variational Autoencoder (DVAE), consisting of a CNN-based encoder-decoder and an LSTM-based transition model, jointly learns latent representations and their evolution under control inputs. For robust real-time inference, a specialized Particle Filter (PF) propagates the latent and LSTM hidden states, preserving process history while suppressing sensor noise. This design is well suited to welding's slow and inertial dynamics. Validation on GTAW and GMAW without process-specific tuning demonstrates the framework's generalizability and robustness.


[40] 2607.16036

Network-Induced Strategic Communication in Opinion Dynamics

Classical opinion dynamics typically assume a fixed mapping from private opinions to public signals, such as linear exchange, saturated signaling, or discrete public actions. In this paper, we show that these communication mappings can be derived from a strategic communication game played on a weighted influence network. Each agent acts as a receiver estimating its neighbors' states and as a sender broadcasting a public signal to influence its audience. We prove that the network's effect on a sender is summarized by a scalar, network-induced exaggeration factor, and the network game decouples into independent scalar cheap talk problems. The communication rules then emerge as behavioral regimes of one model: aligned incentives recover linear averaging; persuasive senders facing naive receivers produce saturated signaling; and persuasive senders facing Bayesian receivers cannot credibly reveal their opinions, so equilibrium communication becomes an interval quantizer, providing a game-theoretic foundation for continuous-opinion discrete-action (CODA) dynamics. Embedded in repeated opinion updates, the resulting strategic-CODA dynamics preserve opinion clustering and exclude extremism. The model predicts that a speaker exaggerates more when their audience is less influenced by them, and that under strong exaggeration, credible public expression collapses to a binary stance even while private opinions remain continuous.


[41] 2607.16084

Pick-to-Learn Calibration of an MPC Policy for an Origin-to-Destination Flight Problem

This paper illustrates the Pick-to-Learn methodology applied to the calibration of a Model Predictive Control policy. While developed around a specific example, the presentation is meant to highlight a methodology of broad applicability. The example concerns an aircraft traveling from an origin point to a destination point in the presence of uncertain crosswinds and a low-connectivity zone that should be avoided. The MPC policy is parameterized by two hyperparameters, which are selected from data by the P2L procedure. Starting from a dataset of 400 wind realizations, also called scenarios, P2L identifies a final compression set containing only two informative scenarios. The resulting MPC policy avoids the low-connectivity zone on all available scenarios and, according to the P2L theory, satisfies a probabilistic risk bound of $4.8\%$ at confidence level $1-10^{-5}$, where the risk is the probability of entering the low-connectivity zone in a future flight under a new wind realization not included in the sample.


[42] 2607.16107

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos

We present Audio-Visual Flamingo (AV-Flamingo), a fully open state-of-the-art audio-visual large language model (AV-LLM) for joint understanding and reasoning over audio, images, and long-form videos. Unlike prior AV-LLMs that primarily focus on short clips, AV-Flamingo is designed for understanding and reasoning over long and complex real-world (audio-visual) videos. To support this, we make three key contributions: (i) Audio-Visual-Skills, a large-scale collection of real-world videos with ~7M caption and question-answer training instances designed to emphasize temporal, compositional, and cross-modal audio-visual reasoning; (ii) a novel three-stage curriculum that progressively trains the model from short-range perception to long-horizon multi-event reasoning; and (iii) Temporal Audio-Visual Interleaved Chain-of-Thought, a reasoning framework that explicitly grounds intermediate reasoning steps to timestamps in long audio-visual streams, improving temporal alignment and interpretability. Extensive experiments across 15+ audio-visual, omni-modal, audio, and vision benchmarks show that AV-Flamingo outperforms similarly sized open models by clear margins and remains highly competitive with, and in some cases surpasses, much larger open-weight and closed models, particularly on long and complex real-world audio-visual understanding and reasoning tasks. Beyond benchmark performance, AV-Flamingo exhibits strong real-world utility and transfers well to unseen tasks, highlighting its robustness and generalization ability.


[43] 2607.16139

A Kalman Filter-Assisted Data-Predictive SAR ADC With Reduced Switching Energy for Low-Power Applications

The proliferation of Internet of Things (IoT) devices and wearable health monitors has created an urgent demand for ultra-low-power analog-to-digital converters (ADCs). Successive approximation register (SAR) ADCs are widely used in such applications, yet their energy efficiency remains constrained by the sequential bit-by-bit switching of the capacitive DAC (CDAC). The high-weight most significant bit (MSB) transitions dominate the total switching energy, and the rigid N -cycle conversion flow imposes a hard lower bound on latency per this http URL paper presents a Kalman filter-assisted data-predictive SAR ADC that replaces the first four comparator-driven decisions with a recursive state estimator. The Kalman filter predicts the 4 MSBs from the complete conversion history before each cycle begins, enabling simultaneous parallel switching of the MSB capacitors. This eliminates redundant CDAC transitions, shortens the quantization cycle by four clock periods, and reduces switching energy by approximately 50%. An optimized 4-bit MSB switching scheme further suppresses residual switching at the hardware level. The ADC, designed in a 180-nm CMOS process, supports configurable dual-mode operation, toggling between a conventional mode and the Kalman-driven predictive mode for robustness under erratic inputs. At 20 MS/s and a 1.8-V supply, the predictive mode reduces total power consumption by 50.3% (from 1.96 mW to 0.975 mW), with a measured SNR/SFDR of 57.88/74.51 dB at 504 kHz, confirming its suitability for energy-constrained wireless sensor networks.


[44] 2607.16172

The Internet of Things for Smart Manufacturing: A Review

The modern manufacturing industry is investing in new technologies such as the Internet of Things (IoT), big data analytics, cloud computing and cybersecurity to cope with system complexity, increase information visibility, improve production performance, and gain competitive advantages in the global market. These advances are rapidly enabling a new generation of smart manufacturing, i.e., a cyber-physical system tightly integrating manufacturing enterprises in the physical world with virtual enterprises in cyberspace. To a great extent, realizing the full potential of cyber-physical systems depends on the development of new methodologies on the Internet of Manufacturing Things (IoMT) for data-enabled engineering innovations. This paper presents a review of the IoT technologies and systems that are the drivers and foundations of data-driven innovations in smart manufacturing. We discuss the evolution of internet from computer networks to human networks to the latest era of smart and connected networks of manufacturing things (e.g., materials, sensors, equipment, people, products, and supply chain). In addition, we present a new framework that leverages IoMT and cloud computing to develop a virtual machine network. We further extend our review to IoMT cybersecurity issues that are of paramount importance to businesses and operations, as well as IoT and smart manufacturing policies that are laid out by governments around the world for the future of smart factory. Finally, we present the challenges and opportunities arising from IoMT. We hope this work will help catalyze more in-depth investigations and multi-disciplinary research efforts to advance IoMT technologies.


[45] 2607.15404

Closed-Loop Bayesian Bandit Encoder with GRAND Receiver for a Bursty Interference Channel

Interleaving mitigates burst errors but introduces decoding delay and removes temporal error structure that a channel-aware decoder could exploit. We consider packet-level selection between a random linear code and the same code used with cross-codeword interleaving, over a channel with an unknown number of on/off interferers. The receiver uses Guessing Random Additive Noise Decoding (GRAND) with a replaceable noise model and feeds aggregate channel statistics back to a Bayesian estimator at the transmitter. Once the interference amplitudes and timing parameters are estimated, the receiver's noise model is replaced: it computes hidden-Markov-model posterior bit-flip probabilities and uses them to order GRAND queries. A discounted Thompson sampler selects between the two transmission modes using a goodput-minus-latency reward whose distribution is endogenously nonstationary: receiver adaptation, rather than channel change, alters the value of each mode. Across five simulation seeds, the interleaved mode is preferred before channel estimation converges. After the learned decoder is activated, the non-interleaved mode becomes preferable because it achieves lower block error rate without interleaving delay. In the reference configuration, the learned noise model reduces block error rate by approximately one order of magnitude relative to ORBGRAND. Using partial channel estimates before full convergence reduces pre-convergence block error rate by up to $4.5\times$. Adding model-predicted utilities as confidence-weighted pseudo-observations reduces post-transition selection of the inferior arm by approximately $65\%$. Under an idealized airtime conversion at a 100~MHz 5G~NR-like symbol rate, the learning transient corresponds to a few milliseconds of occupied symbol time.


[46] 2607.15443

Estimating the Reliability of Dynamic Time Warping Alignments Using Circumstantial Evidence

Recent works have explored ways to handle uncertainty in dynamic time warping (DTW) alignment paths through the use of differentiable variants of DTW like Soft-DTW. In this paper, we approach the issue of uncertainty in DTW alignment paths in a different way. Given a DTW alignment path, we propose a metric that indicates how reliable a local segment of the alignment path is. The intuition for our metric is based on the idea of circumstantial evidence. If DTW has found a very prominent path, then if we re-run the alignment with relaxed boundary conditions, it will still pick the same path. If, on the other hand, DTW has found a "weak" path, then re-running the alignment with relaxed boundary conditions will likely yield a different path. Accordingly, our reliability metric is computed by picking a local section of the DTW alignment path, re-estimating the alignment with FlexDTW (which allows flexibility in the boundary conditions), and then measuring how well the DTW and FlexDTW paths agree. We assess the proposed reliability metric on DTW alignment paths containing both matching and non-matching regions across a range of scenarios on an audio-audio alignment task. We find that the reliability metric correctly identifies reliable regions of the alignment path with an aggregate AUROC of 0.97. This approach provides an unsupervised method for estimating the reliability of a DTW alignment path.


[47] 2607.15475

Segmental DTW: A Parallelizable Alternative to Dynamic Time Warping

In this work we explore parallelizable alternatives to DTW for globally aligning two feature sequences. One of the main practical limitations of DTW is its quadratic computation and memory cost. Previous works have sought to reduce the computational cost in various ways, such as imposing bands in the cost matrix or using a multiresolution approach. In this work, we utilize the fact that computation is an abundant resource and focus instead on exploring alternatives that approximate the inherently sequential DTW algorithm with one that is parallelizable. We describe two variations of an algorithm called Segmental DTW, in which the global cost matrix is broken into smaller sub-matrices, subsequence DTW is performed on each sub-matrix, and the results are used to solve a segment-level dynamic programming problem that specifies a globally optimal alignment path. We evaluate the proposed alignment algorithms on an audio-audio alignment task using the Chopin Mazurka dataset, and we show that they closely match the performance of regular DTW. We further demonstrate that almost all of the computations in Segmental DTW are parallelizable, and that one of the variants is unilaterally better than the other for both empirical and theoretical reasons.


[48] 2607.15478

A Study of Parallelizable Alternatives to Dynamic Time Warping for Aligning Long Sequences

This article investigates several parallelizable alternatives to DTW for estimating the alignment between two long sequences. Whereas most previous work has focused on reducing the total computation and/or memory costs of DTW, our focus is instead on reducing wall clock time by utilizing common hardware like GPUs that are optimized for parallel processing. We propose and study four different parallelizable alignment algorithms: the first three algorithms compute approximations of DTW by breaking the pairwise cost matrix into rectangular regions and processing the regions in parallel, and the fourth algorithm computes an exact DTW alignment by processing the cost matrix along diagonals rather than rows or columns. We characterize the performance of our proposed alignment algorithms on an audio-audio alignment task, and we develop GPU-based implementations for the two best-performing algorithms, which we call weakly-ordered Segmental DTW (WSDTW) and Parallelized Diagonal DTW (ParDTW). Our experiments indicate that ParDTW is the most practical and useful of the four algorithms: it computes an exact DTW alignment and reduces runtime by 1.5 to 2 orders of magnitude on long sequences compared to current alternatives. We present a comprehensive evaluation and study of the alignment accuracy, runtime, and practical limitations of the proposed alignment algorithms.


[49] 2607.15534

Machine Learning-Driven Design of Mixed-Pitch Grating Couplers for Co-Packaged Optics Applications

A mixed-pitch grating coupler which can couple a wide range of wavelengths is preferred in its application in co-packaged optics (CPO). However, the design and optimization of such grating coupler is complex. In this work, we developed software with integrated deep neural network (DNN) model to automatically design the mixed-pitch grating coupler from user-specified peak wavelengths and full-width half-maximum (FWHM) values. We first trained the DNN model with 10,000 rows of grating parameters-power spectrum datasets, where the power spectrum was simulated using finite-difference time domain (FDTD) technique. Upon training, we tested the model using ~1,000 different combinations of peak wavelengths and FWHM values. Among the combinations, 822 attempts have <15% error, while 351 attempts have <5% error when comparing the user-specified and FDTD-verified spectrum. Meanwhile, comparing the user-specified and FDTD-verified peak wavelengths, 844 attempts have peak wavelengths with absolute error (AE) < 2 nm. For FWHMs, 738 attempts have FWHM values with AE < 10 nm. We have also developed a graphical-user interface (GUI) to ease the usage of this software.


[50] 2607.15565

Ask Twice, Look Twice: Prompt Echoing Resolves the Question-First Paradox in Vision-Language Models

Where should the question go in a vision-language model (VLM) prompt: before the image or after it? Intuition says before: knowing what is asked should tell the model where to look. Yet across visual question answering benchmarks, question-first prompting consistently underperforms the image-first ordering recommended for frontier VLMs, a phenomenon we term the question-first paradox. We trace the paradox to a conflict between two stages of VLM computation. Logit-lens and attention probes show the intuition is half right: a question placed before the image genuinely steers perception, moving image patch representations toward question-relevant concepts. The failure lies downstream. Stranded behind hundreds of image tokens, the question is barely attended by the answer token, which instead commits to image-driven (often wrong) answers; a causal attention knockout confirms that the answer reads the question only when the question follows the image. The diagnosis yields a training-free fix: question echoing, restating the question on both sides of the image so that one copy steers perception while the other is read out at answer time. The same division of labor appears in a fifty-year-old finding on human ``adjunct questions'', where repeating a question before and after a passage aids comprehension more than either position alone. Echoing the image as well brings further gains, restoring the whole-image view a causal decoder otherwise loses. The paradox holds across five open VLMs, costing up to 17.5 group-accuracy points. Echoed prompts close it and surpass the best single-pass ordering on NaturalBench, POPE, Winoground, and open-ended VQAv2, by up to 19 Winoground group-accuracy points, with no training, fine-tuning, or architecture change. The paradox reveals a trade-off between steering perception and preserving question access; echoing resolves it through prompt design alone.


[51] 2607.15589

MemoGuard: An Adaptive Runtime for Guarding Against Memory Traps in Communication-Limited Robot Navigation

Communication-limited robots in mission-critical scenarios such as disaster inspection and search-and-rescue must make reliable onboard decisions without access to remote operators or high-capacity reasoning services. Episodic memory reuse is an attractive low-cost fallback, but retrieval similarity does not guarantee execution validity, i.e., a retrieved action may match the current context yet be unsafe due to changed topology, insufficient battery margin, or unreliable prior outcomes. We call such high-similarity but execution-invalid episodes memory traps. This creates a safety-efficiency design space where similarity only reuse minimizes fallback cost but can be unsafe, while always invoking local reasoning improves safety at high computational and energy cost. This paper presents MemoGuard, a lightweight adaptive runtime that validates episodic memories against topology, resource, and outcome contracts before reuse, invoking fallback only when validation fails. In a graph-based corridor-inspection simulator, MemoGuard reduces battery safety violations by 76.6% over similarity-only top-1 reuse while reducing fallback calls by 21.4% over always reasoning. On an NVIDIA Jetson AGX Xavier with local llama3.2:3b fallback reasoning, this corresponds to 3.67 s and 36.97 J of avoided fallback-reasoning overhead per trial. We open-source MemoGuard at this https URL.


[52] 2607.15656

Learning a System-Level Surrogate for Hydraulic Excavators: A Simulation-to-Real LSTM Approach

Developing autonomous hydraulic excavators is constrained by limited access to physical machines and the high cost of real-world experimentation. This paper proposes a simulation-to-real framework for learning a system-level digital surrogate using Long Short-Term Memory (LSTM) networks. Instead of modeling internal dynamics, the excavator is treated as an input-output operator, and the surrogate is trained to reproduce its closed-loop behavior under identical control inputs. The approach is first validated in a MuJoCo simulation environment and then transferred to a real excavator. To address measurement inconsistencies in real-world data, a consistency-aware state estimation method based on adaptive Kalman filtering is introduced. Experimental results demonstrate that the learned surrogate achieves high fidelity in both angular velocity and long-horizon trajectory reproduction under closed-loop autoregressive evaluation. These results confirm that the proposed model can serve as a drop-in surrogate for both simulation and physical systems, enabling scalable and efficient development of excavation automation algorithms.


[53] 2607.15801

Polynomial-Based Solutions to Targeting Problems for Onboard Applications

This paper solves the targeting problem focusing on accuracy, computational efficiency, and reliability. The trajectory optimization problem is first recast as a polynomial optimization problem (POP) by leveraging differential algebra to compute high-order Taylor expansions of the nonlinear dynamics and constraints. Moment-sum-of-squares (SOS) optimization is then utilized to solve this POP. A convex formulation based on a second-order expansion of the dynamics is also proposed. For impulsive targeting, the moment-SOS and convex approaches are compared against traditional nonlinear programming (NLP) solvers and map inversion techniques. Results indicate that the moment-SOS approach provides solutions as accurate as traditional NLP, but with the critical advantage of guaranteeing convergence to the global optimum under mild assumptions. Furthermore, the method excels at handling large maneuvers and long propagation times, conditions in which standard linear approximations rapidly degrade. To demonstrate its versatility, the methodology is extended to a continuous low-thrust station keeping (SK) scenario in the Earth-Moon Circular Restricted Three-Body Problem. The algorithm's performance is then evaluated in the presence of significant state errors. The ability to directly handle non-convex constraints and recast complex, nonlinear dynamics into formulations with reliable convergence properties makes the moment-SOS approach suitable for autonomous onboard applications.


[54] 2607.15853

Maximal quantum leakage: operational interpretation and quantum channel analysis

Maximal quantum leakage quantifies privacy against adversaries with arbitrary intentions. In this work, we prove that computing this leakage is equivalent to minimum-error quantum state discrimination with equal priors. This establishes a computable operational interpretation, addressing the previous difficulty in computing maximal quantum leakage. We further analyze the impact of collective measurements on multiple copies of a state, demonstrating that leakage increases monotonically with the number of copies, which leads to explicitly characterizing the maximal leakage in the asymptotic limit. Extending this framework to quantum channels, we develop an iterative algorithm for the jointly designing of input states and measurements. Numerical examples involving collective measurements and the maximal channel leakage demonstrate our theoretical findings.


[55] 2607.15929

Current Should Not Sneak: Constrained Codes for Reliable Memristor Crossbar Arrays

The approach of squeezing more transistors in the same area in order to speed up computing is no longer effective. Currently, researchers and engineers are searching for novel solutions that offer faster computing. One of these solutions is to compute where you store, known as in-memory computing. Resistive random access memories (ReRAMs), which are based on memristor crossbar arrays, enable in-memory computing. Moreover, ReRAMs offer large storage capacity associated with energy efficiency. In this work, we focus on storing digital data in memristor crossbar arrays. A critical challenge here is the sneak-path problem, occurring when there is a rectangle on the array with three low and one high resistances at the corners. The electric current in this case is prone to sneaking through the low-resistance path upon reading, which results in the high resistance data becoming erroneous. In this paper, we propose effective constrained coding solutions to the sneak-path problem after finding the expected number of sneak paths over a two-dimensional array given their circumferences. In particular, we adopt a literature model where $b$ rows on the crossbar array are read simultaneously while the others are grounded, and we design capacity-achieving non-binary constrained codes for the cases of $b=2$ and $b=3$. We focus more on the sneak paths with shorter circumferences as they are more detrimental. Here, GF refers to Galois field. Our GF$(4)$ codes, for $b=2$, and GF$(8)$ codes, for $b=3$, are a class of lexicographically-ordered constrained (LOCO) codes, and we call them resistive-LOCO (RES-LOCO) codes. RES-LOCO codes operate horizontally, and we also suggest a run-length-limited scheme for coding data on the crossbar array vertically to mitigate the sneak-path problem for $b=4$. We experimentally demonstrate the effectiveness of our RES-LOCO codes for various array setups.


[56] 2607.15939

Strategic Persuasion Through Information Timeliness

We study a dynamic strategic communication problem in which a sender controls the timing of truthful updates from binary continuous-time Markov sources. The receiver chooses between a zero-order-hold estimator that follows the sender's updates and a prior-only default estimator, aiming to maximize a weighted correct-estimation utility. In contrast, the sender seeks to persuade the receiver to estimate the state as 1, regardless of the true state. This misalignment leads to a Stackelberg game in which the sender, as the leader, commits to state-dependent Poisson update rates, and the receiver, as the follower, decides whether to follow the sender's messages. The sender maximizes the long-term average time that the receiver's estimate equals 1, subject to a conditional intensity budget and a participation constraint (PC) ensuring that following the sender's messages does not degrade the receiver's average utility relative to its prior information. For a single source, we show that the sender's optimal policy allocates a minimum state-0 update intensity to the undesired state-0, just enough to satisfy the PC, and the remaining budget to the desired state-1. For multiple sources with heterogeneous minimum state-0 update intensities, we develop a branch-and-bound algorithm that typically avoids exhaustive search. Finally, we extend the solution to multiple receivers over dedicated channels. Our results show that controlling timeliness alone enables the sender to persuade the receiver and increase its utility.


[57] 2607.15986

DebrisTracer: Reliable Tracking in Hypervelocity Impact Fast Imaging

This application paper presents DebrisTracer, a framework for the reliable tracking of debris in hypervelocity impact fast imaging. These noisy and highly specific datasets capture the ejection of a large number of debris fragments after the impact of a projectile launched at hypervelocity into a target material. The reliable estimation of debris mass and speed distributions is of major importance in aerospace applications. We document how to extend an off-the-shelf topology tracking framework based on critical point extraction and matching, in order to incorporate domain knowledge and physical assumptions. Our approach automatically produces an accurate and reliable debris tracking, enabling an interpretable visual analysis of this complex space-time phenomenon. Extensive experiments demonstrate the accuracy improvements provided by our approach over established tools used by domain experts in terms of physical validation, specifically via the prediction of the experimental ejected mass and crater depth profiles. We illustrate the utility of our approach across several use cases (with varying impact angles and physics). We show that our statistical summaries enable the visual identification of distinct regimes within the debris population, corroborating and refining prior expectations of domain experts. Our database and our C++ implementation are available at this address: this https URL.


[58] 2607.15994

A 140-GHz Direct Raised-Cosine Envelope-Shaping Transmitter with Integrated ILO Phase Shifter

A 140-GHz transmitter with direct raised-cosine-like envelope shaping and wide-range phase tuning is presented in 90-nm SiGe BiCMOS. The proposed architecture relaxes conventional baseband pulse shaping and high-speed digital-to-analog converters by directly synthesizing 3-level and 5-level RF envelope states that approximate a raised-cosine waveform, enabled by a 23-dB modulation dynamic range. Measured results demonstrate data rate up to 8 Gbpers with 34-dB sidelobe suppression. To support scalable phased-array applications, an injection-locked-oscillator phase-tuning path is also integrated, providing 22.5deg digital phase resolution together with continuous analog tuning over more than 360 deg. The transmitter achieves up to 2 dBm output power and combines direct RF envelope shaping with wide-range fine phase control for sub-THz wireless links


[59] 2607.16002

A Morphing-Designed Hexarotor Prototype combining Practical Resilience and Efficiency

This work demonstrates experimentally the existence of a hexarotor prototype, termed Opti-Hexa, that simultaneously achieves practical resilience to single-propeller failures and energy efficiency comparable to a standard Star-shaped prototype with the same size, weight, hardware and software. Leveraging a novel open-source morphing platform, we investigate the trade-offs across a continuous range of geometries by varying the angles between adjacent propellers. We study practical efficiency through a data-fitted empirical power model and evaluate practical resilience by comparing the position accuracy and rotational kinetic energy during failure to those observed under nominal hovering conditions. Our experiments confirm the existence of a geometric viability region for this specific morphing platform, where resilience is ensured without the aerodynamic efficiency losses typically associated with practically resilient designs found in the state of the art. The complete hardware and software of the morphing platform are released to support further research.


[60] 2607.16080

Physics-Based Deep Spatiotemporal Hyperlocal Radar Nowcasting with a Multi-Variable U-Net for High-Resolution Precipitation Forecasting

Precipitation nowcasting over the immediate 10-90 min period is important for flood management and real-time decision-making in urban regions. Conventional short-range forecasting with high-resolution numerical weather prediction requires frequent data assimilation, model initialization, and spin-up, introducing computational latency. Machine learning provides an alternative by learning storm evolution directly from high-frequency observations and producing forecasts quickly after training. This is particularly relevant for Mumbai, India, where monsoon convection, land-sea interactions, and localized intense rainfall make short-term prediction difficult. Here, we develop a compact radar-only nowcasting framework that combines multi-elevation reflectivity, Doppler radial velocity, and radial-velocity-gradient proxy features within an encoder-decoder U-Net. Using the most recent radar volume scan, the model predicts 12 future composite reflectivity fields at 7.5-min intervals up to 90 min lead time. The derived velocity magnitude, divergence-like, directional-shear, and vorticity-like channels represent kinematic signatures associated with convergence and boundary interactions without requiring full wind-field retrieval. A high-reflectivity attention module improves sensitivity to convective cores, and physics-guided attribution examines whether the learned sensitivities are meteorologically meaningful. The model is trained using Mumbai Doppler radar observations from May to August 2023 and evaluated on temporally independent events. At 90 min lead time, Critical Success Index values are 0.437, 0.332, and 0.193 for $\geq$10, $\geq$20, and $\geq$30 dBZ thresholds, respectively. Compared with persistence, the model gives lower RMSE and higher spatial correlation at longer lead times. Once trained, it runs on a standard computer, generating nowcasts within seconds for real-time use.


[61] 2607.16085

Controlling Implicit Shortcut Reliance in L2 Spoken English Auto-markers

Increasingly, speech and language processing tasks take either audio or text directly rather than extracting features from these as the input to the classifier or regressor. Often these systems make use of complex, for example transformer-based, processes that have the ability to derive highly non-linear mappings between the input and the output. Unfortunately these systems can also learn ''shortcuts'' where the classifier is overly reliant on particular aspects of the input to yield the output. For the task of language proficiency assessment, this over-reliance can enable learners to increase their score by exploiting the shortcut rather than improving their ability. This paper introduces a novel training criterion that is able to reduce the classifier's reliance on shortcuts, thus for example limiting this option for malpractice in language assessment. This process is illustrated on two forms of assessment system, one based on the audio the other on the speech recognition text. The results show that, for both systems, there is higher correlations with features that could be exploited for malpractice than expected from the human reference, indicating an over-reliance on these features. By introducing the modified training criterion, this correlation can be reduced to be closer to the reference correlation.


[62] 2607.16121

Comparison of Energy System Optimization Software and Evaluation of Selected Frameworks

Optimizing energy systems is a crucial step toward achieving a carbon-neutral future, with software tools playing a major role in the process. However, selecting the most suitable tool for specific optimization challenges can be complex, given the diverse objectives and requirements of various energy systems. In this study, we aim to address this issue by evaluating five preselected software tools-REMix, MTRESS, COMANDO, OEMOF, and HOMER PRO-to identify the scenarios for which they are most suitable. To achieve this, we conducted an extensive review of literature, documentation, tutorials, and example models, and developed a set of comparison criteria that were subsequently used to evaluate these tools. Our analysis suggests that REMix is particularly effective for scenarios where investment path optimization is required. COMANDO excels in systems with atypical components. OEMOF is the best-suited open-source tool for standard optimization problems. HOMER PRO is recommended for users seeking rapid synthesis optimization, especially those with limited programming experience. Due to difficulties in obtaining sufficient information on MTRESS, we were unable to complete the analysis for this tool.


[63] 2607.16130

A Methodology for Auditable Trustworthiness Levels in AI Lifecycle Governance

AI governance increasingly requires judgments about whether an AI system remains adequately trustworthy over time, whether observed changes are tolerable, and how such judgments should be documented in a transparent and contestable way. Yet existing work on AI trustworthiness remains either too high-level to support lifecycle monitoring and reassessment or too narrowly metric-driven to connect with governance needs. We therefore propose a lightweight methodology for auditable trustworthiness levels in AI governance. The methodology has two components: a formal framework for representing and learning trustworthiness levels, and a lightweight AI lifecycle governance procedure for documenting, monitoring, and reassessing them over time. The formal framework models governance-relative trustworthiness through a context-sensitive protocol of measurable dimensions and learns trustworthiness levels as interpretable rules over trustworthiness profiles. Using decision trees as an interpretable proof-of-concept model class, the methodology yields explicit trustworthiness plateaus, readable level transitions, and two simple lifecycle diagnostics: boundary margins and profile drift. The governance procedure embeds these formal objects in a conformity-oriented workflow for design-time labeling, post-deployment monitoring, reassessment, and reporting. It also assigns human responsibilities and control gates for protocol design, validation, monitoring, and reassessment. We illustrate the methodology on synthetic AI lifecycle traces involving degradation, shocks, updates, heterogeneous monitoring cadences, and system comparison. Our methodology does not replace legal or other expert judgment: it supports conformity documentation and lifecycle monitoring by providing an evidential basis for documenting and tracking AI governance-relevant changes over time.


[64] 2607.16154

CLIFE: Camera-LiDAR Fusion Framework for Edge-Deployable Roadside VRU Perception

Reliable roadside perception of vulnerable road users (VRUs) remains challenging under occlusions, variable lighting, and diverse weather conditions, particularly under strict edge-computing and latency constraints. Existing multi-sensor fusion systems rely on cloud or server-grade infrastructure, creating a deployment gap at real-world intersections. We present CLIFE, an edge-native camera-LiDAR fusion framework that integrates targetless online calibration and lightweight late-fusion tracking entirely on a single embedded device, without cloud offloading. CLIFE adaptively refines camera-LiDAR alignment on demand and performs multi-sensor fusion and track association with O(N log N) per-frame cost. We deploy CLIFE across 12 signalized intersections in Chattanooga and conduct an in-depth evaluation at a representative intersection using synchronized camera-LiDAR data that spans diverse daytime, nighttime, and weather conditions. Our experiments demonstrate that the fusion architecture substantially enhances the perceptual range and robustness of the individual sensors under varied environmental and traffic conditions. The late-fusion core operates at 53.2 FPS on the Jetson AGX Thor, ensuring high throughput for real-time intersection-scale applications. By centering perception at the edge, CLIFE provides a deployable foundation for downstream safety applications, while reducing bandwidth and calibration overhead for agencies operating multi-intersection corridors.


[65] 2607.16156

PRISA: Proactive Infrastructure LiDAR Framework for Intersection Safety Assessment

Urban intersections are among the most hazardous locations in road networks, posing significant risks to vehicles and vulnerable road users (VRUs) such as pedestrians and cyclists. The complexity of multi-agent interactions demands continuous, real-time monitoring systems capable of anticipating conflicts before they escalate into crashes. We present PRISA, a modular infrastructure LiDAR framework leveraging privacy-preserving, low-light-robust roadside sensors for long-term traffic observation and real-time risk detection at the edge. The framework comprises two core components: a sensing and perception layer and a plug-and-play risk assessment module. The latter automatically curates site-specific training data from accumulated perception outputs to train a trajectory prediction model without manual annotation. It then deploys the trained model for continuous motion forecasting and dual surrogate safety evaluation, using Time-to-Collision (TTC) for longitudinal conflicts and Predicted Post-Encroachment Time (PPET) for crossing and VRU-involved interactions. PRISA is evaluated on the public R-LiViT dataset and deployed on an NVIDIA Jetson AGX Thor at a live signalized intersection in Chattanooga, Tennessee. PPET-based assessment operates at 194~ms end-to-end latency over a 2.4-second predictive horizon, with TTC-based detection and perception remaining within real-time constraints, demonstrating practical feasibility for proactive multi-agent intersection safety monitoring.


[66] 2607.16171

A Globally Asymptotically Stable Planar Homogeneous Polynomial Vector Field With No Polynomial Lyapunov Function

We disprove the conjecture that every globally asymptotically stable homogeneous polynomial vector field admits a homogeneous polynomial Lyapunov function. The counterexample is a planar homogeneous cubic polynomial vector field with integer coefficients. It admits no positive definite homogeneous polynomial with nonpositive Lie derivative and, more strongly, no real-analytic Lyapunov function even locally. Nevertheless, it has an explicit degree-two homogeneous Lyapunov function that is radially unbounded, continuously differentiable everywhere, and smooth away from the origin. We also provide a machine-checked Lean 4 formalization of the main result.


[67] 2504.20526

Histogram-Probabilistic Multi-Hypothesis Tracking with Integrated Target Existence

The histogram-probabilistic multi-hypothesis tracker (H-PMHT) is a parametric approach to solving the multi-target track-before-detect (TBD) problem, using expectation maximisation (EM). A key limitation of this method is the assumption of a known and constant number of targets. In this paper, we propose the integrated existence Poisson histogram probabilistic multi-hypothesis tracker (IE-PHPMHT), for TBD of multiple targets. It extends the H-PMHT framework by adding a probability of existence to each potential target. For the derivation, we utilise a Poisson point process (PPP) measurement model and Bernoulli targets, allowing for a multi-Bernoulli birth process and an unknown, time-varying number of targets. Hence, integrated track management is achieved through the discrimination of track quality assessments based on existence probabilities. The algorithm is evaluated in a simulation study of two scenarios and is compared with several other algorithms, demonstrating its performance.


[68] 2507.14194

Boosted Enhanced Quantile Regression Neural Networks with Spatiotemporal Permutation Entropy for Complex System Prognostics

This paper presents an integrative prognostic framework that combines Spatiotemporal Permutation Entropy (STPE), Boosted Enhanced Quantile Regression Neural Networks (B-EQRNNs), Gated Temporal Attention, a Spiking Neural Network (SNN) refinement stage, and a Temporal Fusion Transformer (TFT) classifier. The motivation is long-horizon fault prediction in distributed industrial electronic systems, where single-sensor or point-estimate models can miss weak spatially propagating degradation signatures and provide limited uncertainty information. The proposed pipeline first converts 70-channel sensor streams into multiscale STPE descriptors, then learns conditional quantile representations and attention-weighted temporal context before final Normal/Abnormal classification. Evaluation is reported on a nine-system industrial electronic-sensor dataset with 48-, 90-, and 168-hour prediction horizons. The comparison includes a tree-based LightGBM baseline and modern sequence baselines available under the same preprocessing protocol, including LSTM, Autoformer, and TCN models. The full pipeline reaches 81.17% accuracy at the 168-hour horizon and is evaluated with component ablations, computational-cost analysis, and an explicit reproducibility protocol. The contribution is therefore framed as a validated hybrid architecture for uncertainty-aware spatiotemporal prognostics rather than as a new standalone learning theory.


[69] 2508.00307

Acoustic Imaging for UAV Detection: Dense Beamformed Energy Maps and U-Net SELD

We introduce a U-net model for 360° acoustic source localization formulated as a spherical semantic segmentation task. Rather than regressing discrete direction-of-arrival (DoA) angles, our model segments beamformed audio maps (azimuth & elevation) into regions of active sound presence. Using delay-and-sum (DAS) beamforming on a custom 24-microphone array, we generate signals aligned with drone GPS telemetry to create binary supervision masks. A modified U-Net, trained on frequency-domain representations of these maps, learns to identify spatially distributed source regions while addressing class imbalance via the Tversky loss. Because the network operates on beamformed energy maps, the approach is inherently array-independent and can adapt to different microphone configurations and can be transferred to different microphone configurations with minimal adaptation. The segmentation outputs are post-processed by computing centroids over activated regions, enabling robust DoA estimates. Our dataset includes real-world open-field recordings of a DJI Air 3 drone, synchronized with 360° video and flight logs across multiple dates and locations. Experimental results show that U-net generalizes across environments, providing improved angular precision, offering a new paradigm for dense spatial audio understanding beyond traditional Sound Source Localization (SSL). We additionally validate the same beamforming-plus-segmentation formulation on the DCASE 2019 TAU Spatial Sound Events benchmark, showing that the approach generalizes beyond drone acoustics to multiclass Sound Event Localization and Detection (SELD) scenarios.


[70] 2508.10218

On Geometric Asymmetry and Information in Sequential Dimension Reduction

Standard random projection techniques typically operate as a black box, mapping high-dimensional structures directly to a lower-dimensional space where the target dimension must be specified a \textit{priori}. To address scenarios where the optimal ultimate dimension is unknown, this paper investigates the retention of information through a sequential, step-by-step dimension reduction process. We examine a fixed, bounded convex body as it undergoes successive random orthogonal projections, systematically reducing the ambient dimension by one at each step. By demonstrating that this sequence of observed bodies forms a Markov chain, we quantify the information preserved through these reductions using the conditional mutual information between successive projections given the original convex body. We derive a theoretical upper bound on this conditional mutual information, parameterized by the Haar measure of the projection spaces that yield the same observed body. Leveraging the established Markov property, we extend these results to an arbitrary number of iterations, proving that the initial two-step bound characterizes information retention across the entire sequence of projections. Furthermore, by analyzing the projection space under the symmetry group of the initial body, we demonstrate that geometric asymmetry serves as a beneficial asset, resulting in higher overall information retention.


[71] 2509.13600

GNSS Jamming and Spoofing Monitoring Using Low-Cost COTS Receivers

The Global Navigation Satellite System (GNSS) is increasingly vulnerable to radio frequency interference (RFI), including jamming and spoofing, which threaten the integrity of navigation and timing services. This paper presents a methodology for detecting and classifying RFI events using low-cost commercial off-the-shelf (COTS) GNSS receivers. By combining carrier-to-noise ratio (C/N0) measurements with a calibrated received power metric, a two-dimensional detection space is constructed to identify and distinguish nominal, jammed, spoofed, and blocked signal conditions. The method is validated through both controlled jamming tests in Norway and real-world deployments in Poland, and the Southeast Mediterranean which have experienced such conditions. Results demonstrate that COTS-based detection, when properly calibrated, offers a viable and effective approach for GNSS RFI monitoring.


[72] 2510.01475

Comparative Field Deployment of Reinforcement Learning and Model Predictive Control for Residential HVAC

Model Predictive Control (MPC) has demonstrated significant performance improvements over today's control methods for residential Heating, Ventilation, and Air Conditioning (HVAC), but deploying MPC often requires substantial engineering effort. Reinforcement Learning (RL) may offer comparable performance with easier deployment, but its practical application for residential HVAC remains largely undemonstrated, leaving open questions related to occupant comfort and data requirements. To investigate these issues, we deployed one MPC variant and one model-based RL variant for one month each in an occupied house in a cold climate. The controllers adjusted an air-to-air heat pump's thermostat temperature setpoint based on measurements of the indoor temperature and the electric power used for heating. Relative to constant-setpoint operation, MPC saved 18.1\% (95\% confidence interval: 4.4 to 30.9\%) of weather-normalized heat pump energy and RL saved 20.9\% (2.6 to 38.3\%). MPC maintained acceptable occupant comfort. RL kept the house cooler, particularly during an initial adaptation phase, leading to three reports of occupant discomfort. The two algorithms had similar data requirements. We estimate that for a fresh deployment in another house, RL would take about one-third less engineering effort than MPC. While RL reduces deployment effort, it faces difficulties related to safe controller initialization and to mismatches between the modeled and true state and action spaces.


[73] 2510.16495

Performance Evaluation of High Power Microwave Systems Against UAVs A Probabilistic Antenna Propagation Framework with Sensitivity Analysis

We present an uncertainty-aware probabilistic framework for high-power microwave (HPM) counter-UAV performance under stochastic target motion, beam-pointing uncertainty, atmospheric propagation, and uncertain target susceptibility. It couples stochastic UAV kinematics, a jitter-to-gain model, free-space spreading, gaseous absorption, and rain attenuation, and a logistic energy--response model to derive closed-form statistics of received pulse energy and per-pulse and cumulative effectiveness probabilities. Slant-range variability arises from integrated acceleration noise. Received pulse energy is treated as a target-level exposure metric rather than the exact energy absorbed by an internal component. Closed-form moments and a log-normal approximation yield the mean per-pulse probability through Gaussian--Hermite quadrature and a dwell-time expression under an independent-pulse assumption. Analytical predictions closely match Monte Carlo results under matched assumptions. For a vulnerable-target threshold of $E_{\mathrm{th}}=10^{-2}\,\mathrm{J}$, the model predicts $\bar{P}_{\mathrm{kill}}\gtrsim0.4$ per pulse and $P_{\mathrm{kill,tot}}>99\%$ within about $0.1\,\mathrm{s}$ at kilohertz PRF. For a hardened target with $E_{\mathrm{th}}=10^{-1}\,\mathrm{J}$, it predicts $\bar{P}_{\mathrm{kill}}\approx2.2\times10^{-4}$ ($\approx0.02\%$) and $P_{\mathrm{kill,tot}}\approx20\%$ after $1\,\mathrm{s}$ at $1\,\mathrm{kHz}$ under the i.i.d. pulse assumption. Elasticity analysis identifies slant range as dominant ($S_{\bar{R}}\approx-2$), followed by aperture diameter and transmit power; pointing jitter and atmospheric variability are less influential in the evaluated regimes. Within its assumptions, the framework supports system sizing, trade-off analysis, and risk-aware mission planning.


[74] 2510.20067

Semantic Communication for Task Execution and Data Reconstruction in Multi-View Scenarios

Semantic communication has gained significant attention with the advances in machine learning. Most semantic communication works focus on either task execution or data reconstruction, with some recent works combining the two. In this work, we propose a semantic communication system for concurrent task execution and data reconstruction for a multi-view scenario, which we formulate as the maximization of mutual information. To investigate the trade-off between the two objectives, we formulate a joint objective as a convex combination of task execution and data reconstruction. We show that under specific assumptions, the \ac{SSIM} loss can be obtained from the mutual information maximization objective for data reconstruction, which takes human visual perception into account. Furthermore, for constant resource use, we show that by increasing the weight of the reconstruction objective up to a certain point, the task execution performance can be kept nearly constant, while the data reconstruction can be significantly improved.


[75] 2511.16235

Describing Functions and Phase Response Curves of Excitable Systems

The describing function (DF) and phase response curve (PRC) are classical tools for the analysis of feedback oscillations and rhythmic behaviors, widely used across control engineering, biology, and neuroscience. These tools are known to have limitations in networks of relaxation oscillators and excitable systems. For this reason, the paper proposes a novel approach tailored to excitable systems. Our analysis focuses on the discrete-event operator mapping input trains of events to output trains of events. The methodology is illustrated on the excitability model of Hodgkin-Huxley. The proposed framework provides a basis for designing and analyzing central pattern generators in networks of excitable neurons, with direct relevance to neuromorphic control and neurophysiology.


[76] 2601.18333

Integrated Channel Estimation and Sensing for Near-Field ELAA Systems via Low-Rank Tensor Decomposition

In this paper, we study the problem of uplink channel estimation for near-filed orthogonal frequency division multiplexing (OFDM) systems, where a base station (BS), equipped with an extremely large-scale antenna array (ELAA), serves multiple users over the same time-frequency resource block. A non-orthogonal pilot transmission scheme is considered to accommodate a larger number of users that can be supported by ELAA systems without incurring an excessive amount of training overhead. To facilitate efficient multi-user channel estimation, we express the received signal as a third-order low-rank tensor, which admits a canonical polyadic decomposition (CPD) model for line-of-sight (LoS) scenarios and a block term decomposition (BTD) model for non-line-of-sight (NLoS) scenarios. An alternating least squares (ALS) algorithm and a non-linear least squares (NLS) algorithm are employed to perform CPD and BTD, respectively. Channel parameters are then efficiently extracted from the recovered factor matrices. By exploiting the geometry of the propagation paths in the estimated channel, users' positions can be precisely determined in LoS scenarios. Moreover, our uniqueness analysis shows that the proposed tensor-based joint multi-user channel estimation framework is effective even when the number of pilot symbols is much smaller than the number of users, revealing its potential in training overhead reduction. Simulation results demonstrate that the proposed method achieves markedly higher channel estimation accuracy than compressed sensing (CS)-based approaches.


[77] 2602.10936

Indirect data-driven predictive control and the state-space predictor

We define trajectory predictive control (TPC) as a class of indirect data-driven predictive control (DDPC) methods that represent future outputs as linear in past inputs/outputs and future inputs. TPC unifies many DDPC variants with different predictor structures. We introduce a predictor with a state-space representation and show that with it, TPC inherits the mature theory of linear model predictive control. In numerical experiments, the state-space predictor outperforms existing predictors, especially for small training datasets.


[78] 2604.19248

Robust Path Following Control for Vehicles with Uncertain Steering Resistance Using Model Error Compensation

This paper presents a robust path following control method for vehicles that explicitly incorporates steering torque dynamics into the control model. Unlike conventional approaches that treat the steering angle as a direct control input, this study models the steering angle as a state variable driven by a torque-proportional steering command against a speed- and angle-dependent resistance term. Since the resistance coefficient depends on road surface properties and is difficult to determine precisely, it is treated as an uncertain parameter. To compensate for the resulting model error, a Model Error Compensator (MEC) is employed as an add-on compensator that feeds back the discrepancy between the actual plant and a nominal model running in parallel, without requiring an explicit plant inverse or a specific canonical form. The zero dynamics arising from the path following formulation are formally analyzed, and it is shown that they are stable for any positive rear cornering power. Numerical simulations under systematic parameter mismatch conditions (C/C_M=0.5 to 2.0) demonstrate that the proposed method reduces the maximum following error by more than 90\% compared to the conventional method without MEC throughout the tested mismatch range, and maintains practical tracking performance within a mismatch range of 0.75 <= C/C_M <= 1.25. These results confirm that MEC effectively suppresses the influence of steering torque uncertainty, significantly enhancing path following robustness.


[79] 2605.18516

Sparse Channel Estimation for Pixel Antennas: Addressing the Pilot Rank Deficiency

Composed of multiple interconnected pixels controlled by on/off RF switches, the pixel antenna can generate reconfigurable radiation patterns that can be further exploited to construct diverse pilot sequences for effective channel estimation. However, such pilot sequences inherently have rank deficiency, making it difficult to effectively and efficiently acquire the full channel state information (CSI) across all available radiation patterns. To tackle this difficulty, we consider a sparse environment with a limited number of propagation paths for a pixel antenna system, where a user equipped with a pixel antenna transmits only a limited number of pilots to recover the CSI under all radiation patterns. The proposed algorithm exploits the limited number of propagation paths that are invariant with the pixel antenna patterns, and then formulates the full channel estimation as a sparse recovery problem in the angular domain solved by Generalized Approximate Message Passing (GAMP). Moreover, to mitigate the rank deficiency of pilot sequences, we additionally incorporate a Multipath Matching Pursuit (MMP) algorithm for robust initialization. The overall proposed scheme, termed MMP-GAMP, achieves higher estimation accuracy than other algorithm baselines, while requiring lower pilot overhead.


[80] 2605.25878

A Clinically Validated Foundation Model for Comprehensive Lung Pathology Interpretation

Pathological assessment guides lung cancer diagnosis, treatment selection, and prognostic evaluation, yet current CPath approaches rely on task-specific models for isolated objectives. Although pan-cancer foundation models offer versatility, they lack subspecialty-level depth and have not been evaluated across clinical workflows or prospectively validated in real-world settings. We introduce PulmoFoundation, a multi-center, prospectively validated, randomized controlled trial (RCT)-evaluated foundation model for comprehensive lung pathology assessment across pre-operative, intra-operative, and post-operative care. Built upon Virchow2 via subspecialty-specific pretraining using ~40,000 diagnostic H&E-stained whole-slide images (WSIs), PulmoFoundation was systematically evaluated on ~26,000 WSIs across 32 clinically relevant tasks. In addition to accurately predicting molecular markers and patient survival, our model achieves clinical-grade performance in core diagnostic tasks across biopsy, frozen section, and surgical resection slides. In a registered prospective study of 1,357 patients across 11 diagnostic tasks, our model achieved an average AUC of 92.3%. Using pre-specified triage thresholds, PulmoFoundation could reduce additional second-review burden for 68.8% of biopsies and 83.0% of frozen sections, and defer 44.5% of IHC stain orders, with PPVs of 1.000, 0.991, and 0.966. Beyond prospective validation, we conducted a crossover RCT with eight pathologists, in which AI assistance improved diagnostic accuracy across 5,264 case-reader pairs (91.7% w/ AI vs. 83.2% w/o AI). AI assistance also reduced median diagnostic time by 18.3%, increased diagnostic confidence by 9.0%, and improved inter-rater agreement from moderate (kappa = 0.55) to substantial (kappa = 0.76). Together, these evaluations support PulmoFoundation as a clinically validated decision-support system for lung pathology.


[81] 2607.02478

Docking of Autonomous Vehicles with a Stationary Docking Station in 3D Space

In this letter, we present a strategy for autonomous docking of autonomous vehicles in three-dimensional space. Docking is a safety-critical task and requires expert piloting skills. Vehicles with autonomous docking capabilities are highly desirable in various applications, such as marine vehicle docking, aerial vehicle docking, spacecraft docking, and landing. To dock autonomously with the docking station, the vehicle must align itself to a specific desired orientation relative to the docking station and also reduce speed as it approaches. The vehicle achieves near-zero speed to dock successfully and safely without colliding with the docking station. Inspired by the philosophies from the guidance literature, we present a finite-time sliding mode-based strategy to achieve the same. The range and line-of-sight kinematics relations describing the motion of the vehicle with respect to the stationary docking station are used to steer the vehicle to achieve the desired orientation for docking. This docking strategy is validated in MATLAB\textsuperscript{\textregistered} simulations for various initial locations and orientations of both the vehicle and the docking station.


[82] 2504.06479

Holistic Fusion: Task- and Setup-Agnostic Robot Localization and State Estimation with Factor Graphs

Seamless operation of mobile robots in challenging environments requires low-latency local motion estimation and accurate global localization. While most sensor-fusion approaches are designed for specific scenarios, this work introduces a flexible open-source solution for task- and setup-agnostic multimodal sensor fusion distinguished by its generality and usability. Holistic Fusion formulates sensor fusion as a combined estimation problem of i) the local and global robot state and ii) a (theoretically unlimited) number of dynamic variables, including automatic alignment of reference frames; this formulation fits countless real-world applications without conceptual modifications, offering a comprehensive solution beyond hard-coded/task-specific approaches. The proposed factor-graph formulation enables direct fusion of an arbitrary number of absolute, local, and landmark measurements expressed with respect to different frames by explicitly including them as states in the optimization and modeling their evolution as random walks. Moreover, local smoothness and consistency receive particular attention to prevent estimation jumps. Holistic Fusion enables low-latency and smooth online state estimation on typical robot hardware while simultaneously providing low-drift global localization at the IMU measurement rate. The efficacy of this released framework [1] is demonstrated in five real-world scenarios on three robotic platforms with distinct task requirements, highlighting the advantages of fusing multiple absolute measurement types [2]. [1] Code: this https URL [2] Project: this https URL


[83] 2508.03708

Tax reform as a constrained optimization problem: a piecewise-linear framework and software implementation

In many countries, income tax codes have grown into a complex tangle of interacting brackets, benefits, and deductions. Despite widespread calls for systematic reform, successful attempts at reform are rare. Part of the problem is the difficulty of designing viable reform proposals. Politically viable reform must offer hard guarantees on income effects, marginal rates, and budgetary cost. Existing microsimulation tools can evaluate a reform proposal but cannot generate one by themselves. We develop a framework that casts tax reform as a constrained optimization problem. We show that any statutory tax code satisfying four mild assumptions reduces to a finite-dimensional piecewise-linear function for each taxpayer group, so reform becomes a linear or mixed-integer linear program whose decision variables are legislatable parameters: rates, bracket cutoffs, and lump-sum transfers. We are able to recover current tax systems and generate provably optimal reform candidates within the modeled space, or a certificate that no reform satisfying certain policy design constraints exists. Behavioral effects can also be incorporated, producing a nonconvex mixed-integer formulation. We demonstrate the framework through a near-complete reconstruction of the Dutch income tax code, generating reforms that smooth marginal-rate spikes, cap household income losses, and roughly halve the number of active rules through a lexicographic procedure. Developed in close collaboration with the Dutch Ministry of Finance, the methodology is currently in active use there. An open-source software implementation is available as \texttt{TaxSolver}.


[84] 2509.15412

Sym2Real: Symbolic Dynamics with Residual Learning for Data-Efficient Adaptive Control

We present Sym2Real, a fully data-driven framework for highly data-efficient adaptation of low-level controllers. Although symbolic regression is data-efficient, its role in real-world control has been limited due to its sensitivity to measurement noise, which corrupts the equations and leads to model degradation when fitted directly on real-world data. Sym2Real addresses this limitation by 1) learning first from low-fidelity simulation, where noise-free trajectories allow symbolic regression to identify the underlying dynamics, and 2) using a small amount of real-world data for targeted residual adaptation to bridge the sim-to-real gap. Using only about 10 trajectories, we achieve robust control of both a quadrotor and a racecar in the real world, without expert knowledge or simulation tuning. Through experimental validation on both platforms, we demonstrate consistent data-efficient adaptation across 6 out-of-distribution sim2sim scenarios and successful sim2real transfer across 5 real-world conditions. More information can be found at this http URL


[85] 2510.26147

Duality-Based Fixed Point Iteration Algorithm for Beamforming Design in ISAC Systems

In this paper, we investigate the beamforming design problem in an integrated sensing and communication (ISAC) system, where a multi-antenna base station simultaneously serves multiple communication users while performing radar sensing. We formulate the problem as the minimization of the total transmit power, subject to signal-to-interference-plus-noise ratio (SINR) constraints for communication users and mean-squared-error (MSE) constraints for radar sensing. The core challenge arises from the complex coupling between communication SINR requirements and sensing performance metrics. To efficiently address this challenge, we first establish the equivalence between the original ISAC beamforming problem and its semidefinite relaxation (SDR), derive its Lagrangian dual formulation, and further reformulate it as a generalized downlink beamforming (GDB) problem with potentially indefinite weighting matrices. Compared to the classical DB problem, the presence of indefinite weighting matrices in the GDB problem introduces substantial analytical and computational challenges. Our key technical contributions include (i) a necessary and sufficient condition for the boundedness of the GDB problem, and (ii) a tailored efficient fixed point iteration (FPI) algorithm with a provable convergence guarantee for solving the GDB problem. Building upon these results, we develop a duality-based fixed point iteration (Dual-FPI) algorithm, which integrates an outer subgradient ascent loop with an inner FPI loop. Simulation results demonstrate that the proposed Dual-FPI algorithm achieves globally optimal solutions while significantly reducing computational complexity compared with existing baseline approaches.


[86] 2601.12222

Song Aesthetics Evaluation with Multi-Stem Attention and Hierarchical Uncertainty Modeling

Music generative artificial intelligence (AI) is rapidly expanding music content, necessitating automated song aesthetics evaluation. However, existing studies largely focus on speech, audio or singing quality, leaving song aesthetics underexplored. Moreover, conventional approaches often predict a precise Mean Opinion Score (MOS) value directly, which struggles to capture the nuances of human perception in song aesthetics evaluation. This paper proposes a song-oriented aesthetics evaluation framework, featuring two novel modules: 1) Multi-Stem Attention Fusion (MSAF) builds bidirectional cross-attention between mixture-vocal and mixture-accompaniment pairs, fusing them to capture complex musical features; 2) Hierarchical Granularity-Aware Interval Aggregation (HiGIA) learns multi-granularity score probability distributions, aggregates them into a score interval, and applies a regression within the interval to produce the final score. We evaluated on two datasets of full-length songs: SongEval dataset (AI-generated) and an internal aesthetics dataset (human-created), and compared with two state-of-the-art (SOTA) models. Results show that the proposed method achieves stronger performance for multi-dimensional song aesthetics evaluation. The inference code and checkpoint are publicly available at this https URL.


[87] 2602.14913

Coverage Guarantees for Pseudo-Calibrated Conformal Prediction under Distribution Shift

Conformal prediction (CP) offers distribution-free marginal coverage guarantees under an exchangeability assumption, but these guarantees can fail if the data distribution shifts. We analyze the use of pseudo-calibration as a tool to counter this performance loss under a bounded label-conditional covariate shift model. Using tools from domain adaptation, we derive a lower bound on target coverage in terms of the source-domain loss of the classifier and a Wasserstein measure of the shift. Using this result, we provide a method to design pseudo-calibrated sets that inflate the conformal threshold by a slack parameter to keep target coverage above a prescribed level. Finally, we propose a source-tuned pseudo-calibration algorithm that interpolates between hard pseudo-labels and randomized labels as a function of classifier uncertainty. Numerical experiments show that our bounds qualitatively track pseudo-calibration behavior and that the source-tuned scheme mitigates coverage degradation under distribution shift while maintaining nontrivial prediction set sizes.


[88] 2605.07292

Variable Aerodynamic Damping Actuation via Co-Contraction: A Structural Analogy with Variable Stiffness Actuation

This work identifies a passive aerodynamic damping effect induced by co-contraction in antagonistic redundant propulsion. Complementing prior work on aerodynamic promptness, which addressed active wrench-rate authority along constant-wrench fibers, we study the passive side: the local derivative of aerodynamic force with respect to air-relative velocity at a trim. This derivative defines an incremental aerodynamic damping coefficient. We prove that it increases monotonically along constant-force fibers under a mild aerodynamic hardening condition, and derive this property from a first-order Blade Element Theory model exposing the relevant speed-inflow coupling. The resulting mechanism, Variable Aerodynamic Damping Actuation (VADA), is formulated as an antagonistic aerodynamic actuation module and allocation principle, structurally analogous to variable-stiffness actuation at the level of fiber motions and incremental impedance modulation. An impedance-form interpretation clarifies common- and differential-mode roles, while a propeller-data-based assessment using the UIUC Propeller Database shows that the identified damping has practical small-UAV magnitude, is comparable to ordinary low-speed body-drag damping, and depends strongly on low-advance-ratio thrust sensitivity.


[89] 2605.22083

RobustSpeechFlow: Learning Robust Text-to-Speech Trajectories via Augmentation-based Contrastive Flow Matching

While flow-matching text-to-speech (TTS) achieves strong zero-shot speaker similarity and naturalness, it remains susceptible to content fidelity issues, particularly skip and repeat errors from imperfect alignment. We propose RobustSpeechFlow, a training strategy that improves alignment robustness by extending contrastive flow matching with length-preserving repeat and skip latent augmentations. Requiring no external aligners or preference data, our method directly penalizes realistic failure modes and readily integrates into existing pipelines. On Seed-TTS-eval, it reduces the word error rate (WER) from 1.44 to 1.38 using only 0.06B parameters. On our ZERO500 benchmark, it delivers consistent intelligibility improvements across diverse speaker and prosody conditions; at NFE=24, it reduces English character error rate (CER) from 0.48\% to 0.35\% and Korean CER from 0.81\% to 0.57\%. Audio samples: this https URL