📄 arXiv 일일 적재 — 산업 paradigm 시그널
cs.AI · cs.LG · cs.CV · cs.RO · q-bio · physics.optics · econ.GN 등 톱다운 chain 관련. codex 로 chain 매칭·새 paradigm 자동 감지.
🔭 새 패러다임 후보 (8건 — 기존 chain 밖 신호)
StateFlow: Building, Evolving, and Accessing 3D World States for Previsualization
· 2026-08-12 · Vision
- 영상 프롬프트→3D 상태 중심 프리비주얼 전환, 새 아키텍처 신호
- 영화·게임·건축 시각화 도구체인 재편 가능성
- 생성AI 산업화가 콘텐츠 제작 워크플로우로 확장되는 사례
Learning How the World Evolves: Extrapolative Video World Models via Latent Dynamics Reasoning
· 2026-08-10 · Vision
- 픽셀 대신 잠재 동역학을 명시적 적분해 물리법칙 외삽 대폭 개선
- 파라미터 26배 절감·143배 고속화로 world model 실용화 신호
- 물리 정합 비디오 생성은 로보틱스·시뮬레이션 산업 수혜 가능
Macaron-V1: Towards Open Continual Learning with Self-Improvement and Mixture-of-LoRA
· 2026-08-10 · ML
- LoRA 혼합형 지속학습 아키텍처, 배포후 자가개선 신호
- 744B급 초대형 모델+경량 특화 어댑터 조합 확산
- AI 인프라·추론비용·에이전트SW 밸류체인에 간접 수혜
Towards Expert-level Medical AI for Real-time Video Consultations
· 2026-08-10 · AI
- 실시간 화상 진료가 전문의 수준 도달, 텍스트 넘어 오디오·비전 융합 AI
- 원격의료·헬스케어 AI 적용 확장, 진단·문진 자동화 신뢰도 상승
- 구글 제미나이 기반 멀티에이전트, 헬스케어向 AI 인프라 수혜 가능성
BDH-CQ: In-Context Learning with Recurrent Latent Reasoning
· 2026-08-10 · cs.NE
- 언어화 없는 잠재공간 순환추론 신아키텍처 등장
- 150M 소형모델로 ARC 비용효율 신기록, 추론비용 급감 시사
- 저비용 고효율 추론칩·엣지AI 수혜 가능성, 지속검증 필요
A Quantum Circuit Framework for Protein Ensemble-Level Energetics
· 2026-08-06 · Emerging Tech
- 양자회로로 단백질 앙상블 자유에너지 지형 모델링 시도
- 신약설계·펩타이드 안정성 예측에 새 계산패러다임 제시
- 아직 벤치마크 단계, 상용화까진 거리 있어 투자 임팩트 제한적
MASS: Multiplayer World Models with Authoritative Shared State
· 2026-08-06 · Vision
- 월드모델을 상태·렌더링으로 분리하는 새 아키텍처 제시
- 1024명 동시 멀티에이전트 시뮬레이션 확장성 입증
- 게임엔진식 구조로 로보틱스·시뮬레이션 산업 응용 가능성
Chained Recursive Language Models for Multi-Iteration Reasoning
· 2026-08-05 · cs.CL
- LLM 추론을 여러 독립 호출로 체인화, 컨텍스트 오염 문제 완화
- 긴 문서·다단계 추론용 새 추론시점 아키텍처, 모델 재학습 불필요
- 추론 인프라·API 호출량 증가 시사, LLM 서비스업체 수혜 가능
🔗 기존 chain 매칭 논문 30건
- biopharma-cmo How to Spend Your Oracle Budget: Practical Guidance for Protein Structure Prediction Models 2026-08-12
- physical-ai-vlm Learning-Based Behavior Planning for Automated Driving: Real-World Integration and Deployment 2026-08-12
- physical-ai-vlm SCOUT: Unlocking Enhanced Spatial Reasoning via Structured Chain-of-Thought and Multi-Objective Process Reward 2026-08-12
- physical-ai-vlm DreamFly: Causal Memory and Receding-Horizon Diffusion Planning for Aerial Vision-Language Navigation 2026-08-12
- physical-ai-vlm CausalSplat: Towards Comprehensive Hierarchical Reasoning in 3D Gaussian Splatting 2026-08-11
- physical-ai-vlm R4DSG: Relative 4D Scene Graph Memory for Object-Centric Question Answering in Long Egocentric Video 2026-08-11
- physical-ai-vlm Test-Time Self-Evolving GUI Visual Grounding via Reflection-Guided On-Policy Self-Distillation 2026-08-11
- physical-ai-vlm Surgical WAM: A World-Action Model for Data-Efficient Surgical Robot Learning 2026-08-11
- physical-ai-vlm Perception Before Supervision: Self-Contained Visual Distillation from Counterfactual Blind Spots 2026-08-10
- physical-ai-vlm RynnValue: Scaling Robotic Value Foundation Models with Temporal Distance 2026-08-10
- physical-ai-vlm CEAA: A Cognitive Embodied Agents Architecture for Interactive Computing Systems 2026-08-10
- physical-ai-vlm Agentic Harnesses: LLM-Driven Verification Layers for Robot Autonomy 2026-08-10
- physical-ai-vlm Energy-Structured Latent World Models with Neural Time Fields for Physically Constistent Open-World Motion Planning 2026-08-10
- ai-power-bottleneck GENCO - A Unified Neural Solver Embedded in a Development Framework for Steady-State Grid Analysis 2026-08-10
- biopharma-cmo THBKG: A Temporal Biomedical Knowledge Graph for Decision-Aligned Clinical Advancement Prediction 2026-08-06
- ai-substrate PLoRA: An NDP-Enhanced Pooled-Memory System for Cost-Efficient Multi-LoRA Serving 2026-08-06
- ai-substrate Automated Synthesis of Heterogeneous, Hierarchical, Scoped Coherence Protocols 2026-08-06
- physical-ai-vlm Robust-WAM: Bridging Generative Pretraining and Semantic Foresight in World-Action Models 2026-08-06
- physical-ai-vlm GAUGE: A Measurement-Grounded Benchmark for Physical Fidelity in Simulation Engines and Video World Models 2026-08-06
- physical-ai-vlm SkillMemo: Expert-guided Skill Memory Framework for Compositional Embodied Manipulation 2026-08-06
- physical-ai-vlm TRACE: Learned Proprioceptive Odometry for Legged Robots under Unreliable Contact Conditions 2026-08-06
- physical-ai-vlm Beyond Flat Policies: Hierarchical Post-Training for Embodied Agents in Robotic Manipulation 2026-08-06
- physical-ai-vlm Adaptive-WAM: Quality-Guided Early-Exit Planning from Intermediate Video-Diffusion Features 2026-08-06
- physical-ai-vlm Topometric Autonomous Vehicle Localization by Combining Visual Embeddings and Feed-Forward 3D Models 2026-08-06
- physical-ai-vlm IcFuzz: Fuzzing Isaac Sim with Semantic Stage Guidance and Multi-level Mutation 2026-08-06
- physical-ai-vlm VIDP: Variable Impedance Diffusion Policy for Compliant Robot Manipulation from Diverse Demonstrations 2026-08-06
- physical-ai-vlm GeniWorld: A Generalizable Interactive World Model for Robotic Manipulation via Visual Actions 2026-08-06
- physical-ai-vlm A Master-Salve Robot Manipulator for Needle-Based Teleoperation in MRI Chamber 2026-08-06
- physical-ai-vlm DyPES-VLA: Learning Shared Dynamics Priors and Embodiment-Specific Control for Cross-Embodiment Manipulation 2026-08-06
- physical-ai-vlm $ω$-0: A Latent Predictive World Action Model for Concurrent Humanoid Loco-Manipulation 2026-08-06
Organizational Technology Ladders: Remote Work and Generative AI Adoption
This study proposes that firms move along an "organizational technology ladder": adopting one technology transforms hiring and work processes and builds skills and organizational capital that change the cost of adopting subsequent technologies. I study how firms' adoption of remote work technology d…
Gregor Schubert · 📄 PDF
Robustness over efficiency in climate coalitions: a bistable model and a map of architectures
Designs for international climate cooperation face a trade-off between allocative efficiency and robustness to the erosion of institutions by defection, renegotiation, and political turnover. We formalize this trade-off in a stylized coalition-formation game in which membership is driven by two mark…
Juergen Renn · 📄 PDF
Oil price shocks reveal unequal capacities for mobility adaptation
Urban decarbonization often raises the cost of travel, yet which neighbourhoods can adapt remains largely invisible under normal conditions. We leverage the 2026 US-Iran oil shock as a natural experiment, applying a hierarchical panel regression discontinuity design to 1.7 trillion point-of-interest…
Zihao Zhang, Yuanbo Zhang, Xiaolei Ma, Yuan Liao · 📄 PDF
Temporal Coupled-Mode Theory for Quantum-Confined Stark Tuning of Intersubband-Polaritonic Metasurfaces
Electrical tuning of intersubband-polaritonic metasurfaces requires a field-consistent connection between quantum-confined Stark effect (QCSE) modified transitions and measurable optical scattering. We derive a two-oscillator, one-port temporal coupled-mode theory (TCMT) and reduce the complex far-f…
Inyong Hwang · 📄 PDF
Unifying Physical Backpropagation
Physical computing systems exploit device dynamics for computation, but their gradient-based optimization is challenging: backpropagation through a digital twin suffers from model-reality gap. On-device gradient computation could resolve this issue, and a handful of theoretical and experimental stud…
Cyrill Bösch, Yigithan Gediz, Hakan Türeci · 📄 PDF
Towards Terabit/$λ$/s Multidimensional Silicon Photonic Engine
Increasing artificial intelligence (AI) workloads drive co-packaged optics (CPO), which integrates optical engines with electronic components. Optical interconnects can extend transmission distances and reduce latency, allowing distributed clusters in AI factories to operate as a unified computation…
Hao Chen, Zengqi Chen, Wu Zhou, Kaihang Lu, Mingyuan Zhang, Yuxiang Yin, Yiou Cui, Chaoran Huang, Pui-In Mak, Yeyu Tong · 📄 PDF
Near-Unity Excitation and Radiative Efficiencies in Electroluminescence Without External Carrier Injection
Electroluminescence occurring without external charge injection is typically characterized by weak emission and excessive driving voltage, due to low excitation and radiative recombination efficiencies. Here, we demonstrate non-injecting electroluminescence (NI-EL) that challenges this conventional …
Rui Li, Xinrui Li, Jiachen Xie, Pingjun Xu, Wen Li, Congcong Liu, Zheng Ge, Longjia Wu, Xingtong Chen, Zheming Liu, Song… · 📄 PDF
Experimental quantum telecloning across silicon photonic chips
Telecloning -- the combination of quantum teleportation and cloning -- offers a powerful mechanism to disseminate unknown quantum states to multiple spatially separated recipients with optimal fidelity. Despite its conceptual importance for quantum networks, an experimental demonstration of symmetri…
Zicong Wen, Kai Wang, Bochi Wu, Leizhen Chen, Yan-Qing Lu, Shining Zhu, Xiao-Song Ma · 📄 PDF
Single-Cycle Pulses, Transparent Conducting Oxides, Optical Nonlinearity, Quantum Coherence, Thermalization
We present a first-principles study of the nonlinear optical response of transparent conducting oxides at the nanoscale due to excitation by intense, extremely short pulses based on a density matrix framework. We identify a strong ($O(1)$) thermal nonlinearity, which is complemented with stimulated …
Ieng-Wai Un, Subhajit Sarkar, Yonatan Sivan · 📄 PDF
Aggregation-engineered loss-tolerant strong coupling in metallic microcavities
Room-temperature strong coupling in organic microcavities is usually achieved by combining high-quality optical resonators with highly ordered excitonic media, a requirement that limits scalability and processing flexibility. Here we show that this constraint can be relaxed by using molecular aggreg…
Andrea Betti, Eleonora Cara, Giulia Serrano, Lorenzo Poggini, Alessia Valzelli, Natascia De Leo, Paolo Bartolini, Andrea… · 📄 PDF
Spatial coherence enabled sensorless adaptive optical imaging
Optical aberrations degrade imaging performance in label-free microscopy, where the absence of a guide star often necessitates sensorless adaptive optics (AO). Conventional sensorless AO approaches rely on image-quality metrics whose optimal choice depends on both the specimen and the imaging modali…
Pranay Mohta, Shaurya Aarav, Hugo Defienne, Anand K. Jha · 📄 PDF
Two-Stage Deformable-Convolutional Inverse Design of Nanophotonic Absorbers from Optical Spectra
Data-driven inverse design enables efficient generation of nanophotonic structures with prescribed optical responses, but spectrum-to-geometry mapping remains challenging due to non-uniqueness and fine geometric features. This work presents a two-stage deformable-convolutional framework for reconstr…
Waleed Waseer, Muhammad Shahid Jabbar, Muhammad Sohail Ibrahim, Shujaat Khan · 📄 PDF
Programmable vs. Static Beam Shaping in Ultrafast Laser Micromachining: A Critical Review
Beam shaping has become one of the principal determinants of throughput, precision, and process robustness in ultrafast laser micromachining. Despite this, the field is still largely interpreted through a historical distinction between programmable and static optical elements, a framework that incre…
Krystof Kobliha, Peter Hauschwitz · 📄 PDF
Metastable soliton necklaces confined by the boundary of a flattop region
We present quasistationary ring-shaped soliton necklaces in a two-component envelope propagating in a medium with competing cubic-quintic nonlinearity. Metastable propagation of soliton necklaces results from a balance of repulsion between adjacent out-of-phase solitons in one component and confinem…
Dmitry A. Zezyulin · 📄 PDF
State-resolved quantum transport of vortex electrons in accelerators
We develop a density-matrix theory of vortex-electron transport in accelerator lattices. In a periodic round lattice, the protected object is a Lewis-Floquet orbital angular momentum (OAM) invariant rather than the instantaneous kinetic OAM. Ideal transport has a metaplectic lift; stochastic imperfe…
S. S. Baturin · 📄 PDF
Bridging the gap: Using deep learning to reconstruct noise-reduced super-resolved OCT images from gapped spectra
Fourier-domain (FD) optical coherence tomography (OCT) depends on broadband sources to maximize axial resolution and image quality. However, these lasers significantly drive device cost or may be unavailable at desired wavelength and bandwidth ranges. A potential solution lies in integrating multipl…
Jonas Nienhaus, Thomas Schlegl, Wolfgang Drexler, Tilman Schmoll, Rainer A. Leitgeb · 📄 PDF
Spatial and energetic correlations of ultrashort two electron pulses
When two electrons are emitted from a metallic needle tip into a nanometric volume on femtosecond timescales, strong Coulomb correlations arise. While longitudinal correlations, manifested as energy shifts, have been observed both from bare tips and in electron microscopes, transverse correlations r…
Stefan Meier, Jonas Heimerl, Felix López Hoffmann, Tobias Volk, Marco Knipfer, Peter Hommelhoff · 📄 PDF
Beam Routing through Excitons in Transition Metal Dichalcogenide Monolayers
Routing light at the nanoscale typically relies on nanostructured surfaces to imprint directionality on the emission. Using low-temperature, angle-resolved cathodoluminescence spectroscopy, we show that the intrinsic excitonic transitions of a semiconductor can themselves produce routed emission. We…
Yonas Lebsir, Jacob Terndrup Heiden, Jorge Barcia Rodríguez, Maria Papadopoulou, Kenji Watanabe, Takashi Taniguchi, N. A… · 📄 PDF
Twist-Reconfigurable van der Waals Moiré Photonic Crystals
Moiré photonics has emerged as a fascinating concept to design and in situ control of the optical bands. Moiré enabled light localisation arises from the relative twist between periodic layers, rather than from fixed, pre-fabricated cavity features. So far, however, the realisation of practical moir…
Hugo Quard, Jiyun Kim, Anastasiia Zalogina, Xuerong Hu, Evan Williams, Oscar J. Palma Chaundler, Owen R. Wolley, Alexand… · 📄 PDF
Goos-Hanchen-Shift Photonic Sensor for Nanometer-Scale Delayering and Tamper Detection in Semiconductor Packages
We propose a co-packaged photonic tamper sensor that detects progressive delayering and localized drilling through changes in the Goos-Hanchen (GH) shift of a reflected optical beam. Frustrated total internal reflection (FTIR) couples the beam into a high-index sensing layer, where its transverse-wa…
Mia Mohammad Shoaib Hasan, Mohamed Elkabbash · 📄 PDF
SVPLEX: A Nextflow Pipeline for Cohort-level Structural Variant Calling
SVPLEX is a Nextflow pipeline for cohort-level structural variant detection from short-read whole-genome sequencing data. The pipeline implements six different structural variant callers with different strengths and weaknesses, integrating different levels of evidence for SVs, and generates a merged…
Jacob E. Munro, Mark F. Bennett, Melanie Bahlo · 📄 PDF
Uni-SFU: Algorithm-HW Co-Design for Universal SFUs via Mixed-Degree Piecewise Approximation
Nonlinear activation functions are essential to modern deep neural networks (DNNs), but their hardware evaluation places significant pressure on the special-function units (SFUs) of GPUs and custom accelerators. Therefore, piecewise polynomial approximations are commonly used within allowed error bo…
Miao Sun, Yucheng Huang, Mingcong Cao, Jaehyun Park, Partha Pratim Pande, Umit Y. Ogras · 📄 PDF
A Full-Stack Characterization of High-Bandwidth Flash for KV-Centric LLM Serving
High-Bandwidth Flash (HBF) stacks NAND behind a wide, package-local interface, giving flash-scale capacity with far better read latency and bandwidth than an SSD. This makes it tempting to keep an SSD-style Mooncake KV-offloading stack and swap only the backing tier for HBF. We test that substitutio…
Zhuoran Li, Zhuohang Bian, Xin Huang, Yibo Zhao, Guangyu Sun, Youwei Zhuo · 📄 PDF
APEX: Adaptive Expert Prefetching for Memory-Efficient Edge MoE Inference
Mixture-of-Experts (MoE) models are attractive for edge deployment because they provide high model capacity while activating only a small subset of parameters per token, improving compute efficiency. However, MoE inference at the edge is fundamentally limited by memory. Expert parameters are large a…
Alish Kanani, Layan Badawi, Umit Y. Ogras · 📄 PDF
Spec Sheets Are Not Kernels: An ISA- and Source-Level Audit of INT8 Availability on NVIDIA Blackwell Ultra
NVIDIA's published specifications give the Blackwell Ultra GPU (B300) a dense-compute ratio of roughly 30:1 between FP8 and INT8 tensor-core throughput; its predecessors, H200 and B200, both provide 1:1. We audit what this deprioritization means in practice by tracing INT8 W8A8 support through four …
Teng-Ruei Chen · 📄 PDF
Do Not Let CNOTs Overwhelm the Decoder: Scheduling Transversal Gates for Fast FTQC
Transversal CNOT (TCNOT) gates can accelerate fault-tolerant quantum computation (FTQC) in the surface code by reducing the number of syndrome extraction rounds required between logical operations from $O(d)$ to $O(1)$. This is particularly attractive for quantum platforms with long-range connectivi…
Shota Ikari, Yuga Hirai, Yasunari Suzuki, Hiroshi Nakamura, Yosuke Ueno · 📄 PDF
NITRO: High-Performance 3D NAND Flash-Based In-Storage Computing with Enhanced Activation Dataflow
In-storage computing (ISC) is considered a next-generation memory architecture for its ability to relieve the data bottleneck between the host and the memory. While the required resources of large language models (LLMs) have increased significantly in recent years, the memory density has not scaled …
Sanghun Shin, Sangyeon Kim, Gisan Ji, Sungju Ryu · 📄 PDF
Keep the Future, Drop the Rollout: RIFT for World Action Models
World action models (WAMs) condition robot actions on predicted futures, but iterative video rollout increases deployment latency. We ask whether action generation requires the evolving rollout trajectory or only its future representation. Across four WAMs on all 40 LIBERO tasks, paired closed-loop …
Chushan Zhang, Jinguang Tong, Xuesong Li, Yikai Wang, Hongdong Li · 📄 PDF
Repurposing RGB-based Foundation Model for Depth Estimation on Thermal Images Using Hierarchical Supervision
Depth estimation from thermal images is highly valuable for robotic applications in adverse conditions, such as nighttime and rainy weather. Recent studies have sought to transfer knowledge from RGB-based foundation models to thermal modalities, yet the rich hierarchical representations these models…
Jie Hong, Tingtian Li, Xuesong Li, Xiao Li · 📄 PDF
RoadWeaver: Large-Scale Lane-Level HD Map Generation from Scratch for Autonomous Driving Simulation
Autonomous driving simulation requires diverse and scalable lane-level HD maps to support long-horizon evaluation across complex road networks. Existing approaches either rely on handcrafted or reconstructed real-world maps, which limits scalability, or generate only local road structures rather tha…
Yueyuan Li, Zexi Chen, Weijie Xi, Mingyang Jiang, Songan Zhang, Hanyang Zhuang, Ming Yang · 📄 PDF
Video2Track: From Real-World Interaction Videos to Steerable Adversarial Closed-Track Testing for Automated Driving Systems
Closed-track testing plays a fundamental role in the verification and validation of automated driving systems (ADS), particularly for safety-critical scenarios, by enabling reproducible evaluation under controlled conditions. However, most existing approaches still rely on standardized protocols or …
Mengjie Tian, Xinrui Zhang, Tianyu Li, Peizhi Zhang, Guirong Zhou, Haojie Feng, Junpeng Huang, Qixiang Zhang, Lu Xiong · 📄 PDF
IoT-Enabled Autonomous Maritime Navigation in Smart Ports: A Curriculum-Guided Shared Policy Learning Framework
As smart port infrastructures increasingly rely on autonomous maritime devices enabled by the Internet of Things (IoT), ensuring reliable onboard navigation intelligence has become a critical challenge for safe and scalable operations in congested waterways. This paper investigates onboard autonomou…
Yuqing Lin, Rangya Zhang, Kum Fai Yuen · 📄 PDF
Energy-Aware Wind-Resilient Routing for Truck-Assisted Multi-UAV Delivery under Wind Uncertainty
Energy feasibility under wind uncertainty is a critical safety issue for low-altitude air-ground delivery. In truck-UAV systems, UAVs complete assigned deliveries and safely return to a mobile truck or depot, while wind-induced propulsion costs vary online and are only partially observable. Existing…
Tianshun Li, Yanggang Sheng, Hongliang Lu, Zhongzhen Wang, Haoang Li, Xinhu Zheng · 📄 PDF
StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models
Vision-Language-Action (VLA) models can follow instructions and manipulate objects, but their performance often collapses out of distribution (OOD), when the scene, viewpoint, or object differs from training. Adapting to each new situation typically requires collecting more data and fine-tuning. We …
Siyu Xu, Yunke Wang, Zijian Wang, Dihao Zhu, Chenghao Xia, Chengbin Du, Daochang Liu, Tao Huang, Chang Xu · 📄 PDF
ContactIPM: A Structure-Exploiting Interior-Point Solver for Contact-Implicit Trajectory Optimization
Contact-implicit trajectory optimization avoids prescribing contact sequences, but yields mathematical programs with complementarity constraints (MPCCs) whose degeneracy challenges conventional primal--dual solvers. Existing contact-specific methods improve robustness to this degeneracy but do not l…
Yucheng Chen · 📄 PDF
G0.5: One Autoregressive Stream for Robot Reasoning and Action
The prevailing recipe for Vision-Language-Action (VLA) models couples a pretrained VLM with a separately trained flow-matching action expert. This makes the VLM a context encoder rather than a decision-maker. We introduce G0.5, a pretrained autoregressive VLA in which a single transformer decoder em…
Yicheng Liu, Zibin Dong, Baijun Ye, Tianyuan Yuan, Tao Jiang, Anqi Yang, Shicheng Cao, Haonan Liu, Yue Sun, Zihan Guo, X… · 📄 PDF
Policy-Induced Hand Priors in Humanoid Dual-Arm Manipulation: Diagnosing and Mitigating Initial-Pose Dependence
Vision-language-action (VLA) policies are expected to operate robustly across variations in the robot's initial configuration, yet aggregate task success can conceal pose-specific failures and inappropriate hand selection. This work investigates initial-pose dependence in VLA-based humanoid dual-arm…
Chaeyeon Jung, Juyoun Park · 📄 PDF
Enhancing Visual Domain Robustness in Behaviour Cloning via Saliency-Guided Augmentation
In vision-based behavior cloning (BC), conventional image augmentations such as Random Crop and Color Jitter often fall short under substantial visual domain shifts, including changes in shadows, distractors, and backgrounds. Superimposition-based augmentations, which blend in-domain and out-of-doma…
Zheyu Zhuang, Ruiyu Wang, Nils Ingelhag, Ville Kyrki, Danica Kragic · 📄 PDF
D3D-GEN: Robot-Aware Domain-Grounded Interactive 3D World Generation for Social Robotics
Training and validation of Embodied AI for social navigation critically depends on realistic simulation environments, yet many current approaches fail to find a balance between realism and simulability. We propose D3D-GEN, a novel world generation system that combines a domain agent with a retrieval…
Anh Duc Do, Volodymyr Scherbyna, Tai Duc Nguyen, Spaarsh Thakkar, Zhengcheng Shen, Teham Buiyan, Archan Misra, Linh Käst… · 📄 PDF
Scalable Multi-Agent Maze Traversal with Local Communication
Cave networks, pipe systems, and similar maze-like environments pose significant challenges for multi-agent navigation in unknown settings with limited communication. We propose a distributed algorithm that enables agents to collectively traverse an unknown, possibly cyclic graph. Agents enter seque…
Julian Rau, Jahir Argote-Gerald, Grace McFassel, Genki Miyauchi, Paul Trodden, Roderich Groß · 📄 PDF
DaViNCi: A Dataset Towards Outdoor Vision-and-Language Navigation with Continuous Actions and Dynamic Elements
Vision-and-Language Navigation (VLN) has progressively expanded from indoor to outdoor environments. However, existing outdoor VLN datasets still rely on fixed discrete topological graphs for construction. It fails to align with the rapidly changing real-world outdoor environments and impedes the si…
Zihao Xie, Pingrui Lai, Yitong Wu, Hua Yang · 📄 PDF
Learning Loco-Manipulation From SMPC Demonstrations With Sparse Offline-to-Online RL
Integrating locomotion and manipulation is essential for robot autonomy, but scaling standard Reinforcement Learning (RL) to complex tasks is severely bottlenecked by the slow, manual process of dense reward shaping. To bypass this limitation, we leverage Sample-based Model Predictive Control (SMPC)…
Martin Schuck, Maks Sorokin, Simone Manni, Duy Ta, Angela P. Schoellig, Marco Hutter, Simon Le Cleac'H, Jan Brüdigam · 📄 PDF
HandEdit: A Unified Benchmark for Egocentric Human-to-Robot Dexterous Hand Image Editing
Robotic manipulation with dexterous hands is a cornerstone of Embodied AI, yet its progress is stifled by the high cost of collecting embodiment-aware teleoperation data. While abundant egocentric videos of human hands offer a scalable alternative, the profound discrepancies in appearance, articulat…
Zhenjie Yang, Xingyu Jiao, Guopeng Zhong, Shuzhe Yang, Shi Che, Chao Wu, Chenyu Jiang, Dongjie Zhang, Yideng Zhang, Zhen… · 📄 PDF
GenFAR: A generalized representation of brain structure, derived from 49,246 multi-cohort MRIs via deep learning
Deep learning models for neuroimaging have largely been developed for individual tasks, limiting knowledge transfer across applications. Here we introduce GenFAR, a modular deep learning framework that learns general, clinically informed features from brain MRIs. We trained this modular architecture…
Vishnu M. Bashyam, Guray Erus, Junhao Wen, Pratik Chaudhari, Randa Melhem, Sindhuja Govindarajan Tirumalai, Gareth Harma… · 📄 PDF
HSTGFormer: Hyper Spatial-Temporal Graph Transformer for 3D Human Pose Estimation
Transformer-based methods have achieved strong performance in monocular 3D human pose estimation, but most existing approaches organise spatial and temporal reasoning as separate stages, which may weaken unified spatial-temporal interdependencies inherent in human motion and compress frame-level str…
Ruochen Li, Shuang Chen, Wenke E, Farshad Arvin, Amir Atapour-Abarghouei · 📄 PDF
M-Net: Integrating Spectral Features and Physical Field Operators into Deep Learning for Medical Image Segmentation
Purpose: Deep learning-based medical image segmentation has achieved remarkable success, yet purely data-driven approaches often fail to exploit the rich mathematical structure inherent in medical images. We investigate whether explicit mathematical inductive biases, specifically matrix spectral ana…
Jing Zhu, Ye Wang, Fumin Wang · 📄 PDF
GeoFlow: Efficient Driving Video Generation via Geometry-Aligned Priors
Generative models like Diffusion Models and Flow Matching have demonstrated remarkable capabilities in synthesizing high-fidelity driving videos, but are severely constrained by high inference latency due to the requirement of extensive sampling steps. We argue that this inefficiency stems from the …
Jiazheng Liu, Hang Li, Jiawei Zhang, Jiahe Li, Xiaohan Yu, Shengyin Fan, Jin Zheng, Xiao Bai · 📄 PDF
Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction
While Multimodal Large Language Models (MLLMs) have achieved remarkable progress, visual understanding and generation are typically treated as divergent objectives. Existing unified frameworks often rely on discrete visual tokenization or diffusion objectives whose generative targets differ from the…
Zhongbin Guo, Jiahao Xie, Dongling Xiao, Qianle Wang, Ruiqi Lu, Xiaomin He, Wanxuan Sun, Cheng Yang · 📄 PDF
ScaleVid: Geometry-Aware Video Object Scaling with Mesh-Free Inference
Geometry-aware video object scaling aims to anisotropically resize the object along object-centric axes while preserving geometric plausibility, temporal coherence, and background consistency. Existing text-guided methods mainly operate in the 2D image plane, while depth-guided approaches provide co…
Youze Huang, Penghui Ruan, Bojia Zi, Xianbiao Qi, Shihao Zhao, Rong Xiao · 📄 PDF
Automated Borehole Core Analysis with Report-Derived Weak Labels and Supervised Crack Segmentation
Borehole archives commonly contain core tray photographs and corresponding digital log reports, but no native pixel-level crack annotations. We investigate two complementary approaches for extracting defect-spacing information from these archives. First, structured spacing categories recovered from …
Usama Imdad, Ali Khan, Luke Lu, Zubair Khalid, Arif Mahmood · 📄 PDF
XYZFlow:Scaling Multi dimensional Shortcut Flows for Efficient Generative Modeling
High-fidelity image generation faces a trade-off between speed and quality. Diffusion models produce strong visuals but require costly iterative sampling. Existing efficient methods mainly distill pretrained models into few-step samplers, a challenging process that depends heavily on teacher-model q…
Jinxiu Liu, Xuanming Liu, Kangfu Mei, Yandong Wen, Weiyang Liu · 📄 PDF
Curvature-Aware Zeroth-Order Optimization for Memory-Efficient Test-Time Adaptation
Test-time adaptation (TTA) aims to enhance the cross-domain performance of pre-trained models by adapting to unlabeled test data. While most existing TTA methods rely on backpropagation (BP) for finetuning, BP-free methods such as zeroth-order (ZO) methods are more desired in practical on-device sce…
Junming Zhang, Shuyu Yin, Peilin Liu, Rendong Ying, Fei Wen · 📄 PDF
AVA-Encoder: Towards Agent-Native Video Representation Learning
Creative agents still lack an effective way to learn from high-quality human films, limiting their ability to produce cinematic-grade videos. A key challenge is the absence of a structured video representation that is both faithful to film content and directly usable for agentic reasoning and manipu…
Chuyue Li, Jinpeng Yu, Haozhe Wang, Tian Xueyun, Zhijing Zhang, Bingnan Li, Shuqi Gu, Kan Ren, Jiaming Liu, Ruihua Hua · 📄 PDF
StateFlow: Building, Evolving, and Accessing 3D World States for Previsualization
Previsualization is an intermediate layer between ideas and production in film, games, architecture, and urban design. It lets creators iteratively refine scenes, actions, cameras, and spatial-temporal dynamics. Yet existing generative methods rely on simple prompts to jointly control all of these f…
Yuyang Yin, Zixiang Li, Longxuan Deng, Hongkai Li, Shifang Zhao, Junnan Liu, Weirong Huang, Mengyu Wang, Tianxiao Fu, Yi… · 📄 PDF
Beyond Parameter Space: NTK-Guided Personalized Aggregation for Robust Federated Learning
Federated learning (FL) enables collaborative model training across distributed clients while keeping data local. A central challenge is determining which client updates are beneficial for aggregation with respect to each client's target domain. Existing methods typically address this problem in par…
Mirko Konstantin, Stefan Zachow, Anirban Mukhopadhyay · 📄 PDF
The Advective Fisher-Rao Geometry of Deterministic Measure Transport
A novel advective Fisher-Rao metric is introduced for optimization tasks on paths of probability measures governed by the continuity equation. This metric is shown to lead to optimal descent directions. It is then shown that this metric arises naturally from three different perspectives: As the resc…
Benjamin Gess, Johannes Müller · 📄 PDF
Attractor Image-Based Deep Learning of Arterial Pulse Waves for Age Classification
Arterial pulse waveform morphology evolves with age, reflecting structural and functional changes in the cardiovascular system. Thus, vascular age is a valuable surrogate marker of cardiovascular health, and premature vascular ageing can indicate increased disease risk. Pulse wave analysis could sup…
Sara Vardanega, Patrick Segers, Philip Aston, Ernst Rietzschel, Jordi Alastruey, Manasi Nandi · 📄 PDF
Adversarial Resilience of Poisson-Process Submodular Maximization over Matroids: From Robust Offline Optimization to Full-Bandit Learning
We study nonnegative submodular maximization subject to a general matroid when the offline algorithm is given an arbitrary controlled value oracle. Our main result is an adversarial resilience theorem for the Spiteful Greedy Swap Poisson Process (SGS-Poisson): without modifying its Poisson intensity…
Vaneet Aggarwal · 📄 PDF
A corpus-specific clinical RAG system matches or outperforms newer frontier LLMs on HealthBench
General-purpose large language models (LLMs) have recently been reported to match or exceed specialized clinical AI tools on medical benchmarks, but such comparisons draw on a narrow set of systems and on benchmarks developed largely in high-income settings. We evaluate VITA, a retrieval-augmented g…
Praveen Reddy, Charuta Mandke, Suvrankar Datta, Sarah Khan, Siddharth Reddy Anthireddy, Shitij Arora, Vishal Singh · 📄 PDF
FQTree: Fine-grained Quantization and Hardware Generation of Boosted Decision Trees
Boosted decision trees (BDTs) are widely used in latency-critical applications, but efficient hardware deployment remains challenging. Existing designs often rely on uniform or manually tuned fixed-point formats, which can introduce unnecessary hardware cost or accuracy loss. This work presents the …
Zhiqiang Que, Chang Sun, Haiyang Wang, Dinesh Pamunuwa, Roshan Weerasekera, Qijia Tang, Bakhtiar Zadeh, Wayne Luk, Maria… · 📄 PDF
ADEPT: A Unified Framework for Deep Learning Test Adequacy
Over the past decade, many test adequacy metrics have been proposed for deep learning that characterize test dataset adequacy from different perspectives, e.g., neuron activation behavior, latent feature coverage, decision-boundary exploration, etc. However, these metrics are typically released as i…
Yidi Kao, Shawn Burnham, Tommi Rose Fahy, Ali Ghanbari · 📄 PDF
Autonomous Telerehabilitation via Skeletal Motion Prediction and Joint-Level Performance Assessment
Autonomous rehabilitation systems must not only recognize human motion but also provide structured feedback to support users without continuous therapist supervision. This paper presents a telerehabilitation pipeline that integrates skeleton-based exercise quality assessment and short-term motion pr…
Lara Pereira, João Ruivo Paulo, Pedro Santos, Paulo Peixoto · 📄 PDF
How to Spend Your Oracle Budget: Practical Guidance for Protein Structure Prediction Models
Foundation models for protein structure prediction remain unreliable on certain targets. External oracles can flag and correct these failures, but biological oracles are expensive, making oracle budget a critical constraint. Existing guidance methods, such as FK-steering, DPO, and Best K-of-N sampli…
Aleksandra Kalisz, Jack Simons, Krisztina Sinkovics, Noam Ghenassia, Shikha Surana, Henry Moss, Paul Duckworth · 📄 PDF
HYDRA: Hyperbolic Dynamic Representation Architecture for Kolmogorov-Arnold Networks
Kolmogorov-Arnold Networks (KANs) enhance nonlinear function approximation by replacing scalar weights with learnable univariate functions. However, assigning an independent function to every connection results in substantial parameter redundancy, limiting their scalability and efficiency. To reduce…
Zhao Su, Yuxin Xia, Haoran Li, Jun Shen, Qi Zhu, Qingguo Zhou, Binbin Yong · 📄 PDF
Learning-Based Behavior Planning for Automated Driving: Real-World Integration and Deployment
Recent research in machine and deep learning has shown the potential of learningbased motion planning approaches to improve the driving behavior of automated vehicles, especially in complex environments. However, their complex nature and lack of transparency can hinder explainability and trustworthi…
Jean-Pierre Busch, Guido Linden, Jan Bergmann, Lutz Eckstein · 📄 PDF
ScreenShot: A Foundation Model for Few-Shot Combination Drug Screening
Treating patients with combinations of drugs reduces the risk of resistance to any individual drug. Finding effective combinations is difficult because the large search space makes combinatorial screens prohibitively expensive, time consuming, and often technically infeasible. Predictive models can …
Antoine de Mathelin, Christopher Tosh, Wesley Tansey · 📄 PDF
An Efficient Near-Optimal Algorithm for Adversarial $m$-Set Bandits
We study adversarial combinatorial bandits with $m$-set actions, where at each round the learner selects $m$ out of $d$ items and observes only the aggregate loss of the selected items. The resulting action set contains $K=\binom{d}{m}$ elements and can therefore be exponentially large. Nevertheless…
Francesco Bacchiocchi, Tommaso Cesari, Roberto Colomboni · 📄 PDF
Regime-Gated Residual Mixture-of-Experts for Cross-Sectional Volatility Forecasting
Financial volatility is regime dependent, yet incorporating regime information into neural networks can also destabilize training. This paper asks where such information should enter a neural cross-sectional volatility forecasting model. We study five-day realized-volatility forecasts for 1,027 U.S.…
Junyi Ye, Gargi Vijay Borde · 📄 PDF
Calibration Bets on the Past: Post-Training Quantization for Financial Time-Series Forecasting
Financial forecasting models are typically developed in full precision, yet production deployment often requires low-precision inference to reduce memory and computational cost. Post-training quantization (PTQ) enables such deployment without retraining. However, reliable activation quantization req…
Junyi Ye, Ivy Gateri Wanjiku · 📄 PDF
Earth observation embeddings are effective sub-grid descriptors for probabilistic weather downscaling
Global weather reanalyses and forecasts resolve the evolving atmospheric state on coarse grids, but site-specific applications require predictions at arbitrary locations where near-surface conditions also depend on unresolved terrain and land-surface properties. Existing probabilistic downscalers ad…
Pedro Sousa, Will Tebbutt, Sadiq Jaffer, Robin Young, Anil Madhavapeddy, Richard E. Turner · 📄 PDF
A Framework for Designing Reward Functions: From Objectives to Features to Human-Aligned Reward Functions
We present a formal process to enable non-experts to instantiate and iterate on human-aligned reward functions, i.e. reward functions that adhere to a given preference ordering over trajectories. Given a task described in natural language, our process produces a linear reward function in three steps…
Di Yang Shi, W. Bradley Knox · 📄 PDF
Information Abundance Paradox: Long-Context Training Undermines Parametric Knowledge
Large language models are increasingly trained and deployed with long contexts that span documents, code repositories, and interaction histories. This scaling reflects the implicit assumption that training on longer contexts will only help the model by exposing it to richer evidence. We challenge th…
Arda Uzunoglu, Benjamin van Durme, Daniel Khashabi · 📄 PDF
SCOUT: Unlocking Enhanced Spatial Reasoning via Structured Chain-of-Thought and Multi-Objective Process Reward
Existing Vision-Language Models (VLMs) exhibits a critical bottleneck in robust spatial reasoning. Recent reinforcement learning (RL) methods aim to close this gap with verifiable outcomes, yet they suffer from poor credit assignment across intermediate reasoning steps. Concurrently, structured reas…
Zile Zhou, Huining Yuan, Weichen Zhang, Xinlei Chen, Xiao-ping Zhang · 📄 PDF
Domain-Aware Lightweight Spectral-Grouped Convolutions for Hyperspectral Fish Freshness Classification
Hyperspectral imaging (HSI) offers nondestructive assessment of fish freshness by detecting biochemical alterations across spectral bands. However, conventional deep learning approaches do not fully address the particular characteristics of HSI data, such as spectral dominance over spatial textures,…
Kazi Nabiul Alam, Pooneh Bagheri Zadeh, Akbar Sheikh-Akbari · 📄 PDF
Few-Shot Ordinal Learning for Day-Wise Freshness Estimation with Hyperspectral Fish Images
Non-destructive food quality assessment has increasingly benefited from hyperspectral imaging (HSI), which captures spectral signatures linked to biochemical changes during storage. Estimating day-wise freshness, however, remains challenging owing to strong inter-fillet variability and scarce labell…
Kazi Nabiul Alam, Pooneh Bagheri Zadeh, Akbar Sheikh-Akbari · 📄 PDF
How Organizations Use AI: Evidence from ChatGPT
We study how organizations use frontier generative AI by linking ChatGPT Enterprise account records to usage, worker roles, task classifications, and public-company financial data through March 2026. These linked data enable a privacy-preserving analysis of adoption, worker roles, and message-level …
Aaron Chatterji, David Holtz, Neel Rakholia, Prasanna Tambe, Gawesha Weeratunga · 📄 PDF
HAMP-LIC: Hessian-Aware Mixed-Precision Post-Training Quantization for Learned Image Compression
Use this plain-text version for the arXiv abstract field: Learned image compression (LIC) models achieve strong rate-distortion performance but are hindered by high computational complexity and encoding-decoding mismatches across heterogeneous hardware platforms. Uniform fixed-precision quantization…
Yuefeng Zhang · 📄 PDF
VICBench: A Multi-Language Benchmark for Code Vulnerability Detection
Evaluating security vulnerability detection tools requires benchmark datasets with vulnerability-inducing commits (VICs) - the commits that first introduce vulnerabilities into codebases. VICs are essential for determining the full range of vulnerable software versions. Existing vulnerability datase…
Jin Lu, Xuening Han, Yang Zhong, Lin Tan, Kevin Luo, Andrew Gacek, Neha Rungta · 📄 PDF
An Agentic Workflow for Legacy HPC Modernization: Converting the Two-Electron-Integral Core of GAMESS
Modernizing legacy Fortran is a problem of volume: the transformations are individually routine, but the codebases can be enormous, and across much of computational science the work simply goes undone. We propose an agentic workflow that takes this work on at production scale, and we set out to meas…
Yuzhong Shen, Masha Sosonkina, Peng Xu, Mark S. Gordon · 📄 PDF
One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL
Multi-agent reinforcement learning for human-AI interaction typically relies on a single large language model to simulate user behavior. We show that this approach systematically fails to generalize, and trace the failure to simulator collapse: because the simulator LLM is mode-collapsed, an LLM pol…
Simon Yu, Nicholas Tomlin, Marwa Abdulhai, Ximing Lu, Derek Chong, Abe Hou, Dilara Soylu, Sergey Levine, Christopher D. … · 📄 PDF
Diagram-MMU: A Multi-Modal Benchmark for Scientific Diagrams
Multimodal Large Language Models (MLLMs) have been growing the capability for scientific writing and collaboration. For example, OpenAI Prism is a free workspace for scientific writing and collaboration. One important feature in Prism is turning scientific diagrams directly into LaTeX TikZ code. In …
Weihao Bo, Shan Zhang, Yanpeng Sun, Jie Liu, Yongke Yao, Jinhao Du, Wei He, Kai Zou, Zechao Li, Jingdong Wang · 📄 PDF
Convergent Detour Hijacking: Task-Preserving Resource Amplification in Skill-Based LLM Agents
LLM agents increasingly rely on third-party skills, using natural-language descriptions for selection and instruction bodies for planning. This progressive-disclosure design exposes two sequential control points to untrusted publishers: a static skill may steer an otherwise correct task onto an unne…
Junliang Liu, Ruoyu Li, Wenxin Tang, Jingyu Xiao, Zhenyu Liu, Jingheng Xu, Laizhong Cui · 📄 PDF
A Neighborhood Attention Transformer Network for Enhanced 3D Segmentation of the Left Anterior Descending Artery
Background: Accurate segmentation of the Left Anterior Descending (LAD) artery in 3D free-breathing, non-contrast CT is critical for cardiac dose sparing in thoracic radiotherapy. The LAD is extremely small, has poor soft-tissue contrast, and varies substantially across patients; even manual contour…
Rafi Ibn Sultan, Chengyin Li, Yiannos Demetriou, Ahmed I. Ghanem, Joshua P. Kim, Justine Cunningham, Hassan Bagher-Ebadi… · 📄 PDF
Structural Silence: When AI Infrastructure Fails Speakers of Underrepresented Languages
Artificial intelligence tools for education and language support are increasingly framed as scalable responses to access gaps in under-resourced communities. Yet the infrastructure underlying these tools, including training corpora, tokenization schemes, evaluation benchmarks, and deployment archite…
Avijit Roy, Proma Roy · 📄 PDF
VAKRA: Evaluating Multi-Hop Reasoning Across APIs and Retrieval Under Tool-Use Policies
Agents deployed in enterprise settings must reason across structured APIs and document collections, yet existing benchmarks evaluate these capabilities in isolation. We introduce VAKRA (e\textbf{V}aluating \textbf{A}PI and \textbf{K}nowledge \textbf{R}etrieval \textbf{A}gents), a benchmark of over $…
Ankita Rajaram Naik, Anupama Murthi, Benjamin Elder, Siyu Huo, Raavi Gupta, Abhinav Jain, Praveen Venkateswaran, Abdulha… · 📄 PDF
Beyond Trial-and-Error: Agentic Optimization for Image-to-Video Adherence
Modern black-box Image-to-Video (I2V) models offer powerful capabilities in automated content creation, yet their lack of fine-grained control and reliability presents significant challenges in professional workflows. Their inherent stochasticity causes minor variations in textual prompts or hyperpa…
Aman Tyagi, Hemanth Boinpally, Jonathan Chen, Douglas Gebert, Steven Hickson · 📄 PDF
Class Activation Mapping in Explainable Computer Vision: A Method-Centered Review of CNN, Transformer, and Foundation-Model-Era Visual Explanations
Class activation mapping (CAM) is one of the most widely used visual explanation families in explainable artificial intelligence. Its purpose is intuitive: it converts internal model evidence into a heatmap that highlights the image regions, convolutional channels, tokens, or patches that support a …
AmirHossein Eshghi, Hamid Saadatfar, Seyyed Ali Hoseini, AmirMohsen Eshghi, Siavash Arjomand Bigdel · 📄 PDF
Constructing Dynamic Master Logic Models as Knowledge Graphs for Complex System Diagnostics Using Retrieval-Augmented Large Language Models
Dynamic Master Logic (DML) provides a hierarchical framework for representing system behavior by linking functional objectives to underlying structural elements. However, DML construction typically relies on expert interpretation of technical documentation, limiting scalability for complex systems. …
Saman Marandi, Yu-Shu Hu, Mohammad Modarres · 📄 PDF
Redistribution-based Cost Inference Improves Sparse Safe Offline RL
Safe offline RL typically assumes access to dense per-step cost annotations, but in practice supervisors provide only trajectory-level stop-feedback: a binary signal at the first unsafe transition, with no per-step attribution. We frame this as a temporal credit assignment problem and propose the Re…
Ebenezer Gelo, Geraud Nangue Tasse, Steven James, Benjamin Rosman · 📄 PDF
AI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses
Recent work on distillation transfers the capabilities of large models to smaller ones often by updating the latter's parameters, through teacher forcing, on-policy distillation, and related training-time methods. In this paper, we ask whether such transfer can instead occur at test time. We study s…
Cheng Qian, Wenting Zhao, Liangwei Yang, Heng Wang, Jielin Qiu, Heng Ji, Silvio Savarese, Huan Wang, Shelby Heinecke · 📄 PDF
DreamFly: Causal Memory and Receding-Horizon Diffusion Planning for Aerial Vision-Language Navigation
Aerial vision-language navigation (VLN) requires an embodied agent to integrate visual evidence over time, plan future actions, and determine when it has reached a navigation goal under partial observability. Although recent VLA models offer a promising perception-to-action paradigm, adapting them t…
Yan Deng, Fei Xu · 📄 PDF
Do People Follow AI Advice? Evidence from a Pension Portfolio Choice Experiment
We study how differences in AI-generated financial recommendations are transmitted into individual portfolio choices. In an experiment with 400 employed adults enrolled in workplace defined contribution pension plans in South Korea, participants allocate a hypothetical pension balance across eleven …
Hongseok Choi, Jeongbin Kim, Matthew Kovach, Kyu-Min Lee, Euncheol Shin, Hector Tzavellas · 📄 PDF
Technology interactions reshape the economics of China's coal power decarbonization
Decarbonizing existing coal-fired power plants can contribute to near-term climate mitigation, but identifying cost-effective retrofit strategies is complicated by interactions among mitigation technologies. Here we develop an interaction-aware optimization framework that jointly evaluates energy co…
Yun-Long Zhang, Jia-Ning Kang, Xiaoming Kan, Lan-Cui Liu, Zhimin Huang, Song Peng, Biying Yu, Yi-Ming Wei · 📄 PDF
Few-cycle electro-optic light on thin-film lithium niobate
The twin fields of ultrafast optics and nonlinear photonics enable applications ranging from attosecond science [1, 2] and ultrafast electronics [3] to molecular spectroscopy [4, 5], nonlinear optics [6-8], quantum nanophotonics [9] and precision metrology [10]. However, bringing these capabilities-…
Xinyi Ren, Chun-Ho Lee, Ian Christen, Clayton Cheung, Reshma Kopparapu, Yue Yu, Lian Zhou, Zaijun Chen, Mengjie Yu · 📄 PDF
CosMAP: Contrastive Manifold Approximation and Projection for Dimensionality Reduction of Omics and Genealogical Data
Omics datasets, particularly single-cell RNA sequencing data, are high-dimensional, sparse, noisy, and dominated by zero values, making faithful low-dimensional representation challenging. Existing dimensionality-reduction methods may distort local neighbourhoods, global organization, or the cohesio…
Fenosoa Randrianjatovo, Maya Saleh, Simon Girard, Amadou Barry · 📄 PDF
Large-scale AI-Ready Data for Anti-Cancer Drug Response Modeling
Drug response prediction (DRP) models are an active area of research in pharmacogenomics, with growing potential to accelerate the identification of effective anticancer drugs. However, their predictive performance is often constrained by limited dataset scale and insufficient coverages of cancer an…
Vincent Lavelle, Yitan Zhu, Kaitlyn Marlor, Thomas Brettin, Rick Stevens · 📄 PDF
Probing and steering biology across Boltz-1s trunk-diffusion boundary
AlphaFold3-class structure predictors pair a representational trunk, which processes sequence and context, with a diffusion module, which generates atomic coordinates. How biological information changes as it crosses this architectural boundary remains poorly understood. We analyze per-residue activ…
Piotr Jedryszek, Tongmeng Xie, Adam Winnifrith, Alexander Hasson, Weronika Ślesak, George Wicks, Toby Winnifrith, Oliver… · 📄 PDF
A Modular Agentic Framework for Synthetically Constrained Multi-Objective Hit-to-Lead Optimization
Hit-to-lead optimization requires iterative design of hit analogs across competing potency, selectivity, physicochemical, pharmacokinetic, safety, and synthetic constraints. We present SABLE (Synthetically-accessible Agentic Bayesian Ligand Exploration), an open-source framework that employs natural…
Kelvin P. Idanwekhai, Enes Kelestemur, Benjamin Strickland, Matthew Hart, Steini Davidsson, Angelos Angelopoulos, Ron Al… · 📄 PDF
The Fallacy of Independent Ceilings: Characterizing Coupled Load-Branch Stall Interaction
Branch mispredictions and data-cache misses are usually evaluated as separate bottlenecks: studies report perfect-branch or perfect-cache speedups as isolated upper bounds and often treat their product as the joint ceiling. In irregular workloads, however, hard-to-predict branches and cache-missing …
Matthew Constant, Resit Sendag · 📄 PDF
Locomotion Variability and User Experience in Smart Wheelchair Human-Robot Interaction
Human movement is inherently variable, with variability structured according to task relevance: movements are typically more consistent at task-critical points and more flexible elsewhere. In human-robot interaction (HRI), however, model-based assistance strategies commonly assume deterministic huma…
Sean Kille, Adina M. Panchea, Balint Varga, Sören Hohmann · 📄 PDF
Herding End-to-End Autonomous Driving via Neuro-Symbolic Safety Guards
Modern end-to-end driving agents can achieve high average performance yet still violate basic traffic rules that a human driver would never miss. The reason is structural: they learn statistical patterns rather than the physical conditions that guarantee safe driving, leaving their decision-making p…
Simón Patiño Idarraga, Erick Silva, Rehana Yasmin, Ali Shoker · 📄 PDF
Koopman Representation of Nonlinear Virtual Environments in Kinesthetic Haptic Systems
Rendering haptic feedback with nonlinear virtual environments (VEs) is important in many applications that require highly accurate force feedback. This paper considers the use of the Koopman operator to represent a nonlinear VE interacting with a haptic system. Simulation and experimental results de…
Yanting Zhou, Jozsef Kövecses, James Richard Forbes · 📄 PDF
A Note on the Identification Step in "A Semistructural Methodology for Policy Counterfactuals"
The New Keynesian example of Beraja (2023) is not identified at its printed calibration: more than one structure satisfies every condition of its identification step. Two of the paper's six identifying restrictions coincide on equilibrium equations consistent with the printed reduced form, leaving e…
Henri Keränen · 📄 PDF
Mastering Stochastic OLG Models in Continuous Time
We propose a comprehensive framework for solving overlapping-generations (OLG) models in continuous time with both idiosyncratic and aggregate risk. Our general characterization of equilibrium through the master equation operates on the joint distribution over the continuous idiosyncratic states, ag…
Yves Achdou, Johannes Brumm, Lukas Frank · 📄 PDF
Private Correlations Certify Sensing Capability
We show that private correlations in a bipartite quantum state constitute a metrological resource for distributed sensing assisted by a possibly noisy channel from one party to the other. We begin with an example with distillable secret key but poor locally accessible sensing performance and show th…
Yunkai Wang, Peixue Wu, Graeme Smith, Sisi Zhou · 📄 PDF
Causality Sum Rules in Conventional Scattering Matrices
Scattering matrices are the standard experimental and computational description of photonic and electromagnetic devices. Passivity is explicit in the conventional incoming-outgoing matrix, whereas causality sum rules are usually formulated only after transforming the response into auxiliary variable…
Ning Han, Rui Zhao, Shuxing Yang, Mingzhu Li, Hongsheng Chen, Yihao Yang · 📄 PDF
Toward a Thermodynamic Framework for Dissipative Solitons: From Photonics to Turbulence and Bose--Einstein Condensate Analogies
Statistical reasoning in nonlinear optics follows three routes - wave-turbulence kinetics, multimode equilibrium thermodynamics, and mode-locking statistical mechanics -all relying on externally specified modal bases (waveguide, spectral window, or cavity bandwidth). We ask whether a dissipative sol…
Vladimir L. Kalashnikov, Irina T. Sorokina · 📄 PDF
730-nm optical parametric conversion from near- to short-wave infrared band
A record 730 nm parametric conversion in silica fiber from the near-infrared to the short-wave infrared band is reported and analyzed. A parametric gain in excess of 30 dB was measured for a signal at 1300 nm (with corresponding idler at 2030 nm). This conversion was performed in a travelling single…
J. M. Chavez Boggio, J. R. Windmiller, M. Knutzen, R. Jiang, C. Bres, N. Alic, B. Stossel, K. Rottwitt, S. Radic · 📄 PDF
Link-adaptive digital twin for robust physical-layer modeling in hybrid-amplified ultra-wideband optical networks
Accurate physical-layer modeling is increasingly essential for reliable ultra-wideband operation and capacity optimization, especially under the intensified inter-channel stimulated Raman scattering (ISRS) effect. This paper proposes the link-adaptive digital twin (LA-DT) for hybrid-amplified ultra-…
Xiaoxuan Gao, Rentao Gu, Yingchun Wang, Xinyi Liu, Junshi Gao, Yuefeng Ji · 📄 PDF
Dispersion Control of Chiral Exciton-Polariton Transport with Dielectric Metasurfaces
Exciton-polaritons provide a powerful platform for manipulating hybrid light-matter states with low effective masses and strong nonlinearities. Introducing chirality into these quasiparticles enables selective control over their spin and propagation, opening new opportunities for chiral transport an…
Juhyeong Jeon, Kehan Wang, Yu-Chen Wei, Francesca Cussiol, Matthijs Berghuis, Fan Xu, Shunsuke Murai, E. W. Meijer, Jaim… · 📄 PDF
Direct experimental measurement of femtonewton-scale momentum transfer force from electron beams
Electron beams (e-beams) are ubiquitous in imaging, patterning, and propulsion. This prevalence is rooted in the profound mastery of their wave-particle duality and energy-transfer pathways. Yet, a fundamental dimension remains largely unexplored: while the mechanical effect (i.e., the momentum tran…
Chunbo Lin, Xinggang Shang, Xijun Li, Xiaoyu Sun, Kang Zhao, Xujie Wang, Yang Yu, Wenjing Cao, Min Qiu · 📄 PDF
Nanoscale graphitization and defect evolution in silicon-vacancy center-containing nanodiamonds under high-pressure high-temperature annealing
Group-IV color centers, such as the silicon-vacancy (SiV) defect, are highly promising for solid-state quantum technologies. However, nanodiamonds typically exhibit significant lattice strain and structural disorder, which degrade their optical properties and hinder the resolution of the fine spectr…
M. De Feudis, B. Yavkin, L. Henry, K. O. Ho, M. -P. Adam, P. Goldner, F. Benedic, S. Desgreniers, J. -F. Roch · 📄 PDF
Strong, Mode-Selective Evxciton-Photon Coupling Driven by Polariton Scattering in the Mo2 Complexes at Ambient Conditions
Quadruply bonded Mo2 complexes provide a distinctive molecular platform in which multiple two-level electronic transitions interact with quantized scattering modes under ambient conditions. Here we show that, in the Mo2 complexes, the intrinsic photonic modes selectively couple to molecular excitati…
Miao Meng, Ying Ning Tan, Zi Cong He, Yuli Zhou, Chun Y. Liu · 📄 PDF
Stability of Orbital Angular Momentum Modes in Conventional and Ring-Core Optical Fibers
Modal dynamics in multimode optical fibers are fundamentally governed by the inter- play between phase matching and intermodal coupling. In this work, we investigate the fundamental coupling mechanisms in ring-core fibers and identify a geometry-induced cou- pling suppression that significantly redu…
Hassan Asgharzadeh B., Regina Gumenyuk, Marco Ornigotti · 📄 PDF
Revealing time characteristics of optical excitations in dielectric and plasmonic structures through cathodoluminescence interferometry
Cathodoluminescence (CL) spectroscopy provides access to optical excitations with nanometer spatial resolution, but direct time-resolved measurements of optical resonances remain challenging. Here, we demonstrate that CL interferometry provides access to the temporal response, phase behavior, and mo…
Evelijn Akerboom, Hirohsi Sugimoto, Minoru Fujii, Nicolas Pazos-Perez, Ramon A. Álvarez Puebla, A. Femius Koenderink, F.… · 📄 PDF
Quantum Theory of Third-harmonic Generation in Epsilon-Near-Zero Materials
We present a theoretical framework, based on the Green's tensor quantization method, to describe third-harmonic generation in epsilon-near-zero (ENZ) materials and derive analytical, closed-form solutions for the generation efficiency. We validate our model against experimental measurements of wavel…
Sonia Alipour, Riccardo Franchi, Tornike Shubitidze, Smridhi Chawla, Luca Dal Negro, Marco Ornigotti · 📄 PDF
Rf beam steering using on-ohip optical true time delay on 3 micrometer siliocn on insulator platfrom
We present the design, packaging, and system-level experimental demonstration of a reconfigurable 5-bit, 4-channel optical true time delay (OTTD) chip realized on a 3 micrometer thick silicon on insulator (SOI) platform at VTT. The chip enables 32 discrete delay states over a 0 to 101 picoseconds ra…
Somnath Paul, Jussi Saily, Markus Schroder, Mikko Harjanne, Timo Aalto · 📄 PDF
Optical-Memory Transport Imaging: A Transport-History Framework for Finite-Memory Tracers
Finite-memory optical tracers encode upstream transport histories rather than instantaneous local flow velocities. We introduce optical-memory transport (OMT) imaging, a framework in which finite memory couples internal-state relaxation to transport through memory kernels. Under structured illuminat…
Haichun Liu, Jerker Widengren · 📄 PDF
Boosting self hybridized exciton polaritons with metal clad WS2 waveguides
The formation of Fabry Perot and guided wave self hybridized exciton polaritons in two dimensional materials results in long range exciton energy transfer and strong exciton exciton interactions. Here, we demonstrate that the coupling strength between photonic modes and excitons is significantly boo…
Filip Majstorovic, Masoud Taleb, Victor DeManuel-Gonzalez, Kai Rossnagel, Nahid Talebi · 📄 PDF
Plasmonic Fourier Surfaces Revisited: Relating Bandgaps with Bound States in the Continuum
Periodically corrugated metal interfaces supporting surface plasmon polaritons (SPPs) belong to the earliest nanoplasmonic platforms. Even simplest reliefs described by a few harmonics - plasmonic Fourier surfaces - display markedly different far-field signatures depending on the corrugation depth a…
Alexander A. Antonov, Connor Heimig, Hannah Niese, Yannik M. Glauser, Sander J. W. Vonk, David J. Norris, Maxim V. Gorku… · 📄 PDF
Non-resonant laser-driven narrowing of particle velocity distributions
Stark acceleration and deceleration based techniques for generating particle ensembles with low velocity spread are useful in many experimental applications. For a given velocity distribution of a particle ensemble, these techniques accelerate or decelerate a small subset of the total population, wi…
Ashwini Vaishnav, Matthias Rupp, Olivier De Castro, Mikhail N. Shneider, Alexandros Gerakis · 📄 PDF
High-dimensional Supermode Photonics Enabled by Hierarchical Supersymmetric Transformation
Modes provide a fundamental degree of freedom for photonic information processing, yet conventional multimode waveguides exhibit non-equidistant effective-index distributions, making closely spaced modes vulnerable to intermodal crosstalk. Supermode photonics can overcome this limitation by geometri…
Yuan Zhong, Kaile Chen, Qi Lu, Chunxue Wang, Jingchi Li, Yuru Li, Zhaohui Li, Chao Lu, Xinchen Ji, Yikai Su, Lu Sun · 📄 PDF
Impact of strain and dark states on spectroscopic measurements of silicon-vacancy centers in diamond
Negatively charged silicon-vacancy (SiV$^-$) centers in diamond offer an attractive platform for the development of many forms of quantum technology. However, questions remain in connection to how large ensembles of SiV$^-$ centers behave in concert. Here, we develop a computational model designed t…
Tommy Chin, Kesav V. Narayan, Imran Bashir, Kelsey M. Bates, Liam G. Stanton, Ehsan Khatami, Christopher L. Smallwood · 📄 PDF
Floquet Green's functions for lattice electrons driven by Gaussian quantum light
We formulate Floquet Green's functions for noninteracting single-band lattice electrons driven by a reservoir-stabilized single-mode Gaussian quantum light source. The source is prescribed externally and is not updated by the many-electron polarization, while an active electronic probe still conditi…
Atsushi Ono · 📄 PDF
An Information Theory Analysis of Whole Slide Image Pathology AI and Diagnostic Field Selection AI Under Limited Resources
A key issue in using AI for pathology diagnosis is what image information should be given to the AI and how limited analysis resources should be used. This study compares two ways of processing different types of images under limited resources. The first is WSI-AI, in which AI automatically compress…
Tatsuaki Tsuruyama · 📄 PDF
Modeling and Interpreting Correlations, Null Distributions and Significance Levels in Neural Tracking of Natural Stimuli
Neural tracking - the time-locking of neural responses to continuous stimuli such as speech, music, and video - is widely used to study how the brain processes natural input. Tracking strength is typically quantified as the correlation between the recorded neural response and the stimulus, decoded a…
Simon Geirnaert, Alexander Bertrand, Tom Francart, Jonas Vanthornhout · 📄 PDF
Observable-Reduction-Guided Sparse Regression for Partially Observed Active-Quiescent Systems
Active-quiescent switching occurs in biological populations in which growth is confined to a proliferative active state, while cells may reversibly enter a nonproliferative quiescent state. Experiments often observe only part of this process, through active-state markers, aggregate population measur…
Kyle C. Nguyen, Kevin B. Flores · 📄 PDF
Conversational Orchestration for Organic 6G
The Organic 6G vision of a network of networks spanning an edge-cloud continuum complemented by non-terrestrial resources requires, to realize its promise, service provisioning that is simple to operate, scalable across independently administered domains, and agile under domain churn (i.e., domains …
Masoud Shokrnezhad, Tarik Taleb · 📄 PDF
Media-over-Multipath-QUIC for Realtime Video Applications
Multipath transports place a client's WiFi, cellular, and satellite networks under one connection, yet real-time video gains little from them. The scheduler that assigns packets to paths sees only bytes, so it cannot tell a keyframe that anchors a second of video from an enhancement frame whose loss…
Tanya Shreedhar, Zuji Zhou, Nitinder Mohan, Fernando Kuipers · 📄 PDF
CARB: A Characterization-Guided Framework for CNN Inference Cost Prediction and Deployment Screening
Accurate pre-deployment estimation of CNN inference cost--energy, latency, and peak memory--is increasingly critical as models are deployed on resource-constrained GPU platforms. Existing approaches rely on FLOPs, latency measurements, or single-device profiling as energy proxies, overlooking the no…
Linh Nguyen, Zhixin Pan · 📄 PDF
Synthesizing Probabilistic Saturating Counters with Differentially Private Formal Guarantees
Branch predictors improve instruction-level parallelism in modern processors and are commonly modeled using saturating counters. However, classical saturating counters are deterministic and thus vulnerable to side-channel attacks: an attacker can manipulate the counter state and infer the branch dir…
Zhiming Chi, Lutan Zhao, Depeng Liu, Yong Li, Pengfei Yang, Bow-Yaw Wang, Rui Hou, Cheng-Chao Huang, Andrea Turrini, Lij… · 📄 PDF
Adaptive Matrix Multiplication for Dynamic Shapes on Ascend NPUs
Matrix Multiplication (MatMul) faces a "generalization crisis" driven by highly dynamic tensor shapes. This crisis is particularly acute on Ascend NPUs, where explicitly controlled architectures and strict physical constraints render existing GPU-centric optimizations ineffective. To resolve this, w…
Yuhang Zhou, Jiang Peng, Qianyu Jiang, Zhibin Wang, Xinghui Tian, Jianwei Zhou, Songxiang Zhu, Jingyi Zhang, Junsong Wan… · 📄 PDF
You Only Charge Once 2.0 : A End-to-End Analog Computing-in-Memory Architecture with Reconfigurable Switched Capacitors
Analog Computing-in-Memory (ACiM) accelerates deep neural networks by keeping weights inside memory arrays and executing dot products in the analog domain. However, modern ACiM accelerators are often limited by the "ADC wall": analog-to-digital converters consume a large fraction of energy and area,…
Zihao Xuan, Yewen Li, Jia Chen, Wei Xuan, Xiao Huo, Fengbin Tu · 📄 PDF
When and Where Faults Matter: A Study of Transient Errors in CKKS Multiplication
Homomorphic Encryption (HE) is a privacy-preserving encryption paradigm that enables computation directly on encrypted data without requiring decryption. In this paper, we study errors in fully homomorphic encryption (FHE) computations, with a particular focus on server-side homomorphic multiplicati…
Vattana Chan, Matías Mazzanti, Karthik Swaminathan, Augusto Vega, Esteban Mocskos, Radha Venkatagiri · 📄 PDF
On the Sensitivity to Errors in Homomorphic Computing: Single Transient Bit-flip Client-side Error Characterization
Homomorphic Encryption (HE) enables computation on encrypted data without decryption and is a key primitive for privacy-preserving computation in sensitive domains such as healthcare, finance, and government. Its security relies on noise injection, which introduces intrinsic error sensitivity and ra…
Matías Mazzanti, Vattana Chan, Karthik Swaminathan, Augusto Vega, Esteban Mocskos, Radha Venkatagiri · 📄 PDF
When Your State Estimator Has Lost The Plot: Detecting Estimator Failures Via Spectral Analysis
Reliable onboard state estimation is essential for safe robotic operation, yet unmodeled disturbances, such as sensor aliasing or out-of-distribution noise, still cause estimators to degrade or fail completely. While many methods aim to improve estimator robustness, only a few provide introspective …
Christian Lanegger, Helen Oleynikova, Roland Siegwart, Michael Pantic · 📄 PDF
Precise Top-Layer Fabric Segmentation for Fabric Destacking with Edge- and Shape-Aware Deep Networks
Fabric destacking requires precise segmentation of the topmost fabric layer, a task complicated by subtle fabric boundaries and high visual similarity between fabric layers. Existing semantic and edge-based segmentation approaches often struggle with these complexities, limiting the performance of r…
Wenbo Dong, Dipankar Bhattacharya, Akinari Kobayashi, Akira Seino, Fuyuki Tokuda, Xuzhao Huang, Kai Tang, Norman C. Tien… · 📄 PDF
OAA: Three Phases of Vocal Guidance in Human-Drone Teleoperation
Voice-guided teleoperation requires systems that adapt to the evolving dynamics of human guidance. Yet most voice-controlled robot systems treat spoken commands as a stationary stream, ignoring how the guide's communicative behavior changes as the task progresses. Using motion capture and speech dat…
Allan Henry, Christian Graff, Solange Rossato, José-Ernesto Gomez-Balderas, Sylvain Huet · 📄 PDF
Robust Sliding Mode and Admittance Control of Underactuated Aerial Manipulators for Contact-Based Inspection
Contact-based industrial inspection requires aerial platforms to maintain stable interaction while rejecting disturbances. Underactuated aerial manipulators present control challenges due to the dynamic coupling between vehicle attitude and force generation. This paper proposes a robust control fram…
Tareq Aziz Alqutami, Yvan Petillot, Matthew W. Dunnigan, Mustafa Suphi Erden · 📄 PDF
TCAM for Autonomous Deformable Manipulation: The RMC2 Champion System for WBCD 2026 Track 4
This technical report describes the RMC2 Team's champion solution for the WBCD 2026 Track 4: Deformable Manipulation Challenge. The task requires a robot to pick a single T-shirt from a stack, load it onto a printing pallet, align the collar with a target area, and smooth the printing region, a sequ…
Guangrui Shen, Zhili He, Shigang Wang, Yuanjun Sun, Qing Yu · 📄 PDF
Embodied Multimodal Grounding for Open-Vocabulary Mobile Manipulation via Semantic 3D Gaussian Splatting
Embodied mobile manipulation requires language, visual observations, three-dimensional scene structure, and action feasibility to be aligned before execution. We study open-vocabulary target grounding with few-shot manipulation in local household workspaces and present an embodied multimodal groundi…
Huosen Ou, Dongni Song, Yuncong Wang, Tao Zhou, Yiding Ji · 📄 PDF
JEPA-WAM: Stage-Level Joint-Embedding Prediction for World-Action Models in Robot Manipulation
Generalist robot policies aim to map multimodal observations and linguistic task instructions to actions across diverse tasks. However, existing methods typically represent the future as a fixed, short video-action chunk. This short-term future captures local scene evolution for action execution, bu…
Xiao Liu, Yuguang Yang, Xi Wang, Kai Jiang, Cheng Chi, Yong Xu, Wenchao Ding, Yilun Chen, Yan Wang · 📄 PDF
Dual Stress: Runtime Safety Monitoring for Safety-Constrained MPC Navigation
Runtime hazard monitors for autonomous naviga- tion are conventionally built from geometric quantities: predicted clearance, time to collision, and required deceleration. A model-predictive controller that enforces safety through explicit con- straints computes, as a by-product of every control step…
Jamil Chahine, Wenqi Cai, John Abanes, Anthony Tzes · 📄 PDF
AECNav: Active Evidence Consolidation for Efficient Zero-Shot Open-Vocabulary Object Navigation
Zero-shot object-goal navigation (ZSON) in open-vocabulary scenarios is challenging, as it requires a robot to locate an arbitrarily specified object in an unseen environment without task-specific training. Currently, the task still suffers from high latency and limited accuracy due to redundant per…
Guanlin Liu, Shaobin Ling, Renyuan Liu, Zeying Gong, Junjie Hu · 📄 PDF
Neural Introspection Gating for Adaptive KV-Cache Reuse in Vision-Language-Action Models
Vision-Language-Action(VLA) models map camera images and language instructions directly to motor commands through a single autoregressive transformer. In real-time control, they still spend substantial compute recomputing key-value(KV) representations for visual tokens that barely change across neig…
Zhijie Wu, Kento Kawaharazuka, Kei Okada · 📄 PDF
Enabling Scalable Kinesthetic Teaching via Observer-based Hand-guiding with Active Support
Kinesthetic teaching through robot hand-guiding provides a natural interface for collecting demonstrations in imitation learning and programming-by-demonstration. However, extended sessions cause operator fatigue, reducing demonstration quality and limiting scalability. Current industrial hand-guidi…
Anna Tuma, Giuseppe Monetti, Jochen J. Steil, Niels Dehio · 📄 PDF
Flex-$π$: A Multi-Stream World-Action Model with Compute Flexibility
World-action models (WAMs) predict the future to act better, but nearly all of them predict only RGB latents, trained purely for pixel reconstruction, with no explicit signal for the 3D geometry or object semantics manipulation needs. We find a surprising free lunch: the same frozen video-generation…
Ge Yan, Jinghao Liu, Yuzhi Fan, Lei Cai, Minwen Liao, Jesse Zhang, Dieter Fox · 📄 PDF
Robust Safety Filtering for Input-Constrained Underactuated Linear Systems
We present a robust safety-filtering framework for input-constrained underactuated linear systems subject to unknown disturbances. A baseline H-$\infty$ input is derived from a zero-sum differential game, while a disturbance observer supplies an estimate and a transient error bound. The baseline inp…
Muhamad Rausyan Fikri · 📄 PDF
GESTO: Human-Centric Spatio-Temporal Memory for Reasoning in Dynamic Scenes
Robots operating in human environments need memories that capture not only what objects exist and where, but also how people use them over time and how individual interactions compose into goal-directed activities. Existing 4D scene graphs preserve object and place histories but omit activity struct…
Ermanno Bartoli, Buwei He, Dennis Rotondi, Sebastian Koch, Federico Tombari, Kai O. Arras, Patric Jensfelt, Yixi Cai, Io… · 📄 PDF
Aerial Layouting: Design and Control of a Compliant and Actuated End-Effector for Precise In-flight Marking on Ceilings
Aerial robots have demonstrated impressive feats of precise control, such as dynamic flight through openings or highly complex choreographies. Despite the accuracy needed for these tasks, there are problems that require levels of precision that are challenging to achieve today. One such problem is a…
Christian Lanegger, Marco Ruggia, Marco Tognon, Lionel Ott, Roland Siegwart · 📄 PDF
Seeing above the waves: A modular sensing framework for data acquisition at sea
Advancing autonomy for surface vessels requires systematic evaluation of their sensing and perception subsystems. Yet, maritime environments impose unique challenges: sensor installation is constrained by vessel layout, environmental conditions such as fog or sea clutter are difficult to reproduce, …
Jonathan E. Schmidt, Julius Wirbel, P. Nicholas Hansen, Morgan Louédec, Christian L. H. Westerdahl, Dimitrios Dagdilelis… · 📄 PDF
Deployment Is Not Destiny: Robot Recomposition in the Field with Unseen Software, Hardware, and Compute Payloads
The tight coupling of subsystems in most robots, though a natural consequence of their complexity, leads to monolithic designs that are time-consuming and difficult to adapt after initial deployment. To address this challenge, we present a framework and supporting abstractions for recomposition duri…
Steven Swanbeck, Jonathan Salfity, Jeffery Gunawan, Corrie Van Sice, Mitch Pryor, Robert Blake Anderson · 📄 PDF
VIScore: Diagnosing Planning-Relevant Quality in Latent World Models
Regulating the latent space to an isotropic Gaussian distribution provides a stable and information-maximized landscape for world model planning. However, the latent space property and successful planning remain disconnected. We first study this by comparing SIGReg and VISReg, two regularization los…
Haiyu Wu, Randall Balestriero, Morgan Levine · 📄 PDF
Risk-Aware Kinodynamic Motion Planning Under Uncertainty For Safe Navigation on Planetary Environments
For autonomous space exploration, robotic agents need to perform motion planning in which environmental interactions may be unknown. Learning these interactions, such as terrain mechanics for wheeled robots, can introduce uncertainties that lead to risky motion plans and potentially hazardous operat…
Sachin Sunil Kelkar, Tanmay Dokania, Yashwanth Kumar Nakka · 📄 PDF
HUI360: A 360° Egocentric Dataset and Baselines for Human-Robot Interaction Anticipation
As robots increasingly operate in human-populated environments, anticipating human intentions is essential for enabling proactive and socially aware behavior. Automatic anticipation of human-robot interactions is thus emerging as a crucial perception challenge for embodied agents. To this end, we in…
Raphael Lorenzo-Louis, Fabio Amadio, Bertrand Luvison, Serena Ivaldi · 📄 PDF
CapProbe: Evaluating Detailed Image Captions via Full-Scene Dense Question Answering
Evaluating detailed image captions from Vision-Language Models (VLMs) requires going beyond surface-level semantic similarity. Reference-based metrics (e.g., CIDEr and SPICE) and LLM-as-scorer protocols struggle to verify dense factual claims, while existing QA-based alternatives generally offer low…
Mouxiao Huang, Qiangyu Yan, Borui Jiang, Han Shu · 📄 PDF
Static in Frames, Dynamic in Events: Rethinking Features in Event Cameras as Motion Cues
Event cameras capture intensity changes asynchronously with high temporal resolution, requiring novel preprocessing methods for downstream tasks. Unlike static intensity snapshots, event data inherently encode information about scene dynamics and object motion, meaning that features derived from eve…
Hesam Araghi, Jan van Gemert, Nergis Tomen · 📄 PDF
Foundation Model-Enabled Efficient Data Sampling (FEEDS): A label-efficient training strategy for pan-cancer, multi-tracer PET/CT datasets
Automated lesion segmentation in whole-body PET/CT imaging can assist clinicians with cancer detection, staging, and treatment planning across radiotracers and cancer types. However, training lesion segmentation models that capture variations in lesion size, distribution, and appearance requires lar…
Biratal Raj Wagle, Bashirul Azam Biswas, Grant Chau, Matthew E. Maeder, Muhammad Azeem Arshad, Michael S. Leapman, James… · 📄 PDF
Learning Gaussian Structure: Intervention-Guided Density Control for Feed-Forward Driving Reconstruction
Feed-forward Gaussian reconstruction has recently emerged as an efficient approach for driving scene reconstruction. However, prevailing LiDAR-based methods preserve the initial correspondence between observed points and Gaussian primitives, treating the initialized primitive set as the final repres…
Hang Li, Jiahe Li, Meiying Gu, Jin Zheng, Lina Yu, Xiao Bai · 📄 PDF
Every Packet Counts: Dispersing Information for Loss-Resilient Learned Image Compression
Learned image compression (LIC) has achieved impressive rate-distortion performance. However, existing methods remain highly vulnerable to packet loss, a common challenge in satellite and emergency communications. This vulnerability stems from non-uniform information distribution at the packetizatio…
Yuhang Wei, Chuqin Zhou, Yibo Shi, Jing Wang, Guo Lu · 📄 PDF
Is There Really a Camouflaged Object? Towards Realistic Camouflaged Object Detection
Camouflaged object detection (COD) aims to segment objects that are visually concealed in their surroundings and has attracted increasing attention in recent years. However, most existing COD methods are developed under a closed-world assumption, where each input image is assumed to contain a camouf…
Huafeng Chen, Yueming Lyu, Chenyang Si, Wende Tan, Liucheng Guo, Caifeng Shan · 📄 PDF
SAR2Agri: Learning SAR Intensity Representations for Agricultural Monitoring
Agricultural monitoring faces unique challenges, arising from the landscape's complex temporal, phenological, and climate dynamics, yet monitoring them is critical for ensuring food security. Synthetic Aperture Radar (SAR) satellites offer all-weather day-night imaging capability supporting key moni…
Moti Rattan Gupta, Anupam Sobti · 📄 PDF
PRMU: A Corpus-Free Benchmark for Person-Centric Knowledge Unlearning in Multimodal Large Language Models
Multimodal large language models (MLLMs) have demonstrated remarkable capabilities in storing and recalling rich person-related knowledge, raising increasing concerns about reliable knowledge removal. However, existing machine unlearning approaches for MLLMs typically assume access to original forge…
Huafeng Chen, Yueming Lyu, Ziyuan Chen, Wenda Tan, Chenyang Si, Liucheng Guo, Caifeng Shan · 📄 PDF
CausalSplat: Towards Comprehensive Hierarchical Reasoning in 3D Gaussian Splatting
While 3D Gaussian Splatting (3DGS) has advanced open vocabulary scene understanding, existing methods remain confined to explicit queries. They struggle to interpret implicit intents, complex spatial constraints, and commonsense reasoning required for practical embodied interactions. To address this…
Jiayu Ding, Meilu Song, Yun Chen, Wei Gao, Ge Li · 📄 PDF
VidForensics-M1: Meta-Detection Reinforcement Learning with Verifiable Temporal Grounding for AI-Generated Video Forensics
Recent advances in video generation models have significantly improved the realism of synthetic videos, blurring the boundary between generated and authentic content and raising concerns about misinformation. Existing MLLM-based detectors mainly rely on supervised fine-tuning or label-level reinforc…
Bowei Liu, Zheng Lu, Yuhan Bian, Xinchen Zhang, Xingming Shui, Yuesheng Huang, Xuhuan Li, Zihao Liu, Yifan Yang, Jun Zho… · 📄 PDF
Capturing Uncertainty in Human Motion for Representation Learning in Soccer
This paper presents a self-supervised representation learning framework for understanding 3D skeleton-based human motion in soccer, using future motion prediction as the learning objective. Since human motion is inherently uncertain, accounting for multiple plausible futures is essential for capturi…
Yizhou Xu, Lars Bretzner, Tiesheng Wang, Atsuto Maki · 📄 PDF
AdvFD: Boosting Visual Generation via Adversarial Fr'echet Distance Loss
Fréchet distance has recently emerged as an effective distribution-level objective for generator post-training, complementing the conventional sample-level diffusion and flow-matching losses. However, directly optimizing Fréchet objectives can cause Fréchet hacking. The target metrics keep improving…
Mingju Gao, Jingkai Zhou, Kun Gai, Changqian Yu, Hao Tang · 📄 PDF
ReRound: Reconstructive Rounding to Resolve Midpoint Ambiguity in Calibration-Free LLM Quantization
ReRound (Reconstructive Rounding) is a post-training quantization method that addresses the midpoint ambiguity inherent in standard round-to-nearest (RTN) schemes when quantizing weights near the centers of quantization intervals. Starting from a pretrained LLM, ReRound trains a conditional diffusio…
He-Yen Hsieh, H. T. Kung · 📄 PDF
Efficient Hypergradient Descent for Inverse Reinforcement Learning
Inverse reinforcement learning (IRL) aims to recover a reward function under which the resulting policy reproduces the behavior observed in expert demonstrations. A natural approach is to formulate IRL as a bilevel optimization problem, in which the inner level corresponds to policy optimization und…
Nikita Sevriukov, Anna Barabanova, Uliana Gagarina, Karina Ivanova, Sofiia Kasaeva, Ilya Levin, Marina Sheshukova · 📄 PDF
Uncertainty-Aware Deep Learning for Genomics Applications: Insights from an Empirical Study
Deep learning models have emerged as the standard computational tool for a wide range of applications in genomics. Yet, uncertainty quantification (UQ) -- and more specifically, the reliability of different uncertainty estimates in this domain -- has received little systematic attention. This work p…
Sepideh Saran, Mahsa Ghanbari, Uwe Ohler · 📄 PDF
Batch Size or Negatives? A Selection Rule for Memory-Constrained Recommender Training
Large-scale neural recommender systems are typically trained with a softmax cross-entropy objective over the full item vocabulary. For a typical large number of possible items $K$, the final classification layer dominates memory, requiring $O(nK)$ logits and gradients to materialize for a batch of $…
Artyom Sabitov, Daniil Volkov, Alexey Zaytsev · 📄 PDF
A Systematic Sample Size Analysis of ML-Based Path Loss Prediction for LPWAN
Low Power Wide Area Networks like LoRa are increasingly deployed for smart city applications, requiring accurate path loss prediction for effective network planning. Traditional (empirical) propagation models often exhibit limited accuracy in these scenarios. We investigate machine learning models f…
Robert Bitterling, Christian Nettersheim, Jörn Hees, Michael Rademacher · 📄 PDF
Cross-View Feature Matching: Survey, Benchmarking, and Foundation-Model Perspectives
Cross-view feature matching aims to establish reliable correspondences across images with large viewpoint variations. Over the past decade, the field has evolved from task-specific models toward increasingly unified and generalizable correspondence models, with recent progress further driven by the …
Songlin Du, Xiaoyong Lu, Zeyu Wu, Xiaobo Lu, Guobao Xiao, Bin Fan, Jiayi Ma, Takeshi Ikenaga · 📄 PDF
AlbumentationsX: One Augmentation Pipeline for Images and Related Annotations
Augmentation can corrupt a training example when an image and its annotations receive different random changes. A crop must use the same coordinates for the image, mask, boxes, keypoints, stereo views, video frames, or volume. Code paths that choose these values separately can silently misalign the …
Vladimir Iglovikov · 📄 PDF
A Recommendation System Approach for Interference-Robust Sensor Subset Selection
This paper develops a method for sensor-subset selection for tracking. Prior work showed that low-cost acoustic Received Signal Strength Indicator (RSSI) measurements can be used to recommend subsets of sensor nodes whose expensive sensing modalities, such as cameras, can achieve high tracking accur…
Kaan Buyukkalayci, Kyle Pak, Merve Karakas, Christina Fragouli · 📄 PDF
Scheduling Mixed RL Rollouts Beyond Prefix Locality
Modern reinforcement learning (RL) post-training pipelines for large language models (LLMs) increasingly combine rollout workloads across multiple domains and feedback paradigms. Prefix-aware routing improves inference efficiency through cache reuse and load balancing, but it does not control how he…
Zetao Hong, Song Yuan, Yuanhao Ding, Yibo Zhu, Daxin Jiang, Zhibin Wang, Chen Tian · 📄 PDF
DACRI: Decision-Aware Causal Intervention Ranking for Critical Supply Chains
Detecting or attributing a supply-chain disruption is not the same as selecting the intervention that maximizes recoverable net value. We present CriticalSCM-Bench v1, a controlled synthetic benchmark with causal ground truth, paired factual/counterfactual rollouts, and an explicit net-value objecti…
Shiqi Huang, Jiani He, Dingyan Shang, Yihua Xu, Jize Li, Yan Lyu, Lashimi Muraleedharan Nair · 📄 PDF
Conditional Independence Tests for Constraint-Based Causal Discovery: A Survey
Conditional Independence (CI) tests are the statistical engine of constraint-based causal discovery: in algorithms such as PC (Peter-Clark) and FCI (Fast Causal Inference), skeleton pruning and key orientations follow directly from CI decisions. This survey reviews CI testing with emphasis on assump…
Pavel Averin, Theodoros Moysiadis, Ioannis Katakis · 📄 PDF
Hierarchical Empirical-Bayes Naive Bayes: Minimax Smoothing and Calibration with AODE Extension
The Naive Bayes (NB) classifier remains a standard choice for categorical data, yet its widely used smoothing rules, such as Laplace, Lidstone, Krichevsky-Trofimov, and the $m$-estimate, all prescribe a fixed smoothing strength that ignores feature cardinality, sample size, and class imbalance, indu…
Nguyen Thai Anh, Truong Viet Vu, Tran Thien Thanh, Vo Nguyen Quoc Bao, Ngo Hoang Tu · 📄 PDF
MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment
Existing Multimodal Large Language Models (MLLMs) predominantly rely on image-text pairs for modality alignment pretraining, mapping global image representations to long textual descriptions. However, this image-level alignment suffers from referential ambiguity: models struggle to infer the corresp…
Changhao Xiang, Shangyu Xing, Zhen Wu, Jianbing Zhang, Xinyu Dai · 📄 PDF
A Quantum Roadmap for Softmax Attention: Exact Born-Rule Analogs for Softmax Attention on the Probability Simplex
The attention mechanism forms the foundation of many modern AI models such as the Transformer. In one subclass of problems where attention is used, inputs and outputs are bound to the probability simplex so that all outputs sum to one. In this setting, softmax attention admits an exact, component-by…
Eric A. F. Reinhardt, Adam J. Hauser · 📄 PDF
Beyond a Bag of Features: Set-Level Instability in Sparse Autoencoders
Shani et al. (2026) show that LLM representations broadly recover human category boundaries, while failing to reflect fine-grained typicality structure. Their analysis uses cosine similarity over dense model representations. We revisit their approach using overlap over active sparse autoencoder (SAE…
Nikolai Bolik, Lennart Stöpler, Artur Andrzejak · 📄 PDF
R4DSG: Relative 4D Scene Graph Memory for Object-Centric Question Answering in Long Egocentric Video
Long-horizon egocentric video is a rich substrate for wearable AI assistants, but object-centric questions such as where an item was moved, when it last changed state, or why it was relocated remain difficult because caption- and transcript-based memories rarely preserve persistent object identity o…
Ke Ma, Yamin Mao, Weiming Li, Shuai Tan, Yijie Zhong, Hao Chen, Haofen Wang, Meng Wang · 📄 PDF
Workflow Cards: Structured Summaries of Workflow Executions Using Provenance Data
Model Cards and Data Cards have demonstrated the value of structured, human-readable documentation for machine learning artifacts, capturing their context, parameters, limitations, and intended use. However, these practices remain focused on static artifacts (the datasets and trained models themselv…
Nicola Giuseppe Marchioro, Gabriele Padovani, Amal Gueroudji, Rafael Ferreira da Silva, Wesley Brewer, Valentine Anantha… · 📄 PDF
V-FiLLM: Verified Financial LLM Reasoning Benchmark
While existing benchmarks have made substantial progress in evaluating LLMs across STEM domains, financial reasoning over structured data remains comparatively less explored. We introduce V-FiLLM, a framework that generates financial reasoning benchmarks from executable computation trees grounded in…
Alicia Larsen, Victoire Laurent, Aulia Kharis Rakhamsari, Lara Turgut, Nino Antulov-Fantulin · 📄 PDF
Multiclass Sentiment Analysis for Identifying Political Viewpoints
The rapid growth of social media has created vast amounts of political discourse, which provides valuable opportunities to analyze public opinions and identify different political perspectives. Sentiment Analysis (SA) is a core task in Natural Language Processing (NLP) that allows the computational …
Girma Yohannis Bade, Olga Kolesnikova, Jose Luis Oropeza, Grigori Sidorov · 📄 PDF
3D Weighted Geometric Graph Neural Networks for Sheep Facial Pain Assessment
Deep learning systems perform mainly within the 2D for a single image domain and take the face as a single-dimension representation, losing sight of the 3D anatomy of sheep and cross-landmark spatial relationships that are intrinsic to the clinically proven Sheep Pain Facial Expression Scale (SPFES)…
Alam Noor, Luis Almeida, Mohamed Daoudi · 📄 PDF
A Comparative Evaluation of Deep Learning Object Detection Models on a Real-World Multi-Plant Dataset from Africa
The application of computer vision in agriculture has shown significant potential for improving crop monitoring and precision farming. However, many existing approaches rely on controlled datasets that do not adequately represent realworld farming conditions, particularly in underrepresented regions…
Ismail Ismail Tijjani, Sunusi Muhammad Ibrahim, Amina Ibrahim Khaleel, Lanre Olusegun Akinola, Fatima Isa Jibrin, Muhamm… · 📄 PDF
Entropy-Centric Explainable AI for Remote Sensing Image Segmentation
Artificial intelligence (AI) has become a powerful approach to solving complex problems in critical domains. Many concerns arise regarding the decision-making process of its models, mainly due to deep neural networks outperforming their peers at the cost of ambiguity in feature extraction and predic…
Ali Saleh, Abdul Karim Gizzini, Mohamad Ghassany, Ali J. Ghandour · 📄 PDF
Quantum Coordination Advantages in AI State-Tracking Tasks: Semantic Compilation and Latent Memory
We prove inference-time quantum coordination advantages for specified AI state-tracking tasks. A solver compresses semantic history into a future-accessible boundary state and later answers a query. We count communication $B$, persistent instance-dependent memory $M$, and local work $D$; classical r…
Ming Yang · 📄 PDF
SkillZip: Evaluation-Free Skill Compression for Self-Evolving Agents by Discovering Reusable Structure
Self-evolving agents accumulate reusable skills by appending successful procedures and failure fixes. Over time, the same requirement is often restated in several branches, examples, and warnings, while common action sequences are copied rather than reused. The resulting skill becomes expensive to i…
Xiaofan Bai, Hongqiang Lin, Chao Liu, Yantao Zhang, Xuan Jin, Xipeng Cao, Yuhong Li · 📄 PDF
RTSKG: Building a Rail Transit Station Knowledge Graph Dataset
Rail transit systems play a vital role in urban mobility and economic development. As key components of such systems, rail transit stations function as critical transport hubs that enhance urban accessibility and stimulate development in surrounding areas. City-level rail transit station related tas…
Shutong Zhu, Tianxing Wu, Runfeng Liu, Yuang Gu, Xuan He, Yuan Zhu · 📄 PDF
Why Does CLAUDE.md Keep Growing? Catastrophic Remembering in Agentic Coding
Agentic coding READMEs like CLAUDE.md grow without bound in real repositories, stopping only when the repository retires or someone rewrites the file wholesale. We trace this to imperfect recall: appending an instruction is always cheap, but once an instruction's rationale is gone, deleting it witho…
Kushal Chakrabarti · 📄 PDF
Two-stage Odd Residual Flows for Mean-Preserving Probabilistic Time Series Forecasting
Probabilistic forecasting plays an essential role in risk-sensitive decision-making, particularly in long-horizon settings. However, existing approaches often face a fundamental trade-off between distributional flexibility and accurate mean prediction. Traditional parametric methods, such as Mean Va…
Kiran Madhusudhanan, Christian Klötergens, Lars Schmidt-Thieme, Vijaya Krishna Yalavarthi · 📄 PDF
sLTN: Structural Logic Tensor Networks
Logic Tensor Networks (LTN) provide a neurosymbolic framework in which first-order logic is interpreted through tensor operations, enabling logical constraints to be integrated with differentiable learning. However, the original formulation of LTN is primarily suited to data represented as flat coll…
Davide Rinaldi, Luciano Serafini · 📄 PDF
Attention-Path Fragility as an Uncertainty Signal in Large Language Models
We propose that a model's uncertainty about a token is reflected not only in the breadth of its output distribution but also in whether a confident prediction is \emph{fragile} under perturbation of its attention pathways. We instantiate this as ASMI (Attention-Subnetwork Mutual Information), a trai…
Minsoo Kim, Sungyoung Ji, Kisung Moon, Ilyong Yoon · 📄 PDF
From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop
The Workshop on Trustworthy Natural Language Processing (TrustNLP), co-located with major ACL conferences since 2021, has grown from 8 proceedings papers to 41 over six editions, documenting a field-wide transition from post-hoc interpretability of static models to mechanistic understanding and proa…
Rahul Gupta, Abhinav Mohanty, Anaelia Ovalle, Anil Ramakrishna, Anubrata Das, Apurv Verma, Jwala Dhamala, Ninareh Mehrab… · 📄 PDF
How to Verify Consistency of Probabilistic Claims
When a probabilistic predictor answers many conditional-probability queries, are its answers self-consistent, and can this be verified in polynomial time? This problem is of interest for AI safety, where safety is derived from honesty about probabilistic predictions of unwanted outcomes potentially …
Orr Paradise, Oliver Richardson, Yoshua Bengio, Shafi Goldwasser · 📄 PDF
Test-Time Self-Evolving GUI Visual Grounding via Reflection-Guided On-Policy Self-Distillation
GUI Visual Grounding is a fundamental capability for GUI agents. Existing models typically freeze their parameters after deployment, limiting their ability to adapt to unseen interfaces. Although recent methods attempt to adapt models via test-time reinforcement learning, they cannot reflect upon fa…
Shiyu Xuan, Zechao Li · 📄 PDF
Long-Horizon AI Research for Grothendieck Constant: A Case Study in Human-AI Mathematical Collaboration
AI agents are increasingly used in mathematics research, but it is often unclear how to use them effectively. Towards this, we present an extensive case study of how AI was used to improve bounds on the Grothendieck constant $K_G$, which captures the hardness between combinatorial problems and their…
Alan Li, Rahul Saha, Anton Xue, Swarat Chaudhuri, Adam Klivans, Pravesh K Kothari, Raghu Meka · 📄 PDF
ConVAWG: A Retrieval-Grounded Framework for Controlled Synthetic Dialogue Generation in Violence Against Women and Girls
Synthetic dialogue generation offers a way to study conversational dynamics in sensitive domains where real data are difficult to access, release, or annotate. The underlying abuse may occur online or offline: threats and coercion can appear directly in messages, while behaviours such as surveillanc…
Chen Lyu, Xingwei Tan, Simon Cullen, Shelley Wilson, Lois Arthurs, Arshad Jhumka, Gabriele Pergola · 📄 PDF
Surgical WAM: A World-Action Model for Data-Efficient Surgical Robot Learning
Learning reliable surgical manipulation policies is bottlenecked by the scarcity of action-labeled demonstrations: teleoperated surgical robot (e.g., dVRK) trajectories with synchronized kinematics are costly to collect, while surgical tasks demand precise contact handling, long-horizon reasoning, a…
Wenrui Bao, Tianyun Jiang, Zhiben Chen, Ser-Nam Lim, Peter D. Peng, Yuzhang Shang · 📄 PDF
Beyond headcount and human capital: The Effective Cognitive Population as a decomposable capacity unit for AI-era planning
National planning counts population, human capital, and artificial-intelligence preparedness in separate ledgers. Demographic accounting has advanced from headcount to skills-adjusted stocks and still debates how much age structure retains once skills are modeled, yet no existing unit carries the co…
Kwan Soo Shin · 📄 PDF
Environmental and Economic Implications of Artificial Intelligence Data Centers in the United States
In this study, we use electricity demand growth, cooling requirements, and backup system operation to evaluate the environmental and economic implications of artificial intelligence data centers in the United States. Our results indicate that impacts are not determined solely by facility design, but…
Johanna Bolaños-Zuñiga, Alberto J. Lamadrid · 📄 PDF
Flow-based conditional cardiac anatomy generation for virtual cohorts
Cardiac digital twin research is moving from subject-specific anatomical replicas toward virtual cohorts that represent clinically relevant population subgroups. Yet access to representative imaging-derived anatomy datasets remains limited by cohort size, subgroup sparsity, and data-sharing constrai…
Konstantinos Kevopoulos, Beatrice Moscoloni, Benjamin Alheit, Cameron Beeche, Julio A. Chirinos, Alexander Heinlein, Mat… · 📄 PDF
Memory-, Circuit-, and Ansatz-Efficient VQLS for CFD on Hybrid Quantum-HPC Systems
Fluid dynamics workloads are dominated by repeated solves of large, structured linear systems, motivating the search for quantum acceleration. The Variational Quantum Linear Solver (VQLS) is a leading near-term candidate, but practical deployment on hybrid quantum--high--performance computing (HPC) …
Chao Lu, Muralikrishnan Gopalakrishnan Meena, Eduardo Antonio Coello Perez, Kalyana Chakravarthi Gottiparthi, Seongmin K… · 📄 PDF
The Unified Evaluation App for DNA Data Storage Codecs
Background: Deoxyribonucleic acid (DNA) data storage is a paradigm with great potential for ultra-dense and durable information preservation. However, the rapid proliferation of coding schemes, or codecs, each with their own design constraints and reporting practices, has led to a fragmented landsca…
Aleksandar Anžel, Chisom Anyabolu, Leon Wimbes, Luca Staus, Ihsan Tri Heldian, David Sonnabend, Khawla Elhadri, Samuel B… · 📄 PDF
Unsupervised Detection of Groundwater Storage Anomalies in Ghana Using GRACE Satellite Data
Groundwater variability in Ghana remains poorly characterized due to limited long-term in-situ observations. This study investigates groundwater storage anomalies using GRACE-derived data from 2004-2024 combined with statistical analysis and unsupervised machine learning. Groundwater anomalies were …
George Yamoah Afrifa, Theophilus Ansah-Narh, Marcellin Atemkeng · 📄 PDF
Monophonic Audio Synthesizer Using FPGAs
Signal synthesis is used in every aspect of the electronics world, where sinusoidal waveforms are used to perform functions such as clocking, signal transmission, feedback controls, and other applications. Digital synthesis is the method of approximating sinusoidal waveforms using digital logic, whe…
Michael Smith, D. G. Perera · 📄 PDF
SLAC: Access-Driven CPU-to-GPU Side-channel Attacks via System-Level Cache on Apple Silicon
Modern heterogeneous System-on-Chip designs integrate CPU cores and a GPU that share a last-level cache (LLC) or system-level cache (SLC). This sharing exposes a new cross-domain attack surface, and existing attacks on integrated platforms either exploit coarse-grained cache-occupancy contention or …
Tianhong Xu, Saion K. Roy, Ruyi Ding, Aidong Adam Ding, Yunsi Fei · 📄 PDF
FSGen: Agile Fused and Sparse Accelerator Generator with Accurate Power Model for LLM Applications
With the growing demand of artificial intelligence (AI) applications, large language models (LLMs) have become important workloads in many domains. The question of how to efficiently generate optimal AI chip accelerator designs remains unresolved and challenging. Currently, there is a lack of end-to…
Jay Zhe-An Mok, Qijun Zhang, Zhiyao Xie · 📄 PDF
TDMA Based Communications Control Co-Design for Cooperative Carrying: Delay Calibration and Sampling-Rate Optimization
Multi robot teams performing cooperative transportation face a fundamental challenge: maintaining stable control while keeping communications efficient. This paper investigates how adaptive sampling time adjustment informed by measured network delay and strategic leader rotation can distribute wirel…
Zahra Seifaei, Maximilian Luebke, Torsten Reissland, Danial Dehghani, Norman Franchi · 📄 PDF
GenTrack3: Hybrid Stochastic-Deterministic Online Multi-Object Tracking with Cluster-Aware Association
Multi-object tracking (MOT) involves maintaining consistent target identities as objects dynamically enter and leave a scene. Deterministic approaches, such as tracking-by-detection with data association, produce reproducible results and are computationally efficient, but they rely heavily on motion…
Toan Van Nguyen, Rasmus G. K. Christiansen, Dirk Kraft, Leon Bodenhagen · 📄 PDF
FactorDrive: Adaptive Multi-Step Reasoning Driven by Planning-Critical Factors for End-to-End Autonomous Driving
Vision-language models (VLMs) have advanced scene understanding and enabled explicit reasoning in end-to-end autonomous driving. However, existing methods insufficiently integrate spatial-physical evidence into planning reasoning, while reasoning adaptation remains coarse-grained and falls short of …
Guolei Huang, Tengfei She, Yuxuan Lu, Yao Huang, Yuqi Ye, Yongjun Shen · 📄 PDF
Nonlinear Model Predictive Control of a Robotic Soft Esophagus
Strictures caused by esophageal cancer can narrow down the esophageal lumen, leading to dysphagia. Palliation of dysphagia has driven the development of a Robotic Soft Esophagus (RoSE), which provides a novel in vitro platform for esophageal stent testing and food viscosity studies. In RoSE, perista…
Dipankar Bhattacharya, Ryman Hashem, Leo K. Cheng, Weiliang Xu · 📄 PDF
A Semantic Communication Approach to Fiducial Marker Processing in 5G-Enabled Edge SLAM
Autonomous robots increasingly rely on edge computing to offload computationally intensive perception tasks while maintaining real-time operation over 5G networks. However, conventional fiducial marker detection pipelines provide limited opportunities for efficient task partitioning, making them poo…
Boris Radovanovic, Vukan Ninkovic, Katarina Vidojevic, Buda Bajic Papuga, Dejan Vukobratovic · 📄 PDF
Satellite Trajectory Optimization via Proximal Policy Optimization for Space Debris Avoidance
Collision avoidance systems are commonly used to avoid fragmentation events occurring in Low-Earth Orbit (LEO) and Geosynchronous Equatorial Orbit (GEO). However, these events have been growing in frequency as orbital congestion worsens with the launch of megaconstellations. Consequently, conjunctio…
Logan Luna, Juan Ortiz Couder, Raul Alejandro Vargas-Acosta · 📄 PDF
Predictive safety filter enhanced curriculum learning control for efficient vehicle dynamics controller
Recent advances in learning-based control have enabled impressive achievements in solving complex control problems in various domains. However, since learning-based control may not be able to realize safety-guaranties, it is of great importance to enhance safety and robustness while maintaining good…
Baocong Zhang, Siliang Lu, Chenyang Li · 📄 PDF
Removing Infrastructure Barriers in Human-Robot Collaboration Through Wireless Reconfigurable Cells
Human-Robot Collaboration (HRC) plays a vital role in dynamic, high mix, low volume industrial scenarios such as remanufacturing, which frequently face workcell rearrangements. Traditional setups are constrained by power and data cabling, restricting modularity and reconfigurations, while the select…
Emma Takács, Mátyás Hajós, Ádám Juniki, Ádám Fischer, Zoltán Komáromi, Kristóf Abai, Dániel Horváth, Sándor Máthé, Konst… · 📄 PDF
TAMS: Task-Aware Multi-View Adaptive Streaming for Wireless Telerobotic Manipulation
Wireless telerobotic manipulation relies on timely multi-view video feedback, but the available uplink bandwidth is often limited and dynamic. This paper presents Task-Aware Multi-View Adaptive Streaming (TAMS), a system that allocates video bitrate according to the current manipulation phase. TAMS …
Zexin Deng, Zhenhui Yuan, Lu Tian, Subhash Lakshminarayana, Longhao Zou · 📄 PDF
Efficient Real-World Online Reinforcement Learning for Robot Manipulation via Centralized Training and Critic Decomposition
Real-world online reinforcement learning (RL) provides a promising approach for training robotic manipulation policies directly in the physical world, avoiding the sim-to-real gap and enabling continuous policy refinement through human-in-the-loop interaction. Recent methods have demonstrated sample…
Changhao Li, Yifang Zhang, Heng Zhang, Davide Torielli, Damiano Gasperini, Arturo Laurenzi, Luca Muratore, Arash Ajoudan… · 📄 PDF
SLIM-0.5B: Learning Action-Grounded Predictive Latents for Robot Manipulation
Vision-language-action policies rely on large multimodal backbones to jointly perform perception, language conditioning, and action generation at every control step. Much of this capacity supports open-domain semantics, whereas continuous robot manipulation primarily requires compact representations…
Jingkai Wang, Zihan Tang, Gu Zhang, Mingyu Cao, Jiapeng Chen, Jingjiao Zhao, Xiansheng Chen, Pengwei Wang, Lemao Liu, De… · 📄 PDF
RoboSeg: Online Part-Level Semantic Reconstruction for Robotic Manipulation via a Single Eye-in-Hand Camera
Robotic manipulation requires perception systemsthat identify actionable parts such as handles, rims, triggers,and tool tips, not merely object categories or point clouds. This paper presents RoboSeg, a part-level semantic reconstructionsystem that links vision-language model (VLM) functional-partdi…
Zhaochen Lan, Mengxiang Lin · 📄 PDF
WRAP: Wasserstein-Robust Adaptive Plug-in for Robot Localization
Robotic localization under changing sensing conditions can suffer from biased errors and miscalibrated covariances. We present WRAP, an adapter-agnostic Wasserstein-robust plug-in for nonlinear extended Kalman filter (EKF) and error-state Kalman filter (ESKF) stacks. A causal module supplies time-va…
Minhyuk Jang, Astghik Hakobyan, Jungjin Lee, Naira Hovakimyan, Insoon Yang · 📄 PDF
Hierarchical Fast--Slow ReAct Agent for Zero-Shot Object-Goal Navigation
Zero-shot object-goal navigation (ZSON) requires a robot to find a named object category in a building it has never entered. The prevailing approach scores frontiers with a vision--language \emph{value map}: every decision is another argmax over the map as it currently stands, and the evidence behin…
Zhaochen Lan, Zhi Yang, Yuxiang Fu, Mengxiang Lin · 📄 PDF
Entanglement-Free Trajectory Planning for Tethered Mobile Robots with a Slack Tether
In motion planning algorithms for tethered mobile robots, the entanglement state of the tether is a critical aspect to consider during the planning phase. This is particularly important in case of a slack tether, where the shape of the tether is not determined solely by the geometry of the environme…
Gianpietro Battocletti, Dimitris Boskos, Bart De Schutter · 📄 PDF
RoSE: A Robotic Soft Esophagus for Endoprosthetic Stent Testing
Soft robotic systems are well suited for developing devices for biomedical applications. A bio-mimicking robotic soft esophagus (RoSE) is developed as an in vitro testing device of endoprosthetic stents for dysphagia management. Endoprosthetic stent placement is an immediate and cost-effective thera…
Dipankar Bhattacharya, Sherine Jesna V. A., Leo K. Cheng, Weiliang Xu · 📄 PDF
XPolicyLab: A Unified Standard and Open Ecosystem for Robot Policy Evaluation and Deployment
Robot policy evaluation and deployment remain fragmented by model-specific software dependencies, data representations, and runtime interfaces, so that connecting N policies to M evaluation environments requires O(NM) separate integrations. We present XPolicyLab, a unified standard and open ecosyste…
XPolicyLab Community, Tianxing Chen, Yue Chen, Tian Nian, Zijian Cai, Guangyu Chen, Wenwei Lin, Qiwei Liang, Peicheng Xi… · 📄 PDF
Disentangling Co-Occurring Retinal Pathologies with Saliency-Guided Sparse Expert Routing
Retinal fundus images frequently exhibit multiple co-occurring pathologies, yet standard deep learning classifiers apply static, identical computation to every image regardless of the underlying disease distribution. We propose a novel architecture that resolves this via sparse conditional computati…
Nagur Shareef Shaik, Jeongwoo Park, Yeong-Jin Kim, Jaeuk Jung, Hyunjung Oh, Dong Hye Ye · 📄 PDF
REFRAMED: Towards Realistic Audio Description Generation for Movies
Audio Description (AD) is a verbal narration of key visual content in videos, enabling access for visually impaired audiences. Unlike standard video captioning, AD is a structured editorial task: descriptions must be inserted into gaps in dialogue and must convey only what is needed to understand th…
Igor Sterner, Mirella Lapata, Alex Lascarides, Frank Keller · 📄 PDF
NTIRE 2026 Low-light Enhancement: Twilight Cowboy Challenge
This paper presents a review of the NTIRE 2026 Low-light Enhancement: Twilight Cowboy Challenge. The objective of the competition was to merge a set of misaligned smartphone images in the raw domain, captured in low-light conditions, into a single, clean image. Introduced setup simultaneously addres…
Aleksei Khalin, Egor Ershov, Artyom Panshin, Sergey Korchagin, Georgiy Lobarev, Arseniy Terekhin, Sofiia Dorogova, Amir … · 📄 PDF
ADOPD: Reference-Privileged On-Policy Distillation for MLLM-Based Industrial Anomaly Detection
Industrial anomaly detection (IAD) requires identifying fine-grained deviations from normal visual patterns. Multimodal large language models (MLLMs) can improve recognition accuracy by comparing query images with references at inference time, but these benefits rely on additional retrieval and proc…
Jingtai He, Shiyuan Meng, Wenchao Meng, Qinmin Yang · 📄 PDF
Modern Backbones Improve Multi-task DETR for Mammography Classification and Lesion Localization
Joint exam-level prediction and candidate-region localization may improve the usefulness of AI support in mammography. We study this setting using a multi-task DETR framework, where shared representations support both image-level malignancy prediction and lesion localization, and evaluate its perfor…
Dinh Tan Nguyen, Quang-Hien Kha, Le-Hoang Nguyen, Minh-Toan Dinh, Xuan-Huy Nguyen, Dac Phu Ho, Cao Truong Tran, Sai Ho L… · 📄 PDF
From Diagnosis to Correction: Benchmarking and Improving Real-World Table Parsing
Recent document parsers achieve table TEDS scores above 93 on OmniDocBench v1.6, yet community feedback and our audit reveal persistent failures on complex real-world tables. To quantify this gap, we introduce TableParseMap, a diagnostic benchmark of 916 real-world tables organized into five challen…
Jutao Xiao, Yuan Qu, Dongsheng Ma, Fan Wu, Tianyao He, Weihong Li, Jie Yang, Yu Qiao, Bin Wang, Conghui He · 📄 PDF
DistMoE: Private-data Rehearsal-free Routing in Mixture-of-Experts for Distributed Instruction Tuning
Multimodal Large Language Models (MLLMs) have shown strong multimodal instruction-following ability, but adapting them to diverse visual-language domains typically assumes centralized data access and costly joint training. This is restrictive when data is distributed across private, domain-specific,…
Mainak Singha, Niccolò Biondi, Elisa Ricci, Subhankar Roy · 📄 PDF
Beyond Hazard Resemblance: Contrastive Event Adjudication for Training-Free Video Anomaly Detection
Video anomaly detection (VAD) aims to identify and temporally localize abnormal events in videos. Supervised methods learn anomaly decision boundaries from target-domain annotations but require substantial in-domain data. Existing training-free methods leverage the rich semantic knowledge and reason…
Wenti Yin, Xiang Wang, Huaxin Zhang, Hanqing Wang, Hongbo Shao, Changxin Gao, Nong Sang · 📄 PDF
Overcoming Data Scarcity and Confidentiality in Hardware Assurance via Synthetic Generation
Hardware assurance relies on scanning electron microscopy (SEM) to verify nanoscale structures, but assembling the large, high-quality datasets required for automated analysis is impeded by time-intensive acquisition and strict intellectual property (IP) constraints on proprietary designs. We propos…
Gijung Lee, Ronald Wilson, Damon L. Woodard, Domenic Forte · 📄 PDF
Learning How the World Evolves: Extrapolative Video World Models via Latent Dynamics Reasoning
The world evolves following its dynamics, i.e., its laws of motion. However, leading video diffusion models largely fit the pixels without modeling how the pixels transit over time. Thus, they render visually plausible frames but may not accurately obey the laws. To capture the dynamics purely from …
Haodong Li, Shaoteng Liu, Tianyu Wang, Chongjian Ge, Sihui Ji, Jiahan Zhang, Xin Lin, Haolin Lu, Zhe Lin, Manmohan Chand… · 📄 PDF
Perception Before Supervision: Self-Contained Visual Distillation from Counterfactual Blind Spots
Self-improvement for multimodal large language models (MLLMs) is typically driven by reward-based methods that provide only coarse scalar feedback. Distillation offers a richer alternative through dense token-level supervision, but in the visual domain it usually depends on privileged context constr…
Shravan Venkatraman, Omkar Thawakar, Ritesh Thawkar, Abdelrahman Shaker, Rao Muhammad Anwer · 📄 PDF
MoNo: Multiscale Optimal Transport Neural Operator for Solving PDEs on General Geometries
Transformer-based neural operators have achieved substantial progress in solving Partial Differential Equations (PDEs) by projecting spatial observations into compact latent tokens and learning physical interactions in latent spaces. However, we reveal that existing learnable projection mechanisms c…
Zijiang Yang, Xiaomeng Wu, Dongmei Fu · 📄 PDF
ReliableNet: A Chance-Constrained Approach to Trustworthy Classification in Deep Learning
A prediction that is both confident and wrong is a critical reliability failure because it can bypass abstention and human review precisely when the model is mistaken. Empirical risk minimization (ERM) controls average loss but not this failure directly, while calibration, uncertainty estimation, co…
Ange-Clément Akazan, Ineza Remy Mugenga, Abebe Geletu, Jean Medard Ngnotchouye, Issa Karambal · 📄 PDF
C$^2$A: Coupling Spatial Evidence with Clinical Priors via Co-occurrence Aware Class Attention for Multi-Label Chest X-Ray Classification
Thoracic pathologies rarely occur in isolation, yet standard multi-label classifiers rely on shared global descriptors, discarding \emph{where} findings lie and \emph{how} they co-occur. We propose \textbf{C$\mathbf{^2}$A} (Co-occurrence Aware Class Attention), a classification head that explicitly …
Akash Gogineni, Nagur Shareef Shaik, Aasrith Mandava, Adnan Masood, Dong Hye Ye · 📄 PDF
AirFlow: Context Preserving and Multi-Rate State Modeling for Air Quality Forecasting
Accurate air quality forecasting is essential for public health and urban environmental management, but remains challenging because pollutant channels differ in periodicity and distribution drift, while their concentration trajectories contain both multi-scale dependencies and rapid changes. Recent …
Fan Yang, Nan Chen, Yijie Dong, Yuchen Zhang, Wei Zhang · 📄 PDF
Parameter Exploration for RLVR via Variational Learning
Exploration has been a focus of reinforcement learning research for a long time. Recently, there has been growing evidence that it is also an important ingredient in LLM reinforcement learning recipes that can significantly impact downstream performance. Many existing methods control exploration in …
Vatsal Venkatkrishna, Nico Daheim, Iryna Gurevych · 📄 PDF
Macaron-V1: Towards Open Continual Learning with Self-Improvement and Mixture-of-LoRA
Macaron-V1 is an open agent-model family for experiential intelligence: learning from experience in real environments and continuing to learn after deployment. It is organized around two system goals. Adaptation is pursued through recursive improvement of versioned model-harness pairs, where experie…
Mind Lab, :, Vin Bo, Asher Cai, Jingwei Cao, Song Cao, Vic Cao, Amelia Chen, Andrew Chen, Kaijie Chen, Cleon Cheng, Stev… · 📄 PDF
Deep Multimodal Wearable Sensor Fusion for Detection of Body-Focused Repetitive Behaviors
Body-focused repetitive behaviors, such as hair pulling and skin picking, are compulsive motor actions commonly associated with obsessive-compulsive and anxiety disorders. Their early, objective detection remains difficult because the movements are subtle and overlap with ordinary, non-pathological …
Samaneh Rezaeimanesh, Mohsen Behradfar, Mohammad Fili, Guiping Hu · 📄 PDF
RA-FinBERT: Rule-aware LoRA adaptation for low-resource financial sentiment classification
Financial sentiment analysis converts unstructured financial news into quantitative signals that can support market analysis and decision-making. Existing work on resource-efficient financial NLP has largely focused on compressing or adapting pretrained language models, with less attention to combin…
Fan Zhang, Jiaming Li · 📄 PDF
Real-Time Climate Risk Assessment for Supply Chain Resilience: A Data-Driven Nowcasting Framework for Colombian Agriculture
This paper presents a methodological framework for real-time climate risk assessment using data-driven nowcasting techniques to enhance supply chain resilience in Colombian agricultural contexts. Climate variability in Colombia, characterized by irregular rainfall, temperature fluctuations, and recu…
Hernan J. Silva-Sosa · 📄 PDF
RynnValue: Scaling Robotic Value Foundation Models with Temporal Distance
General-purpose reward models are increasingly the bottleneck for scaling robot learning, yet the recipe for learning value-related capabilities from large-scale heterogeneous corpora remains underexplored. Existing approaches tie supervision to task-internal anchors such as preferences or normalize…
Dongchi Huang, Hongyin Zhang, Bohan Hou, Siteng Huang, Zhian Su, Hang Guo, Tong Lu, Zhaofeng Xu, Jiahao Tang, Jianfei Ya… · 📄 PDF
Logarithmic-Free Moment and Generalization Bounds for Uniformly Stable Algorithms
Uniform stability is a classical tool for controlling the generalization error of a learning algorithm. Bousquet, Klochkov, and Zhivotovskiy (2020) showed that the problem can be reduced to a moment inequality for a sum of weakly interacting functions of independent random variables. Their bound con…
Thanh Nguyen-Cung, Binh T. Nguyen · 📄 PDF
Financial Numerical Prediction and Allocation as Token Generation
Financial prediction typically relies on task-specific regression, ranking, or policy heads, separating the language model from the numerical object ultimately evaluated. We investigate whether a causal language model can instead represent forecasts and decisions directly through constrained token g…
Xu Ouyang, Moontae Lee · 📄 PDF
Space-Creating versus Dead Possession: An Off-Ball Possession-Quality Index for Broadcast Football
Ball possession is the most-cited and most-misleading number in football: 60% recycled in one's own half is not 60% spent pinning the opponent back. Existing event-based possession-value frameworks (expected threat, VAEP, on-ball value) price on-ball actions but ignore the off-ball question a steril…
Seongjin Choi · 📄 PDF
Consilience for Verifier-Free Test-Time Scaling
Test-time scaling often uses an external verifier, such as compilers and test cases in coding or trained value functions in robotics applications, to obtain high-quality rollouts. Verifier-free test-time scaling (or VF-TTS) is gaining extensive attention as a mechanism to enhance Large Language Mode…
Lecheng Kong, Like Hui, Haitao Mao, Jun Huan · 📄 PDF
Fairness in Link Prediction Beyond Demographic Parity: A Reproducibility Study
In fair ranked link prediction, demographic parity ($Δ_\mathrm{DP}$) is a common fairness metric. Yet, Mattos et al. (2025) argue that it fails to detect exposure bias because it ignores where links appear in the ranking. In this study, we reproduce this claim by showing that $Δ_\mathrm{DP}$ can ind…
Valentijn Oldenburg, Floris de Kam, Stef de Wildt, Jarno Nilson Balk · 📄 PDF
MedPixel: A Unified Pixel-Language Model for Medical Reasoning and Segmentation
Reliable medical image understanding requires models to connect clinical language and visual reasoning with pixel-level grounding. Yet medical vision-language models often lack precise localization, whereas medical segmenters typically rely on explicit target categories or precise spatial prompts. T…
Haoyu Yang, Meixing Shi, Zengjie Chen, Haoran Sun, Haitao Leng, Xiaoming Shi, Yuxiang Cai, Yankai Jiang · 📄 PDF
Distill Skills into Weights, Not Prompts: Abstract Skills as Privileged Signals for On-Policy Self-Distillation
Reinforcement learning with verifiable rewards yields no group-relative signal when rollout groups are uniformly correct or uniformly wrong, which account for 63.0-68.0% of groups in our experiments. We propose SKALD (Skill-Anchored Latent Distillation), an on-policy self-distillation framework that…
Yubo Jiang, Fengying Xie, Zhiguo Jiang, Haopeng Zhang · 📄 PDF
Multi-Agent AI Safety as an Institutional Design Problem
AI agents increasingly work inside systems that govern how they delegate tasks, move information, execute actions, and use shared resources. Recent work already shows that deployment rules can change collective behavior. Here we ask which parts of an AI institution produce safety and how they do it.…
Abdullah X · 📄 PDF
Mismatch Matters: On-Policy Distillation Beyond Token Agreement
On-policy distillation (OPD) has emerged as a core component of modern LLM post-training pipelines, yet we reveal a failure mode: degenerate agreement, where students exploit repetitive loops to achieve near-perfect token agreement with the teacher despite globally flawed responses. We therefore shi…
Zichao Yu, Chengzhi Yu, Shengze Xu, Yujin Han, Bingqing Jiang, Xu Wang, Difan Zou · 📄 PDF
CEAA: A Cognitive Embodied Agents Architecture for Interactive Computing Systems
The development of embodied Intelligent Virtual Agents (IVAs) that have cognitive capabilities in real-time interactive virtual environments remains a challenge, even with today's advancements in technology. Existing architectures are often focused on either the implementation of low-level reactive …
Aimilios Hadjiliasi, Louis Nisiotis · 📄 PDF
Agentic Auto-Research is Fuzz Testing
Autonomous research agents can generate experiments faster than researchers can validate them. Researchers have responded by scaling the proposer and ranking more samples with a learned judge or human reviewers. We argue that this *generate-and-rank* paradigm misses the problem of sparse feedback. W…
Yifeng He, Jicheng Wang, Yinzhe Zhao, Jiachen Liu, Hao Chen · 📄 PDF
Agentic Harnesses: LLM-Driven Verification Layers for Robot Autonomy
Advances in advanced artificial intelligence tools have sparked research in robot autonomy, but the development of such systems has largely focused on execution rather than verifying the feasibility actions planning models propose. Like general-purpose LLMs, robotics planning models carry risks: bia…
Rohan Bhagra, Mahantesh Halapannavar, Uddhav Bhattarai · 📄 PDF
Towards Expert-level Medical AI for Real-time Video Consultations
Audio-visual interaction is the standard for patient-physician consultations, enabling natural communication and effective assessment of illness through non-verbal cues. While text-based AI has shown promise, it discards essential perceptual dimensions and limits patients who cannot articulate sympt…
Mahvish Nagda, Jihyeon Lee, Matthew Thompson, Chunjong Park, Tim Strother, Valentin Liévin, Roma Ruparel, Akshay Goel, T… · 📄 PDF
Stealing Reasoning Traces from Proprietary LLM APIs
Leading large language model providers now conceal their models' step-by-step reasoning, or chain-of-thought, to protect intellectual property and limit information leakage. Rather than storing these traces server-side, providers return them to the client as blocks of encrypted text, which the clien…
Alexander Panfilov, David Schmotz, Ilia Shumailov, Luca Beurer-Kellner, Joachim Schaeffer, Ameya Prabhu, Jonas Geiping, … · 📄 PDF
Sci-VBench: Evaluating Knowledge- and Reasoning-Intensive Video Generation in Science Domains
We introduce Sci-VBench, a comprehensive benchmark for evaluating knowledge- and reasoning-intensive video generation across scientific domains. It contains 1,253 expert-annotated examples spanning 60 subjects across four core disciplines: Natural Science, Healthcare, Humanities & Social Sciences, a…
Diandian Zhang, Tingyu Song, Lin Fu, Zheyuan Yang, Yilun Zhao · 📄 PDF
ArchAgent v2: A Case Study with the Data Prefetching Championship
Agentic artificial intelligence has shown great promise in automating algorithm design, but scaling similar techniques to computer microarchitecture discovery remains challenging due to vast search spaces, strict hardware budgets, and long simulation times. In this work, we present ArchAgent v2, a f…
Abraham Gonzalez, Raghav Gupta, Akanksha Jain, Hanna Alam, Alexander Novikov, Po-Sen Huang, Matej Balog, Marvin Eisenber… · 📄 PDF
Energy-Structured Latent World Models with Neural Time Fields for Physically Constistent Open-World Motion Planning
Physically consistent motion planning remains a fundamental challenge in embodied AI, as generated trajectories must strictly conform to real-world execution dynamics. While latent world models offer a promising approach by predicting these dynamics, existing methods learn unconstrained future repre…
Yapeng Liu, Yuanzhao Zhai, Bo Ding, Huaimin Wang, Lin Wang · 📄 PDF
SHE: Trajectory-driven Safety Harness Evolution for LLM Agents
The safety of large language model (LLM) agents depends not only on model weights but also on the agent harness that manages context, memory, tools, permissions, and runtime control. Existing safety mechanisms often treat the harness as a fixed deployment artifact, limiting their ability to evolve w…
Wanying Qu, Qinghua Mao, Yu Li, Jiyao Liu, Xin Zhang, Dadi Guo, Yanxu Zhu, Qingyu Liu, Leitao Yuan, Xi Lin, Shanfeng Zhu… · 📄 PDF
BDH-CQ: In-Context Learning with Recurrent Latent Reasoning
We introduce BDH-CQ, a reasoning model that combines in-context learning with recurrent latent reasoning. Inputs presented at inference time continuously update the model's recurrent memory; the model then solves a query through iterative computation in a high-dimensional latent space, without verba…
Björn Engdahl, Adrian Kosowski, Jan Chorowski, Zuzanna Stamirowska, Przemysław Uznański, Junlin Jiang, Rohan Phadke, Rem… · 📄 PDF
Fusion Training for Mathematical Generalization in Large Language Models
Thinking Mode Fusion (TMF) enables large language models to support both concise responses and long-form reasoning by unifying a non-thinking mode and a thinking mode within a single model. However, its training dynamics, including the \emph{data ratio} and \emph{training schedule} between the two m…
Congfeng Cao, Pengyu Zhang, Jelke Bloem · 📄 PDF
DSLE: A Learning Environment for Dark Souls Boss Encounters
We introduce the Dark Souls Learning Environment (DSLE), a containerized platform that presents all 22 boss encounters of Dark Souls: Remastered as game-playing agent benchmarks through a Gymnasium-style interface. DSLE combines real-time combat, high-dimensional visual input, and sparse terminal re…
Derin Gezgin, Jim O'Connor, Tanner Goodwin, Gary B. Parker · 📄 PDF
GENCO - A Unified Neural Solver Embedded in a Development Framework for Steady-State Grid Analysis
Foundation models are transforming business workflows and boosting productivity, yet they remain largely absent from engineering domains such as power system analysis, where strict physical consistency must be enforced. We present GENCO (GEometric Neural Corrective Optimizer), a unified neural solve…
Alban Puech, Matteo Mazzonelli, Tamara R. Govindasamy, Mangaliso Mngomezulu, Héctor Maeso-García, Thomas Tolhurst, Javad… · 📄 PDF
From Values to Benchmarks: Evaluating Large Language Models for Governmental Use in Dutch
Large language models are increasingly being deployed in governmental settings, yet few existing evaluation frameworks jointly reflect the values of public administration and the linguistic requirements of non-English contexts. We present the "Grip on LLMs" framework, a systematic evaluation suite f…
Laurens Samson, Iva Gornishka, Gossa Lô, Yuki M. Asano, Sennay Ghebreab · 📄 PDF
Multimodal Model Diffing for Feature Discovery and Control
Multimodal Large Language Models (MLLMs) exhibit strong visual understanding, yet the internal features that cause these behaviors remain difficult to identify, audit, or control. While applicable to post-hoc inspection, hidden states that are decomposed into interpretable feature directions using s…
Hunar Batra, Lachin Naghashyar, Ashkan Khakzar, Philip Torr, Christian Schroeder de Witt, Constantin Venhoff, Ronald Cla… · 📄 PDF
Beyond Naturalness: Probing Automated Text-To-Speech Evaluators on Linguistically Grounded Dimensions
Automated Text-to-Speech (TTS) evaluation methods (Mean Opinion Score (MOS) predictors and Audio Large Language Models (Audio-LLM) judges) are expected to reflect human perception, yet it is unclear how well they capture the distinct aspects of speech that listeners actually perceive. We deconstruct…
Oluwanifemi Bamgbose, Simon Rosen, Jash Shah, Lindsay Devon Brin, Hoang H Nguyen, Anke Koelzer, Rachel Hansen, Tara Boga… · 📄 PDF
ARMOR: Accelerating RTL Simulation by Mitigating the Front-End Bottleneck Using Node Compression
RTL simulation is indispensable in chip design. High-performance simulators typically lower each node in the RTL graph into an instruction sequence. Although this per-node lowering enables aggressive compiler optimizations, it dramatically increases the code footprint, severely exceeding instruction…
Jiaping Tang, Jianan Mu, Zhiteng Chao, Jingzhong Wen, Jing Ye, Huawei Li · 📄 PDF
ReVolt: Power Delivery Network-Aware Voltage Droop Control for 2.5D PIM Chiplet Architectures
Processing-in-memory (PIM)-based 2.5D multi-chiplet platforms are enablers for machine learning (ML) workloads. However, their performance is affected by the power delivery network (PDN), where varying chiplet-level current demand induces spatially and temporally varying voltage droop. These droop e…
Vibhanshu Sharma, Alish Kanani, Miao Sun, Janardhan Rao Doppa, Umit Y. Ogras, Partha Pratim Pande · 📄 PDF
C2C-Explorer: An Exploration Framework for Chip-to-Chip Interconnect Architectures in LLM Cloud Computing Systems
The scaling-up of large language models (LLMs) necessitates computing systems to have multi-processor-chip architectures, elevating the importance of chip-to-chip (C2C) communication. However, designing efficient C2C hardware architectures for LLM workloads faces three key challenges: generating rea…
Jiayi Li, Di Wu, Qingxu Li, Hongxiao Zhao, Jiaqi Yang, Anjunyi Fan, Wenbin Zhang, Boqiang Wu, Shuting Liu, Shifeng Fang,… · 📄 PDF
Eco-SoC: A Sustainable VLSI Architecture for Energy-Proportional Artificial Intelligence
In an era defined by escalating climate change and the pervasive deployment of edge intelligence, the environmental cost of semiconductor manufacturing and operation has reached a critical threshold. As Deep Learning (DL) accelerators dominate System-on-Chip (SoC) die area, achieving true sustainabi…
Jatin Chopra · 📄 PDF
IDRAAK: From Multi-Agent NLP to Few-Shot Prompting for Semantic Drift Detection in Technical Requirements
Translating technical requirements across languages can introduce semantic drift, altering numerical constraints, polarities, modalities, or other specification-critical meaning. IDRAAK is presented as an interpretable framework for detecting such drift using a language-independent Semantic Requirem…
Shiva Ahir · 📄 PDF
Benefits of Shifting Passenger Traffic from Air to Rail: A Case Study of California High-Speed Rail
This study provides a method to quantify the benefits of shifting passenger traffic from air to high-speed rail from the perspective of flight-delay cost reduction. We first estimate the number of flight reductions for airport origin-destination pairs based on the high-speed rail ridership forecasts…
Kaijing Ding, Lu Dai, Mark Hansen · 📄 PDF
Large-Market Discipline in Combinatorial Double Auctions: No Assembly, Bundle Selection, and Complementarities
We study double auctions for markets in which goods are valuable in bundles, such as data, model weights, and fine-tuned AI assets. A key friction in such markets is No Assembly: a platform may be unable, for legal or technical reasons, to combine components supplied by different sellers into a sing…
Konstantinos E. Zachariadis, Yongxin Yang · 📄 PDF
Inverse mask design for interference lithography using automatic differentiable wave propagation
Interference lithography (IL) is powerful for fabricating high-resolution periodic nanostructures, but designing masks to produce non-periodic patterns remains challenging. We introduce a gradient-based optimization framework for binary IL mask design using automatic differentiation. The forward mod…
Chuntian Cao, Jangwoon Sung, Jack Griffiths, Yuan Gao, Xi Yu, Paul Baity, Nikhil Tiwale, Zhitian Shi, Juhong Ahn, Shinja… · 📄 PDF
Generating two-mechanical mode entangled cat states, and steady-state entanglement, in cavity optomechanics in the presence of dissipation
We investigate a dissipation-engineering approach to produce a phase-dependent collective-mode Schrödinger cat state involving two modes. Our model features a single cavity mode that interacts with two spectrally identical mechanical oscillators. Both oscillators are coupled through a phase-dependen…
Sanket Das, Jason Twamley · 📄 PDF
Dual-polarization control of broadband nonreciprocal thermal radiation by combining local and nonlocal metasurfaces
Nonreciprocal thermal radiation offers a route to decouple spectral directional absorptivity and emissivity, thereby enabling new paradigms in thermal-photonic systems. However, in magneto-optical platforms, the intrinsic gyroelectric response generally confines observable nonreciprocity to transver…
Shuang Xia, Mengqi Liu, Wenjian Wan, Jialong Wang, Weihao Yang, Chaoran Wang, Huiqin Ma, Jun Qin, Hua Li, Yuan Wang, Lei… · 📄 PDF
Spectrally flat and broadband doublepumped fiber optical parametric amplifiers
We study theoretically and experimentally spectrally flat and broadband double-pumped fiber-optical parametric amplifiers (2P-FOPAs). Closed formulas are derived for the gain ripple in 2P-FOPAs as a function of the pump wavelength separation and power, and the fiber non-linearity and fourth order di…
J. M. Chavez Boggio, J. D. Marconi, S. R. Bickham, H. L. Fragnito · 📄 PDF
Net electron spin rotation in a plane-wave pulse: Holonomy set by the anomalous magnetic moment
We compute the spin rotation that survives after a relativistic electron has crossed a plane-wave laser pulse of finite duration. In the interaction picture built on the exact $g=2$ evolution, the Thomas-Bargmann-Michel-Telegdi equation becomes parallel transport by a connection with constant coeffi…
N. S. Akintsov, A. P. Nevecheria, S. N. Andreev, Qing-Hua Qin · 📄 PDF
Engram-E2VID: Reference-Based Event-to-Video Reconstruction via Generative Activation of Appearance Engrams
Reference-based event-to-video reconstruction aims to recover target RGB frames from a reference frame and the event stream captured over the reference-to-target interval. Although events provide fine-grained temporal cues, they encode sparse and asynchronous log-intensity changes rather than absolu…
Feiyu Ji, Xiang Li, Hao Ma, Tianxiang Huang, Qingxin Lu, Mengqi Ji, Lei Han, Xiaokang Yang, Xiaoyun Yuan · 📄 PDF
Transverse quantum-state characterization of programmable electron optics
Programmable electron optics -- electronically controlled phase plates -- underpin proposals from dose-efficient phase imaging to shaped-electron X-ray sources, nearly all assuming a pure, fully coherent delivered wave whose purity has never been measured. Here we reconstruct the transverse density …
Shengbo You, Paolo Rosi, Enzo Rotunno, Alberto Roncaglia, Luca Belsito, Amir H. Tavabi, Rafal E. Dunin-Borkowski, Vincen… · 📄 PDF
Abruptly autofocusing waves enter space-time
Whereas conventional Gaussian focusing gradually concentrates optical energy around the focal plane, abruptly autofocusing waves maintain a low peak intensity over most of their evolution before undergoing a sudden, high-contrast intensity surge at a prescribed focus. Since their introduction in 201…
Nikolaos K. Efremidis, Demetrios N. Christodoulides · 📄 PDF
Harnessing thermo-optic dynamics for frequency-agile soliton microcombs
Dissipative Kerr soliton microcombs enable compact and scalable frequency comb sources for precision metrology, spectroscopy, communications and coherent LiDAR, where broad and reliable frequency tuning is essential. Thermo-optic response can support thermal locking during soliton operation, enablin…
Yang Liu, Suwan Sun, Yueguang Zhou, Yanjing Zhao, Chaochao Ye, Xinda Lu, Yi Zheng, Leif Kastuo Oxenløwe, Kresten Yvind, … · 📄 PDF
OPERA: Operator-residual feedback for reliable autonomous optical experiments with language-model agents
Autonomous agents choose actions using scores that may not reflect experimental success. We developed OPERA, an operator-residual framework for optical experiments. It represents experimental actions as optical operators and evaluates their outcomes using physically interpretable residuals. Operator…
Ning Xu, Xiang Zheng, Fuqiang Zhong, Huadong Wang, Xiaolong Wu, Zhiyuan Liu, Hui Ning · 📄 PDF
Twist-angle Control of Nonlinear Interference in a ZnO Nanowire/Monolayer WSe$_2$ Hybrid Structure
Nanoscale devices that integrate materials of different dimensionalities (0D, 1D, and 2D) hold great potential for advanced applications in photonics and optoelectronics. A fundamental requirement for the development of such devices is the engineering and control of light-matter interactions beyond …
Maximilian Tomoscheit, Benedikt Mathes, Alexander Zaunick, Moritz Willems, Edwin Eobaldt, Priyanka S. Prakash, Eva Perlt… · 📄 PDF
Noise-driven pseudovorticity multipoles in self-focusing beams with quintic saturation
We investigate pseudovorticity generation in Gaussian beams undergoing self-focusing under amplitude and phase noise, using the cubic-quintic nonlinear Schrödinger equation. Pseudovorticity, defined as the curl of the optical momentum flux, characterizes local rotational flow in the absence of phase…
Chengbo Zhang, Xiaohui Gao · 📄 PDF
Pulse-Duration Control of Subcycle Multiband Electron Dynamics Extends the High-Harmonic Cutoff in a Light-Driven Insulator
We demonstrate pathway-selective control of extreme-ultraviolet high-harmonic generation by jointly tuning laser pulse duration ($5$ - $29$ fs) and intensity ($0.8$ - $74$ TW/cm$^2$). Many-cycle pulses at moderate intensities, $\sim 6$ TW/cm$^2$, promote cumulative carrier transfer over successive o…
Hortense Allegre, Simon V. B. Jensen, Joseph J. Broughton, Tim Klee, Yan Li, Jon P. Marangos, Nicolas Tancogne-Dejean, A… · 📄 PDF
Differentiable eigendecomposition-free rigorous coupled-wave analysis for general photonic structures
For periodic photonic structures with full permittivity and permeability tensors and arbitrarily oriented optical axes, we present, to the best of our knowledge, the first anisotropic rigorous coupled-wave analysis framework that supports automatic differentiation and efficient GPU acceleration. The…
Enbo Yang, Qiang Song, Weiwei Cai · 📄 PDF
Numerical Model of a Multiple-Input-Multiple-Output Distributed Acoustic Sensing System with Joint Phase and Birefringence Estimation
In this work, we introduce and experimentally validate a numerical model for a Multiple-Input-Multiple-Output Distributed Acoustic Sensing (MIMO-DAS) system that accounts for dynamic perturbations of fiber birefringence and of the common optical phase of the backscattered signal (or polarization-ave…
Diane Prato, Mehran Mokhtari Sheramin, Renaud Gabet, Elie Awwad · 📄 PDF
Structured coherence: A modern perspective on optical coherence as a resource
Optical coherence is a well-established branch of physical optics in which the statistical properties of fluctuating optical fields are described in terms of correlation functions over continuous spatial and temporal degrees of freedom (DoFs). Nevertheless, in any practical setting, only discrete Do…
Ayman F. Abouraddy, Bahaa E. A. Saleh · 📄 PDF
EpiBench: Can LLMs Understand Epitopes for Antibody Drug Discovery?
Epitopes determine where antibodies bind antigens and shape downstream therapeutic properties such as functional blockade and escape resistance, making epitope understanding central to antibody drug discovery. Although large language models (LLMs) have shown strong biomedical reasoning ability, it r…
Zirui Wang, Jiaqi Wang, Qinghan Wang, Yuzhi Xu, Gang Du, Tingjun Hou, Odin Zhang · 📄 PDF
A Low-Power Wearable Respiratory Sensor for Non-Invasive Stress Monitoring
Respiration provides a continuously available window into physiological state and behavior. However, monitoring it outside controlled settings remains challenging because a wearable system must capture small body deformations while remaining comfortable, low power, and robust to changes in posture a…
Mohammad Hosseini, Hamed Khatounabadi, Mohammad Fakharzadeh · 📄 PDF
Multiparametric MRI Radiomics and Machine Learning Framework for Predicting Treatment Response in Glioblastoma
Distinguishing True Progression (TP) from Pseudo-Progression (PsP) after chemoradiotherapy remains a major diagnostic challenge in GBM, as both entities present near-identical appearances on conventional contrast-enhanced post-treatment MRI. This distinction carries substantial clinical weight, sinc…
Suchibrata Patra · 📄 PDF
Curriculum Multiple Shooting for Robust Training of Neural and Universal Differential Equations
Neural ordinary differential equations (NODEs) and universal differential equations (UDEs) provide flexible and popular frameworks for learning interpretable dynamical systems from noisy time-series data. However, training these models remains challenging, and versatile methods that robustly handle …
Sebastian Persson, Giacomo Fabrini, Branwen Snelling, Fabian Fröhlich · 📄 PDF
THBKG: A Temporal Biomedical Knowledge Graph for Decision-Aligned Clinical Advancement Prediction
Inadequate target--disease linkage accounts for 40--50\% of Phase~II efficacy failures, so anticipating which programmes will advance would let sponsors back the hypotheses most likely to reach patients. What a programme can be judged on is the evidence that supported its linkage \emph{when it enter…
Pui Chung Siu, Claudia Cabrera, Mani Mudaliar, Arkaitz Zubiaga · 📄 PDF
PyOMES: an open-source framework for biochemical process modelling
PyOMES is a Python-based, Open-source Modelling Environment for (bio)chemical process Simulation that aims to simplify the modelling of dynamic (including steady state) processes. This is done in a generalied, modular way to facilitate modelling a broad range of biological, chemical, and biochemical…
Ethan Errington, Tom Vinestock, Jaewook Lee, Miao Guo · 📄 PDF
A Quantum Circuit Framework for Protein Ensemble-Level Energetics
Proteins occupy heterogeneous free-energy landscapes in which high-entropy ensembles converge toward compact, low-energy basins with multiple sub-states. Molecular dynamics can access these landscapes at atomic resolution, but exhaustive sampling remains computationally demanding. Meanwhile, most qu…
Pratik Patil, Bhushan Bonde, Bhaskar Choubey · 📄 PDF
Viveka: Context-Aware Sensing for Energy Efficiency in Smart Wearables
The proliferation of multi-sensor Internet of Things (IoT) systems, from Body Sensor Networks (BSNs) to industrial monitoring, is increasingly constrained by strict energy budgets and limited on-device storage. Continuous high-fidelity sensing leads to rapid battery depletion and data gaps that comp…
Nikhil Sreekumar, Abhishek Chandra · 📄 PDF
LC-Implicit-QAOA: Active-Workspace-Capped Exact Objective-and-Gradient Evaluation for Training over Bounded QUBO Light Cones
QAOA training repeatedly queries an objective and all shared gradients, making exact evaluation a feasibility bottleneck even when QUBO terms have bounded causal cones. Building on established causal-cone restriction and adjoint differentiation, LC-Implicit-QAOA profiles cone structure and induced-e…
Chih-Chung Hsu · 📄 PDF
RASP-QAOA: Resource-Aware Per-Instance Selection for Exact QAOA Simulation
Exact QAOA simulation spans several computational representations whose useful regions differ sharply across graph structure, circuit depth, precision, and available memory. Choosing only a backend name hides these differences: an executable choice also fixes the representation, adapter, precision m…
Chih-Chung Hsu · 📄 PDF
Balanced Routing for Symmetric Quantum Circuits
Mapping quantum programs to restricted physical chips requires SWAP operations, incurring depth and error penalties. In symmetric programs, this routing overhead breaks theoretical symmetry because identical logical roles experience unequal shuffling. While often attributed to hardware topology alon…
Samuel Punch · 📄 PDF
PLoRA: An NDP-Enhanced Pooled-Memory System for Cost-Efficient Multi-LoRA Serving
Multi-LoRA serving is how one base model becomes thousands of specialized variants, one adapter per user, task, or agent, and the deployments can hold 1000-plus adapters. Serving them is hard because the workload inverts what GPUs provide: terabytes of memory against only tens of TFLOPS, and because…
Zhongkai Yu, Ohm Rishabh Venkatachalam, Zheng Wang, Yikai Li, Yichen Lin, Zihao Yu, Yuke Wang, Liu Liu, Xulong Tang, Shu… · 📄 PDF
Zero-Instruction Sensor Reads: Register-Mapped Peripherals and Hardware PWM on a Five-Stage Soft Processor
We present a case study in application-driven specialization of a five-stage soft processor, evaluated on the inner control loop of a reaction-wheel self-balancing bicycle. Starting from a custom 32-bit RISC core in the MIPS tradition, we specialize the design in two ways. First, two frequently acce…
Nathanael Ren · 📄 PDF
An Open-Source Power Measurement Platform for System-Level Semiconductor Testing
Accurate power measurement is not only essential for evaluating the energy efficiency of modern embedded and semiconductor systems, but power draw is an important proxy during stress testing. Industrial semiconductor test equipment, however, is often expensive and difficult to integrate into flexibl…
Linus Bantel, Sarah Rottacker, Dirk Pflüger · 📄 PDF
Automated Synthesis of Heterogeneous, Hierarchical, Scoped Coherence Protocols
Processor design is converging on a new model of cache-coherent shared memory characterized by heterogeneity, hierarchy, and scopes. Protocols like CXL or AMBA CHI are used as global protocols to combine multiple clusters, each with its own cluster-level coherence protocols. Manually designing shims…
Fletch Rydell, An Qi Zhang, Nicolai Oswald, Andres Goens, Vijay Nagarajan, Daniel Sorin · 📄 PDF
Breaking Memory Bottlenecks in Quantum Control Systems for More Precise Experiments and Higher Throughput Computing
As quantum computing continues to demonstrate promise and attract growing attention, there is an increasing need for more precise experiments to advance the development of quantum devices, as well as higher circuit throughput to validate more domain applications. However, this need is hindered by a …
Yicheng Guang, Neel Vora, Yilun Xu, Yueqi Chen, Gang Huang · 📄 PDF
Shape-Aware Oriented Bounding Box (OBB) to Horizontal Bounding Box (HBB) Conversion
Accurate object detection in aerial and satellite imagery is dependent upon the bounding box representation. This is especially true for spatially oriented objects such as ships or aircrafts. Oriented Bounding Boxes (OBB) have a tighter fit and more robust non-max suppression compared to Horizontal …
Badha Rathna Sabhapathy, Gotam Dahiya, Vishesh Vatsal · 📄 PDF
Robust-WAM: Bridging Generative Pretraining and Semantic Foresight in World-Action Models
Mainstream World-Action Models (WAMs) adapt pretrained video generation models (VGMs) for robot control, transferring their learned dynamics prior for action prediction. These VGMs are typically trained in a variational autoencoder (VAE) latent space. However, the VAE latent space is optimized for p…
Haodong Yan, Junfeng Li, Junjie He, Zhide Zhong, MingMing Yu, Wenxuan Song, Jiaguan Zhu, Yangyang Zheng, Yuqiao Du, Jiad… · 📄 PDF
GAUGE: A Measurement-Grounded Benchmark for Physical Fidelity in Simulation Engines and Video World Models
Physics engines facilitate large-scale training and evaluation for embodied intelligence, while generative video world models are emerging as implicit simulators of future states and interactions. However, existing evaluations of physical fidelity are often conducted in isolation and rely heavily on…
Shuai Wang, Yaxin Feng, Xuekun Jiang, Shihan Tian, Ningyu Yan, Xing Shen, Chaoyang Lyu, Hui Wang, Yunsong Zhou, Hanqing … · 📄 PDF
SkillMemo: Expert-guided Skill Memory Framework for Compositional Embodied Manipulation
Embodied visuomotor models, including Diffusion Policy (DP) and Vision-Language-Action (VLA) models, have demonstrated promising performance on robotic manipulation benchmarks. However, their potential remains fundamentally constrained by the scarcity of large-scale embodied trajectory datasets, lea…
Changyuan Wang, Chubin Zhang, Zhenyu Wu, Runhao Li, Angyuan Ma, Ke Chao, Yinan Liang, Xiuwei Xu, Ziwei Wang, Yansong Tan… · 📄 PDF
TRACE: Learned Proprioceptive Odometry for Legged Robots under Unreliable Contact Conditions
In this paper, we present TRACE (Tokenized Robust Attention for Contact-Aware Estimation), an end-to-end learned proprioceptive odometry estimator for legged robots under unreliable contact conditions. The proposed estimator directly predicts relative displacement, relative rotation, and body-frame …
Taehyeon Kong, Woojin Kim, Jemin Hwangbo · 📄 PDF
Observation-Grounded Self-Predictive Reinforcement Learning for Visual Continuous Control
Sample-efficient policy learning from pixels is a long-standing challenge in reinforcement learning (RL). Recent dynamics-based representation learning methods have significantly improved the sample efficiency of model-free visual RL by learning dynamics-aware representations through auxiliary predi…
Xinwei Liu, Junyuan Liang, Jianting Zhang, Wuhui Chen · 📄 PDF
Beyond Flat Policies: Hierarchical Post-Training for Embodied Agents in Robotic Manipulation
Vision-language-action (VLA) models have demonstrated remarkable capabilities in robotic manipulation by leveraging pretrained vision-language models. However, existing post-training methods predominantly optimize VLA models as flat policies, making it difficult to explicitly model task progression …
He Kong, Zengjue Chen, Qi Wang, Qianli Xing, Runliang Niu, Peidong Liu, Jiawei Li, Shiqi Wang, Yi Chang · 📄 PDF
Adaptive-WAM: Quality-Guided Early-Exit Planning from Intermediate Video-Diffusion Features
Large video diffusion models provide rich spatiotemporal priors for autonomous driving, but existing world-action models often inherit the cost of iterative future-video generation even though deployment only requires an ego trajectory. We ask a more basic question: how much of a video diffusion mod…
Sining Ang, Yuguang Yang, Yan Wang · 📄 PDF
Topometric Autonomous Vehicle Localization by Combining Visual Embeddings and Feed-Forward 3D Models
Effective Visual Localization (VL) requires a map of the environment that combines compactness for efficient scalability with robustness against visual appearance changes and metric precision. Through low-dimensional image embeddings, Visual Place Recognition (VPR) is able to successfully meet the f…
Eulogio Quemada-Torres, Alberto Jaenal, Francisco-Angel Moreno, Javier Gonzalez-Jimenez · 📄 PDF
IcFuzz: Fuzzing Isaac Sim with Semantic Stage Guidance and Multi-level Mutation
Robotics simulators serve as a foundational infrastructure for embodied AI, facilitating safe and scalable robotic system development. NVIDIA Isaac Sim has emerged as one of the most popular simulators, distinguished by its GPU-accelerated physics engine and photorealistic rendering, which enable hi…
Zhixiang Chen, Zhuangbin Chen, Ruoxi Jia, Zeqin Liao, Wei Li, Jinyang Liu, Zibin Zheng · 📄 PDF
ErgoSurf: Ergodic Control for the Coverage of Unknown Surfaces
Contact-centric tasks on surfaces, ranging from inspection and cleaning to sanding and polishing, require robots to systematically cover the surface while maintaining stable contact. Ergodic control generates trajectories that spend time at a location proportional to a desired, task-specific spatial…
Stefan Schneyer, Timo Bachmann, Maged Iskandar, Korbinian Nottensteiner, Alin Albu-Schäffer, Freek Stulp, João Silvério · 📄 PDF
VIDP: Variable Impedance Diffusion Policy for Compliant Robot Manipulation from Diverse Demonstrations
Contact-rich manipulation requires precise tracking and mechanical compliance, where variable impedance control can improve robustness in task success, whereas static compliance cannot adapt to varying contact constraints. Variable impedance skills can be learned from demonstrations, avoiding comple…
Hisham Khalil, Neil Fernandes, Thomas M. Kwok, Hsiu-Chin Lin, Yue Hu · 📄 PDF
Design and Evaluation of a Touchscreen-Based Teleoperation Interface for Robotic Manipulators
Intuitive teleoperation interfaces are crucial for the safe and effective operation of robotic manipulators in challenging environments. In the nuclear industry, surface contact tasks such as swab sampling require precise path and force tracking, obstacle avoidance, and sustained operator attention,…
Juan José García Cárdenas, Alperen Kenan, Hamidreza Raei, Paul Bremner, Manuel Giuliani, Arash Ajoudani, Adriana Tapus · 📄 PDF
Robot Learning from Human Demonstrations: Handwritten Alphabet Trajectories and Human-Likeness Evaluation
Learning from demonstration (LfD) provides a developmental framework through which robots can develop motor skills by observing and imitating human dynamics, reducing reliance on explicit programming to teach a skill to a robot. The resulting human-like robot motion is recognised as a key factor in …
Alperen Kenan, Paul Bremner, Manuel Giuliani · 📄 PDF
GeniWorld: A Generalizable Interactive World Model for Robotic Manipulation via Visual Actions
Generalist robot policies exhibit strong capabilities, but their robustness in complex and unseen environments remains limited. Scaling robot learning and evaluation in diverse real-world environments remains costly and challenging. Action-conditioned world models offer a promising alternative, but …
Chenghao Gu, Hanyang Yu, Jingbo Zhang, Haitao Lin, Wenyao Zhang, Jinghe Wang, Hanglei Jin, Shuzhao Xie, Jingyan Jiang, Z… · 📄 PDF
A Master-Salve Robot Manipulator for Needle-Based Teleoperation in MRI Chamber
We present a MR safe, master-slave robot manipulator for abdominal interventions in the MRI chamber. A human operated 2+1-DoF master controller manipulator transmits motion and force to a 2+1-DoF slave manipulator via fluid transmission. Jointly, a digital master controller provides multimodal contr…
Omar Curiel, Jing-Yuan Huang, Po-Chih Chen, Ji Ma, Qing Dai, Wenqi Zhou, David Lu, Holden H. Wu, Tsu-Chin Tsao · 📄 PDF
DyPES-VLA: Learning Shared Dynamics Priors and Embodiment-Specific Control for Cross-Embodiment Manipulation
Vision-Language-Action (VLA) models have become a powerful paradigm for robot manipulation, but training a single generalist policy for heterogeneous robot embodiments remains an open problem. Existing methods have two main limitations. First, they underuse dynamics priors shared across diverse visu…
Junfeng Li, Junjie He, Zhide Zhong, Yangyang Zheng, Pingyue Sheng, Jiayu Dong, Ruixin Li, Haodong Yan, Jiaguan Zhu, Tian… · 📄 PDF
$ω$-0: A Latent Predictive World Action Model for Concurrent Humanoid Loco-Manipulation
Humanoid household tasks often require concurrent loco-manipulation, where the robot must move, adjust posture, maintain balance, and manipulate objects as a single coordinated behavior. Yet existing humanoid policies typically decompose locomotion and manipulation, while recent world-action models …
Zhe Li, Zhenzhe Zhang, Yangyang Wei, Wenjie Zhang, Xichen Yuan, Peiyuan Zhi, Gen Li, Xinying Guo, Fengjie Gao, Jianfei Y… · 📄 PDF
Patient Pose Assessment Using a CT-Based Framework for Synthetic Data Generation
An adequate diagnostic quality of radiographs is essential for reliable diagnoses and treatment planning. The patient's pose during radiography is one of the most important factors determining the diagnostic quality. Since patient positioning is difficult and not standardized, an automated AI-based …
Manuel Laufer, Dominik Mairhöfer, Malte Sieren, Hauke Gerdes, Fabio Leal dos Reis, Arpad Bischof, Thomas Käster, Erhardt… · 📄 PDF
Learning visual representations for compositional analysis of artworks and photographs
Composition, the deliberate arrangement of visual elements, is central to how meaning, emotion, and aesthetic quality are conveyed in artwork, yet it remains among the least formalized dimensions of visual understanding. Prior work highlights a persistent gap in learning meaningful compositional rep…
Fatemeh Behrad, Tinne Tuytelaars, Johan Wagemans · 📄 PDF
CogVis: Must Open-Vocabulary Change Detection Perceive the Scene Anew for Every Query?
Earth-surface monitoring requires change detection models capable of recognizing arbitrary semantic categories. Open-Vocabulary Change Detection (OVCD) addresses this need. However, existing methods often entangle temporal perception, semantic discrimination, and region verification, causing unstabl…
Zijie Wang, Chen Zhong, Wei He · 📄 PDF
Visual Grounding in Zero-Shot Vision-Language Control
Vision-language models (VLMs) are increasingly used as zero-shot controllers, but successful trajectories do not necessarily show that decisions are grounded in visual input: simulator dynamics and conservative action priors can produce favourable scores without meaningful perception. We investigate…
J. de Curtò, Dayani Plasencia, Diego Sánchez, I. de Zarzà · 📄 PDF
BendTwin: Robust Dense-to-Sparse Physical Reconstruction with Bending-Aware Differentiable Spring-Mass Models
Reconstructing objects with mechanical properties from video observations enables physically consistent dynamic prediction, benefiting robotics planning and interaction. Existing spring--mass based physical driven reconstruction approaches offer efficient and differentiable physical reconstruction, …
Yixiong Jing, Qi Wang, Lin Chen, Junwei Jiang, Guangming Wang, Haibing Wu, Olaf Wysocki, Wanli Ma, Brian Sheil · 📄 PDF
Prior-SG: Task and Prior Driven Region Segmentation for Scene Graphs in Arbitrarily-Structured Environments
Hierarchical 3D scene graphs are a promising representation for high-level spatial reasoning in autonomous mobile platforms. However, existing extraction frameworks typically rely on purely local visual clustering or strict geometric heuristics, such as wall-separated rooms, which fail in open-plan …
Giorgio Tonetti, Laurent Kneip, Abel Gawel, Marco Hutter · 📄 PDF
Support Operation Factorization: Compositional Readout of Frozen Vision Encoders under Controlled Interventions
Compositional analysis of frozen vision encoders should determine both what changed and where it changed. Standard factor probes score these axes separately, however, and can reward multiple operations that reuse the same predicted slot. We call this failure operation laundering. We introduce an inj…
Zhongyao Wang, Wanli Ouyang, Taoyong Cui, Pheng Ann Heng · 📄 PDF
EvReflection: Event-Driven Micro-Dynamics for Reflection Removal
Despite remarkable progress in reflection removal, current methods primarily exploit static image priors from a single frame and still suffer from severe residual artifacts due to the inherent ambiguity between the reflection and transmission layers. In this paper, we propose leveraging event signal…
Jiaxiao Wang, Dachun Kai, Huyue Zhu, Quanquan Hu, Zhenyang Xu, Xiaoyan Sun · 📄 PDF
HOPE: Hand-Object Pressure Estimation from Monocular Videos
Estimating physical pressure from vision is essential for understanding contact-rich hand-object interaction. However, prior vision-based pressure estimation methods are largely limited to planar surfaces and single image input, making them difficult to apply to dynamic hand-object interaction with …
Subin Jeon, Byungjun Kim, Hanbyul Joo · 📄 PDF
CFGPNet: Cross-Attention-Based Fused Gradient Programmed Network Framework for Multispectral Object Detection
RGB--T object detection exploits the complementary strengths of visible and infrared imagery, supporting robust perception in low-light, adverse-weather, and complex multi-scale environments. However, existing methods still suffer from insufficient cross-modal interaction, unstable fusion from modal…
Nima Hatami, Karim Faez, Saeed Sharifian, Hamidreza Amindavar · 📄 PDF
Reversible Unlearnable Examples: Towards the Copyright Protection in Deep Learning Era
Significant advancements in deep learning have been made possible by the utilization of large datasets, underscoring the critical importance of copyright protection. Adding meticulously designed perturbations to examples, making them unlearnable has become a crucial approach for safeguarding data co…
Binze Wang, Jinyu Tian, Xingrun Wang, Xiaochen Yuan, Jianqing Li · 📄 PDF
EmoWorld: A Decoupled Affective Field for Controllable Emotional Video Generation
Emotion shapes how viewers interpret a scene, yet existing video generators entangle global atmosphere, affect-bearing semantic cues, and temporal progression within a single text condition. We present EmoWorld, a framework that decouples these factors within a frozen flow-matching video diffusion t…
Bingyuan Wang, Baistan Zhyldyzbekov, Kunyu Feng, Zeyu Wang · 📄 PDF
Depth-Guided Video Object Counting in Crowded Scenes
Our primary objective is to advance video object counting in crowded scenes, aiming to robustly count all instances of a target category based on given text or visual prompts. Existing methods rely on RGB information, limiting their discriminative ability in crowded and occluded conditions. To addre…
Yuanjing Xu, Xinyan Liu, Weidong Chen, Zixuan Zou, Linhao Zhang, Zhuangzhe Meng, Antoni B. Chan, Weigang Zhang · 📄 PDF
PRISM: Distribution-Gated Flow Matching for Controllable Unpaired Image Translation
Unpaired image-to-image translation must decide, per image, what to change and what to preserve without paired supervision. Many diffusion-based unpaired translators control preservation through a single global noise or guidance value applied across the image, which cannot separate content to keep f…
Elad Yoshai, Natan T. Shaked · 📄 PDF
MASS: Multiplayer World Models with Authoritative Shared State
Current video world models struggle in multiplayer environments because they entangle world state with view-dependent visual latents, leading to redundant compute, view inconsistencies, and poor scalability. We propose MAS (Multiplayer world models with Authoritative Shared State) to resolve this li…
Ziqi Cai, Siqi Yang, Yimu Wang, Zixian Gao, Yunheng Liu, Shuchen Weng, Erwin Wu, Kaipeng Zhang, Boxin Shi · 📄 PDF
TLNM: Externally Validated Tooth Detection, Numbering and Segmentation from Smartphone Photographs Using Mask R-CNN
Oral health issues affect billions globally, but the cost and limited access to professional dental care hinder preventive oral healthcare. Research relies on clinical-grade radiographs or intraoral camera images, unavailable for public self-screening. This study introduces a tooth localisation and …
Arash Nedaei, Henna Tiensuu, Elina Väyrynen, Saujanya Karki, Jaakko Suutala · 📄 PDF
UQ-Loc: Uncertainty-Aware LiDAR Scene Coordinate Regression
LiDAR-based Scene Coordinate Regression (SCR) maps point clouds directly to 3D scene coordinates, enabling precise 6-DoF localisation without explicit map retrieval. However, existing methods produce deterministic predictions, discarding aleatoric uncertainty that could improve robustness and downst…
Jacek Komorowski · 📄 PDF
Minimax Optimal Early-Stopped Gradient Descent for Gaussian Mixture Classification
In overparameterised classification, training data can be linearly separable even when the underlying distribution is not. In this setting, gradient descent (GD) on the logistic loss diverges in norm while converging in direction to a max-margin interpolating classifier, whose implicit bias can be s…
Alex Buna, Shirley Xiaoqi Liu, Patrick Rebeschini · 📄 PDF
MetaboLLM: a metabolomics-specialized large language model for biochemical knowledge integration and predictive metabolite graph construction
Metabolomics knowledge is distributed across heterogeneous resources and remains difficult to translate into predictive representations. We developed MetaboLLM, a metabolomics-specialized large language model adapted through continual pretraining, supervised fine-tuning, and structured retrieval, to…
Dohyun Ku, Min Gu Kwak, Francisco J. Pasquel, Jing Li · 📄 PDF
RxnCLF: Contrastive Transformation-Aware Reaction Foundation Model for Improved Reactivity Prediction
Reaction yield prediction remains challenging because labeled data are scarce and reaction space is both combinatorially large and sparsely populated, limiting the generalization of existing reaction representations. String-, fingerprint-, and graph-based reaction encodings only partially capture ch…
Yiting Zheng, Cheng Fang, Anthony Donofrio, Haote Li · 📄 PDF
Hypothesis Testing with Conditional Queries: Learnability and the Value of Interaction
Model evaluations may fix all tests before observing any responses or select later tests using earlier responses. We study this choice in a conditional-query model on a finite outcome space $\mathcal{X}$ with $|\mathcal{X}|=N$. We first ask which pairs of distribution classes can be reliably disting…
Zonghuan Xu · 📄 PDF
OTLesMix: Wasserstein Barycenter and Optimal Transport Map for Synthetic Lesion Generation with Diverse Shapes and Locations
The development of deep learning over the past decade has revolutionized medical imaging segmentation, allowing the extraction of precise descriptors from large volumes to characterize pathologies. Data augmentation is a technique widely regarded as a way to improve model training. It includes simpl…
Robin Trombetta, Carole Lartizien · 📄 PDF
Stochastic Dynamics on Persistence Diagram Space via Reinforcement Learning
Persistence diagrams (PDs) provide stable and interpretable summaries of multiscale topological structure. While substantial progress has been made in the statistical analysis of PDs, existing literature often treats diagrams as static objects and provide limited frameworks for probabilistic modelin…
Farzana Nasrin · 📄 PDF
The Tamed Subgradient Unadjusted Langevin Algorithm beyond Convexity
We study the problem of sampling from target distributions whose potentials are simultaneously non-smooth, subject to superlinear gradient growth, and non-convex. We introduce the Subgradient Tamed Unadjusted Langevin Algorithm (SG-TULA), a discretisation of the Langevin diffusion that operates dire…
Iosif Lytras, Nikolaos Makras, Sotirios Sabanis · 📄 PDF
Surv-IPTB: An Attention-Based Model for Estimating Individual Probability of Treatment Benefit with Survival Data
This work presents a novel attention-based framework for estimating the Individual Probability of Treatment Benefit (IPTB) in survival analysis contexts. The proposed model, called Surv-IPTB, directly quantifies the probability that a specific patient will experience extended survival time under tre…
Lev V. Utkin, Stanislav K. Kogan, Andrei V. Konstantinov · 📄 PDF
On-Policy Self-Distillation without Any Supervision
On-policy (Self-)Distillation (OPD / OPSD) has shown strong potential for post-training large language models (LLMs). However, existing methods still rely heavily on external supervision, including ground-truth signals, environmental feedback, or guidance from larger models, and therefore fall short…
Yijiang Li, Bingyang Wang, Yijun Liang, Yunjie Tian, Di Fu, Nuno Vasconcelos · 📄 PDF
RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction
Recent advances in reward modeling show a paradigm shift from discriminative reward models to generative reward models. However, despite their strong capabilities in response ranking, generative reward models have not realized their potential in reinforcement learning (RL). Our analysis reveals that…
Chenglong Wang, Ziming Zhu, Yifu Huo, Bei Li, Qiaozhi He, Yan Ding, Xiaoyang Hao, Yuxin Gao, Tianhua Zhou, Xiaojia Chang… · 📄 PDF
Optimal Rates for Learning with Monotone Adversaries
A monotone adversary observes an i.i.d. labeled sample and appends a finite number of further examples of its choice, every one of them labeled correctly by the target hypothesis. The learner sees a uniform shuffle of the combined sample and is scored on the original distribution. Every example is c…
Anay Mehrotra · 📄 PDF
Scalable estimation of VARMA models
Vector autoregressive moving-average (VARMA) models have long been considered impractical beyond moderate dimensions: the likelihood is non-convex, the parametrization is identified only up to equivalence, and every evaluation costs a pass over the entire series. Yet their moving-average term captur…
Daniel Paulin, Victor Elvira · 📄 PDF
CalibForge: Adversarial Solver Calibration for Scaling Learnable Terminal Tasks
Training terminal agents requires executable and verifiable tasks that are not merely solvable, but appropriately challenging for learning. Executable validation establishes feasibility, yet does not reveal how a task behaves relative to a given solver setting. In this paper, we present CalibForge, …
Fanzhe Meng, Guoxin Chen, Jiale Zhao, Shuang Sun, Zhiyu Lin, Wayne Xin Zhao, Ruihua Song, Ji-Rong Wen, Kai Jia · 📄 PDF
Toward Deployable Bangla Sign Language Recognition with Expert-Validated Data and a Lightweight Attention-Based Model
Deaf and hard-of-hearing people in Bangladesh communicate mainly through Bangla Sign Language (BdSL). Automatic BdSL recognition on personal devices could widen access to education and services. Existing systems use controlled-setting datasets without expert verification and heavyweight pretrained b…
Saad Ahmed, Md Khalid Syfullaha · 📄 PDF
Improving the Realism of Synthetic Clinical Benchmarks Under Utility Constraints
Synthetic clinical benchmarks for enterprise AI agents can pass existing utility checks and still remain structurally unrealistic, especially in privacy-sensitive healthcare settings where operational data are hard to access. We study how to improve such benchmarks without breaking the downstream ut…
Omid Bazgir, Md Nasir, Jacob Hoffman, Yang Yang, Manu Agrawal, Anusua Trivedi, Vinay Rao Dandin, Chris Gibbons, Christin… · 📄 PDF
The Illusion of Visual Tool-Use: A Causal Audit of Thinking with Images
The "thinking-with-images" paradigm equips multimodal LLMs with active visual operations such as crop-and-zoom. However, models using these operations often achieve only marginal or negative gains over direct inference at substantially higher token cost. They may also repeatedly crop irrelevant regi…
Zhiheng Wang, Bo Peng, Lai Wei, Chaochao Lu · 📄 PDF
BaKron: Efficient Quantization with Kronecker-Factored Hessians
We accelerate a family of algorithms for neural network quantization whose geometry is informed by any Kronecker-factored approximation of the Hessian. GPTQ-style adaptive rounding typically uses one-sided information derived from input activations. Two-sided Kronecker-factored Hessian approximation…
Johann Birnick, Rayan Saab · 📄 PDF
QuanTiMedAI: Quantum-Enhanced Time-Series Model guided by Agentic AI for Cardiac Arrest Mortality Prediction
Cardiac arrest remains one of the most lethal conditions encountered in intensive care units. Despite the growing availability of electronic health record data, existing mortality prediction studies in this population largely depend on static summaries derived from early admission. Such approaches i…
Mutasim Fuad Sarker, Adiba Rahman Namira, Wafa Binte Alam, Md Adnan Arefeen, Mahzabeen Emu, Sumaiya Tabassum Nimi · 📄 PDF
Bias Analysis of L2 Speaking Assessment Systems Using Concept Activation Vectors
Automatic speaking assessment systems are increasingly deployed in high-stakes settings to mark second language (L2) learners' speaking tests, making it critical to show that their scores depend on speaking proficiency rather than irrelevant speaker attributes such as first language (L1) or age. Tra…
Arya Labroo, Mengjie Qian, Kate Knill · 📄 PDF
HarnessOpt-Bench: Evaluating LLMs at Harness Optimization
As LLMs are increasingly deployed within agentic systems, their capabilities depend not only on the model weights but also on the harness: the prompts, tools, control flow, memory, and orchestration code surrounding them. This makes automated harness optimization -- the iterative and evaluation-guid…
Varun Ursekar, Apaar Shanker, Yash Maurya, Shehab Yasser, Vijay S. Kalmath, Veronica Chatrath, Yuan Xue · 📄 PDF
Beyond Top-K: Replacing Black-Box Retrieval with Interpretable Agentic Operations
Retrieval-augmented generation over long documents is dominated by one design: chunk the text, embed the chunks, and surface the top-k nearest neighbours of the query. We argue that for an important class of documents -- financial statements, audit reports, regulatory returns -- this design is struc…
Sagar Tamang, Ayush Vyas, Tabarakul Hazarika · 📄 PDF
Does FLAIR super-resolution erase or hallucinate small white-matter lesions?
White matter hyperintensities (WMH), bright regions on Fluid-attenuated Inversion Recovery (FLAIR) scans are associated with cerebrovascular pathology and neurodegeneration. FLAIR is usually acquired with thick slices in clinical settings, giving it poor through-plane resolution. Super-resolution (S…
Zahra Khodakarami, Yue Li, Pulkit Khandelwal, John Detre, Sandhitsu Das, Christopher Brown, David Wolk, Paul Yushkevich · 📄 PDF
Benchmarking the Benchmarks: Evaluating Benchmarks for Conversational Agents
Task-oriented conversational agents are evaluated using curated or automatically generated benchmarks, yet benchmark quality is rarely assessed. Poor benchmarks may contain inconsistent tasks, simplistic scenarios, or limited policy coverage, leading to unreliable evaluations. We introduce a referen…
Noam Koren, Roy Bar-Haim, Abigail Goldsteen · 📄 PDF
Tytan: Interactive Neurosymbolic Construction of Analytic Semantic Schemas from Relational Data
From natural-language query interfaces to automated report generation, data analysis tools need a description of the data: the real-world entities it contains, which columns function as measures or identifiers, and how tables connect into units of analysis. Today, this semantic layer is usually writ…
Donna Hooshmand, Shubham Shahi, Cameron Barrie, Abhratanu Dutta, Marko Sterbentz, Harper Pack, Kristian J. Hammond · 📄 PDF
TRAJDEBUG: Tracing Error Lifecycle to Identify Critical Failures in Long-Horizon Agent Trajectories
LLM-based agentic systems have shown remarkable capabilities in complex domains, while suffering from cascading errors and difficulty in debugging. Critical error detection aims to locate the earliest error step in a failed trajectory that is responsible for the final failure. However, progress face…
Yunjia Qi, Zehua Yin, Xintong Shi, Hao Peng, Songyuanyi Lu, Yixian Liu, Richeng Xuan, Yuhong Liu, Zhichao Hu, Xiaozhi Wa… · 📄 PDF
Challenges in Evaluating Explanation Methods for Static and Evolving Data
This paper addresses the limitations of Explainable Artificial Intelligence (XAI) with respect to insufficient evaluation. They are illustrated through the DetoxAI image recognition system for bias detection and concept unlearning. Then, an example of a human-grounded evaluation of methods for expla…
Jerzy Stefanowski · 📄 PDF
Resourced Authority A Mechanism-Design Model for Participatory Governance of Deployed AI Agents
We give a formal mechanism design model for the continuous participatory governance of a deployed AI agent. The mechanism is built on the principle that governance should control an AI agent through resource allocation so as to make authorization self enforcing via compute budgets. The mechanism see…
Praphul Chandra, Sujit Gujar, Ganesh Ghalme · 📄 PDF
The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping
Real-world video benchmarks provide broad coverage, but their fixed clips entangle event count, rate, duration, and visual complexity, making failure modes hard to isolate. While existing programmatic benchmarks offer better control, they score only the final answer rather than auditing reported eve…
Sarvesh Baskar, Zikui Cai, Shayan Shabihi, Anirudh Satheesh, Muhammad R. Islam, Udari Madhushani Sehwag, Tom Goldstein, … · 📄 PDF
AV-AIVAT: 74x Cheaper Agent Evaluation with Certified Anytime-Valid Stopping in Imperfect-Information Games
Deciding which of two agents is stronger means playing games until skill outweighs luck, and every game costs money, model inference, or expert time. Since the number of games needed is unknown, fixed-budget evaluations either keep paying after the result is settled or stop before the agents can be …
Boning Li, Yu Chen, Longbo Huang · 📄 PDF
An Optimal Agnostic PAC Algorithm
Let $H\subseteq\{-1,+1\}^X$ be a class of finite VC dimension $d\ge1$. Writing $L$ for the binary risk and $L^*=\min_{h\in H}L(h)$, we construct a learner achieving the statistically optimal risk bound: from an i.i.d.\ sample of size $n$, for every $0<δ\le 1/2$, with probability at least $1-δ$, \[ L…
Markus Engelund Mathiasen, Jian Qian, Nikita Zhivotovskiy · 📄 PDF
Investigating Artificial Intelligence Digital Sovereignty in Mobile Shopping Apps: A Case Study of Nigeria
The use of e-commerce mobile applications is expanding in Nigeria, creating both opportunities and risks, including fraud and reduced user control over digital technologies, raising concerns about digital sovereignty. This research examines how Artificial Intelligence (AI) in Nigerian mobile applica…
George Grispos, Sajda Qureshi · 📄 PDF
Tracing the Heart: An Evidence-Linked Pipeline for Heart-Failure Feature Engineering
Electronic health record (EHR) feature engineering is a major bottleneck in clinical research and AI, accounting for 39-45% of data scientists' workload. This is especially pronounced in heart failure, which affects an estimated 6.7 million U.S. adults and requires integrating fragmented EHR data wi…
Soorya Ram Shimgekar, Michelle Hu, Dorisa Shehi, Daniel Kang, Roy Ka-Wei Lee, Koustuv Saha, Christian Poellabauer, Chris… · 📄 PDF
Learning When to Trust via Selective Context Preference Optimization
Language models increasingly condition their answers on external signals, and a single misleading one can turn a correct answer wrong. The obvious remedy, training models to resist such signals, hides a failure mode: a model that ignores all context looks robust yet is useless when the context is wo…
Xian Sun, Wei Chow, Yingshuo Wang, Junhao Liu, Wei Gao, Qing Wu, Lingdong Kong · 📄 PDF
Legal aid eligibility and court outcomes: a design-based double-machine-learning approach
Equality before the law is a human right, and access to high-quality legal aid for indigent defendants is essential to enforce it. In a context where all defendants have access to a lawyer, I study the impact of denying legal aid on court outcomes. I combine double machine learning and a new adminis…
Fabio Italo Martinenghi · 📄 PDF
Overcoming Scattering in High-Cell-Density Tomographic Volumetric Bioprinting Using Computational Light Optimization
Tomographic volumetric additive manufacturing has emerged as a transformative 3D printing technology for rapidly fabricating complex geometries. It offers significant advantages for bioprinting due to its contactless and short process time (a few tens of seconds). However, the presence of high cell …
Qianyi Zhang, Felix Wechsler, Viola Sgarminato, Christophe Moser, Riccardo Rizzo · 📄 PDF
Comparative analysis of fiber Bragg grating filter losses inscribed by continuous wave UV and femtosecond-IR lasers for astrophotonics
Fiber Bragg grating (FBG) filters have been demonstrated as promising components in astrophotonic instrumentation for near-infrared ground-based observations. Given the photon-starved nature of astronomical applications, it is critical to minimize insertion losses across astrophotonic components. In…
Aashia Rahman, Ria G. Krämer, Abani Shankar Nayak, Julius Göhring, Piyamas Choochalerm, Anna Maria Weiß, Samuel L. Döpfn… · 📄 PDF
High Spectral Energy Density All-Fiber Nanosecond Pulsed 1.7 $μ$m Light Source for Photoacoustic Microscopy
We present a high spectral energy density all-fiber nanosecond pulsed 1.7 $μ$m light source specifically designed for photoacoustic microscopy (PAM). The system targets the first overtone absorption of C-H bonds near 1720 nm within the near-infrared-III (NIR-III) window, where lipids exhibit strong …
Seongjin Bak, Sang Min Park, Yuon Song, Jeesu Kim, Tae Won Nam, Dong-Wook Han, Chang-Seok Kim, Soon-Woo Cho, Brett E. Bo… · 📄 PDF
Frozen but Not Always Accessible: A Representation Analysis of Genomic Language Models
Genomic foundation models are increasingly reused as frozen feature extractors for downstream sequence prediction, offering a compute-efficient alternative to full fine-tuning. However, it remains unclear when biological information encoded by these models is accessible without task-specific adaptat…
Nirjhor Datta, Swakkhar Shatabda, M. Sohel Rahman · 📄 PDF
IL-10 rs1800896 polymorphism predicts biochemical remission in IBD patients undergoing biologic therapy
Background. Genetic factors, including single-nucleotide polymorphisms (SNPs), may modulate disease course and therapeutic efficacy in patients with inflammatory bowel disease (IBD). Aim. We investigated the association between four SNPs in cytokine genes and clinical phenotype as well as the respon…
Michela Helga Falzone, Davide Giuseppe Ribaldone, Martina Buglione, Irene Cottone, Marta Vernero, Demis Pitoni, Angelo A… · 📄 PDF
Scalable Circuit Cutting: A Framework for Combined Gate and Wire Cuts Using Gate Groups
Quantum circuit cutting enables the execution of large circuits on devices with a limited number of qubits by partitioning circuits into independent subcircuits. However, this introduces a sampling overhead, which grows exponentially with the number of cuts, rendering the choice of cut placements cr…
Fiona Jiali Fröhler, Yannick Stade, Christian Ufrecht, Daniel D. Scherer, Robert Wille · 📄 PDF
Filtered Vector Search in a Disaggregated Lakehouse: Composing Table-Format Pruning with Per-File ANN
Approximate nearest-neighbor (ANN) search increasingly runs alongside structured data - "find the 10 nearest documents where tenant='acme' AND lang='en'" - yet similarity and filtering are usually bolted together: a specialized vector index for one, a separate filter step for the other. We ask what …
Rakesh Jain, Thomas Griffin, Syed Zawad · 📄 PDF
EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding
On-device deployment of Large Language Models (LLMs) has become essential for personalized edge applications. A primary bottleneck is external memory access (EMA) in feed-forward network (FFN) layers. Speculative decoding and mixture-of-experts (MoE) are promising solutions. Speculative decoding red…
Sangwoo Ha, Hyunwoo Seo, Yurim Jo, Youngjin Moon, Hoi-Jun Yoo · 📄 PDF
PowerScope: ML-based Intra-Cycle Power Estimation
Power estimation at sub-clock-cycle temporal resolutions is critical for tasks such as power delivery network (PDN) design, dynamic voltage droop analysis, and pre-silicon power side-channel security evaluation. Designers commonly rely on commercial post-layout gate-level power analysis tools for th…
Jayanth Balasubramanian, Sujay Pandit, Radha Vaidya, Anand Raghunathan · 📄 PDF
Quantifying Different Gains from Trade in Quality
In this paper, I study to what extent countries differ in their preferences for quality and their technologies for improving quality. The paper also quantifies the contribution of those differences to the differences in gains from trade across countries. I adopt ANTONIADES, which allows endogenous q…
Yuting Chen · 📄 PDF
From Long to Short: How Interest Rates Shape Life Insurance Markets
This paper explores how financial institutions pass interest rate risk through to product markets using the life insurance industry as a setting. We show theoretically that it is optimal for insurers to distort product issuance across maturities to offset duration gaps. We examine insurers exogenous…
Ziang Li, Derek Wenning · 📄 PDF
The Role of Risk Sharing in Attenuating Business Cycles Within Currency Unions
The United States is a currency union where multiple risk-sharing mechanisms--- migration, fiscal transfers, income diversification and credit markets---buffer consumption from local income fluctuations. We show that risk sharing not only directly smooths consumption but also indirectly stabilizes i…
Alberto Pavia, Christian Proebsting · 📄 PDF
Quantum optoelectronics in semiconductor solar cell materials and devices
We analyze the integration of quantum optical phenomena, such as cavity quantum electrodynamics (CQED), Fabry Perot resonances, and strong light-matter coupling, into the design and engineering of next generation photovoltaic systems. We examine how these phenomena can be harnessed through photonic …
Xi Liu, Wenxi Fang, Ken Perlin · 📄 PDF
All-Optical Field-Resolved Spectroscopy With Interferometric Nonlinear Cross-Correlations
Direct time-domain measurements of electric fields enable sub-cycle spectroscopy of light-matter interactions, but established techniques such as electro-optic sampling are constrained in their bandwidth by gate-pulse duration and phase-matching limitations. Alternative approaches have emerged in re…
Felix Ritzkowsky, Gian Luca Dolso, Benjamin Mazur, Matthew Yeung, Phillip D. Keathley · 📄 PDF
Photogalvanic second harmonic generation in Si3N4 for 1 Hz level on-chip metrology and spectroscopy
The coherent photogalvanic (PG) effect induces an effective $χ^{(2)}$ nonlinearity in natively $χ^{(3)}$ silicon nitride integrated photonics, unlocking pathways toward chip-scale precision spectroscopy and optical clockworks via second harmonic generation (SHG). While quasi-phase-matched PG-SHG usi…
Andrei Diakonov, Roy Zektzer, Xiyuan Lu, Kartik Srinivasan, Liron Stern · 📄 PDF
Design of a thermal loading resilient optical enhancement cavity for operation at 515 nm for X-ray production through inverse Compton scattering in an energy recovery linac
A fast and simple method to optimize a high-power optical enhancement cavity is proposed. It is applied to a four-mirror bow-tie cavity operating at 515 nm, which is to be implemented within the PERLE ERL for the production of X-rays through inverse Compton scattering. The optimized figure of merit …
Alice Renaux, Aurelien Martens, Yann Peinaud, Ronic Chiche, Kevin Dupraz, Marie Jacquet, Daniele Nutarelli, Fabian Zomer · 📄 PDF
Universal Function Approximation via Diffractive Optical Processors: Physical Limits, Error Bounds, and Learnability
We present a unified theoretical framework connecting classical universal approximation theory, Fourier-feature approximation, and diffractive optical processors. We show that phase-encoded diffractive processors implement finite Fourier-feature expansions whose mathematical completeness follows fro…
Md Sadman Sakib Rahman, Che-Yung Shen, Aydogan Ozcan · 📄 PDF
Time-resolved THz Stark spectroscopy of molecules in water
Stark spectroscopy is a powerful method for probing molecular dipole moment changes, charge transfer dynamics, and polarizability under applied electric fields. Time-Resolved Terahertz Stark Spectroscopy (TRTSS), which employs intense single-cycle terahertz (THz) pulses to induce transient Stark shi…
Elnaz Zyaee, Vladislav Slama, Seyyed Jabbar Mousavi, David Rohrbach, Ursula Rothlisberger, Thomas Feurer · 📄 PDF
Freeform super-oscillatory optics for CMOS-integrated THz super-resolution imaging
The diffraction limit fundamentally constrains the spatial resolution of far-field imaging systems. While near-field techniques can circumvent this limit, their inherently short working distances (WD) severely restrict practical applications. Super-oscillatory lenses (SOLs) offer a far-field alterna…
Jin Chen, Liang Gao, Hao Guo, Zhi Chao Chen, Kang Jie Lin, Kam Man Shum, Ka Fai Chan, Chi Hou Chan · 📄 PDF
Photonic-chip-based generation of sub-100-femtosecond optical frequency combs
Sub-100-fs optical pulses and frequency comb sources have been revolutionizing a wide range of applications, from ultrafast optical science to optical frequency standard and measurement. To date, the leading techniques for generating such pulses in practical systems rely on tabletop mode-locked lase…
Weiqiang Xie, Zhengshun Lei, Zeyu Xiao, Yudi Zhao, Xing Zou, Wenqi Wei, Zihao Wang, Ting Wang, Jianjun Zhang, Bofang Zhe… · 📄 PDF
Mask-free fast patterning of organic light-emitting diode pixels using laser-assisted close-space sublimation
Existing patterning processes for organic light-emitting diode displays offer micrometer-scale precision but are constrained by long processing times for large-area substrates. In this work, we study a fast growth method for patterned organic film deposition, aimed at applications including active-m…
Subhamoy Sahoo, Jain Jose, Mani R, Arghya Saha, Kanimozhi V, Dhruvajyoti Barah, R. Bairava Ganesh, Amitava Majumdar, Jay… · 📄 PDF
Parameter identification for predator-prey system with sparse data
Parameter identification from observations of dynamical systems is a fundamental problem in population biology. Mechanistic models of ecological systems rely on optimization methods that require accurate initial guesses to guarantee convergence. In ecological applications, datasets contain observati…
Eduard Campillo-Funollet, James Van Yperen · 📄 PDF
Toward Blockage-Resilient 6G-V2X Connectivity: Semi-Distributed Bandit with Dynamic Arm Set for mmWave HetNets
The vision for 6G vehicle-to-everything (V2X) communications demands reliable, adaptive connectivity for fully autonomous driving across complex dynamic environments. Millimeter-wave (mmWave) user association (UA) in heterogeneous vehicular networks presents a particularly demanding instance of this…
Weiqi Chi, Bo Qian, Hanlin Wu, Donghui Li, Haibo Zhou, Manabu Tsukada · 📄 PDF
Deltoris: Enabling Real-time VLA Inference in Embodied AI via Bit-level Sparsity and Speculative Inference
Vision-language-action (VLA) models have emerged as a key component in embodied AI. Among existing approaches, diffusion-based VLA models achieve superior motion quality and generalization. However, diffusion-based VLA models are compute-intensive and must run at high control frequency, e.g., 50-200…
Zheng Liu, Zeyu Guo, Zihan Liu, Anbang Wu, Han Zhao, Fangxin Liu, Zhezhi He, Yinhe Han, Jingwen Leng, Minyi Guo, Yiming … · 📄 PDF
MCHA: A Memory-Centric Hierarchical Architecture for Parallel-Sequential Computing
Emerging workloads, such as Multi-Agent Reinforcement Learning (MARL), large-scale neuromorphic computing, and probabilistic graphical models, intrinsically exhibit parallel-sequential computing patterns. While these tasks demand massive parallelism to achieve high throughput, they are severely bott…
Daijing Shi, Hongxiao Zhao, Yihan Fu, Zhan Chen, Jiayi Li, Yihang Zhu, Anjunyi Fan, Yaoyu Tao, Yuchao Yang, Bonan Yan · 📄 PDF
Architectural Implications of Agentic AI Workflows
Agentic AI is emerging in datacenters, but its architectural implications remain unexplored. We organize agentic workflows in a taxonomy and present its first architectural characterization with a production study at Microsoft Azure and a controlled study of open-source frameworks. We show that agen…
Jirong Yang, Peizhe Liu, Chaojie Zhang, Jovan Stojkovic · 📄 PDF
A Systolic Array Architecture for Nonlinear Activation Functions and Softmax Computation using Chebyshev Polynomials
Neural Network Accelerators have gained popularity in recent years due to their greater efficiency than CPU-based platforms. Often, these accelerators utilize different hardware units for univariate activation functions, such as tanh, and the multivariate softmax, thereby missing opportunities for r…
Benedikt Schaible, Anirudh Suresh Bharadwaj, Ulf Schlichtmann, Jiang Hu · 📄 PDF
LLM-Assisted Detection and Repair of Hardware Security Vulnerabilities in Verilog Designs
Hardware designs, like software, are susceptible to bugs that can introduce security vulnerabilities and create opportunities for malicious exploitation. Unlike software vulnerabilities, however, hardware flaws become permanently embedded in silicon after fabrication, making them difficult or imposs…
Ethen Santana, Gabriel Gyaase, Hao Zheng · 📄 PDF
Kerckhoffs-Compliant Watermarking for Physical Design IP Protection: From Placement to Routing
Physical design (PD) intellectual property (IP) is a valuable artifact of modern VLSI implementation. It includes optimized cell placement, clock distribution, and routing decisions produced by carefully tuned PD flows. As access to PD tools expands, unauthorized reuse of placed-and-routed databases…
Andrew B. Kahng, Yiting Liu · 📄 PDF
Differential 6-DOF Pose Estimation with Provable First-Order Immunity to Camera Calibration Errors
Accurate six-degree-of-freedom (6-DOF) motion estimation is essential for robotic manipulation, autonomous systems, and structural displacement monitoring. Conventional 3D-2D methods estimate absolute camera poses independently at each time and recover platform motion through camera-to-platform extr…
Yueqiang Zhang, Liang Deng, Yi Zhang, Baoqiong Wang, Wenjun Chen, Shuixin Pan, Yulan Guo, Qifeng Yu · 📄 PDF
Suppression Sticks, Locality Is Fragile: A Closed-Loop Target-and-Control Audit of Task-Vector Negation in VLA Policies
Task-vector arithmetic offers a closed-form way to modify a model, yet its behavioral locality remains unclear in closed-loop robot control. We present a target-and-control audit of per-skill task-vector subtraction from multitask vision-language-action (VLA) policies. Across all ten LIBERO-Goal ski…
Shaoguang Wang, Weiyu Guo, Rushi Dai, Yiren Zhao, Yandong Guo, Hui Xiong · 📄 PDF
A Multi-Sensor Dataset for Monitoring the Operational Environment of Rail Vehicles
Reliable environment monitoring is essential for the safe and efficient operation of automated railway systems, covering all Grades of Automation (GoA), from partially automated (GoA2) to fully automated operation (GoA4). Artificial Intelligence (AI) plays a central role in enabling these systems to…
Claudio Diotallevi, Rodrigo Gudiño, Zaharia Pachalieva, Philipp Neumaier, Patrick Naumann, Erik Bochinski, Volker Eisele… · 📄 PDF
Enabling Urgency-aware Robot Swarm Intralogistics using Smart IoT Tags
Warehouse items differ in how urgently they must be moved: perishable goods, pharmaceutical shipments, and just-in-time production materials must be delivered sooner than the rest of the stock. Decentralised robot swarms suit warehouses that cannot justify fixed automation infrastructure, but curren…
Youssef Alboraei, Murray Groves, Shane Wen, Wenda Zhao, Senhui Qiu, Mohammud J. Bocus, Robert Piechocki, Sabine Hauert, … · 📄 PDF
A Vision-based Control Framework for Real-time Autonomous UUV Operations
This paper presents a fully integrated vision-based framework for real-time and robust localization, autonomous navigation, and mapping for unmanned underwater vehicles (UUVs) in dynamic, visually challenging environments. The proposed pipeline enables both net-relative and global localization while…
Erik Tjærand Frøland, Marco Job, Md Ether Deowan, Eleni Kelasidi · 📄 PDF
A GitOps-Driven Annotation Catalog for Fully Automatic Railway Operations
Automatic train operation (ATO) at grade of automation 3 and above (GoA3-GoA4) requires robust AI-based perception systems capable of reliably detecting obstacles and railway-specific objects under real-world conditions. The effectiveness of these modern artificial intelligence approaches depends he…
Martin Köppel, Tobias Cronauer, Zekiye Ilknur-Öz, Sebastian Dubiel, Patrick Naumann, Philipp Neumaier · 📄 PDF
Toward Integrating Adaptive Experience Replay and Online Uncertainty Estimation in Safe Actor-Critic Optimal Control
Safe actor-critic control often treats barrier filtering, uncertainty estimation, and experience replay as separate modules, even though each changes the data used for learning and control. We develop an integrated architecture in which the uncertainty estimate updates the obstacle geometry used by …
Mahshad Rastegarmoghaddam, Davoud Nikkhouy, Shima Samadzadeh · 📄 PDF
Explicit Language Memory for Long-Horizon Planning in Vision-Language-Action Models
Vision-language-action (VLA) models provide a unified paradigm for connecting visual perception, language understanding, and robotic control. However, existing VLA models still face major challenges in long-horizon tasks: sparse expert demonstrations constrain cross-task compositional generalization…
Houze Xu, Jizhong Li, Ziyi Ye · 📄 PDF
From Transparent Labware Segmentation to Collision Avoidance: A Real-Time Edge-Aware Perception Pipeline
This paper presents an edge-aware instance segmentation framework that enables real-time robotic collision avoidance with transparent laboratory glassware using purely visual perception. Transparent vessels defy conventional segmentation due to refraction, specular reflection, and the absence of sta…
Shijun Ding, Chen Qian, Weiwei Shang, Junlin Xiong · 📄 PDF
Deliberate Before You Fly: Vision-Guided Spatial Deliberation for UAV See-and-Reach Navigation
UAV see-and-reach navigation requires an aerial agent to approach a language-specified target visible in its initial view and stop reliably near it. Existing methods typically map vision-language representations directly to action outputs without explicitly modeling intermediate fine-grained spatial…
Fanfu Xue, En Yu, Bohang Liu, Hongjun Wang, Yang Yang, Xindi Wang, Jiande Sun · 📄 PDF
RORA: Realistic Object Reconstruction with Articulation
Replicating real-world environments into simulation by realistic visual representation like NeRF and 3D Gaussian Splatting (3DGS) has emerged as an effective strategy to reduce the sim-to-real gap in robot learning. However, implementing object articulation during the real-to-sim process is still a …
Hyesung Lee, Youngseon Lee, Kyutae Lee, Dongjun Lee, Yongseok Lee · 📄 PDF
PRIMAL3: Pathfinding via Reinforcement and Imitation Multi-Agent Learning - Leveraging LaCAM3
We present PRIMAL3, an ultra-large-scale learning-based framework for multi-agent pathfinding (MAPF) that integrates reinforcement learning, topology-aware communication, LaCAM3-guided training, and PIBT-based action refinement. PRIMAL3 targets failures at topologically critical states, where agents…
Chengyang He, Tanishq Duhan, Gadiel Sznaier Camps, Fangyuan Wang, Yuhong Cao, Jiankai Sun, Ge Sun, Mac Schwager, Guillau… · 📄 PDF
Mimir: A Neuro-Symbolic Memory System with Dynamic Grounding for Embodied Agents in Interactive Environments
Long-horizon embodied task requires agents to act under partial observability while preserving both scene belief and execution progress. Flat histories or implicit policy states may contain past observations, but they do not provide an explicit interface for deciding which world facts support the cu…
Haoming Xu, Zhenlin He, Hengyi Wang, Jiafeng Xu, Hao Dong · 📄 PDF
DreamWAM: Beyond RGB Future Prediction for World Action Models
World Action Models (WAMs) learn action-relevant representations by predicting how the observed world will evolve. Most existing WAMs define this future in RGB space, where task-relevant state transitions are entangled with nuisance variations in texture, illumination, background, and viewpoint. We …
Shanglin Yuan, Weiheng Zhao, Xin Shi, Haoyi Jiang, Xianda Guo, Liu Liu, Wenyu Liu, Wei Sui, Xinggang Wang · 📄 PDF
Optimal Constrained sc-LTL Planning in MDPs via Switching Policies
We study the synthesis of optimal policies for planning problems on Markov decision processes with both objectives and safety constraints specified in co-safe linear temporal logic (sc-LTL). Our problems are inherently non-Markovian due to the complexity of the sc-LTL specification and may require p…
Zetong Xuan, Yu Wang · 📄 PDF
BridgeVLA++: A Data-Efficient, Generalizable, and Memory-Augmented Vision-Language-Action Framework for 3D Manipulation
Leveraging pre-trained vision-language models (VLMs) to construct vision-language-action (VLA) models has emerged as a promising paradigm for 3D robot manipulation. However, existing 3D VLA methods remain data-hungry, exhibit limited generalization under distribution shifts, and lack explicit memory…
Peiyan Li, Yuze Zhu, Yixiang Chen, Qisen Ma, Yuan Xu, Jiabing Yang, He Guan, Yan Huang, Hongtao Wu, Xiao Ma, Tao Kong, L… · 📄 PDF
Exact Model-Free Policy Iteration for Co-safe LTL Planning
This work studies model-free reinforcement learning for co-safe linear temporal logic (sc-LTL) objectives in finite Markov decision processes, which can be reduced to maximal reachability objectives via the standard product construction. For this problem, direct sample-based bootstrap methods (e.g.,…
Zetong Xuan, Yu Wang · 📄 PDF
SpikingNav: Robust Embodied Navigation with Spiking Neural Policies
Embodied navigation requires an agent to make sequential decisions from egocentric observations in a physical environment. Existing Artificial Neural Network (ANN)-based navigation models have achieved strong performance, yet they often rely on dense computation and may degrade under visual corrupti…
Jiahong Zhang, Sijun Shen, Dehua Wu, Yifan Lin, Xuechen Xia, Xu Chu, Youhui Zhang, GuoqiLi · 📄 PDF
AI-based single-shot structured-light depth reconstruction for real-time laparoscopic surgical guidance
Significance. Accurate intraoperative depth perception is important for autonomous and semi-autonomous robotic laparoscopic surgery. Conventional fringe projection profilometry can achieve millimeter-scale accuracy but often requires multi-shot acquisition, digital-micromirror-device projection, and…
Wayne Wonseok Rodgers, Xiangyi Le, Seonghoon Jang, Shuwen Wei, Justin Opfermann, Michael Kam, Axel Krieger, Jin U. Kang · 📄 PDF
UG-UMRE: Uncertainty-Guided Modality Augmentation and Distributional Calibration for Unified Multimodal Relation Extraction
Unified Multimodal Relation Extraction (UMRE) aims to identify intra-modal and cross-modal relations between textual entities and visual objects. However, existing UMRE studies still encounter two critical issues: ignoring inherent aleatoric uncertainty causes noise propagation, and deep-seated hete…
Bo Kong, Liruiz Jia, Yi Liang, Chao Liu, Dongfang Han, Tianwei Yan, Yuan Liu, Shengquan Liu · 📄 PDF
Towards Valid B-Rep Generation: Training-Free Wireframe Anomaly Detection and Repair
Multi-stage boundary representation (B-Rep) generation leverages intermediate wireframes to synthesize CAD models. However, geometric and topological risks in these wireframes -- such as self-intersections, edge collapses, and disconnected vertices -- can propagate to invalid final B-Reps. Mitigatin…
Jingyu Wu, Youcheng Cai, Tengyu Luo, Ligang Liu · 📄 PDF
ContextMaster: Interactive Multi-Shot Video Creation via Fixed-Budget Sparse Context Routing
Recent video models increasingly support generation, reference conditioning, and editing within a single model, yet typically expose them as separate operations over fixed inputs. Practical creation unfolds across multiple shots, requiring one model to generate from text, follow a reference, or edit…
Xu Guo, Zhengxuan Wei, Xinghui Li, Hanzhuo Huang, Xinyu Liu, Xiangyang Luo, Min Wei, Yiran Zhu, Qiulin Wang, Yulong Xu, … · 📄 PDF
Promptable Animal Pose Tracking Across Species
Animal pose estimation and tracking is important for wildlife monitoring and conservation research, and with limited expert time for labelling automated approaches are imperative. While human pose estimation and tracking has seen rapid progress thanks to large annotated datasets, animal pose remain …
Le Li, Daniela Ivanova, Nicolas Pugeault · 📄 PDF
Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes
Vision offers a critical axis for advancing foundation models, driving a shift towards natively unified multimodal pretraining. Despite this momentum, the design space and the fundamental mechanisms of how modalities interact during unified training remain underexplored. We provide empirical clarity…
Junlin Han, Shengbang Tong, David Fan, Minghao Chen, Philip Torr, Filippos Kokkinos, Mike Lewis · 📄 PDF
OmniEdit-Bench: A Comprehensive Benchmark for Instruction-based Video Editing
Instruction-based video editing (IVE) is an emerging field with broad applications, yet evaluating editing models remains challenging. Existing benchmarks suffer from two major limitations: limited task coverage inherited from image editing, which overlooks video-specific dimensions, and inadequate …
Chenxuan Miao, Yutong Feng, Yi Lu, Yunfeng Yan, Donglian Qi, Shiwei Zhang, Yu Liu, Xi Chen, Hengshuang Zhao · 📄 PDF
Beyond Reprojection Error: Camera Calibration with 3D Targets
In 3D reconstruction, camera calibration is an essential element for achieving high fidelity and accuracy of the reconstructed geometry. While existing approaches rely upon 2D planar calibration, this work proposes a framework tailored for 3D reconstruction that is based on predicting scene rays, wh…
Dennis Ruppel, Hasan Kutlu, Kai A. Neumann, Martin Knuth, Pedro Santos, Andreas Weinmann, Arjan Kuijper · 📄 PDF
HelloWorld: Enabling Socially Interactive Characters in Video World Models
Despite the remarkable recent progress of video world models, social interaction between users and the characters within these worlds remains unsupported. To fill this gap, we present HelloWorld, a video world model that enables social interaction with in-world characters. With a single button press…
Liangyang Ouyang, Ruicong Liu, Xuangeng Chu, Kaipeng Zhang, Yoichi Sato · 📄 PDF
Bag-of-Visual-Words for Spatial Mapping of Lung Adenocarcinoma Growth Patterns
Spatial mapping of lung adenocarcinoma (LUAD) growth patterns across whole slide images (WSIs) requires resolving architectural context at the region level, yet existing methods operate at the individual tile level and produce generic morphological clusters rather than clinically defined pattern map…
Darya Ardan, Valentin Oreiller, Henning Müller · 📄 PDF
Lesion Detection in CT with Frozen Self-Distilled Features: SALT, a Spatially Adaptive Label-Guided Temperature
Self-supervised pretraining objectives are spatially uniform: the teacher temperature and the per-patch loss weight are identical everywhere in the image, so a lesion a few patches wide contributes no more to the training signal than the surrounding parenchyma. Prior work biases the views toward ann…
Mahmut S. Gokmen, Evan W. Damron, Mitchell A. Klusty, Caroline N. Leach, Emily B. Collier, V. K. Cody Bumgardner · 📄 PDF
HexMIL: Hierarchical Attention MIL for Ante-Hoc Explainable Detection of AI-Manipulated CT Volumes
The emergence of medical deepfakes, i.e., medical images manipulated by deep generative models, poses a significant threat to clinical workflows. However, existing detectors suffer from two critical limitations: poor generalization to unseen generative architectures for manipulation detection and la…
Orazio Pontorno, Luca Guarnera, Zahid Akhtar, Sebastiano Battiato · 📄 PDF
IRIS: A Visual Cortex-Inspired Framework for Analyzing Orientation Selectivity in Vision Transformers
Vision transformers (ViTs) have become the de facto standard for image encoding across many perception tasks. Despite their empirical success, it remains mechanistically unclear how they encode low-level features, given their lack of inductive biases: ViTs process information globally rather than re…
Vaishnavi B Mohan, Vijayakrishna Naganoor, Yashas Annadani, Shashank Hegde · 📄 PDF
SmartMage: Dynamic Modality Orchestration for 3D Scene Understanding
Understanding 3D scenes is fundamental to embodied intelligence, requiring joint reasoning over heterogeneous information from multiple modalities, including visual and geometric cues. However, the relevance of these modalities often varies across queries. Existing Multimodal Large Language Models (…
Yue Zhang, Yingzhao Jian, Yunqiu Xu, Xiaoxiao Sun, Hehe Fan · 📄 PDF
Objects as Audio-Visual Modal Sound Fields
While modern 3D reconstruction excels at modeling object geometry and appearance, it largely ignores the rich acoustic cues revealed through physical interaction. Object impact sounds convey material, stiffness, and structural properties that complement vision, yet existing impact sound modeling app…
Zisen Shao, Zihao Wei, Derong Jin, Ruohan Gao · 📄 PDF
CoCo-IR: Contextual Composed Image Retrieval
Current instruction-based image retrieval systems are powerful but limited to single-turn interactions, failing to capture the iterative nature of complex, real-world visual searches. To overcome this limitation, we introduce Contextual Composed Image Retrieval (CoCo-IR), a novel task that enables u…
Shengcao Cao, Tanmaya Shekhar Dabral, Zhongli Ding, Madhuri Shanbhogue, Kaifeng Chen, Zhe Li, Mojtaba Seyedhosseini, Yu-… · 📄 PDF
Optimizing What Policies Learn From: Recoverability-aware Rollout Intervention Learning
Critic-free group-based reinforcement learning has become a scalable approach for post-training large language models. However, most existing methods allocate the same number of rollouts to every task and trajectory state, even though some rollouts provide much more useful learning signals than othe…
Zheyuan Zhang, Manqing Mao, Hong Wang, Zhuoer Wang, Samson Koelle, Jie Yuan, Yanjun Lin, James Feng, Nikki Lijing Kuang,… · 📄 PDF
Learning When to Stop: Prefix-Optimal Dynamic Diffusion Policies for Continuous Control
Diffusion policies are a powerful policy class for continuous control, but their iterative denoising process creates a substantial computational bottleneck. Reducing this cost requires adapting the number of denoising steps to the difficulty of each action while preserving task performance. We intro…
Rohit Kumar Salla, Manoj Saravanan, Simon Stepputtis · 📄 PDF
MALT: Lightweight Curvature-Aware Muon via Diagonal Preconditioning
Muon has recently emerged as a promising alternative to AdamW for language model pretraining by orthogonalizing momentum matrices using Newton-Schulz iterations. Although Muon mitigates gradient anisotropy, it does not explicitly account for the curvature geometry of the loss landscape and may there…
Tongle Wu, Huanyu Dong, Ying Sun, Ziye Ma · 📄 PDF
Multimodal Spatiotemporal Atmospheric Data Assimilation with Latent Flow-matching
Data assimilation (DA) uses Bayesian inference to update the state of a numerical forecast model with observed data. In this study, we propose a fundamentally different, unified approach to atmospheric data assimilation. We use latent video flow-matching to sample temporally consistent trajectories …
Dibyajyoti Chakraborty, Romit Maulik · 📄 PDF
BnBERT-iPET: Sparse Few-Shot Language Modeling for Bengali via Lottery Ticket Pruning
Deep neural networks have shown impressive success in NLP tasks owing to their complex structure and huge number of edges. Achieving state-of-the-art performance in natural language processing with a large pre-trained model such as BERT is expensive and time-consuming, carries a large carbon footpri…
Sajib Hossain, Md Kamrus Samad, Anan Ghosh, Labib Imam Chowdhury, Nabeel Mohammed · 📄 PDF
Reward Structure Shapes the Interaction Between Episodic Exploration and Neural Memory in Reinforcement Learning
In partially observable reinforcement learning, agents face a dual bottleneck: they must explore to encounter rewarding states and retain that experience in memory to optimize their policies. Exploration bonuses and memory architectures are traditionally evaluated in isolation, leaving their interac…
Jai Malegaonkar, Rohan Patil, Henrik I. Christensen · 📄 PDF
Stable Density Ridges: Consistency and Convergence of Subspace Constrained Mean Shift
The Subspace Constrained Mean Shift (SCMS) algorithm is a popular nonparametric method for extracting density ridges, which serve as a low-dimensional representation of high-dimensional data. It is a widely held belief in the literature that SCMS trajectories converge to the classical density ridge,…
Wanli Qiao · 📄 PDF
DASyR-LLM: Domain-Aware Symbolic Regression with LLMs for Kinetic Model Discovery
Kinetic model discovery is a central challenge in chemical engineering, as accurate rate expressions are essential for understanding and controlling chemical and biological processes. Symbolic regression (SR) has emerged as a powerful data-driven approach for identifying interpretable kinetic models…
Roberto Aliaga Medina, Paulina Quintanilla, Antonio del Rio Chanona · 📄 PDF
Predicting Brain Morphometry with MT-GNN: Mesh Evolution in Continuous Time with Graph-Based Metric Tensor Embeddings
Predicting how a subcortical structure's shape will evolve from a few prior scans could support prognosis and clinical-trial enrichment. Existing longitudinal mesh predictors either extrapolate shape trajectories via high-dimensional embeddings or regress vertex deformations directly. We instead pre…
Hao Ding, Daniel Semchin, Paul M. Thompson, Boris Gutman · 📄 PDF
The Loss Does Not See the Basis, but Adam Does
Gradient descent on a factored model $W = UV^\top$ is implicitly biased toward low-rank solutions, while Adam, starting from the same small initialization, is not. We trace the difference to the gauge symmetry of the loss, its invariance under $(U, V) \mapsto (UQ, VQ)$. Gradient flow's low-rank mech…
Devender Singh · 📄 PDF
Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning
Long-horizon reasoning in recent LLMs demands that the model switch between distinct skills inside a reasoning chain, such as first doing a math derivation, then using the result to plan a schedule. We call such problems cross-skill long-horizon tasks: multi-step tasks whose steps require different …
Yinghui He, Ling Yang, Jiarui Liu, Yongjin Yang, Lechen Zhang, Yingcheng Wu, Zhenfei Yin, Mengdi Wang, Sanjeev Arora · 📄 PDF
The Effect of Perceived Race and Gender on Police Language Use: Experimental Evidence from VR Simulations
Against the backdrop of violence in police interactions with the U.S. public, we explore how deferentially police officers speak to virtual characters depicted as Black adult males in vir- tual reality (VR) simulations. We evaluate the effect of seeing and communicating with these characters through…
Sandra C. Sandoval, Navita Goyal, Rashawn Ray, Long Doan, Rachel Rudinger, Hal Daumé · 📄 PDF
MarsCast: Transfer Learning of AI Weather Foundation Models to Planetary Atmospheres
We investigate the transferability of Earth weather foundation models to planetary atmospheres by adapting the GraphCast graph neural weather forecasting model to Mars. While GraphCast achieves state-of-the-art performance for terrestrial forecasting, its applicability to non-Earth environments rema…
M. L. Carroll, J. Li, S. D. Guzewich, G. Villanueva, J. A. Caraballo-Vega, M. J. Frost · 📄 PDF
RepairFormer: Automated Repair of Structured Inputs Using Transformers
Structured input files such as JSON, DOT, OBJ, INI, S-expression, and TinyC are widely used in software systems, but small corruptions can cause parsers to reject otherwise useful data. Repairing such inputs is important because malformed configuration, program, and data files can interrupt testing,…
Ovi Paul, Tom J King, Ali Shokri · 📄 PDF
Hardware Design and Security in the Era of Chiplets and LLMs
The semiconductor industry is undergoing a dual revolution: the shift toward heterogeneous 2.5D chiplet systems and the integration of Large Language Models (LLMs) into Electronic Design Automation (EDA) flows. While these paradigms offer unprecedented benefits in yield, modularity, design productiv…
Johann Knechtel, Ozgur Sinanoglu, Paul V. Gratz, Ramesh Karri · 📄 PDF
Provable Limits and Certified Deferral for Verbalized Uncertainty in Small Language Models
Small open-weight language models increasingly run in private, offline, and cost-sensitive settings, where the key deployment question is not only what a model answers but when it should defer to a human. We study whether verbalized confidence can support risk-controlled deferral, evaluating eleven …
Jianru Shen · 📄 PDF
VQ-VAD: Vector-quantized Motion Representation Learning for Human-centric Video Anomaly Detection
Video Anomaly Detection (VAD) is inherently challenging due to the scarcity of anomalies and the large visual variability in surveillance footage, including changes in lighting, viewpoint, and human appearance. To mitigate visual noise and address privacy concerns, recent work has shifted to pose-ba…
Narges Rashvand, Ghazal Alinezhad Noghre, Shanle Yao, Gabriel Maldonado, Hamed Tabkhi · 📄 PDF
MultiPathFormer: Towards a Foundation Model for Multipath Wireless Propagation
Recent advances in machine learning have enabled training of wireless foundation models, which aim to support tasks such as channel estimation, beam prediction, and localization based on wireless signals. Existing wireless foundation models typically pretrain on channel tensors using masked reconstr…
Blessed Guda, Kayley Sze, Carlee Joe-Wong · 📄 PDF
Capability-Gated Planning: Cost-to-Goal Discovery and the Limits of Myopic Experiment Selection
Systems that automate scientific discovery must repeatedly decide which experiment to run, which hypothesis to test, which tool to build, and when to stop. Many systems make these decisions by maximizing a myopic score such as expected information gain per unit cost or a learned plausibility score. …
Ahmed Hassoon, Mark Dredze · 📄 PDF
Item Response Theory for AI Safety
Language models differ in how safely they behave and these differences are measured by safety benchmarks. But aggregated benchmark scores are hard to trust and interpret, because benchmarks duplicate one another, correlate heavily, and models may sandbag when they detect evaluation. To address these…
Joshua Fonseca Rivera, Neil Shah, David Demitri Africa, Konstantinos Voudouris · 📄 PDF
Hierarchical Graph Memory for LLM Agents with Path-level Localization and Rewrite
Agents for long term reasoning require a memory that can be efficiently and effectively updated over time, as new facts and external feedback continue to arrive. Recently, graph memory has been adopted to offer structural organization for multi-hop retrieval and reasoning. However, existing methods …
Xiawei Yue, Boran Wang, Xiaoqing Zhang, Shuxin Zheng, Ziwei Zhang · 📄 PDF
ABSeeker: Training Long-Horizon Search Agents via Answer-Backtracked Credit Assignment
Long-horizon search agents must make multiple sequential actions (steps) to search, retrieve, verify, and integrate evidence to reach a final answer. However, existing methods for training these agents typically treat all steps within a trajectory uniformly during both supervised fine-tuning (SFT) a…
Yijun Lu, Rui Ye, Jiajun Wang, Yuwen Du, Tian Jin, Songhua Liu, Siheng Chen · 📄 PDF
CoPlan: A Trustworthy Co-Intelligence Interface for Care Planning through Role-Based Contestable Argument Graphs
AI-supported care planning can help clinicians, patients, caregivers, and care teams coordinate complex decisions across clinical, functional, psychosocial, and environmental needs. However, many AI systems present recommendations as fixed outputs, limiting stakeholders' ability to inspect, challeng…
Hung Truong Thanh Nguyen, Hélène Fournier, Piper Jackson, Makoto Itoh, Shannon Freeman, Rene Richard, Hung Cao · 📄 PDF
Representational separation between unitary and channel quantum generative models via shared classical randomness at shallow depth
Near-term quantum hardware limits circuit depth and often imposes geometrically local connectivity for quantum generative models, restricting the output distributions accessible to shallow unitary Born models. Introducing stochasticity into a unitary quantum Born model can improve the empirical gene…
Arunava Majumder, Marius Krumm, Hendrik Poulsen Nautrup, Hans J. Briegel · 📄 PDF
Robust and Efficient Motion Reasoning for Privacy-Aware Classroom Incident Recognition
Can computer vision help make classrooms safer? In this pilot study, we investigate privacy-aware and computationally efficient classroom incident recognition from CCTV-style observations. This setting remains underexplored, with limited benchmarks and few methods designed for the privacy, efficienc…
Paritosh Parmar, Landy Lan, Hong Yang, Chen Yi, Chiat Pin Tay · 📄 PDF
Chained Recursive Language Models for Multi-Iteration Reasoning
Long context reasoning in large language models (LLMs) is usually constrained by the fact that a single inference trajectory has to simultaneously explore the context, store intermediate state, verify evidence, and produce the final answer. This becomes particularly difficult in tasks that require e…
Purbesh Mitra, Sennur Ulukus · 📄 PDF
SSTQ:Privacy-Preserving Vector Quantization via Subsampled Stochastic TurboQuant
Achieving local differential privacy in distributed optimization while maintaining low communication cost remains challenging. Existing vector quantization methods, such as vqSGD, use high-dimensional geometric constructions but incur unfavorable dimension-dependent variance. In this work, we propos…
Adel Javanmard, David P. Woodruff, Vahab Mirrokni · 📄 PDF
OPD-V: Visual On-Policy Self-Distillation with Modality Balance
On-Policy Self-Distillation (OPSD) has become a standard post-training approach for improving visual reasoning in multimodal large language models (MLLMs). Existing methods draw privileged information from diverse input sources to guide self-distillation. Yet these designs overlook Modality Imbalanc…
Aniri, Jinhe Bi, Peng Liao, Zengjie Jin, Volker Tresp, Fei Shen, Yunpu Ma, Tat-Seng Chua · 📄 PDF
Teaching Nemotron Greek: Mining a Corpus, Adapting Retrieval, and Grounding Generation for Modern Greek across Specialist Domains
Modern Greek is absent from NVIDIA's Nemotron retrieval models and from major multilingual retrieval benchmarks, despite being important for retrieval-augmented generation (RAG) in legal, energy, financial, and medical applications. We present an end-to-end adaptation of the Nemotron retrieval stack…
Ayoub Kirouane, Christos Petrocheilos · 📄 PDF
OctoLong: Mid-Training On Cross-Repository Code Contexts Enhances Long-Context Modeling
Context lengths of language models (LMs) have dramatically increased, driven by the demands for in-context learning, self-improvement, and long-horizon agentic workflows. Existing long-context corpora, however, are dominated by books, academic articles, and code repositories, which are finite resour…
Indraneil Paul, Falko Helm, Goran Glavaš, Iryna Gurevych · 📄 PDF
Argus: A General-Purpose Agentic Runtime for Long-Horizon Reasoning
Long-horizon reasoning requires an agentic runtime that can persist when evidence supports its current approach and pivot when measurements reveal failure, hidden constraints, or a misspecified objective. We present Argus, a persistent, self-evolving runtime in which Manager, Planner, Engineer, and …
Boxiu Li, Zimo Wen, Yijia Fan, Junxiang Lei, Sufeng Guo, Jiaao Wu, Ruize Tang, Mukai Li, Yifei Shen, Xiaoyu Chen, Wanbo … · 📄 PDF
Does the Gender Wage Gap Originate at Labor Market Entry? Evidence from South Korea
When in the lifecycle does a large gender wage gap emerge? South Korea has the largest gender pay gap in the OECD, 29%. Among recent college graduates the conditional gap is only 4.3% over 2008-2019, falling from 5.0% to 3.0%, and correcting for differential selection into full-time wage employment …
Dongwoo Kim · 📄 PDF
Digital State Capacity
Digital State Capacity is the ability of governments to deploy ICT infrastructure and information systems to implement policy. This paper introduces a new measure of government ICT capacity based on an observable stock of deployable public-sector network infrastructure: public IPv4 address space hel…
Patrick Healy, Simon D. Angus, Paul Raschky, Klaus Ackermann, Nathan Lane, Weijia Li, Cynthia Huang · 📄 PDF
Stochastic Choice with Advertising
We study how advertised products (e.g., Top Picks, Recommended, Featured) affect consumer choice on digital platforms and retail interfaces by extending the Luce (1959) (or multinomial logit) model. A consumer either focuses on the advertised items or considers the full menu, then chooses among the …
Henrik Petri, Kai Wang · 📄 PDF
Cities and political violence in West Africa
This paper examines how political violence in West Africa is distributed between urban agglomerations and their surrounding areas using spatially disaggregated conflict and population data. Covering seventeen countries from 2012 to mid-2025, the analysis shows that political violence remains clearly…
Steven M. Radil, Olivier J. Walther · 📄 PDF
Transnational political violence in African borderlands
This paper examines the relationship between borderlands and political violence in Africa. Using spatiotemporal data on conflict events from 1997 to 2024 alongside the innovative OECD Spatial Conflict Dynamics indicator (SCDi), it suggests that borderlands see more conflict than non-borderlands. Pol…
David G. Russell, Olivier J. Walther · 📄 PDF
Preying on Leveraged ETFs
We argue that the extreme volatility of the Korean market in 2026 was driven by arbitrageurs preying on the closing rebalance of leveraged exchange-traded funds (LETFs). An LETF must trade in the direction of the day's move at a close that also measures it, so its demand rises in price, and arbitrag…
Yinhong Zhao · 📄 PDF
Synthetic supply networks
A good representation of the population of firms and households is essential for large-scale economic models. While there exist good methods to create synthetic populations of households, creating synthetic populations of firms and, crucially, their supply chain links, is typically much harder. Here…
Galvin Ng, Luca Mungo, Damien Bertrand, François Lafond · 📄 PDF
Does generative AI narrow education-based productivity gaps? Evidence from a randomized experiment
Does generative artificial intelligence (AI) widen or narrow productivity gaps across workers? We study this in a randomized online experiment with 1,174 adults aged 25-45 who completed a workplace-style problem-solving task with or without a generative AI assistant, followed by an unassisted module…
Guillermo Cruces, Diego Fernandez Meijide, Sebastian Galiani, Ramiro Galvez, Maria Lombardi · 📄 PDF
Fully Distributed Fiber-Optic Sensing Enabled by Kalman Filtering
Signal fading creates points along the fiber where phase cannot be extracted, so they are conventionally discarded. Instead, we propose a Kalman-based solution for φ-OTDR full-fiber monitoring. Experiments demonstrate phase and temperature estimation with approximately 15 times better spatial densit…
Juan M. Marin, Roman Ermakov, Florian Azendorf, André Sandmann, Francesco Da Ros, Darko Zibar · 📄 PDF
AFLOW-EMERALD: ElectroMagnetic modes EngineeRing in Advanced LayereD materials
Layered and periodically patterned heterostructures underpin advanced optical, photonic, and plasmonic (meta)materials, whose rational design demands electromagnetic solvers that are both numerically robust and tightly linked to the underlying material properties. Here, we present AFLOW-EMERALD (Ele…
Stefano Campanaro, Luca Bursi, Nicholas H. Anderson, Stefano Curtarolo, Arrigo Calzolari · 📄 PDF
Self-Focusing Control for Depth-Precise Wafer Slicing of 4H-SiC in Femtosecond Laser Processing
4H-SiC has emerged as a third-generation chip material because its superior thermal conductivity and high breakdown field enable the material to achieve high power density and higher switching frequencies in power-electronics applications. As chip architectures evolve toward 3D and heterogeneous int…
Dong Hee Kang, Jaeseung Lim, Mishfaqur Rahman, Seongheum Han, Jae-Hak Lee, Seungman Kim, Jihoon Jeong · 📄 PDF
Joint spectral characterization of SPDC photon pairs near 2 $μ$m in (Al)GaAs-on-insulator waveguides
Integrated photon-pair sources are a core component of chip-based quantum computing, communication, and metrology. Although such sources have been demonstrated at conventional telecom wavelengths, the 2 $μ$m band remains comparatively less explored, despite offering advantages for free-space quantum…
Alexandre Z. Leger, Emil Z. Ulsig, Samuel E. Fontaine, Dileep V. Reddy, Eric J. Stanton, Lynden K. Shalm, Richard P. Mir… · 📄 PDF
Modulation in degree of cross-polarization at Young's interferometer illuminated by non-uniformly polarized electromagnetic fields
The degree of cross-polarization (DoCP) and the electromagnetic degree of coherence (EM DoC) of an electromagnetic beam are investigated at the observation points for incoherent and non-uniformly polarized, i.e., different degree of polarization with respect to space at the two pinholes in Young's i…
Rajneesh Joshi, Gyaprasad · 📄 PDF
Can crystal symmetry reshape ENZ photonics?: Opinion
Over the last decade, epsilon-near-zero (ENZ) photonics has been driven by the search for lower losses and stronger nonlinear responses. Here, I ask a different question: can crystal symmetry also be used to shape the ENZ response? I focus on low-symmetry conductors, where the geometry of the electr…
Mario G. Silveirinha · 📄 PDF
Refraction laws in spatio-temporal media
We study the time-dependent Maxwell system, formulated in the sense of distributions, for electromagnetic waves propagating through media with temporal and spatial material interfaces. Under explicit trace and regularity assumptions on the permittivity and permeability, we derive the jump conditions…
Cristian E. Gutiérrez, Eric Stachura · 📄 PDF
Structured light under turbulence
Structured light has emerged as a promising resource for high-capacity and secure free-space optical communication, where atmospheric turbulence remains a major source of signal degradation. In this work, we investigate the resilience of different transverse mode structures of an optical beam with r…
Guilherme S. Barros, Lucas C. Céleri, Antonio Zelaquett Khoury, André L. S. Santos Junior, Rafael M. Gomes, Guilherme L.… · 📄 PDF
An optical-fibre-integrated buffer for packet-switched quantum networks
Packet-switched quantum networks require buffers that can delay qubit payloads while routing information is read out in real time. Previous approaches have not provided this functionality in a fully fibre-integrated architecture compatible with telecom infrastructure. Here we demonstrate an optical-…
Daniel Spegel-Lexne, Joakim Argillander, Martin Clason, Åsa Claesson, Kenny Hey Tow, Gustavo Lima, João M. B. Pereira, G… · 📄 PDF
Analysis of Nonlinear Phase Noise in Coherent Fiber-Optic Systems Based on Phase Shift Keying
Analytical expressions for the phase variance in a nonlinear fiber optic system based on phase-shift keying are developed. The Gauss-Hermite functions are used as the orthogonal basis to represent the noise field. Number of degrees of freedom (DOF) to accurately model the phase variance is estimated…
Shiva Kumar · 📄 PDF
Spatial proteomics guided by H&E-based AI reveals recurrence-risk niches in triple-negative breast cancer
Deep learning models can predict cancer recurrence from H&E stained slides, but the localized molecular states underlying these predictions remain largely obscured. Here, we developed an outcome informed spatial pathology framework in TNBC that integrates AI generated recurrence risk heatmaps with m…
Yesung Cho, Ji Hwan Park, Chanil Kim, Hyewon Kim, Honglan Li, Yumin Lee, Geongyu Lee, Sujeong Hong, Seong Min Park, Yoon… · 📄 PDF
Persistent homology broadens the controllable subspace in human structural connectomes
Network control theory applied to structural connectomes typically ranks brain regions as candidate driver nodes by their structural connectivity strength, and evaluates performance through scalar control energy. We test whether this framing captures the most relevant information about how driver-no…
Carter Sale, Marco Coraggio, Mengsen Zhang, Michael J. Richardson · 📄 PDF
The Cost of Binarizing Survival Outcomes in Clinical Prognostic Modeling
Survival analysis is an established framework for analyzing time-to-event data, yet many clinical machine learning studies still binarize the outcome before model training. This practice excludes censored patients, collapses temporal information into a single threshold, and can affect which features…
Shashank Yadav, David M. Routman, Andrew Y. K. Foong · 📄 PDF
Stochastic partial differential equation model for environmental DNA dynamics in river environments
Environmental DNA (eDNA) has emerged as a novel tool for quantifying the seasonal abundance of aquatic species in water bodies; however, its mathematical modeling is still at a germinating stage because of its mechanistic uncertainties. We propose a first-step mathematical and computational framewor…
Hidekazu Yoshioka · 📄 PDF
Population Structures with Positive Feedback and Asymmetric Division
In bioreactor experiments, budding yeast can manifest stable metabolic oscillations. In some instances these oscillations involve the cell cycle and are associated with a dynamical phenomenon called temporal clustering. Since yeast divide asymmetrically, mother cells may be able to divide sooner tha…
Gabriel Dooley, Camden Kilton, Brynley Needham, Graham Walther, Todd R. Young · 📄 PDF
Enactive Artificial Intelligence: A Decision-Centric Architecture for Complex Systems
As artificial intelligence (AI) continues to evolve and mature, recent AI practices have moved beyond large language models (LLMs) and text or image generation tasks, increasingly integrating tools, agents, and harnesses to solve real business and industrial problems. However, the power of AI is not…
Zuojun Max Shen, Yuan Qu, Pujun Zhang, Anbang Liu, Yunhao Liang · 📄 PDF