Paper Plaza - AcademicHub

01.

arXiv (CS.LG) 2026-06-18 DOI: arXiv:2606.18390

MOLAR: Learning Multimodal Molecular Representations from Noisy Labels

Authors:

Yingxu Wang ↗Kunyu Zhang ↗Nan Yin ↗Yu Li ↗Eran Segal ↗

arXiv:2606.18390v1 Announce Type: new Abstract: Motivation: Noisy labels are a common challenge in molecular property prediction because molecular annotations are often obtained from assays, curated databases, or weak annotation pipelines rather than directly observed clean biological states. Treating recorded labels as reliable supervision can cause models to memorize corrupted observations and learn misleading molecular evidence. In multimodal molecular representation learning, this issue can be amplified by graph-text fusion or alignment, which may propagate label-induced errors across modalities. Results: We propose MOLAR, a noise-aware framework for learning multimodal molecular representations from noisy labels. MOLAR separates latent clean-property inference from recorded-label observation: graph and text views contribute residual evidence to a clean-property distribution, and a categorical label-observation channel maps this distribution to recorded labels for training. This formulation derives posterior label reliability and modality-specific molecular evidence from the model. Experiments on naturally noisy molecular benchmarks and controlled label-flipping benchmarks show that MOLAR consistently outperforms representative baselines. Visualization analyses further show that MOLAR provides interpretable reliability and modality-evidence diagnostics.

Read & Discuss → View Source →

02.

arXiv (CS.CV) 2026-06-16 DOI: arXiv:2604.18866

HMR-Net: Hierarchical Modular Routing for Cross-Domain Object Detection in Aerial Images

Authors:

Pourya Shamsolmoali ↗Masoumeh Zareapoor ↗Michael Felsberg ↗Nick Pears ↗Huiyu Zhou ↗Yue Lu ↗

Despite advances in object detection, aerial imagery remains a challenging domain, as models often fail to generalize across variations in spatial resolution, scene composition, and semantic label coverage. Differences in geographic context, sensor characteristics, and object distributions across datasets limit the capacity of conventional models to learn consistent and transferable representations. Shared methods trained on such data tend to impose a unified representation across fundamentally different domains, resulting in poor performance on region-specific content and less flexibility when dealing with novel object categories. To address this, we propose a novel modular learning framework that enables structured specialization in aerial detection. Our method introduces a hierarchical routing mechanism with two levels of modularity: a domain routing layer that uses latent geographic embeddings to assign inputs to domain-specialized expert modules, and a scene routing mechanism that allocates image subregions to scene-specific expert modules. This allows our method to specialize across datasets and within complex scenes. Additionally, the framework contains a conditional expert module that uses external semantic information (e.g., category names or textual descriptions) to enable detection of novel object categories during inference, without the need for retraining or fine-tuning. By moving beyond monolithic representations, our method provides an adaptive framework for remote sensing object detection. Comprehensive evaluations on four datasets highlight improvements in multi-dataset generalization, region-level specialization, and open-category detection.

Read & Discuss → View Source →

03.

medRxiv (Medicine) 2026-06-23 DOI: HASH:ae6e54ff4a79aeb6b871813dbdf74fe4

Multidimensional motivation in aging: a person-centred framework spanning goal-directed behaviour, social reward and pleasure

Authors:

Warren ↗S. L ↗Somasundaram ↗Horne ↗Robinson ↗G. A ↗Behenska ↗Nguyen ↗Goh ↗A. M. Y ↗Jeon ↗Y.-H ↗…

Motivational changes are determinants of healthy aging, social engagement, and functional independence, and may signal early neurodegenerative risk. Existing assessment approaches in aging typically treat motivation as a unitary construct. Here, we introduce MotDem, an age-appropriate measure of motivation co-designed with people living with dementia, carers, and clinicians. Across a broad adult lifespan sample (18-80 years), MotDem revealed a robust three-domain motivational architecture encompassing goal-directed behaviour, social reward, and pleasure, with a fourth satiety factor retained as exploratory. This structure was replicated in an independent older cohort (45-80 years) from a different national context. MotDem showed strong convergence with established measures of apathy and anhedonia, alongside more modest associations with depressive symptomatology. Together, these findings show that motivational aging is multifaceted and poorly captured by traditional unitary assessment. MotDem provides a multidimensional framework for measuring distinct motivational drivers of heterogeneous aging trajectories, with implications for resilience, wellbeing, and neurodegenerative risk.

Read & Discuss → View Source →

04.

bioRxiv (Bioinfo) 2026-06-22 DOI: HASH:2a5258a58f620cde242d30fc3478b73c

Complex-valued representations of time-series gene expression profiles for network analysis

Authors:

Sun ↗Cao ↗Ikumi ↗Shimizu ↗K. K ↗Sese ↗

Time-series RNA sequencing provides a powerful framework for studying dynamic gene regulation, yet conventional analyses usually represent gene expression profiles as real-valued vectors in Euclidean space and quantify similarity using correlation or distance. Inspired by quantum information theory, we present a framework for encoding time-series gene expression profiles as complex-valued vectors comprising amplitude and phase components in Hilbert space. We designed multiple encoding models to represent gene expression in the amplitude of complex-valued vectors, encode temporal differences in the phase, and extend the phase representation to incorporate the direction of local expression changes. Gene-gene similarity was then quantified using fidelity, which measures the overlap between two encoded vectors. Evaluation using time-series RNA-seq datasets across diverse species and biological contexts showed that different encoding models produced distinct fidelity distributions that were related to, but distinct from, conventional correlation measures. We then constructed gene-gene networks using pairwise fidelity values and detected communities containing genes with similar temporal profiles. Although fidelity distributions differed across encoding models, the resulting communities captured major temporal expression programs, and functional annotations based on gene ontology and Kyoto encyclopedia of genes and genomes pathway analyses provided exploratory biological context. The detected communities were comparable to those obtained using conventional methods, including weighted correlation network analysis and fuzzy c-means clustering. Furthermore, as a proof-of-concept, we performed SWAP-test circuit simulations to mimic fidelity computation on a quantum computer; under noise-aware conditions, these simulations produced less accurate fidelity estimates with higher computational cost than classical computation. As a proof-of-concept, this study provides a complementary view of temporal transcriptome organization, rather than a uniformly superior alternative to conventional methods.

Read & Discuss → View Source →

05.

medRxiv (Medicine) 2026-06-24 DOI: HASH:30feaee165a4b33910b7c1e0e43d2619

Uncovering the fitness of endemically circulating Zika virus strains

Authors:

Chen ↗Fritz ↗Clapham ↗H. E ↗Lefrancq ↗Salje ↗BUZZ study team ↗

Zika virus (ZIKV) is an arbovirus that usually causes few symptoms and has circulated endemically in Asia for decades. However, a large outbreak in South America in 2015 uncovered the serious risk of congenital Zika syndrome in infants born from ZIKV infected mothers. It is unknown whether a lineage with distinct pre-existing fitness advantage emerged from Asia to cause the South American outbreak, and whether there is ongoing evolution that can result in future globally fit strains. Here we used 107 sequences from a single setting (Thailand) collected over an 18 year period (2006-2023). We used novel analytical tools to identify distinct lineages that have circulated in the population and estimated their relative epidemiological fitness. We found there have been six lineages circulating sequentially in the country, with regular emergence and replacement of lineages showing higher fitness than their predecessors. We identified 15 lineage-defining amino acid changes, including four well-documented fitness-enhancing mutations, and two UTR substitutions. The lineage that emerged in South America was evolutionarily linked to the highest-fitness lineage in Thailand, carrying seven of our lineage-defining substitutions acquired during endemic circulation there, and subsequently accumulating four additional changes. After the global pandemic, endemic ZIKV in Thailand continued to evolve, with newly emerged lineages showing novel mutations and increased fitness. Our findings have key implications for the monitoring of ZIKV and can help identify the pathway to increased transmissibility of this globally important pathogen.

Read & Discuss → View Source →

06.

medRxiv (Medicine) 2026-06-16 DOI: HASH:19331ad18dcc6f5ed3e4a02042594f6c

Infections and suicide and self-harm: a population-based matched cohort study

Authors:

Fanjul Iglesias ↗Gore-Langton ↗G. R ↗Cadogan ↗S. L ↗Mansfield ↗K. E ↗Douglas ↗I. J ↗Fazel ↗Tazare ↗Morton ↗…

Background Infections have been associated with adverse mental health outcomes, including suicide, but evidence beyond severe or central nervous system infections is limited. We investigated associations between a range of acute infections and subsequent suicide/self-harm outcomes. Methods We conducted six infection-specific matched cohort studies using English primary care records from the Clinical Practice Research Datalink Aurum (2007-2024), linked to hospital admissions and mortality data. Adults ([≥]18 years) with a primary care record of infection (gastroenteritis, lower respiratory tract [LRTI], skin/soft-tissue [SSTI], urinary tract [UTI], sepsis, meningitis/encephalitis [positive control]) were matched (age, sex, practice, calendar period) to up to five comparators without infection. We estimated hazard ratios (HRs) for suicide/self-harm outcomes using Cox regression, stratified by matched set and implicitly adjusting for matching factors, with additional adjustment for deprivation, lifestyle factors, and comorbidities. We examined whether associations varied over time, by infection severity, antimicrobial treatment, sex, and prior mental health conditions. Findings Cohorts ranged from 18,192 individuals with meningitis/encephalitis (matched to 90,915 without) to 398,099 with SSTI (matched to 1,743,747). After adjustment, individuals with infection had a higher hazard of suicide/self-harm outcomes than comparators across all cohorts: sepsis (HR 1.79, 95% CI 1.65-1.93), gastroenteritis (1.62, 1.55-1.70), meningitis/encephalitis (1.56, 1.32-1.84), UTI (1.41, 1.33-1.50), SSTI (1.37, 1.31-1.43), and LRTI (1.37, 1.31-1.44). Risk was highest in the year post-infection, attenuating over time, and was higher among severe infections and those without prior mental health conditions. Interpretation Common acute infections recorded in primary care are associated with increased risk of suicide and self-harm, particularly following severe infections and in the year post-infection. Findings support suicide risk monitoring following acute infection, particularly among individuals without prior mental health conditions, and highlight infection prevention as a potentially modifiable strategy in vulnerable populations. Funding Wellcome and La Caixa. Copyright This work is licensed under a Creative Commons Attribution (CC BY) licence.

Read & Discuss → View Source →

07.

medRxiv (Medicine) 2026-06-12 DOI: HASH:1d1667ee1dd0931385986c6ab2063254

High coverage, persistent gaps: quality of Antenatal Care and its determinants in Zambia based on the 2024 Demographic and Health Survey.

Authors:

Tukamuhebwa ↗P. M ↗Nuwabaine ↗

Abstract Background Evaluating antenatal care (ANC) quality is critical to reducing maternal and neonatal mortality. In Zambia, despite high basic ANC attendance, comprehensive national evidence on the clinical content and quality of services remains limited. This study assessed the coverage of WHO-recommended ANC interventions and identified factors associated with care quality using the latest national data. Methods A cross-sectional analysis was conducted using data from the 2024 Zambia Demographic and Health Survey. The final analytic sample comprised 4,829 women aged 15-49 with a live birth in the preceding 5 years. A composite index of 15 selected, equally weighted WHO-recommended components evaluated clinical assessment, counseling/screening, preventive interventions, and utilization. Survey-weighted Poisson regression estimated adjusted incidence rate ratios (aIRRs) for the count of ANC components received. Results The mean ANC quality score was 12.5 out of 15 (95% CI: 12.4-12.6), and 78.5% (95% CI: 77.0-80.0) of women achieved adequate ANC ([≥] 12/15 components). While individual clinical and counseling coverage generally exceeded 90%, only 47.2% (95% CI: 45.3-49.0) of women initiated care during the first trimester, and just 4.8% (95% CI: 4.1-5.6) achieved [≥] 8 ANC contacts. Maternal education was the strongest and most stable predictor of quality across all models. Compared to no education, higher education was associated with an 8.0% higher expected quality score (aIRR = 1.080, 95% CI: 1.051-1.110). Lower ANC quality was significantly associated with unwanted pregnancies (aIRR = 0.970, 95% CI: 0.956-0.993) and with residence in Western (aIRR = 0.923, 95% CI: 0.897-0.951) and North Western (aIRR = 0.966, 95% CI: 0.937-0.996) provinces. Absence of distance barriers and residence in Eastern, Luapula, and Copperbelt provinces were associated with higher quality scores. Conclusion While average ANC component coverage in Zambia is high, critical gaps persist in early initiation and total contact frequency. Care adequacy is strongly influenced by maternal education, relationship status, pregnancy intention, and regional inequities. These findings underscore the need for interventions targeted at uneducated women, preventing unintended pregnancies, and underserved regions such as Western and North Western Provinces. Keywords: Antenatal care quality, ANC content, Zambia, maternal education.

Read & Discuss → View Source →

08.

arXiv (CS.CL) 2026-06-17 DOI: arXiv:2606.17299

Examining the Limits of Word2Vec with Toki Pona

Authors:

Daniel Zhenhan Huang ↗Hongchen Wu ↗

Word2Vec's effectiveness at generating semantic embeddings has been widely validated, yet it has been tested almost exclusively on languages with large vocabulary inventories. This study examines whether Word2Vec can successfully capture semantic relationships within an extremely reduced vocabulary using data from Toki Pona, a constructed language with approximately 130 words. We sourced 1.4 million sentences (7.95 million tokens) from the Toki Pona community for training. Approximately 23% of sentences in the corpus contain non-Toki Pona tokens such as named entities, loanwords, and neologisms. To investigate whether this linguistic noise enhances or hinders performance – a topic rarely addressed in word embedding literature – we trained two distinct models: one retaining these incidental tokens and another filtering them out completely. Evaluation was conducted using quantitative methods measuring word proximity to semantic category centroids, automated silhouette scores via agglomerative clustering, and qualitative analysis utilizing representational similarity matrices compared against English. The results indicate that while sparse, non-core tokens do not affect the relative structure of the learned embeddings, they actually draw similar words closer together in the vector space. Importantly, Word2Vec's effectiveness depends more on distributional patterns than lexicon size even at this extreme lower bound.

Read & Discuss → View Source →

09.

arXiv (CS.CV) 2026-06-12 DOI: arXiv:2606.12913

Selecting Samples on Graphs: A Unified Dataset Pruning Framework for Lossless Training Acceleration

Authors:

Dongyue Wu ↗Zilin Guo ↗Xiaoyu Li ↗Jiajia Liu ↗Jingdong Chen ↗Nong Sang ↗Changxin Gao ↗

The rapid growth of modern training datasets has significantly increased computational cost, motivating dataset pruning~(DP) methods which retain only a subset of informative samples to reduce training cost. Existing pruning criteria typically rely on either intrinsic signals that assess samples independently or extrinsic signals that promote diversity via pairwise relations. While effective in their own specific regimes, each captures only one aspect of sample utility and lacks robustness across different pruning ratios or data distribution. In this work, we present a unified graph-based DP framework. By modeling the dataset as a weighted graph, where node weights encode intrinsic value and edge weights encode extrinsic value, DP can be cast as a Maximum Weight Clique Problem (MWCP). Although MWCP is NP-hard, its structure admits a principled greedy solution based on sample-wise marginal gains. Under a few mild conditions, we further prove that this unified objective enjoys a formal approximation guarantee, which applies to a broad family of importance metrics and provides practical design guidelines. Extensive experiments show that our method outperforms existing DP methods while substantially reducing training cost, reducing training time by over 40\% without sacrificing accuracy on ImageNet-1k with ResNet-50.

Read & Discuss → View Source →

10.

PLOS Computational Biology 2026-06-10 DOI: HASH:73c8f5c625ba663a84c7b5b533c272be

Ten simple rules for teaching data science

Authors:

Tiffany A. Timbers ↗

by Tiffany A. Timbers, Mine Çetinkaya-Rundel

Read & Discuss → View Source →

11.

arXiv (CS.AI) 2026-06-15 DOI: arXiv:2603.10444

The Curse and Blessing of Mean Bias in FP4-Quantized LLM Training

Authors:

Hengjie Cao ↗Zhendong Huang ↗Mengyi Chen ↗Yifeng Yang ↗Fang Dong ↗Anrui Chen ↗Ruijun Huang ↗Xin Zhang ↗Mingzhi Dong ↗Yujiang Wang ↗Jinlong Hou ↗Qin Lv ↗…

arXiv:2603.10444v2 Announce Type: replace-cross Abstract: FP4 training promises substantial memory and compute savings for large language models, but remains fragile because blockwise quantization is dictated by extreme activation magnitudes, which inflate dynamic range and compress long-tail signals. We identify a counterintuitive source of this failure: dominant activation outliers are not merely arbitrary sparse events, but are largely induced by a coherent rank-one mean bias, whose direction aligns with the leading anisotropic spectral component. This mean component strengthens during training, is amplified and reshaped by attention and FFN operators, and increasingly dominates top activation magnitudes. Crucially, this discovery reveals that a seemingly complex outlier-suppression problem admits a truly simple solution: isolate the coherent mean before quantization. We therefore propose Averis, a mean-residual splitting quantization method that separates the mean component using only reductions and elementwise subtractions before FP4 quantization. Across Qwen3 0.6B Dense trained on 100B tokens and Qwen3 7B A1.5B MoE trained on 50B tokens, Averis enables robust W4A4G4 FP4 training, reducing BF16 loss gaps to 1.19%/0.81% versus 2.05%/1.10% for NVIDIA's recently released Hadamard-based outlier-smoothing method, while limiting downstream gaps to 0.89/0.71 points. With only 2.20% end-to-end overhead over vanilla NVFP4, about 30% of NVIDIA's Hadamard-based design, Averis provides a hardware-efficient path to stable low-bit LLM training. Complementary to Hadamard, Averis further reduces the Qwen3-0.6B loss and downstream gaps to 0.94% and 0.73 points when combined. Code is available at: https://anonymous.4open.science/r/averis-504D.

Read & Discuss → View Source →

12.

arXiv (CS.AI) 2026-06-19 DOI: arXiv:2606.20208

Beyond Accuracy: Measuring Logical Compliance of Predictive Models

Authors:

Guillaume Olivier Delplanque ↗Pierre Genev\`es ↗Nabil Laya\"ida ↗Zephirin Faure ↗

arXiv:2606.20208v1 Announce Type: new Abstract: Machine learning models are predominantly evaluated through predictive performance metrics such as ranking quality, prediction error, or classification accuracy. While these metrics effectively quantify how closely predictions match the ground truth, they do not assess whether model outputs respect predefined logical or domain-specific constraints. In high-stakes applications, including healthcare, finance, and autonomous systems, logical consistency can be as critical as predictive accuracy, yet no standard metric captures this dimension. We introduce the Rule Violation Score (RVS), a complementary evaluation metric that quantifies the extent to which a predictive model respects a given set of logical rules, independently of predictive accuracy. RVS treats hard rules (strict constraints) and soft rules (statistical regularities) differently, can be evaluated on any dataset and on any predictive model expressed over a relational vocabulary, and can be computed using SQL queries that are automatically generated for Horn rules. Beyond evaluating models, RVS can also evaluate the logical consistency of training datasets and help identify poorly defined rules. We evaluate RVS on three benchmarks covering knowledge graph link prediction and relational regression, including rule-based, embedding-based, and neuro-symbolic predictive models. Our results demonstrate that two models achieving comparable predictive accuracy can exhibit substantially different levels of logical compliance, revealing differences in model behavior that standard metrics fail to capture.

Read & Discuss → View Source →

13.

arXiv (CS.CV) 2026-06-11 DOI: arXiv:2606.11645

Motion Reinforces Appearance: RGB-Skeleton Gated Residual Fusion for Micro-Gesture Online Recognition

Authors:

Jialin Liu ↗Xinwen He ↗Pengyu Liu ↗Jiale Shi ↗Huaijuan Zang ↗Yanbin Hao ↗

Micro-gesture analysis attracts increasing attention for inferring spontaneous emotion from subtle body movements. Micro-gesture online recognition, which localizes and classifies each gesture instance in untrimmed videos, is a core task in the 4th EI-MiGA-IJCAI Challenge. Compared with typical temporal action detection, MGR emphasizes the localization and classification of actions, requiring the model to output the start time, end time, and category of each micro-gesture. Moreover, since micro-gestures are highly spontaneous, relying solely on a single modality makes it difficult to capture the complete and accurate multi-modal cues. In this work, we propose DyFADet+, which extends DyFADet into a dual-stream RGB-skeleton framework. In our model, both modalities are projected into shared multi-scale temporal embeddings and fused through a gated residual module, which adaptively injects skeleton motion into the RGB representation rather than using naive concatenation. Finally, these fused features are decoded by a Dynamic TAD head for online classification and boundary regression. On the SMG dataset, our method achieves an F1 score of 40.88, ranking 2nd in the Micro-gesture Online Recognition track.

Read & Discuss → View Source →

14.

arXiv (CS.AI) 2026-06-24 DOI: arXiv:2606.23983

Maestro Order: A Model-Agnostic Orchestration Harness

Authors:

Hidayet Aksu ↗

arXiv:2606.23983v1 Announce Type: cross Abstract: A single forward pass of a capable model is a fast, fluent, and unreliable problem-solver: it is right often enough to be useful and wrong often enough to be dangerous; in language models, such confident errors are known as hallucinations. We present Maestro Order, a model-agnostic orchestration harness that turns unreliable solvers into reliable problem-solving systems by composing them according to four structural primitives (decompose, ensemble, verify, and recurse) and a budget-aware controller that decides where to spend compute. The harness treats any model as a black-box base solver behind a uniform interface, layers a verifier ensemble whose discrimination is measured online, and allocates verification and voting to the stages with the highest marginal reliability per unit cost. We give the architecture, the message and state schema, the controller algorithm, and the engineering that makes it deterministic, observable, and fault-tolerant. We then specify an evaluation methodology (reliability at fixed cost, coverage, calibration, and ablations) and report results from a faithful Monte Carlo simulation of the harness over a parameterized solver/verifier model. The simulation reproduces the predicted laws quantitatively: verification amplifies reliability geometrically (e.g. $0.55\to0.98$ with two gates, $\to0.999$ with four), voting helps only above chance and is limited by shared errors, and a budget-aware controller reaches a target reliability at a small fraction of the cost of voting alone by selecting the cheapest mechanism for each regime. We close with failure modes (verifier gaming, correlated errors, and decomposition error compounding) and concrete guidance: build robust checkers, diversify solvers, and let the controller put compute where the information is.

Read & Discuss → View Source →

15.

arXiv (CS.CV) 2026-06-16 DOI: arXiv:2511.12024

Null-Space Diffusion Distillation Unlocks Speed, Fidelity and Realism in Lensless Imaging

Authors:

Jose Reinaldo Cunha Santos A V Silva Neto ↗Hodaka Kawachi ↗Yasushi Yagi ↗Tomoya Nakamura ↗

Lensless imaging reconstructs scenes from highly multiplexed measurements, resulting in a severely ill-posed inverse problem. In this work, we identify a fundamental trade-off between measurement consistency, perceptual quality, and inference speed across lensless reconstruction paradigms. Traditional methods favor consistency but produce perceptually degraded results, supervised approaches achieve high-quality reconstructions with fast inference but may violate physical constraints, and diffusion-prior methods achieve high perceptual quality and consistency–particularly when structured constraints such as range-null decomposition are used–but remain slow due to iterative sampling. Motivated by this observation, we propose Null-Space Diffusion Distillation (NSDD), a single-pass reconstruction model that distills structured diffusion-prior inference into an efficient feed-forward network. NSDD learns to produce high-quality reconstructions that preserve measurement consistency while avoiding costly iterative sampling. Experimental results demonstrate that NSDD achieves perceptual quality and consistency competitive with diffusion-prior methods, while providing significantly faster inference and offering a favorable balance across all three objectives. Furthermore, ablation experiments show that distilling the range–null decomposition improves reconstruction quality and robustness over unstructured full-reconstruction distillation, including on unseen real scenes. These results highlight the potential of structure-aware distillation for efficient lensless imaging. Code is available at github.com/JRCSAVSN/NullSpaceDiffusionDistillation.

Read & Discuss → View Source →

16.

arXiv (CS.CL) 2026-06-24 DOI: arXiv:2606.24083

CAVEWOMAN: How Large Language Models Behave Under Linguistic Input and Output Compression

Authors:

Morayo Danielle Adeyemi ↗Ryan A. Rossi ↗Franck Dernoncourt ↗

"Talk short. Drop grammar. Save token." This caveman style is widely promoted as a way to cut inference cost, but whether it actually saves anything depends on which channel (the user's prompt or the model's response) is being compressed. We present Cavewoman, a two-channel evaluation protocol that scores every generation on task accuracy, realized per-item cost, and reference-text agreement against the model's unconstrained reference. We evaluate eight models on five datasets at five reduction levels, with both channels measured on the same items. Output compression cuts realized cost on most API models (1.4-2.4x per model, up to 3x in the best case) and on all four open-weight models under public-tier pricing. Input compression has the opposite effect, a strict lose-lose: it raises net cost rather than lowering it (~1.15x on the five-benchmark mean, up to 1.8x on the worst dataset and 2.7x under stronger compression), because models compensate with longer responses even as accuracy collapses. Under the same setting, surface text diverges from the unconstrained reference: on the non-reasoning models, roughly half of all generations are correct yet their surface text no longer entails the model's own unconstrained baseline generation. The divergence survives length-controlled re-scoring, multiple-comparisons correction, and replication under complementary semantic measures. Code and data are available at https://github.com/danielle34/cavewoman.

Read & Discuss → View Source →

17.

arXiv (quant-ph) 2026-06-16 DOI: arXiv:2505.16606

Fast and high-fidelity transfer of edge states via dynamical control of topological phases and effects of dissipation

Authors:

Yuuki Kanda ↗Yusuke Fujisawa ↗Kousuke Yakubo ↗Norio Kawakami ↗Hideaki Obuse ↗

arXiv:2505.16606v2 Announce Type: replace-cross Abstract: Topological edge states are robust against symmetry-preserving perturbations and noise, making them promising for quantum information and computation, particularly in topological quantum computation through the braiding operations of Majorana quasiparticles. Realizing these applications requires fast and high-fidelity dynamic control of edge states. In this work, we theoretically propose a high-fidelity protocol for transferring topological edge states by dynamically moving a domain wall between two regions with different topological numbers in one dimension. This protocol fundamentally relies on Lorentz invariance and relativistic effects, because moving the domain wall at a constant speed is described by a mass term with the uniform linear motion in the Dirac equation. We demonstrate the effectiveness of our protocol in transferring edge states with high fidelity using a one-dimensional quantum walk with two internal states, which is feasible with current experimental technology. We also investigate how bit-flip and dephasing dissipation to the environment affect transfer efficiency. Remarkably, bit (dephasing) dissipation does not affect the fidelity at the slow (fast) transfer limit, which can be explained by the relativistic effects on the edge states.

Read & Discuss → View Source →

18.

medRxiv (Medicine) 2026-06-24 DOI: HASH:ca4ad85af17b2482e6764f61655e4de9

Five-Year Breast Cancer Risk Prediction From Screening Breast Ultrasound Using Deep Learning

Authors:

Chen ↗Yang ↗Xu ↗Soni ↗Heacock ↗Lis ↗Stanek ↗Puto ↗Lewin ↗A. A ↗Moy ↗Schnabel ↗…

Objective: To develop and evaluate a deep learning model for five-year breast cancer risk prediction from screening breast ultrasound (BUS) examinations. Methods: This retrospective study included 295,298 breast ultrasound examinations from 122,072 women imaged between 2012 and 2020. Patients were split into training, validation, and test sets; the test set included screening examinations only. BUS-Risk-Net aggregated image features using attention-based multiple instance learning and combined them with age and ultrasound-estimated breast density to predict 2- to 5-year risk. Performance was compared with the full Tyrer-Cuzick model in a matched case-control cohort and with a reduced Tyrer-Cuzick model in the held-out test set. Risk stratification was evaluated within BI-RADS density categories. Results: In the matched case-control cohort (n = 240 women), BUS-Risk-Net achieved a 5-year AUC of 0.632 (95% CI, 0.562-0.702), versus 0.514 for the full Tyrer-Cuzick model (95% CI, 0.440-0.588; p = 0.04). Among 19,548 examinations from 9,015 women eligible for 5-year evaluation in the test set, BUS-Risk-Net achieved an AUC of 0.679 (95% CI, 0.653-0.706), versus 0.594 for the reduced Tyrer-Cuzick model (95% CI, 0.564-0.623; P < .001). Observed 5-year cancer incidence increased across AI-defined risk tiers within each BI-RADS density category, ranging from 0.0% to 5.8% after AI stratification, compared with 2.1% to 3.6% across density categories alone. Discussion: Deep learning models applied to screening breast ultrasound could enable long-term breast cancer risk prediction and stratify risk beyond breast density alone. External and prospective validation is needed before clinical use.

Read & Discuss → View Source →

19.

medRxiv (Medicine) 2026-06-23 DOI: HASH:6cfb10c2f2558e76de3f2b39e6040e03

Changes in hierarchical brain dynamics of rumination following mindfulness-based cognitive therapy for depression

Authors:

Dagnino ↗P. C ↗van der Velden ↗A. M ↗Ruhe ↗H. G ↗Kuyken ↗Kringelbach ↗M. L ↗Vohryzek ↗Deco ↗

Major depressive disorder (MDD) is a leading cause of disability worldwide with risk of onset and recurrence linked to depressive ruminative thought patterns. Mindfulness-based cognitive therapy (MBCT) is an evidence-based treatment for depression that targets the ability to recognise, decenter, and disengage from ruminative thought patterns. Elucidating how MBCT impacts hierarchical brain organisation may be key to understanding the processes by which MBCT can modulate ruminative tendencies. In a randomised controlled functional magnetic resonance imaging (fMRI) trial on individuals with MDD (N=80) before and after MBCT in addition to treatment as usual (TAU), we investigated changes in hierarchical brain organisation during resting-state and rumination. We built whole-brain models to obtain generative connectivity (GEC) matrices per patient and quantified brain hierarchy by measuring the global directedness and regional trophic levels in each GEC, in which greater directedness reflects more directional information flow and less recurrence. Global directedness in MBCT+TAU compared to TAU increased during rumination, with no changes during resting-state. Furthermore, increased regional breadth of hierarchy during rumination was related to improvements in clinical and behavioural outcomes following MBCT+TAU. Increased brain hierarchy during rumination following mindfulness training may be consistent with a shift away from self-reinforcing negative mental loops towards more differentiated and less coupled cognitive and bodily cycles, supporting MBCT's ability to interrupt ruminative processes. Hierarchical brain dynamics may hold promise as a treatment-sensitive marker and a potential mechanism of therapeutic change in MBCT for depression.

Read & Discuss → View Source →

20.

PLOS Computational Biology 2026-06-23 DOI: HASH:4383fad5758df55addcde28b8b235773

A novel biclustering algorithm for mining m<sup>6</sup>A co-methylation patterns based on beta-binomial distribution and data screening strategy

Authors:

Zhaoyang Liu ↗

by Zhaoyang Liu, Yuteng Xiao, Dao Xiang, Hao Shi, Kaijian Xia Studies have shown that m6A plays a key role in different life processes such as RNA metabolism, physiology and pathology. However, due to the complexity of life processes, its specific regulatory details are still not revealed. The computational approach based on co-methylation pattern mining of m6A sequencing data can assist in revealing its mechanism and save time and economic cost, however, the current algorithms suffer from the problems of insufficient robustness to low signal-to-noise data and unreliable performance. Based on this, this paper proposes an enhanced beta-binomial distribution biclustering algorithm (EBBM) based on data screening strategy. This algorithm is based on the framework of Bayesian, adopts Gibbs sampling method for parameter inference, and introduces the data screening strategy in the process of parameter inference, which effectively removes the problem that the low signal-to-noise data in the original sequencing data of m6A affects the reliability of the clustering results. The simulation experiment results show that this algorithm can effectively deal with the interference of low signal-to-noise data and accurately mine the co-methylation patterns pre-planted in the data, which is significantly better than the current mainstream biclustering algorithm. In real human m6A sequencing data with 32 samples, this algorithm mined two effective co-methylation patterns, which were enriched to different biological processes, such as negative regulation of phosphorylation and peptidyl lysine methylation, etc. The scoring results of GEO_Score indicate that the results of this algorithm are more biologically meaningful than the clustering results of current mainstream m6A co-methylation pattern mining algorithms.

Read & Discuss → View Source →

21.

medRxiv (Medicine) 2026-06-23 DOI: HASH:335255d85209159e2b3b2dd7c26774c0

Respiratory support with Continuous Positive Airway Pressure in preterm neonates: an analysis of coverage and quality of care in 66 neonatal units in Kenya, Malawi, Nigeria and Tanzania implementing with the NEST360 Alliance

Authors:

Shemwell ↗Wainaina ↗Lawn ↗J. E ↗Salim ↗Penzias ↗R. E ↗Malla ↗Johari ↗Tillya ↗Bohne ↗C. A ↗…

Background: Prematurity is the leading cause of child deaths worldwide, with the highest neonatal mortality in sub Saharan Africa. Respiratory distress syndrome (RDS) is the leading mortality pathway in preterm neonates, but continuous positive airway pressure (CPAP) has high impact. This analysis reports CPAP coverage and quality of care for preterm neonates admitted to 66 neonatal units in Kenya, Malawi, Nigeria and Tanzania. Methods: Analyses used individually linked neonatal inpatient data and cross-sectional health systems data. All admitted neonates were eligible for inclusion (January 2021 through December 2024). Service readiness for CPAP delivery and mean CPAP coverage were described for CPAP eligible newborns (weighing 1500g). Quality of care cascades were constructed to illustrate key indicators. Survival among CPAP eligible neonates was analysed using regression models, stratified by clinical severity scores. Results: 375,255 newborn admissions were analysed in 66 neonatal units. Functional CPAP availability varied with median 16% of days (IQR: 4 to 47%) classified as high demand (>1.5 eligible newborns per CPAP). Of 64,761 CPAP eligible neonates, 22,006 (34%, 95% CI 33 to 34%) received CPAP. All countries showed improvement in CPAP coverage, with Tanzanian hospitals recording 63% increase in mean coverage (p-value=0.001) over time. Quality of care cascades showed treatment was initiated 1 day for 42% (95% CI 41 to 43%) of eligible neonates receiving CPAP. Only 10% of neonates

Read & Discuss → View Source →

22.

arXiv (CS.CL) 2026-06-16 DOI: arXiv:2606.16307

State-Grounded Multi-Agent Synthetic Data Generation for Tool-Augmented LLMs

Authors:

Rahul Khedar ↗Eshita ↗Sneha Teja Sree Reddy Thondapu ↗Mayank Malhotra ↗Arup Das ↗Jitesh Chandra ↗Yun-Shiuan Chuang ↗Chaitanya Kulkarni ↗Arun Menon ↗Linsey Pang ↗Avinash Karn ↗Mouli V ↗…

Training tool-augmented LLM agents requires large corpora of multi-turn, tool-grounded conversational data that is expensive to annotate, privacy-constrained in production settings, and largely absent from public datasets. We present StateGen, a synthetic data generation platform that produces scored, reasoning-trace-rich training conversations by orchestrating a four-role LLM loop: a persona-conditioned user simulator, an agent under test, a state-grounded tool simulator, and a multi-axis LLM judge. The key architectural contribution is an authoritative state manager that maintains a structured world-state object across turns, enforcing a backend-is-truth invariant that eliminates the dominant class of tool-call hallucinations by construction. StateGen extends naturally to hierarchical multi-agent settings by declaring sub-agents as tools, all sharing a single state object. We report results on 64,698 evaluated conversations across three production corpora: tool-call hallucination scores reach 9.66/10, the system supports persona-driven variation via a 23-dimensional trait vector, and a cleanly separated train and golden evaluation set split confirms the data is not memorization bait (per-criterion gap analysis). Comparison with eight external systems shows that no single publicly available platform combines multi-turn generation, state-grounded tool simulation, hierarchical multi-agent support, and built-in judge scoring.

Read & Discuss → View Source →

23.

medRxiv (Medicine) 2026-06-16 DOI: HASH:84ae281699f52b48aaec31f22d4fb0f9

Cross-sectional study of the association between depressive symptoms and attentional bias to emotional stimuli in patients with acute stroke: Study protocol

Authors:

Yamashita ↗Takizawa ↗Koizumi ↗Hamaguchi ↗

Post-stroke depression affects approximately 30% of patients after stroke and is associated with delayed recovery in activities of daily living, reduced rehabilitation effectiveness, and poorer quality of life. Attentional bias modification may provide a low-burden, nonpharmacological approach for patients in the acute phase of stroke. However, before such an intervention can be implemented in clinical practice, it is necessary to clarify whether attentional bias is present in patients with acute stroke and depressive symptoms, whether cognitive function influences the manifestation of this bias, and which task and stimulus formats are most appropriate for assessment. This multicenter, cross-sectional observational study will enroll patients with acute stroke between 7-30 days after stroke onset. Depressive symptoms will be assessed using the depression subscale of the Hospital Anxiety and Depression Scale. Attentional bias will be measured under four task conditions based on the dot-probe task and the cue-target task, using face and word stimuli. Secondary assessments will include cognitive function, anxiety symptoms, activities of daily living, health-related quality of life, and clinical background variables. The aims of this study are to investigate the association between depressive symptoms and attentional bias in patients with acute stroke, compare attentional bias characteristics across task and stimulus types, and examine the potential influence of cognitive function on this association. The findings are expected to provide an empirical basis for designing future attentional bias modification protocols targeting post-stroke depression in the acute phase. This study has been registered with the UMIN Clinical Trials Registry (UMIN000059166).

Read & Discuss → View Source →

24.

arXiv (quant-ph) 2026-06-24 DOI: arXiv:2406.07884

Reinforcement Learning to Disentangle Multiqubit Quantum States from Partial Observations

Authors:

Pavel Tashev ↗Stefan Petrov ↗Matthew T. Diaz ↗Friederike Metz ↗Alaina M. Green ↗Norbert M. Linke ↗Marin Bukov ↗

arXiv:2406.07884v3 Announce Type: replace Abstract: Using partial knowledge of a quantum state to control multiqubit entanglement is a largely unexplored paradigm in the emerging field of quantum interactive dynamics with the potential to address outstanding challenges in quantum state preparation and compression, quantum control, and quantum complexity. We present a deep reinforcement learning (RL) approach using an actor-critic algorithm for constructing short disentangling circuits for states with up to 16 qubits. With access to only two-qubit reduced density matrices, our agent decides which pairs of qubits to apply two-qubit gates on; requiring only local information makes it directly applicable on modern NISQ devices, as we demonstrated experimentally on a trapped-ion quantum computer. Utilizing a permutation-equivariant transformer architecture, the agent can autonomously identify qubit permutations within the state, and adjusts the disentangling protocol accordingly. Once trained, it provides circuits from different initial states without further optimization. We demonstrate the agent's ability to identify and exploit the entanglement structure of multi-qubit states. We analyze the disentangling circuits constructed by the agent for 4- and 5-qubit Haar-random states, and observe strong correlations between consecutive gates and among the qubits involved. Through extensive benchmarking, we show the efficacy of the RL approach to find disentangling protocols with minimal gate resources. We explore the resilience of our trained agents to noise, highlighting their potential for real-world quantum computing applications. Analyzing optimal disentangling protocols, we report a general circuit to prepare an arbitrary 4-qubit state using at most 5 two-qubit (10 CNOT) gates.

Read & Discuss → View Source →

25.

arXiv (CS.CV) 2026-06-18 DOI: arXiv:2606.18707

PEFT-MedSAM: Efficient Fine-Tuning of Medical Foundation Models for Explainable Skin Lesion Segmentation

Authors:

Asad Channa ↗Abdullah Khan ↗Asghar Ali Chandio ↗Aamir Akbar ↗Shahzad Memon ↗Aqib Hussain ↗Ameer Hamza ↗

Automated segmentation of skin lesions using deep learning models for dermoscopic images can be very helpful in finding melanomas earlier than they would normally be detected. However, most deep learning methods available do not perform well. The aim of this paper is to present a parameter-efficient fine-tuning method called PEFT-MedSAM for adapting the Medical Segment Anything Model (MedSAM) to automatically segment dermoscopic skin lesions. The PEFT-MedSAM method uses only the lightweight mask decoder for training the model while keeping the pre-trained image encoder and prompt encoder frozen. The experiments performed on the ISIC 2018 benchmark dataset shows that PEFT-MedSAM obtains a dice coefficient of .9411 and an intersection over union value of .8918 when compared to both a fully trained U-Net baseline (.8715 dice coefficient) and zero-shot MedSAM inference (.8997 dice coefficient). The external validation of the model using PH2 dataset shows .9467 dice coefficient with +/- .0310 standard deviation. Supportive evidence for these claims include a p-value less than .0001 for Wilcoxon signed rank tests comparing the two datasets and bootstrap-estimated 95% confidence intervals of [.9364,.9447] that represent the estimated range of possible values for the average dice coefficient obtained by repeating the test. To increase clinical trustworthiness, we used Grad-CAM explainability along with a pointing game based evaluation methodology to evaluate the CNN baseline model on the validation set. The results showed that we had an accuracy rate of 98.27% on the validation set of 519 images and confirmed that the model classified regions containing skin lesions.

Read & Discuss → View Source →

Explore the Frontier of Global Academia