论文广场 - AcademicHub

01.

bioRxiv (Bioinfo) 2026-06-11 DOI: HASH:5dea537644017f9516a0d147dad369fb

DLDN-Bench: A Benchmark Framework for Deep Learning de Novo Peptide Sequencing in Proteomics

作者:

Schneider ↗Hartwig ↗Chadt ↗Lehr ↗Al-Hasani ↗Turewicz ↗

De novo peptide sequencing is an essential approach for analyzing mass spectrometry data because it enables the identification of novel peptides without relying on protein sequence databases. Recent advances in deep learning have substantially improved the performance of de novo sequencing methods, but the rapid emergence of new models has led to heterogeneous evaluation practices and limited comparability. To address this, we introduce DLDN-Bench, a benchmark framework including a set of benchmark datasets derived from human muscle biopsy mass spectrometry data retrieved from PRIDE and annotated through consensus across multiple widely used database search engines. Using these datasets, we systematically benchmark recent deep learning-based de novo sequencing tools alongside traditional approaches. Performance is assessed using established metrics, including precision and coverage relative to a pseudo-ground truth defined by cross-engine agreement. To demonstrate the utility of DLDN-Bench, we benchmark four recent deep learning models and make all results publicly available. This benchmark framework provides a standardized basis for comparing state-of-the-art methods and offers an extensible resource for evaluating future tools in de novo peptide sequencing.

阅读与讨论 → 访问原文 →

02.

arXiv (quant-ph) 2026-06-17 DOI: arXiv:2606.17264

Coupled-Mode Equations with Arbitrary Mode Combinations for Kinetic-Inductance Superconducting Traveling-Wave Parametric Devices: Theory and Experimental Validation

作者:

F. Patricio Mena ↗Camilo Espinoza ↗Ryan O. Berriel ↗Ricardo Finger ↗David J. Thoen ↗

arXiv:2606.17264v1 Announce Type: cross Abstract: The coupled-mode equations (CMEs) have proven very successful in describing parametric processes in nonlinear optics. More recently, the same formulation has been used to model microwave superconducting parametric amplifiers and frequency multipliers. However, when applied to the microwave regime, not all assumptions remain valid and losses play a more dramatic role. Here, we revisit the CMEs applied to traveling-wave superconducting amplifiers to include losses and provide a formulation that enables their systematic derivation for any combination of traveling waves. As examples, we discuss the impact of unwanted harmonics and intermodulation products on parametric amplification, as well as harmonic generation. We verify that, if not properly accounted for, device performance can deviate considerably from the ideal case. Furthermore, using a superconducting CPW-based artificial transmission line and combining an independent experimental determination of its nonlinear parameter $I'_*$ with simulations of its linear properties, we obtain a parameter-free validation of this formulation. The nonlinear parameter was determined to be $I'_* \approx 27$ mA which, surprisingly, scales with the theoretical depairing current and not with the much smaller critical current of the device. For the validation, we measured multiple-harmonic generation and found excellent agreement between theory and experiment. The fact that $I'_* \gg I_C$ has direct implications for device design.

阅读与讨论 → 访问原文 →

03.

medRxiv (Medicine) 2026-06-24 DOI: HASH:1da7d8a123616938049cc598c4046884

Evaluation of Corneal Subbasal Nerve Plexus Alterations in ARSACS and SPG7 by In Vivo Corneal Confocal Microscopy

作者:

Guleser ↗U. Y ↗Akkaya ↗Kesim ↗Cakmak ↗O. O ↗Karslioglu ↗M. Z ↗Basak ↗A. N ↗Ertan ↗Hasanreisoglu ↗…

Purpose: To investigate corneal subbasal nerve plexus alterations using in vivo corneal confocal microscopy (IVCM) in patients with Autosomal Recessive Spastic Ataxia of Charlevoix-Saguenay (ARSACS) and Spastic Paraplegia Type 7 (SPG7). Methods: This cross-sectional pilot study included eight ARSACS patients, five SPG7 patients, and twenty age- and sex-matched healthy controls. All participants underwent neurological and ophthalmological examination followed by central corneal imaging using IVCM. Quantitative corneal nerve parameters were analyzed with automated software, and correlations with clinical severity scales were assessed. Results: The mean age was 34.2 +/- 3.4 years in controls, 34.5 +/- 0.7 years in the ARSACS group, and 38.2 +/- 3.5 years in the SPG7 group. Corneal nerve branch density (CNBD) and corneal nerve total branch density (CTBD) were significantly lower in ARSACS and SPG7 patients compared with healthy controls. CNFD, CNFL, CNFA, CNFW, and CNFrD were lower in ARSACS and SPG7 patients compared with healthy controls; however, these differences did not reach statistical significance. No statistically significant differences in IVCM parameters were detected between ARSACS and SPG7 patients. Spearman correlation analysis did not show significant correlations between corneal nerve parameters and FARS, SARA, ADL scores, or disease duration. Conclusion: IVCM revealed reduced corneal nerve branching parameters in patients with ARSACS and SPG7. These findings indicate involvement of the corneal subbasal nerve plexus and support the potential role of corneal confocal microscopy as a non-invasive ocular imaging modality for evaluating peripheral neural alterations in hereditary spastic ataxias.

阅读与讨论 → 访问原文 →

04.

arXiv (CS.LG) 2026-06-16 DOI: arXiv:2510.01175

On the Benefits of Weight Normalization for Overparameterized Matrix Sensing

作者:

Yudong Wei ↗Liang Zhang ↗Bingcong Li ↗Niao He ↗

arXiv:2510.01175v2 Announce Type: replace Abstract: While normalization techniques are widely used in deep learning, their theoretical understanding remains relatively limited. In this work, we establish the benefits of (generalized) weight normalization (WN) applied to the overparameterized matrix sensing problem. We prove that WN with Riemannian optimization achieves linear convergence, yielding an exponential speedup over standard methods that do not use WN. Our analysis further demonstrates that both iteration and sample complexity improve polynomially as the level of overparameterization increases. To the best of our knowledge, this work provides the first characterization of how WN leverages overparameterization for faster convergence in matrix sensing.

阅读与讨论 → 访问原文 →

05.

arXiv (CS.CV) 2026-06-25 DOI: arXiv:2511.20651

RubricRL: Simple Generalizable Rewards for Text-to-Image Generation

作者:

Xuelu Feng ↗Yunsheng Li ↗Ziyu Wan ↗Zixuan Gao ↗Junsong Yuan ↗Dongdong Chen ↗Chunming Qiao ↗

Reinforcement learning (RL) has recently emerged as a promising approach for aligning text-to-image generative models with human preferences. A key challenge, however, lies in designing effective and interpretable rewards. Existing methods often rely on either composite metrics (e.g., CLIP, OCR, and realism scores) with fixed weights or a single scalar reward distilled from human preference models, which can limit interpretability and flexibility. We propose RubricRL, a simple and general framework for rubric-based reward design that offers greater interpretability, composability, and user control. Instead of using a black-box scalar signal, RubricRL dynamically constructs a structured rubric for each prompt–a decomposable checklist of fine-grained visual criteria such as object correctness, attribute accuracy, OCR fidelity, and realism–tailored to the input text. Each criterion is independently evaluated by a multimodal judge (e.g., o4-mini), and a prompt-adaptive weighting mechanism emphasizes the most relevant dimensions. This design not only produces interpretable and modular supervision signals for policy optimization (e.g., GRPO or PPO), but also enables users to directly adjust which aspects to reward or penalize. Experiments with an autoregressive text-to-image model demonstrate that RubricRL improves prompt faithfulness, visual detail, and generalizability, while offering a flexible and extensible foundation for interpretable RL alignment across text-to-image architectures.

阅读与讨论 → 访问原文 →

06.

arXiv (CS.CL) 2026-06-16 DOI: arXiv:2604.25853

G-Loss: Graph-Guided Fine-Tuning of Language Models

作者:

Aditya Sharma ↗Vinti Agarwal ↗Rajesh Kumar ↗

Traditional loss functions, including cross-entropy, contrastive, triplet, and su pervised contrastive losses, used for fine-tuning pre-trained language models such as BERT, operate only within local neighborhoods and fail to account for the global semantic structure. We present G-Loss, a graph-guided loss function that incorporates semi-supervised label propagation to use structural relationships within the embedding manifold. G-Loss builds a document-similarity graph that captures global semantic relationships, thereby guiding the model to learn more discriminative and robust embeddings. We evaluate G-Loss on five benchmark datasets covering key downstream classification tasks: MR (sentiment analysis), R8 and R52 (topic categorization), Ohsumed (medical document classification), and 20NG (news categorization). In the majority of experimental setups, G-Loss converges faster and produces semantically coherent embedding spaces, resulting in higher classification accuracy than models fine-tuned with traditional loss functions.

阅读与讨论 → 访问原文 →

07.

arXiv (CS.AI) 2026-06-18 DOI: arXiv:2606.18733

SWE-Future: Forecast-Conditioned Data Synthesis for Future-Oriented Software Engineering Agents

作者:

Qiao Zhao ↗JianYing Qu ↗Jun Zhang ↗Yehua Yang ↗Hanwen Du ↗Zhongkai Sun ↗

arXiv:2606.18733v1 Announce Type: cross Abstract: Realistic coding-agent benchmarks often replay public GitHub issues and pull requests, making them vulnerable to overlap with model pretraining, fine-tuning, synthetic-data generation, or benchmark-driven model selection. Fully synthetic tasks avoid direct historical replay, but can drift away from real repository needs. We propose SWE-Future, a forecast-conditioned data synthesis method for future-oriented coding tasks. Given a forecast snapshot at time $T_0$, the method uses only pre-$T_0$ repository evidence to forecast future feature implementation/enhancement, bugfix, and refactor task families. We first validate this forecasting step retrospectively: after forecasts are fixed, later pull requests are used only to measure whether the predicted task families match future repository work. In an 80-repository study, the forecaster achieves 58.1\% future-work relevance under the main semantic matching metric. We then use validated forecast families as conditioning signals to synthesize a 200-task coding-agent dataset across 61 repositories from a task-generation snapshot, rather than replaying the later pull requests used for validation. SWE-Future shows that repository-evolution forecasts can guide realistic, future-oriented coding-task synthesis while reducing direct dependence on historical pull-request replay.

阅读与讨论 → 访问原文 →

08.

arXiv (CS.CL) 2026-06-25 DOI: arXiv:2504.09910

Learning to Erase Private Knowledge from Multi-Documents for Retrieval-Augmented Large Language Models

作者:

Yujing Wang ↗Jinwen Chen ↗Hainan Zhang ↗Liang Pang ↗Yongxin Tong ↗Binghui Guo ↗Hongwei Zheng ↗Zhiming Zheng ↗

Retrieval-Augmented Generation (RAG) is a promising technique for applying LLMs to proprietary domains. However, retrieved documents may contain sensitive knowledge, posing risks of privacy leakage in generative results. Thus, effectively erasing private information from retrieved documents is a key challenge for RAG. Unlike traditional text anonymization, RAG should consider: (1) the inherent multi-document reasoning may face de-anonymization attacks; (2) private knowledge varies by scenarios, so users should be allowed to customize which information to erase; (3) preserving sufficient publicly available knowledge for generation tasks. This paper introduces the privacy erasure task for RAG and proposes Eraser4RAG, a private knowledge eraser which effectively removes user-defined private knowledge from documents while preserving sufficient public knowledge for generation. Specifically, we first construct a global knowledge graph to identify potential knowledge across documents, aiming to defend against de-anonymization attacks. Then we randomly split it into private and public sub-graphs, and fine-tune Flan-T5 to rewrite the retrieved documents excluding private triples. Finally, PPO algorithm optimizes the rewriting model to minimize private triples and maximize public triples retention. Experiments on four QA datasets demonstrate that Eraser4RAG achieves superior erase performance than GPT-4o.

阅读与讨论 → 访问原文 →

09.

arXiv (CS.CL) 2026-06-19 DOI: arXiv:2606.19815

Clusters are All You Need: Pre-Training the Tsetlin Machine with Semantic Clusters from Language Models for Interpretability

作者:

Jiechao Gao ↗Rohan Kumar Yadav ↗Yuangang Li ↗Yuandong Pan ↗Jie Wang ↗Ying Liu ↗Michael Lepech ↗

Pre-trained language models such as BERT achieve strong text classification performance but lack transparency, limiting their use in high-stakes settings. The Tsetlin Machine (TM) offers fully interpretable, clause-based reasoning but captures little semantic information, and prior attempts to bridge the two rely on static word embeddings that miss contextual meaning. We propose a semantic pre-training framework that transfers knowledge from a pre-trained language model into a TM without using embeddings. Text samples are grouped into semantically coherent clusters with K-means or Top2Vec, and the resulting cluster-sample pairs pre-train a non-negated TM with enhanced Type I feedback. The TM thereby learns interpretable semantic keywords that are fine-tuned on downstream tasks. Across five datasets, our method substantially outperforms vanilla and embedding-based TMs and reaches performance competitive with BERT while remaining interpretable.

阅读与讨论 → 访问原文 →

10.

Nature (Science) 2026-06-24 DOI: HASH:92f530f86cd5c8206b5e7968a24e7612

The mutational landscape of STING-induced immunity

作者:

Bing Zhang ↗

Stimulator of interferon genes (STING) is an evolutionary conserved immune signalling protein with key roles in host defence, cancer, senescence and inflammation1–3. Downstream of STING, type I interferon, inflammatory cytokine signalling and non-canonical autophagy are governed by a multilayered mechanism integrating ligand-induced structural transitions, protein–protein interactions and coordinated intracellular trafficking4–13. Despite its central role in immunity and relevance as therapeutic target14, the sequence elements that govern STING (in)activation in cells remain incompletely understood. Here we developed a massively parallel assay to systematically chart the sequence-function landscape of STING. Profiling thousands of single amino-acid variants, we identified structural and functional determinants that shape the immunostimulatory capacity of STING and its ability to translate ligand recognition into distinct signalling outputs. Cryogenic-electron microscopy structures of select STING hyperactive variants revealed new regulatory principles dictating conformational transition from inactive to signalling-competent states of STING. Mutational effects are widespread across the functional landscape and can sensitize STING towards the natural ligand 2′3′-cGAMP15–18 or decouple interferon induction from non-canonical autophagy, demonstrating a diversity of possible responses that can be accessed through single point substitutions. Finally, our data showed the clinical and evolutionary relevance of naturally occurring STING protein variants. Collectively, these findings define molecular principles that tune STING activity and chart the landscape of its functional potential across immune contexts. A massively parallel assay systematically charts the sequence-function landscape of the STING signalling protein, and the findings define molecular principles that tune STING activity and show its functional potential across immune contexts.

阅读与讨论 → 访问原文 →

11.

arXiv (CS.LG) 2026-06-16 DOI: arXiv:2406.06855

Design and Scheduling of an AI-based Queueing System

作者:

Jiung Lee ↗Hongseok Namkoong ↗Yibo Zeng ↗

arXiv:2406.06855v3 Announce Type: replace-cross Abstract: To leverage prediction models to make optimal scheduling decisions in service systems, we must understand how predictive errors impact congestion due to externalities on the delay of other jobs. Motivated by applications where prediction models interact with human servers (e.g., content moderation), we consider a large queueing system comprising of many single server queues where the class of a job is estimated using a prediction model. By characterizing the impact of mispredictions on congestion cost in heavy traffic, we design an index-based policy that incorporates the predicted class information in a near-optimal manner. Our theoretical results guide the design of predictive models by providing a simple model selection procedure with downstream queueing performance as a central concern, and offer novel insights on how to design queueing systems with AI-based triage. We illustrate our framework on a content moderation task based on real online comments, where we construct toxicity classifiers by finetuning large language models.

阅读与讨论 → 访问原文 →

12.

medRxiv (Medicine) 2026-06-24 DOI: HASH:5fafad4e92788ff52f63cf1a4c7cfff8

CerViX-Net: A Multi-Branch Fusion of Vision Transformer and Convolutional Neural Networks for Cervical Cancer Detection using Cytology Images

作者:

De ↗

Cervical cancer represents a pressing global health challenge, emphasizing the critical need for accurate and timely diagnostic methods to facilitate effective treatment and improve survival rates. In response to this challenge, the study presents CerViX-Net, an innovative classification framework designed to advance cervical cancer detection through enhanced computational efficiency and diagnostic accuracy. The development of CerViX-Net is motivated by the limitations of traditional diagnostic models, particularly in handling the computational and memory demands of large-scale data, while ensuring precise feature extraction and classification. CerViX-Net employs a hybrid deep learning architecture that combines the capabilities of ResNet50, EfficientNet-B0, and a Modified Vision Transformer (ViT) module. The ResNet50 branch extracts hierarchical features through stacked convolutional and identity blocks. In another path, the modified ViT module transforms image patches via linear projection, augments them with positional and class embeddings, and processes them using Parallel Transformer Encoder layers to model contextual relationships. Concurrently, EfficientNet-B0 utilizes MBConv blocks to extract multi-scale representations. The feature outputs from all three branches are integrated and passed through a classification head consisting of dropout layers and dense layers to ensure robust and accurate predictions. The proposed framework is rigorously evaluated on the Mendeley LBC dataset, achieving exceptional performance metrics with an accuracy of 99.69%, precision of 99.28%, recall of 99.48%, and an F1-score of 99.52%. The robustness of CerViX-Net is further validated on the SIPaKMeD and Herlev Pap Smear datasets, where it demonstrates comparable excellence, underscoring its efficacy and adaptability across diverse cytology datasets. Statistical validation using Friedman's test further reinforces its superiority over competing methods.

阅读与讨论 → 访问原文 →

13.

arXiv (CS.AI) 2026-06-17 DOI: arXiv:2605.26195

CyberEvolver: Structured Self-Evolution for Cybersecurity Agents On the Fly

作者:

Yihe Fan ↗Changyi Li ↗Lichen Xu ↗Xudong Pan ↗Jiarun Dai ↗Hong Geng ↗Min Yang ↗

arXiv:2605.26195v2 Announce Type: replace-cross Abstract: LLM-based agents are increasingly used for cybersecurity tasks, but most existing systems rely on fixed, human-designed scaffolds that struggle to adapt across diverse targets and failure modes. We introduce \textsc{CyberEvolver}, a self-evolving cybersecurity agent framework that iteratively revises its own scaffold based on experience from failed execution attempts. Self-evolution in cybersecurity is challenging because the space of possible scaffold changes is largely unstructured, execution feedback is sparse and often obscured by the environment, and low-diversity updates can cause errors to compound over repeated iterations. \textsc{CyberEvolver} addresses these challenges with a four-layer evolvable agent architecture that decomposes scaffold optimization into structured components, a trace-to-diagnosis mechanism that converts noisy execution logs into actionable revision signals, and a population-based beam search strategy that preserves diverse agent variants during evolution. We evaluate \textsc{CyberEvolver} on CTF challenges, vulnerability exploitation, and penetration-testing tasks using four open-source LLMs. Across these settings, \textsc{CyberEvolver} improves the seed agent's success rate by $13.6$\,\% on average, and outperforms six human-designed cybersecurity agents as well as two self-improvement methods adapted from other domains. These results suggest that scaffold self-evolution is a promising direction for building adaptive LLM agents for security testing.

阅读与讨论 → 访问原文 →

14.

arXiv (CS.CV) 2026-06-25 DOI: arXiv:2507.16863

Position: Reasoning After Perception Means Reasoning Without Vision

作者:

Hongcheng Gao ↗Zihao Huang ↗Jingyi Tang ↗Lin Xu ↗Xinhao Li ↗Haoyang Li ↗Yue Liu ↗Minhua Lin ↗Xinlong Yang ↗Taihang Hu ↗Ge Wu ↗Balong Bi ↗…

A common belief in multimodal research is that the perceptual weaknesses of vision–language models can be compensated by stronger language reasoning (e.g., chain-of-thought, in-context learning, or external tools). We challenge this assumption. We argue that for a broad class of visual tasks hard to specify in language, failures stem from a structural fatality where the temporal decision of when to reason strictly dictates the spatial constraint of where reasoning takes place. When visual reasoning is deferred to language generation, current architectures do not merely delay computation; they displace it from the continuous visual representation to a discrete textual space. Consequently, the sequential ``Perception-then-Reasoning'' paradigm degenerates perception into a passive, one-off feature encoding process, rendering it functionally equivalent to ``Reasoning-in-Text-Space'', where task-critical spatial signals are collapsed before reasoning begins. We substantiate this claim with the Turing Eye Test (TET): tasks that must be resolved in visual space and are hard to verbalize; results show text-only reasoning cannot remedy these perceptual failures. Our findings suggest rethinking the architectural divide: shifting from reasoning about perception to reasoning within perception. This facilitates actively reasoning-driven perception that operates directly on pixel-level visual representations, rather than within a collapsed textual space.

阅读与讨论 → 访问原文 →

15.

arXiv (quant-ph) 2026-06-24 DOI: arXiv:2606.24572

The Vector and Canonical Components of the Momentum Operator in 3D Euclidean Space Spanned by General Curvilinear Coordinates

作者:

M. S. Shikakhwa ↗

arXiv:2606.24572v1 Announce Type: new Abstract: We construct the Hermitian vector and canonical components of the momentum operator in 3D Euclidean space spanned by general curvilinear coordinates (GCC's) using a simple, natural and unified approach based on identifying the momentum operator in any coordinate system as mass times the velocity operator. When this latter is calculated by applying the Heisenberg equation of motion, it returns ($-i\hbar$ times) the gradient operator plus an additional zero-valued sum, which when distributed among the components of the gradient, it makes each the Hermitian vector component of the momentum operator in GCC's. The canonical components follow immediately upon symmetrizing each of these vector components in the corresponding base vector. For accessability by wider audiences, we first develop the formalism for the simple polar coordinates and then we develop the case for GCC's.

阅读与讨论 → 访问原文 →

16.

bioRxiv (Bioinfo) 2026-06-15 DOI: HASH:c331f6fb311a5ab72548d4ceaa8e6d66

AliceDB database and pipeline for identification of natural protein variants based on mass spectrometry measurement data

作者:

Thiel ↗Rozycka ↗Puchalski ↗Oldziej ↗

The natural variation that distinguishes living organisms within a single species is currently being studied intensively, primarily at the genetic level. Unfortunately, studies of natural variants at the level of protein gene products are not very common, mainly due to the lack of appropriate databases and bioinformatics tools. The main research technique used to study proteomes/peptidomes is mass spectrometry (MS). A classic method for interpreting raw mass spectrometry data in proteomic/peptidomic studies involves the use of databases containing representative (canonical) sequences that define the proteome of the organism under study. In this paper, we present the AliceDB database, which contains information on over 7 million natural variants of protein sequences described in the scientific literature for Homo sapiens. The data contained in the AliceDB database can be utilized using widely available and commonly used software for interpreting proteomic data. Test results regarding the use of the AliceDB database for the interpretation of proteomic data indicate that accounting for the presence of natural variants increases both the number and quality of identified proteins. Furthermore, it is easy to identify protein sequence variants that may, for example, be of significance in medicine.

阅读与讨论 → 访问原文 →

17.

arXiv (CS.CL) 2026-06-17 DOI: arXiv:2606.17688

LLMs Infer Cultural Context but Fail to Apply It When Responding

作者:

Yisong Miao ↗Jian Zhu ↗Vered Shwartz ↗

Recent work has shown that LLMs overrepresent dominant cultures, particularly Western ones, while marginalizing others. We investigate whether this affects models' ability to generate culturally adapted responses by evaluating their use of local measurement units based on the user's perceived cultural background. We introduce Cultural and Pragmatic Response Inference (CAPRI), a dataset of conversations with varying levels of cultural cues. Experiments with state-of-the-art LLMs show that models can infer cultural background and recall relevant conventions, but often fail to utilize the information to adapt their answers to the relevant cultural conventions, unless explicitly prompted to perform the tasks sequentially. We further evaluate adaptation to the interpretation of time and quantity expressions, two subjective language grounding dimensions that are affected by culture. We find that models increasingly adapt their answers as cultural cues accumulate, but their priors are not culture-neutral, sometimes aligning with the model's country of origin. Overall, CAPRI provides a resource for future research aimed at narrowing the gap between cultural knowledge and culturally adaptive language generation.

阅读与讨论 → 访问原文 →

18.

arXiv (CS.CL) 2026-06-11 DOI: arXiv:2606.11613

Factions Within, Uncertain Across: Within-Document Reader Sub-Groups in Social Highlighting

作者:

Kazuki Nakayashiki ↗Keisuke Watanabe ↗

When many people highlight the same document, is the crowd a single consensus, or is it internally structured into reader sub-groups that mark different things – and is that structure a stable property of a reader or of the document? Building on prior work showing an individual's within-document highlighting signal is a whisper while individuality lives in selection, we ask the group-level question on a co-readership platform using a margin-preserving curveball null. Experiment 1: within a document, readers form strong sub-groups – pairs agree far beyond what shared salience, mark density, and sentence popularity predict (nearest-neighbour agreement z=+6.3, significant in 88% of documents). Under an eight-block region-preserving null, shared engagement with the same coarse regions of the document accounts for about 40% of this excess; the majority survives as finer reader-specific agreement (z=+3.6, 77% significant). So the within-document crowd is, in a descriptive sense, factional. Experiment 2: is that grouping a stable reader trait? Here we are honest about power. The cross-document split-half reproducibility of a pair's agreement is near zero pooled (+0.078 and 0.000 in two separately drawn samples), and a power calibration shows the test is informative only for pairs that co-read many documents. In the only informative high-overlap subset (k>=4), point estimates are positive but small-sample, imprecise across the separately drawn samples, never significant, and attenuate under the region-preserving null. We therefore leave cross-document stability unresolved: the data is consistent with anything from situational grouping to a weak-to-moderate stable reader trait. The crowd is factional within a document; whether its factions follow the reader across documents is, honestly, beyond our reach.

阅读与讨论 → 访问原文 →

19.

Nature (Science) 2026-06-10 DOI: HASH:f8a664017587ebba32576c5bffcc957a

Artificial intelligence shines a light on hidden global migration flows

作者:

Francisco Lara-García ↗

Training a neural network to collate data from several sources provides a high-resolution view of how people are moving around the world. Training a neural network to collate data from several sources provides a high-resolution view of how people are moving around the world.

阅读与讨论 → 访问原文 →

20.

arXiv (CS.CV) 2026-06-24 DOI: arXiv:2606.24144

Geometry-Aware Style Transfer in 3D Gaussian Splatting

作者:

Min Hyeok Bang ↗Jun Hyeong Kim ↗Seung-Wook Kim ↗Se-Ho Lee ↗

In this paper, we present a novel geometry-aware style transfer framework for 3D Gaussian splatting (3DGS) that simultaneously transfers appearance attributes and geometric structures. Unlike prior works that primarily focus on color-based stylization and often overlook structural adaptation, our method explicitly incorporates geometry adaptation through a decoupled optimization scheme that alternately updates color and geometry parameters. This strategy alleviates potential interference between color and geometry updates, leading to stable and consistent scene-level geometry transformation. The decoupled optimization is enabled by the proposed geometry-aware contrastive feature matching (GCFM). GCFM integrates RGB, depth, and edge cues into a contrastive objective and is employed in both optimization phases to effectively transfer structural characteristics from style images to Gaussian primitives. Extensive experiments show that our approach achieves superior performance in both qualitative fidelity and quantitative metrics, significantly outperforming existing 3DGS-based stylization methods. Our code is available at \href{https://github.com/oweixx/gast}{https://github.com/oweixx/gast}.

阅读与讨论 → 访问原文 →

21.

medRxiv (Medicine) 2026-06-22 DOI: HASH:79bfa8e43bebe84179d3aa5324c087d2

Symptom-based phenotype discovery in motor neuron disease using natural language processing of electronic health records

作者:

Abdulle ↗Dinu ↗Wu ↗Kim ↗Budhdeo ↗Yao ↗Tomlinson ↗Al-Chalabi ↗Dobson ↗Iacoangeli ↗

Background: Motor neuron disease (MND) is a fatal neurodegenerative condition with significant clinical heterogeneity that is incompletely captured by existing phenotype classifications based on onset site. Electronic health records (EHRs) contain detailed symptom documentation in clinical narratives that may enable data-driven discovery of clinically meaningful patient subgroups. Methods: We developed a natural language processing (NLP) pipeline using MedCAT to extract symptoms from clinical notes of 2,361 people with a confirmed diagnosis of MND at a tertiary neurology center. MND cohort confirmation used three complementary methods: clinic attendance records, text-based diagnosis detection, and NLP extraction with negation detection. Extracted symptoms were filtered to Unified Medical Language System semantic type T184 (Sign or Symptom) with removal of negated concepts. Patients were clustered using latent class analysis on binary symptom profiles. Survival differences were assessed using Kaplan-Meier analysis, log-rank tests, and Cox proportional hazards regression. Results: From the first clinical notes, we identified four clusters of symptoms among 872 patients and 76 symptoms: Motor-Bulbar (n=373), Motor-Tremor (n=154), Sensory-Pain (n=222), and Motor-Respiratory (n=123). When extended to all clinical notes (n=2,065; 184 symptoms), these reorganized into three clusters: Autonomic-Respiratory (n=472), Nocturnal-Respiratory (n=338), and Classic Motor (n=1,255). Survival differences were significant across all clusters in both the first notes and all notes analyses (log-rank p < 0.001). Conclusions: NLP-based symptom extraction from EHRs identifies clinically meaningful MND subgroups that extend beyond traditional onset-site classifications. Autonomic-respiratory symptom burden is associated with poorer survival while a newly identified Sensory-Pain subtype with a better prognosis. These data-driven phenotypes may improve prognostication and inform targeted supportive care.

阅读与讨论 → 访问原文 →

22.

arXiv (quant-ph) 2026-06-11 DOI: arXiv:2606.12338

Entanglement generation between field modes mediated by a fluctuating conducting wall

作者:

Luca Giovanni Cammarata ↗Tommaso Fazio ↗Roberto Passante ↗Lucia Rizzuto ↗

arXiv:2606.12338v1 Announce Type: cross Abstract: We consider a movable conducting plate of finite mass, between two fixed ones, whose mechanical degrees of freedom are treated quantum-mechanically and bound to its equilibrium position by a harmonic potential. The movable wall is thus subjected to quantum fluctuations of its position. This creates a system of two sub-cavities separated by the movable fluctuating plate, and two massless one-dimensional scalar fields, one in each sub-cavity. This system is described by an appropriate generalization of the Law Hamiltonian. The presence of the movable wall yields an effective plate-fields interaction, as well as an effective interaction between the field modes. We obtain, at the second order in perturbation theory, the ground state of the interacting system and the reduced density operator of the fields in each sub-cavity by tracing out the wall's degrees of freedom. We calculate the entanglement between two field modes, one in each cavity, by evaluating analytically the negativity; we then evaluate numerically also the total multimode negativity. Our results show that in both cases the fields in the two sub-cavities are entangled, in contrast to the case in which the wall is fixed in space. We discuss the amount of the field entanglement present as a function of relevant physical parameters of the system such as the mass and oscillation frequency of the movable wall, its distance from the fixed walls and the frequencies of the field modes considered.

阅读与讨论 → 访问原文 →

23.

arXiv (CS.LG) 2026-06-12 DOI: arXiv:2606.12709

Smarter Saboteurs, Better Fixers: Scaling & Security in Linear Multi-Agent Workflows

作者:

Timothy McAllister ↗Sina Abdidizaji ↗Ivan Garibay ↗Ozlem Ozmen Garibay ↗

arXiv:2606.12709v1 Announce Type: cross Abstract: As LLM-based multi-agent systems (MAS) are deployed in the wild, the resilience of their collaboration structures against adversarial compromise becomes a critical safety concern. Attackers may leverage prompt-injection or jailbreaking to sabotage individual agents within MAS workflows, but the interaction between model scaling and system-level resilience remains poorly understood. This paper investigates how model scale affects the security of linear multi-agent workflows. Our experiments across scales of two open-weight model families on the HumanEval benchmark reveal a compliance-correction symmetry: larger models are far more likely to faithfully execute malicious instructions, with the control-to-malicious performance drop reaching 53.7pp at 27B in uncorrected pipelines. However, appending a lightweight terminal Fixer stage collapses this to 0.6pp and restores statistical parity with control-level performance, demonstrating that strictly linear collaboration structures can be viable and resilient to adversaries at this scale, and suggesting that the brittleness previously attributed to linear topology may stem from a lack of correction.

阅读与讨论 → 访问原文 →

24.

arXiv (CS.LG) 2026-06-19 DOI: arXiv:2606.19412

Spectral Retrieval-Augmented Time-Series Forecasting

作者:

Huu Hiep Nguyen ↗Minh Hoang Nguyen ↗Dung Nguyen ↗Hung Le ↗

arXiv:2606.19412v1 Announce Type: new Abstract: Time series forecasting leverages historical patterns to predict future values, but traditional methods face challenges when dealing with complex, non-stationary patterns that are difficult to memorize during training. Retrieval-augmented approaches have emerged as promising solutions by retrieving similar historical patterns to enhance predictions. However, existing retrieval methods suffer from two fundamental limitations: spectral blindness, which overlooks critical frequency-domain characteristics that capture underlying periodic structures, and temporal recency, which treats all historical data equally without emphasizing recent, more relevant patterns. In this paper, we propose SpecReTF, a novel retrieval method that addresses these issues by converting time series into windowed frequency representations, measuring similarity with a combined metric that captures both amplitude and phase information. To balance recency and historical context, we apply an exponential moving average weighting scheme that emphasizes recent windows. Extensive experiments on benchmark datasets demonstrate that SpecReTF outperforms time-domain retrieval methods, achieving superior forecasting accuracy across diverse, non-stationary time series.

阅读与讨论 → 访问原文 →

25.

arXiv (math.PR) 2026-06-12 DOI: arXiv:2512.13829

Conditional means, vector pricings, amenability and fixed points in cones

作者:

Nicolas Monod ↗

arXiv:2512.13829v4 Announce Type: replace Abstract: We develop a generalization of conditional probability for arbitrary ordered vector spaces. A related problem is that of assigning a numerical value to one vector relative to another. We characterize the groups for which these generalized probabilities can be stationary, respectively invariant. Our results deviate from the setting of classical probability and lead to a new criterion for amenability and for fixed points in cones.

阅读与讨论 → 访问原文 →

探索全球前沿学术脉络