Paper Plaza - AcademicHub

01.

Nature (Science) 2026-06-11 DOI: HASH:f22fb8081d205679b6fb7f676903179e

My diverse academic background is affecting my PhD studies — what do I do?

Authors:

A PhD student in China who has changed his focus twice since obtaining his undergraduate degree is struggling to keep up. A PhD student in China who has changed his focus twice since obtaining his undergraduate degree is struggling to keep up.

Read & Discuss → View Source →

02.

arXiv (CS.LG) 2026-06-18 DOI: arXiv:2606.18420

Measurement noise limits the advantage of nonlinear models over linear models in biomedical prediction

Authors:

Marc-Andre Schulz ↗Kerstin Ritter ↗

arXiv:2606.18420v1 Announce Type: new Abstract: On biomedical tabular data, flexible models such as deep networks, gradient-boosted trees, and kernel methods are repeatedly matched or beaten by linear and logistic regression given the same features. The usual reaction is to treat this as a model-side shortfall, to be fixed with more data, a better architecture, or tuning, on the assumption that the nonlinear structure is there and the model has failed to capture it. We argue that these fixes cannot help when the binding limit is the measurement rather than the model, as it frequently is in biomedicine. Additive noise blurs the population-optimal predictor, and because blurring removes a function's fine, rapidly varying detail before its broad shape, it erases nonlinear structure faster than linear structure. A degree-$k$ interaction is attenuated by the $k$-th power of feature reliability, while the linear part is attenuated only once. At the reliabilities typical of biomedical measurement, the nonlinear advantage can vanish even when the underlying biology is strongly nonlinear, and what the noise removes cannot be recovered by a larger cohort or a more flexible model, only by better measurement. The nonlinearity is hidden, not absent, and a tie between linear and flexible models is not by itself a verdict on the biology. These pieces are classical, drawn from measurement-error statistics, psychometrics, and Gaussian analysis, and we assemble them into an exact excess-risk identity. Measurement reliability is one of three conditions, alongside sample size and feature representation, that must align for a flexible model to help, and together they leave only a narrow window that most biomedical tasks fall outside. Across 140 UK Biobank tasks, the gap between flexible and linear models, where it exists, carries the predicted noise signature, and the three conditions can be separated by intervention but not by a benchmark alone.

Read & Discuss → View Source →

03.

arXiv (CS.CL) 2026-06-17 DOI: arXiv:2606.00024

ART: Attention Run-time Termination for Efficient Large Language Model Decoding

Authors:

Chen Qiu ↗Guozhong Li ↗Cristian McGee ↗Aritra Dutta ↗Panos Kalnis ↗

Long-context decoding in Large Language Models (LLMs) is constrained by the cost of accessing and processing the Key-Value (KV) cache. Despite evidence that attention outputs depend jointly on keys and values, most existing KV management methods rely on key-only pruning, since incorporating values incurs prohibitive overhead. In this paper, we propose Attention Run-time Termination (ART), a lightweight run-time mechanism that tracks accumulated attention outputs during kernel execution and terminates subsequent KV block accesses once further contributions become negligible. Rather than replacing KV selection, ART dynamically terminates redundant KV traversal on top of existing dense or sparse attention policies. We introduce a stability-based criterion that monitors both magnitude and directional changes of intermediate attention outputs and provideds a theoretical characterization of the resulting truncation error. Experiments on the LongBench and RULER Needle-in-a-Haystack tasks show that ART increases the generation throughput of existing KV-cache methods by up to 20%, without compromising the result quality.

Read & Discuss → View Source →

04.

arXiv (CS.LG) 2026-06-18 DOI: arXiv:2606.18527

Toward Simultaneously Optimal Regret in U-Calibration

Authors:

Rafael Frongillo ↗Haipeng Luo ↗Nishant A. Mehta ↗Jon Schneider ↗

arXiv:2606.18527v1 Announce Type: cross Abstract: U-calibration studies online forecasting algorithms whose predictions can be consumed by any unknown downstream agent, guaranteeing sublinear regret simultaneously for all proper loss functions. Existing U-calibration algorithms achieve worst-case optimal $O(\sqrt{T})$ regret for every bounded proper loss, but they fail to adapt to easier losses: as we show, even for smooth losses such as squared loss, they incur $\Omega(\sqrt{T})$ regret instead of the optimal $O(\log T)$ regret. In this work, we show that this limitation is not inherent. Specifically, we design a single forecast algorithm that simultaneously achieves $\tilde O(\sqrt{T})$ regret for every bounded proper loss and $O(\log T)$ regret for every bounded smooth proper loss. More generally, our algorithm also attains logarithmic regret for losses that are smooth relative to the log-barrier, which include several non-Lipschitz examples. Our approach is based on a novel variant of Follow-the-Perturbed-Leader (FTPL) in which perturbations are applied directly in the prediction space using self-concordant noise. The resulting analysis also departs substantially from prior FTPL analyses due to the complex nature of this noise and may be of independent interest.

Read & Discuss → View Source →

05.

medRxiv (Medicine) 2026-06-24 DOI: HASH:55a1a083101466764c6b763543394ff0

Beyond Nodal Status: Interactions Between Molecular Subtype, Tumor Burden, and Survival in 12,225 Patients with Breast Cancer

Authors:

Akrami ↗Tavakolian ↗Arianpour ↗Moosazadeh ↗Rajabi ↗A. H ↗Keumarsi ↗Ghoddusi Johari ↗Zangouri ↗Talei ↗

Background Lymph node status and molecular subtype are among the most established prognostic factors in breast cancer. However, the extent to which their prognostic effects vary across different tumor size categories and clinical subgroups remains incompletely understood. We investigated the interplay between nodal status, molecular subtype, and tumor size in a large real world breast cancer cohort and developed a prognostic nomogram for individualized survival prediction. Methods A total of 12,225 women with invasive breast cancer from the Shiraz Breast Cancer Registry were analyzed. Patients were stratified according to tumor size, lymph node status, and molecular subtype. Overall survival (OS) and disease free survival (DFS) were evaluated using Kaplan Meier analyses and subgroup comparisons. Logistic regression was performed to identify predictors of lymph node involvement, while Cox regression was used to determine independent prognostic factors. A nomogram was subsequently developed and internally validated for prediction of 3-year and 5-year OS. Results Of 12,225 patients, 41.7% had lymph node positive disease. Across nearly all tumor size categories and molecular subtypes, nodal involvement was associated with significantly worse OS and DFS. Notably, the survival disadvantage associated with nodal positivity was more pronounced among patients with larger tumors and among those with HER2 positive and triple negative breast cancer (TNBC). Although TNBC demonstrated the lowest rate of lymph node involvement among molecular subtypes (adjusted OR 0.54, 95% CI 0.46-0.63), it appeared to show one of the largest survival gaps between node positive and node negative disease. In the overall cohort, survival outcomes generally ranked from best to worst as Luminal A, Luminal B, HER2 positive, and TNBC. However, survival differences among molecular subtypes were not consistently observed across all tumor size and nodal status subgroups. When significant differences were present, Luminal A and Luminal B tumors consistently showed superior outcomes compared with HER2 positive and TNBC tumors. Multivariable analysis identified lymph node status, tumor size, molecular subtype, lymphovascular invasion, tumor necrosis, type of surgery, radiotherapy, hormone therapy, and adjuvant chemotherapy as independent prognostic factors. A nomogram integrating clinicopathological and treatment variables demonstrated good predictive performance, with time dependent AUCs of 0.749 and 0.751 for 3 year and 5 year OS, respectively, and showed good calibration. Conclusions The prognostic impact of lymph node status is not uniform across breast cancer subgroups and appears particularly pronounced in larger tumors and biologically aggressive subtypes. Despite a lower likelihood of nodal involvement, TNBC showed substantial outcome deterioration when nodal metastasis was present. These findings highlight the importance of jointly considering nodal status, molecular subtype, and tumor burden in prognostic assessment.

Read & Discuss → View Source →

06.

arXiv (CS.AI) 2026-06-18 DOI: arXiv:2606.18272

Mitigating Anchoring Bias in LLM-Based Agents for Energy-Efficient 6G Autonomous Networks

Authors:

Hatim Chergui ↗Claudia Carballo Gonz\'alez ↗Farhad Rezazadeh ↗Merouane Debbah ↗

arXiv:2606.18272v1 Announce Type: cross Abstract: This paper presents an autonomous agentic resource negotiation framework designed to enable zero-touch network slicing in 6G architectures using Large Language Model (LLM) agents. While LLMs offer powerful reasoning capabilities, we demonstrate that such agents inherently suffer from anchoring bias, rigidly adhering to initial heuristic proposals and causing severe network over-provisioning. To systematically mitigate this cognitive bias, we propose a novel randomized anchoring strategy modeled via a Truncated 3-Parameter Weibull distribution. This mathematically bounded approach seamlessly integrates with burst-aware Digital Twins (DTs) employing Conditional Value at Risk (CVaR) to rigorously guarantee strict Service Level Agreement (SLA) tail-latencies. To validate our methodology, we introduce and prove the Bimodal Constraint-Avoidance Utility Theorem, demonstrating that while feasible negotiations follow classical convex bounds, highly constrained scenarios undergo a phase transition governed by an inverse rational decay envelope. Empirical results generated using a locally hosted 1B-parameter model (\texttt{otel-llm-1b-it}) confirm these dual-regime bounds. Our cognitive de-biasing successfully dismantles rigid negotiation patterns, forcing agents into active exploration to safely ride SLA boundaries and boost system energy savings up to 25\%. Crucially, the lightweight 1B LLM achieves sub-second inference latencies (0.95s mean), ensuring our multi-agent framework is compatible with the operational timescales of the O-RAN non-Real-Time RAN Intelligent Controller (non-RT RIC)\footnote{Our source code is available for non-commercial use at https://github.com/HatimChergui.

Read & Discuss → View Source →

07.

arXiv (CS.CL) 2026-06-15 DOI: arXiv:2410.15051

Automatic identification of diagnosis from hospital discharge letters via weakly supervised Natural Language Processing

Authors:

Vittorio Torri ↗Elisa Barbieri ↗Anna Cantarutti ↗Carlo Giaquinto ↗Francesca Ieva ↗

Identifying patient diagnoses from hospital discharge letters is essential for large-scale cohort selection and epidemiological research, but traditional supervised approaches require extensive manual annotation, which is often impractical for large textual datasets. We present a weakly supervised Natural Language Processing (NLP) pipeline for classifying Italian discharge letters without document-level manual annotation. The method extracts diagnosis-related sentences, generates semantic embeddings using a transformer model further pre-trained on Italian medical documents, and applies a two-level clustering procedure to derive weak labels that are then used to train a document-level classifier. The approach was evaluated in a case study on bronchiolitis using 33,176 discharge letters of children admitted to 44 emergency rooms or hospitals in the Veneto Region, Italy, between 2017 and 2020. The best weakly supervised model achieved an AUROC of 77.68% ($\pm4.30\%$), an AUPRC of 73.13% ($\pm4.93\%$), and an F1-score of 78.14% ($\pm4.89\%$) against manually annotated data. Performance surpassed unsupervised baselines and approached fully supervised models, while reducing the need for manual annotation by more than 1,500 hours for a dataset of this size. Similar model rankings were observed in a secondary validation on a smaller bronchitis dataset (3,188 discharge letters, 2020-2025), where the best weakly supervised model achieved an AUPRC of 76.72% ($\pm 5.02\%$). These results suggest the potential of weakly supervised NLP methods for scalable disease identification from clinical discharge letters.

Read & Discuss → View Source →

08.

arXiv (CS.CV) 2026-06-19 DOI: arXiv:2606.19495

LooseControlVideo: Directorial Video Control using Spatial Blocking

Authors:

Shariq Farooq Bhat ↗Niloy J. Mitra ↗Kalyan Sunkavalli ↗

Precise 3D spatial orchestration in text-to-video generation remains a significant challenge, particularly for multi-object scenes where semantic layout and temporal dynamics are often entangled. While existing depth-conditioned models achieve good structural fidelity, they necessitate dense, frame-accurate guidance that is labor-intensive to author for dynamic events involving deformable objects. We present LooseControlVideo, a framework that enables intuitive and expressive control by using sparse, oriented 3D boxes as a "blocking" proxy. This allows users to author high-level layout and trajectory while leveraging a video generative model to generate realistic occlusions, dynamics and interactions. We achieve this by fine-tuning a Wan 2.2 backbone on a video dataset annotated with DNOCS, a novel encoding for 3D size, orientation and depth-ordered occlusions. Furthermore, our method allows for localized refinement, such as adjusting a jump trajectory or adding an interaction, with minimal disruption to the global scene context. Extensive evaluations on the nuScenes, HO-3D, and BEHAVE benchmarks demonstrate that LooseControlVideo significantly outperforms existing 2D-box and flow-based baselines. Our findings indicate a 1.2x to 3x improvement in Trajectory Error; 2x improvement in Rigid Motion Consistency; and a 1.5x to 2x increase in Occlusion Accuracy over current state-of-the-art layout-conditioned models, demonstrating that oriented 3D primitives provide good geometric prior for complex, multi-agent video authoring.

Read & Discuss → View Source →

09.

arXiv (CS.LG) 2026-06-12 DOI: arXiv:2606.13614

Majority-of-Three is Optimal

Authors:

Divit Rawal ↗Nikita Zhivotovskiy ↗

arXiv:2606.13614v1 Announce Type: cross Abstract: We give a short proof that the majority vote of three independent consistent classifiers is an optimal learner in the realizable PAC setting. This proves optimality for the simplest voting scheme, while simplifying both the algorithmic structure and the probabilistic analysis of previous voting learners, including the algorithm of S. Hanneke and the analysis of bagging by K. Green Larsen.

Read & Discuss → View Source →

10.

arXiv (CS.LG) 2026-06-19 DOI: arXiv:2606.19569

On the QUEST for Uncertainty Quantification via Highest Density Regions

Authors:

Sam Goring ↗Tom Kuipers ↗Nicola Paoletti ↗David S. Watson ↗

arXiv:2606.19569v1 Announce Type: new Abstract: Uncertainty quantification (UQ) is essential for reliable decision-making in safety-critical applications in probabilistic machine learning. For regression problems, dominant scalar UQ approaches - notably, those based on proper scoring rules - measure uncertainty via pointwise predictive risk. This can lead to counterintuitive results when the target statistic is not the conditional expectation. We propose an alternative framework, in which uncertainty is characterised by the volume of the most probable subset of a distribution's support. QUEST (Quantifying Uncertainty via highest dEnSiTy regions) is a novel approach to UQ based on the concentration of Lebesgue measure at a distribution's peak(s), evaluated at one or more values of a robustness parameter $\alpha$. We establish connections between our measures and classical statistics from information theory and economics. We show that, unlike popular alternatives based on proper scoring rules, QUEST measures of epistemic and aleatoric uncertainty satisfy a set of axioms adapted from the UQ literature, including monotonicity under distributional spread and invariance to location shifts. Selective prediction benchmarks confirm that QUEST performs favourably against standard measures such as variance and differential entropy.

Read & Discuss → View Source →

11.

arXiv (CS.CL) 2026-06-11 DOI: arXiv:2606.11681

UR-BERT: Scaling Text Encoders for Massively Multilingual TTS Through Universal Romanization and Speech Token Prediction

Authors:

Sangmin Lee ↗Eekgyun Ahn ↗Woongjib Choi ↗Hong-Goo Kang ↗

We propose UR-BERT, a Romanized transcription-based text-to-speech (TTS) encoder for massively multilingual TTS systems. Conventional grapheme-to-phoneme (G2P)-based approaches are limited to around 100 languages due to the availability of reliable G2P resources. In contrast, UR-BERT scales to 495 languages by unifying diverse writing systems into a shared Romanization representation. To further enhance phonetic fidelity and text-speech alignment, we introduce a speech token prediction objective during training, which encourages the encoder to learn speech-aware phonetic representations in a data-efficient manner. Experiments show that TTS systems built on UR-BERT consistently outperform recent text encoder baselines across a wide range of languages and resource conditions, and demonstrate strong generalization to unseen languages.

Read & Discuss → View Source →

12.

arXiv (CS.LG) 2026-06-16 DOI: arXiv:2606.16411

Not all Jensen-Shannon Divergence Estimators are Equal

Authors:

Alba Garrido ↗Alejandro Almod\'ovar ↗Mar Elizo ↗Patricia A. Apell\'aniz ↗Santiago Zazo ↗Juan Parras ↗

arXiv:2606.16411v1 Announce Type: new Abstract: The Jensen-Shannon divergence is widely reported as a scalar measure of fidelity for synthetic tabular data. Yet, in practice, it is estimated from finite samples using protocols that are often underspecified. This creates a measurement problem. Although the population divergence is well defined, the empirical value depends on the estimator family, sampling protocol, calibration, dimensionality, and class balance. We show that different protocols can yield non-comparable values: marginal-based estimators ignore dependencies in the joint distribution and can severely underestimate divergence, while classifier-based estimators capture joint structure but exhibit strong estimator dependence. We systematically study this behavior across controlled settings with reference divergences and real-world synthetic tabular benchmarks. Our analysis reveals dependence blindness in marginal estimators, prior-shift bias under class imbalance, and estimator sensitivity in high dimensions. To address prior shift, we derive a closed-form posterior correction for classifier-based Jensen-Shannon estimation. Our results show that empirical Jensen-Shannon divergence values are inherently protocol-dependent, making explicit specification of the estimation procedure necessary for meaningful comparison. We provide practical guidelines and an open-source tool for estimator-aware Jensen-Shannon evaluation.

Read & Discuss → View Source →

13.

arXiv (CS.CV) 2026-06-19 DOI: arXiv:2606.19804

HypOProto: Hyperbolic Ordinal Prototypes for Left Ventricular Filling Pressure Classification

Authors:

Victoria Wu ↗Nima Hashemi ↗Hooman Vaseli ↗Christina Luong ↗Purang Abolmaesumi ↗Teresa S. M. Tsang ↗

Echocardiography (echo) is a widely used imaging modality for assessing cardiac function, with Left Ventricular Filling Pressure (LVFP) serving as a critical physiological marker for conditions such as heart failure. Standard LVFP classification into normal vs elevated categories relies on the Doppler-derived $E/e'$ ratio, which is operator-dependent and often unavailable in resource-limited settings, motivating methods that infer LVFP directly from B-mode echo. Existing deep learning approaches achieve high performance but remain largely black-box, limiting clinical interpretability. We propose HypOProto, a hyperbolic, ordinal prototype-based framework for interpretable LVFP classification using a frozen, explainable foundation model backbone. HypOProto arranges prototypes along the physiological $E/e'$ scale, placing borderline cases near the hyperboloid root where small angular differences separate similar cases, while normal and elevated cases occupy outward positions reflecting increasing diagnostic certainty. This hyperbolic geometry encodes clinically meaningful ordinal relationships and improves interpretability. We also introduce a novel Hyperbolic Prototype Angular Separation (HyperPAS) loss, enforcing inter-class prototype separation in hyperbolic space. HypOProto achieves SOTA performance while maintaining transparency, and highlights clinically relevant regions in visualizations. This work represents the first prototype-based framework for LVFP classification in echo. Our code can be found at https://github.com/DeepRCL/HypOProto.

Read & Discuss → View Source →

14.

arXiv (CS.AI) 2026-06-16 DOI: arXiv:2606.16190

Embedded Arena: Iterative Optimization via Hardware Feedback

Authors:

arXiv:2606.16190v1 Announce Type: cross Abstract: Embedded devices from wildlife monitoring stations to clinical wearables require local AI inference due to latency, communication, or privacy constraints. Optimizing models for heterogeneous microcontrollers (MCUs) requires simultaneously satisfying hard physical constraints on memory, power, and temperature while preserving accuracy, a multidimensional optimization that is today performed manually by experts. We ask whether an LLM agent can autonomously navigate this complex, multi-turn pipeline guided by real hardware feedback, and introduce a hardware-in-the-loop agent arena in which the agent iteratively refines both model and firmware – compiling, flashing, and measuring on real hardware – to enable closed-loop optimization. Frontier models, including Claude Opus 4.7 and Gemini 3.1 Pro, fail entirely without hardware feedback (0% deployment success), whereas our hardware-in-the-loop formulation achieves the first successful deployment within three iterations and can surpass human expert results within seven. This agentic co-optimization achieves 250x compression for vision models with

Read & Discuss → View Source →

15.

arXiv (CS.CV) 2026-06-25 DOI: arXiv:2606.24944

A Leakage-Aware Comparative Benchmark of Machine Learning, Deep Learning, and Transformer Models for Reliable Leukemia Detection

Authors:

Nisreen Albzour ↗

Automated classification of acute lymphoblastic leukemia (ALL) from peripheral blood smear images has often reported near-perfect performance on the C-NMC 2019 dataset. We show that such results can be inflated by patient-level data leakage caused by random image-level partitioning, where cells from the same subject may appear in both training and test folds. We establish a leakage-aware benchmark under a strict subject-disjoint protocol, comparing LightGBM, RBF-SVM, EfficientNet-B0, EfficientNet-B1, and ViT-Tiny. Models are developed using three subject-disjoint folds from 73 subjects and evaluated on an external preliminary-phase test set of 1,867 images from 28 unseen subjects with zero patient overlap. Beyond discrimination, we assess calibration using expected calibration error, Brier score, and temperature scaling. Under honest evaluation, EfficientNet-B1 achieves the best performance, with AUROC 0.913, sensitivity 0.87, specificity 0.80, and calibrated ECE 0.024. Frozen-feature classifiers and ViT-Tiny show high sensitivity but poor specificity, indicating a tendency to over-predict the malignant class. A random-versus-subject-disjoint ablation shows that random splitting inflates AUROC by about 0.04 even in the conservative frozen-feature setting. These findings caution against image-level evaluation on C-NMC 2019 and provide a reproducible, calibration-aware benchmark for future work.

Read & Discuss → View Source →

16.

arXiv (math.PR) 2026-06-25 DOI: arXiv:2606.25501

An RDT based approach to large deviations of Wishart and Wigner matrices spectral edges

Authors:

Mihailo Stojnic ↗

arXiv:2606.25501v1 Announce Type: new Abstract: We present a novel methodology for studying large deviations principles (LDPs) of random matrices. By utilizing a partially lifted variant of random duality theory (RDT), we develop a generic LDP framework that completely circumvents traditional random matrix theory (RMT) methods. To demonstrate the framework's simplicity and accuracy, we apply it to the Wishart and Wigner GOE classical statistical ensembles. In both cases, we obtain elegant LDP characterizations of the upper and lower spectral edges that fully match the results achieved through traditional Coulomb gas methodologies in [85,95].

Read & Discuss → View Source →

17.

arXiv (CS.AI) 2026-06-24 DOI: arXiv:2606.24808

Large-Language-Model Discovery of Quantum LDPC Codes through Structured Concept Evolution

Authors:

Zidu Liu ↗Florian Marquardt ↗

arXiv:2606.24808v1 Announce Type: cross Abstract: Quantum computers could outperform classical machines on important problems, but only if the errors that pervade quantum hardware can be corrected at scale. Quantum low-density parity-check (qLDPC) codes offer a promising route to this goal by combining sparse parity checks with finite encoding rate and growing distance, but their construction remains a challenging discrete design problem. Here we introduce structured concept evolution (SCE), a search framework that pairs a large language model with a structured algebraic mutation grammar to discover lifted-product code families, a class of CSS qLDPC codes. Instead of asking the LLM to design codes from first principles, SCE evolves structured concepts consisting of algebraic specifications paired with executable programs that realize them, using hierarchical mutations that modify the group algebra, protograph geometry, or base space. Running SCE, we discover a diverse set of competitive code families, ranging from abelian constructions to families over non-abelian groups beyond those underlying standard designs such as bivariate-bicycle codes, and characterize them under code-capacity depolarizing noise with BP+OSD decoding. These results are obtained with lightweight models (GPT-5.4-mini and GPT-5.4-nano).

Read & Discuss → View Source →

18.

arXiv (CS.AI) 2026-06-17 DOI: arXiv:2606.18098

IsabeLLM: Automated Theorem Proving Applied to Formally Verifying Consensus

Authors:

Elliot Jones ↗William Knottenbelt ↗

arXiv:2606.18098v1 Announce Type: new Abstract: Advances in Artificial Intelligence (AI) have led AI for Theorem Proving to become a promising means of formally verifying computer systems. Whilst formal verification is traditionally reserved for safety-critical systems due to the required amount of expertise and effort, AI can help to automate a large amount of this workload and make it far more accessible. Blockchain-based systems are becoming increasingly popular and are frequently targeted by malicious actors, often resulting in huge financial losses, highlighting the need to better verify these systems and mitigate vulnerabilities. Arguably the most important component of these systems is the consensus protocol, which allows nodes to agree on decisions in a potentially adversarial environment. In this paper, we improve upon IsabeLLM, the automated theorem proving tool in Isabelle. Namely, we implement a Retrieval-Augmented Generation framework, Error tracing and counterexample generation for improved context supplied to the Large Language Model. Compatibility with the latest version of Isabelle and Sledgehammer is also implemented for improved efficiency. We compare the performance of the two versions of IsabeLLM in their ability to complete the verification of Bitcoin's Proof of Work consensus.

Read & Discuss → View Source →

19.

arXiv (CS.CV) 2026-06-16 DOI: arXiv:2606.16188

teasr: training-efficient any-step diffusion transformer for real-world image super-resolution

Authors:

Xiang Gao ↗Chenxin Zhu ↗Yushun Fang ↗Qiang Hu ↗Xiaoyun Zhang ↗

Diffusion models excel in Real-World Image Super-Resolution (Real-ISR) due to their powerful generative priors but suffer from slow iterative sampling. Although existing one-step distillation methods accelerate inference, they typically require auxiliary teacher models that inflate training memory and restrict scalability to large-scale architectures. Furthermore, these fixed-step models lack the flexibility to trade off speed for quality. In this paper, we propose TEASR, a training-efficient any-step diffusion framework for Real-ISR that enables both one-step and multi-step restoration within a unified model. Our key idea is to perform self-adversarial distillation within a single diffusion model, eliminating the need for auxiliary teachers or discriminators. Specifically, we propose a timestep-aware rectification strategy that stabilizes one-step generation across noise levels. These two designs further enables the distillation of 20B-parameter diffusion models on a single GPU, significantly improving training efficiency. Moreover, we introduce a dual-branch diffusion transformer with decoupled timestep condition to separate the current noise state and the denoising target to enhance sampling quality. Extensive experiments demonstrate that TEASR supports seamless any-step sampling and consistently outperforms state-of-the-art methods across multiple datasets.

Read & Discuss → View Source →

20.

arXiv (CS.AI) 2026-06-17 DOI: arXiv:2606.18191

DRFLOW: A Deep Research Benchmark for Personalized Workflow Prediction

Authors:

Md Tawkat Islam Khondaker ↗Raymond Li ↗Muhammad Abdul-Mageed ↗Laks V. S. Lakshmanan ↗Issam H. Laradji ↗

arXiv:2606.18191v1 Announce Type: new Abstract: Deep research (DR) systems are increasingly used for complex information-seeking tasks, but existing works mainly focus on generating reports and summaries. In contrast, many enterprise tasks instead require an agent to identify concrete workflows which is a sequence of action-steps. For example, rather than summarizing budgeting policies, an agent should be able to determine the steps needed to answer a question such as: "How do I request new headcount given a fixed budget?". Therefore, we introduce DRFLOW, a benchmark for evaluating personalized workflows predicted by agents from heterogeneous sources. Each task requires the agent to identify relevant evidence from scattered sources, then use that evidence to predict the correct action-step sequence for the user's task. DRFLOW contains 100 tasks across five domains, with 1,246 reference workflow steps grounded in more than 3,900 sources. We define seven diagnostic metrics covering factual grounding, step recovery, structural ordering, condition resolution, and personalization. We further present DRFLOW-Agent (DRFA), a workflow-oriented reference agent to predict personalized workflow. We show that although DRFA improves over strong baseline agents (upto 10.02% average F1 score), there is substantial room for improvement remains across these workflow metrics, indicating that predicting complete and correct personalized workflows remains a challenging frontier for deep research.

Read & Discuss → View Source →

21.

arXiv (CS.CV) 2026-06-24 DOI: arXiv:2606.22371

ZeroGVC: Zero-Shot Generative Video Compression with Autoregressive Diffusion Priors

Authors:

Yixin Gao ↗Xiaohan Pan ↗Lin Liu ↗Xin Li ↗Zhibo Chen ↗Qi Tian ↗

Recent generative video compression methods leverage powerful generative priors to achieve perceptually pleasing reconstructions. However, most existing approaches require additional training to adapt generative models to produce realistic reconstructions from compact representations. In this paper, we propose ZeroGVC, a zero-shot generative video compression framework that leverages pretrained autoregressive diffusion priors for low-delay video reconstruction. ZeroGVC encodes the first frame of each group of pictures (GOP) with an image codec and represents subsequent P-frames through Codebook-Guided Autoregressive Latent Compression. This design is motivated by our observation that the compression scheme of denoising diffusion codebook models is effective in few-step consistency sampling. By selecting compact combinations of reproducible codebook noise vectors, ZeroGVC steers the latent denoising trajectory toward the target P-frame while allowing the decoder to reproduce the same trajectory in only a few denoising steps. In addition, we design an optional bidirectional reference mode that mitigates error propagation by leveraging the next I-frame context without introducing any additional bitrate overhead. Extensive experiments on standard video compression benchmarks demonstrate that ZeroGVC achieves superior perceptual reconstruction quality at ultra-low bitrates without any additional training.

Read & Discuss → View Source →

22.

arXiv (quant-ph) 2026-06-24 DOI: arXiv:2606.23951

M{\o}lmer-S{\o}rensen gates in trapped-ions chains in the presence of correlated noise

Authors:

D. V. Donchenko ↗E. A. Anikin ↗O. Lakhmanskaya ↗K. Lakhmanskiy ↗

arXiv:2606.23951v1 Announce Type: new Abstract: We analyze the impact of correlated laser frequency noise on M{\o}lmer-S{\o}rensen gates in qubit registers based on trapped-ion chains. Using perturbation theory, we calculate gate fidelities in the presence of noise with arbitrary power spectral density for different chain lengths and ion positions in the chain. With our approach, we account for simultaneous excitation of multiple phonon modes during gate operation. We find out that the impact of medium-frequency laser noise depends considerably on the positions of the ions in the chain. In contrast, low-frequency noise has similar effect for different chain lengths and ion positions.

Read & Discuss → View Source →

23.

arXiv (CS.CL) 2026-06-18 DOI: arXiv:2606.18543

CEO-Bench: Can Agents Play the Long Game?

Authors:

Haozhe Chen ↗Karthik Narasimhan ↗Zhuang Liu ↗

Language model agents are becoming proficient executors at isolated, short-horizon tasks such as software engineering and customer service. Yet real-world challenges require a combination of sophisticated skills that remain largely untested in agents: (1) navigating long horizons amid uncertainty; (2) acquiring information in noisy environments; (3) adapting to a changing world; (4) orchestrating multiple moving parts toward a coherent goal. We introduce CEO-Bench, which evaluates these capabilities together by simulating a representative real-world task: operating a startup for 500 days. An agent manages pricing, marketing, budgeting, and many other aspects of a fictional company through a programmable Python interface, operating in the same environment and facing the same challenges as a human CEO. Success demands analyzing noisy, interconnected business databases, translating signals into sound strategy, and coordinating many decisions with programming. The strongest agents write sophisticated code that simulates customer cohorts to forecast future cash and mines negotiation history to uncover hidden customer preferences. Even so, most state-of-the-art models struggle in this environment. Only Claude Opus 4.8 and GPT-5.5 finish above the $1M starting balance, and neither consistently turns a profit. CEO-Bench takes a first step toward measuring the intelligence required to drive sustained, adaptive progress over time.

Read & Discuss → View Source →

24.

arXiv (CS.CL) 2026-06-18 DOI: arXiv:2606.18979

Mitigating Scoring Errors and Compensating for Nonverbal Subtests in Speech-Based Dementia Assessment

Authors:

Franziska Braun ↗Christopher Witzl ↗Andreas Erzigkeit ↗Hartmut Lehfeld ↗Thomas Hillemacher ↗Tobias Bocklet ↗Korbinian Riedhammer ↗

Early detection of cognitive impairment relies on neuropsychological tests to minimize subjectivity by assessing multiple cognitive domains. Speech-based evaluation can support diagnostics and improve accessibility, but transcription errors and the omission of nonverbal subtests (e.g., motor skills) limit accuracy. Beyond conventional test scores, speech-derived features can provide additional insights into cognitive status. This study investigates the speech-based evaluation of the German "Syndrom-Kurz-Test," a standardized dementia screening test comprising verbal and motor subtests. We train models that integrate transcript-derived scores and Whisper embeddings per verbal subtest to reduce scoring errors. To compensate for missing motor subtests, we then leverage these fused representations to approximate expert overall ratings. Despite omitting subtests, our models strongly correlate with expert ratings and efficiently and accurately discriminate between cognitive status groups.

Read & Discuss → View Source →

25.

arXiv (CS.LG) 2026-06-24 DOI: arXiv:2606.21228

Sakana Fugu Technical Report

Authors:

Yujin Tang ↗Edoardo Cetin ↗Jinglue Xu ↗Qi Sun ↗Stefan Nielsen ↗Vincent Richard ↗Haruto Goda ↗Iaroslav Tymchenko ↗Nhan Nguyen ↗Hyunin Lee ↗Mari Ashiga ↗Shashank Kotyan ↗…

arXiv:2606.21228v2 Announce Type: replace Abstract: The capabilities of frontier Large Language Models (LLMs) continue to advance, with different providers increasingly specializing in distinct domains. This raises a natural next objective: how to combine the individual specializations of various LLMs into a collectively intelligent system. To this end, we report the development of Sakana Fugu, a family of orchestrator models that harness and amplify the capabilities of an LLM agent team. Fugu models are themselves language models trained to understand user queries and dynamically devise agentic scaffolds to solve them. Through these adaptive scaffolds, Fugu accesses performance beyond any individual LLM agent, achieving state-of-the-art results compared to other publicly accessible models across a range of challenging tasks, including SWE-Bench Pro, Terminal Bench, LiveCodeBench, GPQA-Diamond, Humanity's Last Exam, and CharXiv Reasoning. We release two models: Fugu, which balances performance with latency for everyday use, and Fugu-Ultra, which prioritizes answer quality on the hardest problems. We describe our training paradigm, which encompasses large-scale fine-tuning, evolutionary algorithms, and reinforcement learning approaches, along with the infrastructure and core design principles that turn these methods into a production system. We hope this report encourages further research into multi-agent systems and dynamic, query-adaptive agentic scaffolds as a path toward the next frontier of AI capabilities, accessed through collective intelligence.

Read & Discuss → View Source →

Explore the Frontier of Global Academia