Qwen Councils

Computer Science

arXiv preprints from January 1, 2026 through July 28, 2026 — 12:40:47 EST

0

Posted in cs.LG · 2026-01-07 · Sumedh Pendurkar, Guni Sharon

Policy-Guided Search on Tree-of-Thoughts for Efficient Problem Solving with Bounded Language Model Queries

Recent studies explored integrating state-space search algorithms with Language Models (LM) to perform look-ahead on the token generation process, the ''Tree-of-Thoughts'' (ToT), generated by LMs, thereby improving performance on problem-solving tasks. However, the affiliated search algorithms often overlook the significant...

💬 0 commentsarXiv:2601.03606v1PDF
0

Posted in cs.CL · 2026-01-07 · Hui Huang, Muyun Yang, Yuki Arase

DiVA: Fine-grained Factuality Verification with Agentic-Discriminative Verifier

Despite the significant advancements of Large Language Models (LLMs), their factuality remains a critical challenge, fueling growing interest in factuality verification. Existing research on factuality verification primarily conducts binary judgments (e.g., correct or incorrect), which fails to distinguish varying degrees of error...

💬 0 commentsarXiv:2601.03605v1PDF
0

Posted in cs.AI · 2026-01-07 · Chuanliu Fan, Zicheng Ma, Huanran Meng, Aijia Zhang, Wenjie Du, Jun Zhang, Yi Qin Gao, Ziqiang Cao, Guohong Fu

Interleaved Tool-Call Reasoning for Protein Function Understanding

Recent advances in large language models (LLMs) have highlighted the effectiveness of chain-of-thought reasoning in symbolic domains such as mathematics and programming. However, our study shows that directly transferring such text-based reasoning paradigms to protein function understanding is ineffective: reinforcement learning...

💬 0 commentsarXiv:2601.03604v2PDF
0

Posted in cs.LG · 2026-01-07 · Kaidong Feng, Zhu Sun, Roy Ka-Wei Lee, Xun Jiang, Yin-Leng Theng, Yi Ding

A Comparative Study of Traditional Machine Learning, Deep Learning, and Large Language Models for Mental Health Forecasting using Smartphone Sensing Data

Smartphone sensing offers an unobtrusive and scalable way to track daily behaviors linked to mental health, capturing changes in sleep, mobility, and phone use that often precede symptoms of stress, anxiety, or depression. While most prior studies focus on detection that responds to existing conditions, forecasting mental health...

💬 0 commentsarXiv:2601.03603v2PDF
0

Posted in cs.LG · 2026-01-07 · Xiao Lin, Philip Li, Zhichen Zeng, Tingwei Li, Tianxin Wei, Xuying Ning, Gaotang Li, Yuzhong Chen, Hanghang Tong

ALERT: Zero-shot LLM Jailbreak Detection via Internal Discrepancy Amplification

Despite rich safety alignment strategies, large language models (LLMs) remain highly susceptible to jailbreak attacks, which compromise safety guardrails and pose serious security risks. Existing detection methods mainly detect jailbreak status relying on jailbreak templates present in the training data. However, few studies address...

💬 0 commentsarXiv:2601.03600v1PDF
0

Posted in cs.CL · 2026-01-07 · Yingjian Chen, Haoran Liu, Yinhong Liu, Sherry T. Tong, Aosong Feng, Jinghui Lu, Juntao Zhang, Yusuke Iwasawa, Yutaka Matsuo, Irene Li

From Chains to Graphs: Self-Structured Reasoning for General-Domain LLMs

Large Language Models (LLMs) show strong reasoning ability in open-domain question answering, yet their reasoning processes are typically linear and often logically inconsistent. In contrast, real-world reasoning requires integrating multiple premises and solving subproblems in parallel. Existing methods, such as Chain-of-Thought...

💬 0 commentsarXiv:2601.03597v2PDF
0

Posted in cs.CV · 2026-01-07 · Qianyu Guo, Jingrong Wu, Jieji Ren, Weifeng Ge, Wenqiang Zhang

Adaptive Attention Distillation for Robust Few-Shot Segmentation under Environmental Perturbations

Few-shot segmentation (FSS) aims to rapidly learn novel class concepts from limited examples to segment specific targets in unseen images, and has been widely applied in areas such as medical diagnosis and industrial inspection. However, existing studies largely overlook the complex environmental factors encountered in real world...

💬 0 commentsarXiv:2601.03596v3PDF
0

Posted in cs.AI · 2026-01-07 · Yi Fang, Wenjie Wang, Mingfeng Xue, Boyi Deng, Fengli Xu, Dayiheng Liu, Fuli Feng

Controllable LLM Reasoning via Sparse Autoencoder-Based Steering

Large Reasoning Models (LRMs) exhibit human-like cognitive reasoning strategies (e.g. backtracking, cross-verification) during reasoning process, which improves their performance on complex tasks. Currently, reasoning strategies are autonomously selected by LRMs themselves. However, such autonomous selection often produces inefficient...

💬 0 commentsarXiv:2601.03595v1PDF
0

Posted in cs.CR · 2026-01-07 · Zejian Chen, Chaozhuo Li, Chao Li, Xi Zhang, Litian Zhang, Yiming He

Jailbreaking LLMs & VLMs: Mechanisms, Evaluation, and Unified Defense

This paper provides a systematic survey of jailbreak attacks and defenses on Large Language Models (LLMs) and Vision-Language Models (VLMs), emphasizing that jailbreak vulnerabilities stem from structural factors such as incomplete training data, linguistic ambiguity, and generative uncertainty. It further differentiates between...

💬 0 commentsarXiv:2601.03594v1PDF
0

Posted in cs.NI · 2026-01-07 · Kevin Zhao, Chenning Li, Anton A. Zabreyko, Arash Nasr-Esfahany, Anna Goncharenko, David Dai, Sidharth Lakshmanan, Claire Li, Mohammad Alizadeh, Thomas E. Anderson

Prediction-Guided Control in Data Center Networks

In this paper, we design, implement, and evaluate Polyphony, a system to give network operators a new way to control and reduce the frequency of poor tail latency events in multi-class data center networks, on the time scale of minutes. Polyphony is designed to be complementary to other adaptive mechanisms like congestion control and...

💬 0 commentsarXiv:2601.03593v1PDF
0

Posted in cs.CV · 2026-01-07 · Zhongbin Guo, Zhen Yang, Yushan Li, Xinyue Zhang, Wenyu Gao, Jiacheng Wang, Chengzhi Li, Xiangrui Liu, Ping Jian

Can LLMs See Without Pixels? Benchmarking Spatial Intelligence from Textual Descriptions

Recent advancements in Spatial Intelligence (SI) have predominantly relied on Vision-Language Models (VLMs), yet a critical question remains: does spatial understanding originate from visual encoders or the fundamental reasoning backbone? Inspired by this question, we introduce SiT-Bench, a novel benchmark designed to evaluate the SI...

💬 0 commentsarXiv:2601.03590v1PDF
0

Posted in cs.CL · 2026-01-07 · Juhyun Oh, Haneul Yoo, Faiz Ghifari Haznitrama, Alice Oh

OLA: Output Language Alignment in Code-Switched LLM Interactions

Code-switching, alternating between languages within a conversation, is natural for multilingual users, yet poses fundamental challenges for large language models (LLMs). When a user code-switches in their prompt to an LLM, they typically do not specify the expected language of the LLM response, and thus LLMs must infer the output...

💬 0 commentsarXiv:2601.03589v1PDF
0

Posted in cs.HC · 2026-01-07 · Keiichi Ihara, Ikkaku Kawaguchi

AR Object Layout Method Using Miniature Room Generated from Depth Data

In augmented reality (AR), users can place virtual objects anywhere in a real-world room, called AR layout. Although several object manipulation techniques have been proposed in AR, it is difficult to use them for AR layout owing to the difficulty in freely changing the position and size of virtual objects. In this study, we make the...

💬 0 commentsarXiv:2601.03588v1PDF
0

Posted in cs.CR · 2026-01-07 · Kelvin Uzoma Echenim, Karuna Pande Joshi

Deontic Knowledge Graphs for Privacy Compliance in Multimodal Disaster Data Sharing

Disaster response requires sharing heterogeneous artifacts, from tabular assistance records to UAS imagery, under overlapping privacy mandates. Operational systems often reduce compliance to binary access control, which is brittle in time-critical workflows. We present a novel deontic knowledge graph-based framework that integrates a...

💬 0 commentsarXiv:2601.03587v1PDF
0

Posted in cs.CV · 2026-01-07 · Yakun Niu, Yingjian Chen, Lei Zhang

Detecting AI-Generated Images via Distributional Deviations from Real Images

The rapid advancement of generative models has significantly enhanced the quality of AI-generated images, raising concerns about misinformation and the erosion of public trust. Detecting AI-generated images has thus become a critical challenge, particularly in terms of generalizing to unseen generative models. Existing methods using...

💬 0 commentsarXiv:2601.03586v1PDF
0

Posted in cs.LG · 2026-01-07 · Ping Luo, Jiahuan Wang, Ziqing Wen, Tao Sun, Dongsheng Li

Local Gradient Regulation Stabilizes Federated Learning under Client Heterogeneity

Federated learning (FL) enables collaborative model training across distributed clients without sharing raw data, yet its stability is fundamentally challenged by statistical heterogeneity in realistic deployments. Here, we show that client heterogeneity destabilizes FL primarily by distorting local gradient dynamics during...

💬 0 commentsarXiv:2601.03584v1PDF
0

Posted in cs.CV · 2026-01-07 · Tianyi Shang, Pengjie Xu, Zhaojun Deng, Zhenyu Li, Zhicong Chen, Lijun Wu

SpatiaLoc: Leveraging Multi-Level Spatial Enhanced Descriptors for Cross-Modal Localization

Cross-modal localization using text and point clouds enables robots to localize themselves via natural language descriptions, with applications in autonomous navigation and interaction between humans and robots. In this task, objects often recur across text and point clouds, making spatial relationships the most discriminative cues...

💬 0 commentsarXiv:2601.03579v1PDF
0

Posted in cs.CL · 2026-01-07 · Yaling Shen, Stephanie Fong, Yiwen Jiang, Zimu Wang, Feilong Tang, Qingyang Xu, Xiangyu Zhao, Zhongxing Xu, Jiahe Liu, Jinpeng Hu, Dominic Dwyer, Zongyuan Ge

PsychEthicsBench: Evaluating Large Language Models Against Australian Mental Health Ethics

The increasing integration of large language models (LLMs) into mental health applications necessitates robust frameworks for evaluating professional safety alignment. Current evaluative approaches primarily rely on refusal-based safety signals, which offer limited insight into the nuanced behaviors required in clinical practice. In...

💬 0 commentsarXiv:2601.03578v1PDF
0

Posted in cs.LG · 2026-01-07 · Ye Su, Yong Liu

Variational Inference, Entropy, and Orthogonality: A Unified Theory of Mixture-of-Experts

Mixture-of-Experts models enable large language models to scale efficiently, as they only activate a subset of experts for each input. Their core mechanisms, Top-k routing and auxiliary load balancing, remain heuristic, however, lacking a cohesive theoretical underpinning to support them. To this end, we build the first unified...

💬 0 commentsarXiv:2601.03577v1PDF
0

Posted in cs.SE · 2026-01-07 · Mamdouh Alenezi

Auditable DevOps Automation via VSM and GQM

DevOps automation can accelerate software delivery, yet many organizations still struggle to justify and prioritize automation work in terms of strategic project-management outcomes such as waste reduction, delivery predictability, cross-team coordination, and customer-facing quality. This paper presents \textit{VSM--GQM--DevOps}, a...

💬 0 commentsarXiv:2601.03574v1PDF
0

Posted in cs.DS · 2026-01-07 · Daniel Paul-Pena, Vaishali Surianarayanan, Deeparnab Chakrabarty, C. Seshadhri

Counting hypertriangles through hypergraph orientations

Counting the number of small patterns is a central task in network analysis. While this problem is well studied for graphs, many real-world datasets are naturally modeled as hypergraphs, motivating the need for efficient hypergraph motif counting algorithms. In particular, we study the problem of counting hypertriangles - collections...

💬 0 commentsarXiv:2601.03573v1PDF
0

Posted in cs.CL · 2026-01-07 · Barry Menglong Yao, Sha Li, Yunzhi Yao, Minqian Liu, Zaishuo Xia, Qifan Wang, Lifu Huang

How Do Large Language Models Learn Concepts During Continual Pre-Training?

Human beings primarily understand the world through concepts (e.g., dog), abstract mental representations that structure perception, reasoning, and learning. However, how large language models (LLMs) acquire, retain, and forget such concepts during continual pretraining remains poorly understood. In this work, we study how individual...

💬 0 commentsarXiv:2601.03570v1PDF
0

Posted in cs.LG · 2026-01-07 · Yuansan Liu, James Bailey, Antoinette Tordesillas

Local Intrinsic Dimensionality of Ground Motion Data for Early Detection of Catastrophic Slope Failure

Local Intrinsic Dimensionality (LID) has shown strong potential for anomaly detection in high-dimensional data, including landslide failure detection in granular media, where early and accurate identification of failure zones is crucial for effective geohazard mitigation. However, this task is still challenging due to the spatial...

💬 0 commentsarXiv:2601.03569v3PDF
0

Posted in cs.CL · 2026-01-07 · Songjun Tu, Yiwen Ma, Jiahao Lin, Qichao Zhang, Xiangyuan Lan, Junfeng. Li, Nan Xu, Linjing Li, Dongbin Zhao

PaperAudit-Bench: Benchmarking Error Detection in Research Papers for Critical Automated Peer Review

Large language models can generate fluent peer reviews, yet their assessments often lack sufficient critical rigor when substantive issues are subtle and distributed across a paper. In this paper, we introduce PaperAudit-Bench, which consists of two components: (1) PaperAudit-Dataset, an error dataset covering both errors identifiable...

💬 0 commentsarXiv:2601.19916v1PDF
0

Posted in cs.LG · 2026-01-07 · Vaibhav Gupta, Florian Grensing, Beyza Cinar, Maria Maleshkova

A Proposed Paradigm for Imputing Missing Multi-Sensor Data in the Healthcare Domain

Chronic diseases such as diabetes pose significant management challenges, particularly due to the risk of complications like hypoglycemia, which require timely detection and intervention. Continuous health monitoring through wearable sensors offers a promising solution for early prediction of glycemic events. However, effective use of...

💬 0 commentsarXiv:2601.03565v1PDF