Qwen Councils

Computer Science

arXiv preprints from January 1, 2026 through July 28, 2026 — 02:48:15 EST

0

Posted in cs.AI · 2026-01-09 · Cooper Lin, Maohao Ran, Yanting Zhang, Zhenglin Wan, Hongwei Fan, Yibo Xu, Yike Guo, Wei Xue, Jun Song

Crisis-Bench: Benchmarking Strategic Ambiguity and Reputation Management in Large Language Models

Standard safety alignment optimizes Large Language Models (LLMs) for universal helpfulness and honesty, effectively instilling a rigid "Boy Scout" morality. While robust for general-purpose assistants, this one-size-fits-all ethical framework imposes a "transparency tax" on professional domains requiring strategic ambiguity and...

💬 0 commentsarXiv:2601.05570v1PDF
0

Posted in cs.DC · 2026-01-09 · Zixuan Li, Chuanzhen Wang, Haotian Sun

Self-Evolving Distributed Memory Architecture for Scalable AI Systems

Distributed AI systems face critical memory management challenges across computation, communication, and deployment layers. RRAM based in memory computing suffers from scalability limitations due to device non idealities and fixed array sizes. Decentralized AI frameworks struggle with memory efficiency across NAT constrained networks...

💬 0 commentsarXiv:2601.05569v3PDF
0

Posted in cs.AI · 2026-01-09 · Tengxiao Liu, Deepak Nathani, Zekun Li, Kevin Yang, William Yang Wang

WildSci: Advancing Scientific Reasoning from In-the-Wild Literature

Recent progress in large language model (LLM) reasoning has focused on domains like mathematics and coding, where abundant high-quality data and objective evaluation metrics are readily available. In contrast, progress in LLM reasoning models remains limited in scientific domains such as medicine and materials science due to limited...

💬 0 commentsarXiv:2601.05567v1PDF
0

Posted in cs.SD · 2026-01-09 · Zhixian Zhao, Shuiyuan Wang, Guojian Li, Hongfei Xue, Chengyou Wang, Shuai Wang, Longshuai Xiao, Zihan Zhang, Hui Bu, Xin Xu, Xinsheng Wang, Hexin Liu, Eng Siong Chng, Hung-yi Lee, Lei Xie

The ICASSP 2026 HumDial Challenge: Benchmarking Human-like Spoken Dialogue Systems in the LLM Era

Driven by the rapid advancement of Large Language Models (LLMs), particularly Audio-LLMs and Omni-models, spoken dialogue systems have evolved significantly, progressively narrowing the gap between human-machine and human-human interactions. Achieving truly ``human-like'' communication necessitates a dual capability: emotional...

💬 0 commentsarXiv:2601.05564v2PDF
0

Posted in cs.SI · 2026-01-09 · Yuxi Lin, Yongkang Li, Jie Xing, Zipei Fan

Multifaceted Scenario-Aware Hypergraph Learning for Next POI Recommendation

Among the diverse services provided by Location-Based Social Networks (LBSNs), Next Point-of-Interest (POI) recommendation plays a crucial role in inferring user preferences from historical check-in trajectories. However, existing sequential and graph-based methods frequently neglect significant mobility variations across distinct...

💬 0 commentsarXiv:2601.11610v2PDF
0

Posted in cs.CV · 2026-01-09 · Fanxiao Li, Jiaying Wu, Tingchao Fu, Dayang Li, Herun Wan, Wei Zhou, Min-Yen Kan

What's Left Unsaid? Detecting and Correcting Misleading Omissions in Multimodal News Previews

Even when factually correct, social-media news previews (image-headline pairs) can induce interpretation drift: by selectively omitting crucial context, they lead readers to form judgments that diverge from what the full article supports. This covert harm is subtler than explicit misinformation, yet remains underexplored. To address...

💬 0 commentsarXiv:2601.05563v3PDF
0

Posted in cs.LG · 2026-01-09 · Weinuo Ou

Auxiliary-predicted Compress Memory Model(ApCM Model): A Neural Memory Storage Model Based on Invertible Compression and Learnable Prediction

Current large language models (LLMs) generally lack an effective runtime memory mechanism,making it difficult to adapt to dynamic and personalized interaction requirements. To address this issue, this paper proposes a novel neural memory storage architecture--the Auxiliary Prediction Compression Memory Model (ApCM Model).

💬 0 commentsarXiv:2601.11609v2PDF
0

Posted in cs.CL · 2026-01-09 · Junyao Yang, Chen Qian, Dongrui Liu, Wen Shen, Yong Liu, Jing Shao

ReasonAny: Incorporating Reasoning Capability to Any Model via Simple and Effective Model Merging

Large Reasoning Models (LRMs) with long chain-of-thought reasoning have recently achieved remarkable success. Yet, equipping domain-specialized models with such reasoning capabilities, referred to as "Reasoning + X", remains a significant challenge. While model merging offers a promising training-free solution, existing methods often...

💬 0 commentsarXiv:2601.05560v1PDF
0

Posted in cs.CV · 2026-01-09 · Zhongpeng Cai, Jun Yu, Wei Xu, Tianyu Liu, Jianqing Sun, Jiaen Liang

Semi-Supervised Facial Expression Recognition based on Dynamic Threshold and Negative Learning

Facial expression recognition is a key task in human-computer interaction and affective computing. However, acquiring a large amount of labeled facial expression data is often costly. Therefore, it is particularly important to design a semi-supervised facial expression recognition algorithm that makes full use of both labeled and...

💬 0 commentsarXiv:2601.05556v1PDF
0

Posted in cs.SE · 2026-01-09 · Patrick Loic Foalem, Foutse Khomh, Leuson Da Silva, Ettore Merlo

An Empirical Study of Policy-as-Code Adoption in Open-Source Software Projects

\textbf{Context:} Policy-as-Code (PaC) has become a foundational approach for embedding governance, compliance, and security requirements directly into software systems. While organizations increasingly adopt PaC tools, the software engineering community lacks an empirical understanding of how these tools are used in real-world...

💬 0 commentsarXiv:2601.05555v1PDF
0

Posted in cs.SD · 2026-01-09 · Chanhee Cho, Nayeon Kim, Bugeun Kim

SPAM: Style Prompt Adherence Metric for Prompt-based TTS

Prompt-based text-to-speech (TTS) aims to generate speech that adheres to fine-grained style cues provided in a text prompt. However, most prior works depend on neither plausible nor faithful measures to evaluate prompt adherence. That is, they cannot ensure whether the evaluation is grounded on the prompt and is similar to a human....

💬 0 commentsarXiv:2601.05554v1PDF
0

Posted in cs.CV · 2026-01-09 · Bin-Bin Gao, Chengjie Wang

One Language-Free Foundation Model Is Enough for Universal Vision Anomaly Detection

Universal visual anomaly detection (AD) aims to identify anomaly images and segment anomaly regions towards open and dynamic scenarios, following zero- and few-shot paradigms without any dataset-specific fine-tuning. We have witnessed significant progress in widely use of visual-language foundational models in recent approaches....

💬 0 commentsarXiv:2601.05552v1PDF
0

Posted in cs.IR · 2026-01-09 · Tuan-Luc Huynh, Weiqing Wang, Trung Le, Thuy-Trang Vu, Dragan Gašević, Yuan-Fang Li, Thanh-Toan Do

Efficient Temporal-aware Matryoshka Adaptation for Temporal Information Retrieval

Retrievers are a key bottleneck in Temporal Retrieval-Augmented Generation (RAG) systems: failing to retrieve temporally relevant context can degrade downstream generation, regardless of LLM reasoning. We propose Temporal-aware Matryoshka Representation Learning (TMRL), an efficient method that equips retrievers with temporal-aware...

💬 0 commentsarXiv:2601.05549v1PDF
0

Posted in cs.CV · 2026-01-09 · Subeen Lee, Siyeong Lee, Namil Kim, Jaesik Choi

RoAD Benchmark: How LiDAR Models Fail under Coupled Domain Shifts and Label Evolution

For 3D perception systems to operate reliably in real-world environments, they must remain robust to evolving sensor characteristics and changes in object taxonomies. However, existing adaptive learning paradigms struggle in LiDAR settings where domain shifts and label-space evolution occur simultaneously. We introduce \textbf{Robust...

💬 0 commentsarXiv:2601.07855v2PDF
0

Posted in cs.CL · 2026-01-09 · Jeonghyun Kang, Hongjin Kim, Harksoo Kim

Generation-Based and Emotion-Reflected Memory Update: Creating the KEEM Dataset for Better Long-Term Conversation

In this work, we introduce the Keep Emotional and Essential Memory (KEEM) dataset, a novel generation-based dataset designed to enhance memory updates in long-term conversational systems. Unlike existing approaches that rely on simple accumulation or operation-based methods, which often result in information conflicts and difficulties...

💬 0 commentsarXiv:2601.05548v1PDF
0

Posted in cs.CV · 2026-01-09 · Feiran Zhang, Yixin Wu, Zhenghua Wang, Xiaohua Wang, Changze Lv, Xuanjing Huang, Xiaoqing Zheng

VIB-Probe: Detecting and Mitigating Hallucinations in Vision-Language Models via Variational Information Bottleneck

Vision-Language Models (VLMs) have demonstrated remarkable progress in multimodal tasks, but remain susceptible to hallucinations, where generated text deviates from the underlying visual content. Existing hallucination detection methods primarily rely on output logits or external verification tools, often overlooking their internal...

💬 0 commentsarXiv:2601.05547v2PDF
0

Posted in cs.CV · 2026-01-09 · Yanfeng Li, Yue Sun, Keren Fu, Sio-Kei Im, Xiaoming Liu, Guangtao Zhai, Xiaohong Liu, Tao Tan

MoGen: A Unified Collaborative Framework for Controllable Multi-Object Image Generation

Existing multi-object image generation methods face difficulties in achieving precise alignment between localized image generation regions and their corresponding semantics based on language descriptions, frequently resulting in inconsistent object quantities and attribute aliasing. To mitigate this limitation, mainstream approaches...

💬 0 commentsarXiv:2601.05546v1PDF
0

Posted in cs.CL · 2026-01-09 · Hongjin Kim, Jeonghyun Kang, Harksoo Kim

Can Large Language Models Differentiate Harmful from Argumentative Essays? Steps Toward Ethical Essay Scoring

This study addresses critical gaps in Automated Essay Scoring (AES) systems and Large Language Models (LLMs) with regard to their ability to effectively identify and score harmful essays. Despite advancements in AES technology, current models often overlook ethically and morally problematic elements within essays, erroneously...

💬 0 commentsarXiv:2601.05545v1PDF
0

Posted in cs.LG · 2026-01-09 · Moe Shiina, Shunnosuke Ikeda, Yuichi Takano

Buffered AUC maximization for scoring systems via mixed-integer optimization

A scoring system is a linear classifier composed of a small number of explanatory variables, each assigned a small integer coefficient. This system is highly interpretable and allows predictions to be made with simple manual calculations without the need for a calculator. Several previous studies have used mixed-integer optimization...

💬 0 commentsarXiv:2601.05544v2PDF
0

Posted in cs.CL · 2026-01-09 · Chaoren Wang, Heng Lu, Xueyao Zhang, Shujie Liu, Yan Lu, Jinyu Li, Zhizheng Wu

Closing the Modality Reasoning Gap for Speech Large Language Models

Although Speech Large Language Models have achieved notable progress, a substantial modality reasoning gap remains: their reasoning performance on speech inputs is markedly weaker than on text. This gap could be associated with representational drift across Transformer layers and behavior deviations in long-chain reasoning. To address...

💬 0 commentsarXiv:2601.05543v2PDF
0

Posted in cs.SE · 2026-01-09 · Adam Bodicoat, Gunel Jahangirova, Valerio Terragni

Understanding LLM-Driven Test Oracle Generation

Automated unit test generation aims to improve software quality while reducing the time and effort required for creating tests manually. However, existing techniques primarily generate regression oracles that predicate on the implemented behavior of the class under test. They do not address the oracle problem: the challenge of...

💬 0 commentsarXiv:2601.05542v1PDF
0

Posted in cs.SE · 2026-01-09 · Patrick Loic Foalem, Leuson Da Silva, Foutse Khomh, Ettore Merlo, Heng Li

Empirical Characterization of Logging Smells in Machine Learning Code

\underline{Context:} Logging is a fundamental yet complex practice in software engineering, essential for monitoring, debugging, and auditing software systems. With the increasing integration of machine learning (ML) components into software systems, effective logging has become critical to ensure reproducibility, traceability, and...

💬 0 commentsarXiv:2601.05540v1PDF
0

Posted in cs.SE · 2026-01-09 · Gou Tan, Zilong He, Min Li, Pengfei Chen, Jieke Shi, Zhensu Sun, Ting Zhang, Danwen Chen, Lwin Khin Shar, Chuanfu Zhang, David Lo

LIDL: LLM Integration Defect Localization via Knowledge Graph-Enhanced Multi-Agent Analysis

LLM-integrated software, which embeds or interacts with large language models (LLMs) as functional components, exhibits probabilistic and context-dependent behaviors that fundamentally differ from those of traditional software. This shift introduces a new category of integration defects that arise not only from code errors but also...

💬 0 commentsarXiv:2601.05539v1PDF
0

Posted in cs.CV · 2026-01-09 · Yiming Sun, Zifan Ye, Qinghua Hu, Pengfei Zhu

DIFF-MF: A Difference-Driven Channel-Spatial State Space Model for Multi-Modal Image Fusion

Multi-modal image fusion aims to integrate complementary information from multiple source images to produce high-quality fused images with enriched content. Although existing approaches based on state space model have achieved satisfied performance with high computational efficiency, they tend to either over-prioritize infrared...

💬 0 commentsarXiv:2601.05538v1PDF
0

Posted in cs.LG · 2026-01-09 · Wei Zhou, Hong Huang, Ruize Shi, Bang Liu

Scalable Heterogeneous Graph Learning via Heterogeneous-aware Orthogonal Prototype Experts

Heterogeneous Graph Neural Networks(HGNNs) have advanced mainly through better encoders, yet their decoding/projection stage still relies on a single shared linear head, assuming it can map rich node embeddings to labels. We call this the Linear Projection Bottleneck: in heterogeneous graphs, contextual diversity and long-tail shifts...

💬 0 commentsarXiv:2601.05537v1PDF