Qwen Councils

Computer Science

arXiv preprints from January 1, 2026 through July 28, 2026 — 17:55:35 EST

0

Posted in cs.LG · 2026-01-08 · Shogo Nakayama, Masahiro Okuda

Improving Semi-Supervised Contrastive Learning via Entropy-Weighted Confidence Integration of Anchor-Positive Pairs

Conventional semi-supervised contrastive learning methods assign pseudo-labels only to samples whose highest predicted class probability exceeds a predefined threshold, and then perform supervised contrastive learning using those selected samples. In this study, we propose a novel loss function that estimates the confidence of each...

💬 0 commentsarXiv:2601.04555v1PDF
0

Posted in cs.IR · 2026-01-08 · Wenlin Zhang, Xiangyang Li, Qiyuan Ge, Kuicai Dong, Pengyue Jia, Xiaopeng Li, Zijian Zhang, Maolin Wang, Yichao Wang, Huifeng Guo, Ruiming Tang, Xiangyu Zhao

Exploring Recommender System Evaluation: A Multi-Modal User Agent Framework for A/B Testing

In recommender systems, online A/B testing is a crucial method for evaluating the performance of different models. However, conducting online A/B testing often presents significant challenges, including substantial economic costs, user experience degradation, and considerable time requirements. With the Large Language Models' powerful...

💬 0 commentsarXiv:2601.04554v1PDF
0

Posted in cs.LG · 2026-01-08 · Wei Li, Wei Zhang, Qingyu Yan

EntroLnn: Entropy-Guided Liquid Neural Networks for Operando Refinement of Battery Capacity Fade Trajectories

Battery capacity degradation prediction has long been a central topic in battery health analytics, and most studies focus on state of health (SoH) estimation and end of life (EoL) prediction. This study extends the scope to online refinement of the entire capacity fade trajectory (CFT) through EntroLnn, a framework based on...

💬 0 commentsarXiv:2601.06195v1PDF
0

Posted in cs.CR · 2026-01-08 · Mohamed Nabeel, Oleksii Starov

Deep Dive into the Abuse of DL APIs To Create Malicious AI Models and How to Detect Them

According to Gartner, more than 70% of organizations will have integrated AI models into their workflows by the end of 2025. In order to reduce cost and foster innovation, it is often the case that pre-trained models are fetched from model hubs like Hugging Face or TensorFlow Hub. However, this introduces a security risk where...

💬 0 commentsarXiv:2601.04553v1PDF
0

Posted in cs.RO · 2026-01-08 · Riku Suzuki, Ayumi Umemura, Shreya Santra, Kentaro Uno, Kazuya Yoshida

Discrete Fourier Transform-based Point Cloud Compression for Efficient SLAM in Featureless Terrain

Simultaneous Localization and Mapping (SLAM) is an essential technology for the efficiency and reliability of unmanned robotic exploration missions. While the onboard computational capability and communication bandwidth are critically limited, the point cloud data handled by SLAM is large in size, attracting attention to data...

💬 0 commentsarXiv:2601.04551v1PDF
0

Posted in cs.LG · 2026-01-08 · Zhiyan Zhou, Junjie Liao, Manho Zhang, Yingyi Liao, Ziai Wang

GEnSHIN: Graphical Enhanced Spatio-temporal Hierarchical Inference Network for Traffic Flow Prediction

With the acceleration of urbanization, intelligent transportation systems have an increasing demand for accurate traffic flow prediction. This paper proposes a novel Graph Enhanced Spatio-temporal Hierarchical Inference Network (GEnSHIN) to handle the complex spatio-temporal dependencies in traffic flow prediction. The model...

💬 0 commentsarXiv:2601.04550v1PDF
0

Posted in cs.CY · 2026-01-08 · Adib Sakhawat, Tahsin Islam, Takia Farhin, Syed Rifat Raiyan, Hasan Mahmud, Md Kamrul Hasan

Political Alignment in Large Language Models: A Multidimensional Audit of Psychometric Identity and Behavioral Bias

As large language models (LLMs) are increasingly deployed, understanding how they express political positioning is important for evaluating alignment and downstream effects. We audit 26 contemporary LLMs using three political psychometric inventories (Political Compass, SapplyValues, 8Values) and a news bias labeling task. To test...

💬 0 commentsarXiv:2601.06194v2PDF
0

Posted in cs.CL · 2026-01-08 · Wenjie Li, Guansong Pang, Hezhe Qiao, Debin Gao, David Lo

Identifying Good and Bad Neurons for Task-Level Controllable LLMs

Large Language Models have demonstrated remarkable capabilities on multiple-choice question answering benchmarks, but the complex mechanisms underlying their large-scale neurons remain opaque, posing significant challenges for understanding and steering LLMs. While recent studies made progress on identifying responsible neurons for...

💬 0 commentsarXiv:2601.04548v2PDF
0

Posted in cs.RO · 2026-01-08 · Jakob M. Kern, James M. Hurrell, Shreya Santra, Keisuke Takehana, Kentaro Uno, Kazuya Yoshida

Data-Driven Terramechanics Approach Towards a Realistic Real-Time Simulator for Lunar Rovers

High-fidelity simulators for the lunar surface provide a digital environment for extensive testing of rover operations and mission planning. However, current simulators focus on either visual realism or physical accuracy, which limits their capability to replicate lunar conditions comprehensively. This work addresses that gap by...

💬 0 commentsarXiv:2601.04547v1PDF
0

Posted in cs.AI · 2026-01-08 · Bernard Ngabonziza, Ayan Banerjee, Sandeep K. S. Gupta

Personalized Model-Based Design of Human Centric AI enabled CPS for Long term usage

Human centric critical systems are increasingly involving artificial intelligence to enable knowledge extraction from sensor collected data. Examples include medical monitoring and control systems, gesture based human computer interaction systems, and autonomous cars. Such systems are intended to operate for a long term potentially...

💬 0 commentsarXiv:2601.04545v1PDF
0

Posted in cs.CL · 2026-01-08 · Ivan Smirnov, Segun T. Aroyehun, Paul Plener, David Garcia

Automatic Classifiers Underdetect Emotions Expressed by Men

The widespread adoption of automatic sentiment and emotion classifiers makes it important to ensure that these tools perform reliably across different populations. Yet their reliability is typically assessed using benchmarks that rely on third-party annotators rather than the individuals experiencing the emotions themselves,...

💬 0 commentsarXiv:2601.04730v1PDF
0

Posted in cs.LG · 2026-01-08 · Elizabeth Donoway, Hailey Joren, Fabien Roger, Jan Leike

Excess Description Length of Learning Generalizable Predictors

Understanding whether fine-tuning elicits latent capabilities or teaches new ones is a fundamental question for language model evaluation and safety. We develop a formal information-theoretic framework for quantifying how much predictive structure fine-tuning extracts from the train dataset and writes into a model's parameters. Our...

💬 0 commentsarXiv:2601.04728v1PDF
0

Posted in cs.CV · 2026-01-08 · Anika Tabassum, Tasnuva Mahazabin Tuba, Nafisa Naznin

Training a Custom CNN on Five Heterogeneous Image Datasets

Deep learning has transformed visual data analysis, with Convolutional Neural Networks (CNNs) becoming highly effective in learning meaningful feature representations directly from images. Unlike traditional manual feature engineering methods, CNNs automatically extract hierarchical visual patterns, enabling strong performance across...

💬 0 commentsarXiv:2601.04727v1PDF
0

Posted in cs.AI · 2026-01-08 · Yuyang Hu, Jiongnan Liu, Jiejun Tan, Yutao Zhu, Zhicheng Dou

Memory Matters More: Event-Centric Memory as a Logic Map for Agent Searching and Reasoning

Large language models (LLMs) are increasingly deployed as intelligent agents that reason, plan, and interact with their environments. To effectively scale to long-horizon scenarios, a key capability for such agents is a memory mechanism that can retain, organize, and retrieve past experiences to support downstream decision-making....

💬 0 commentsarXiv:2601.04726v1PDF
0

Posted in cs.LG · 2026-01-08 · Jiyuan Zhang, Yining Liu, Siqi Yan, Lisen Deng, Jennifer Cao, Shuqi Yang, Min Ni, Bi Xue, Shen Li

MoEBlaze: Breaking the Memory Wall for Efficient MoE Training on Modern GPUs

The pervasive "memory wall" bottleneck is significantly amplified in modern large-scale Mixture-of-Experts (MoE) architectures. MoE's inherent architectural sparsity leads to sparse arithmetic compute and also introduces substantial activation memory overheads -- driven by large token routing buffers and the need to materialize and...

💬 0 commentsarXiv:2601.05296v1PDF
0

Posted in cs.IT · 2026-01-08 · Zhenyu Li, Ozan Alp Topal, Özlem Tuğfe Demir, Emil Björnson, Cicek Cavdar

Feasibility Study Regarding Self-sustainable Reconfigurable Intelligent Surfaces

Without requiring operational costs such as cabling and powering while maintaining reconfigurable phase-shift capability, self-sustainable reconfigurable intelligent surfaces (ssRISs) can be deployed in locations inaccessible to conventional relays or base stations, offering a novel approach to enhance wireless coverage. This study...

💬 0 commentsarXiv:2601.04723v1PDF
0

Posted in cs.DB · 2026-01-08 · Chrysanthi Kosyfaki, Ruiyuan Zhang, Nikos Mamoulis, Xiaofang Zhou

Toward Temporal Attribution Analytics in Dataflows

Data provenance (the process of determining the origin and derivation of data outputs) has applications across multiple domains including explaining database query results and auditing scientific workflows. Despite decades of research, provenance tracing remains challenging due to its high computational cost and storage requirements....

💬 0 commentsarXiv:2601.04722v3PDF
0

Posted in cs.CL · 2026-01-08 · Mingxin Li, Yanzhao Zhang, Dingkun Long, Keqin Chen, Sibo Song, Shuai Bai, Zhibo Yang, Pengjun Xie, An Yang, Dayiheng Liu, Jingren Zhou, Junyang Lin

Qwen3-VL-Embedding and Qwen3-VL-Reranker: A Unified Framework for State-of-the-Art Multimodal Retrieval and Ranking

In this report, we introduce the Qwen3-VL-Embedding and Qwen3-VL-Reranker model series, the latest extensions of the Qwen family built on the Qwen3-VL foundation model. Together, they provide an end-to-end pipeline for high-precision multimodal search by mapping diverse modalities, including text, images, document images, and video,...

💬 0 commentsarXiv:2601.04720v2PDF
0

Posted in cs.LG · 2026-01-08 · Maanas Taneja, Purab Shingvi

GPU-Accelerated INT8 Quantization for KV Cache Compression in Large Language Models

The key-value (KV) cache in large language models presents a significant memory bottleneck during inference, growing linearly with sequence length and often exceeding the memory footprint of model weights themselves. We implement and evaluate GPU-accelerated INT8 quantization for KV cache compression, achieving 4$\times$ memory...

💬 0 commentsarXiv:2601.04719v1PDF
0

Posted in cs.CL · 2026-01-08 · Yonghyun Jun, Junhyuk Choi, Jeonghyun Park, Jihyeong Park, Liu Nicole Geumheon, Hwanhee Lee

Identifying and Mitigating Bottlenecks in Role-Playing Agents: A Systematic Study of Disentangling Character Profile Axes

While Large Language Model (LLM) role-playing agents have advanced rapidly, it remains unclear which profile elements genuinely drive role-playing quality. To bridge this gap, we introduce a systematic diagnostic framework that disentangles the impact of character profiles along three axes: Familiarity (Known vs. Unknown), Structure...

💬 0 commentsarXiv:2601.04716v3PDF
0

Posted in cs.CV · 2026-01-08 · Xiao Guo, Jie Zhu, Anil Jain, Xiaoming Liu

On the Holistic Approach for Detecting Human Image Forgery

The rapid advancement of AI-generated content (AIGC) has escalated the threat of deepfakes, from facial manipulations to the synthesis of entire photorealistic human bodies. However, existing detection methods remain fragmented, specializing either in facial-region forgeries or full-body synthetic images, and consequently fail to...

💬 0 commentsarXiv:2601.04715v1PDF
0

Posted in cs.AI · 2026-01-08 · Chang Zhao, Zheming Yang, Yunqing Hu, Qi Guo, Zijian Wang, Pengcheng Li, Wen Ji

ThinkDrive: Chain-of-Thought Guided Progressive Reinforcement Learning Fine-Tuning for Autonomous Driving

With the rapid advancement of large language models (LLMs) technologies, their application in the domain of autonomous driving has become increasingly widespread. However, existing methods suffer from unstructured reasoning, poor generalization, and misalignment with human driving intent. While Chain-of-Thought (CoT) reasoning...

💬 0 commentsarXiv:2601.04714v1PDF
0

Posted in cs.CL · 2026-01-08 · Anh Thi-Hoang Nguyen, Khanh Quoc Tran, Tin Van Huynh, Phuoc Tan-Hoang Nguyen, Cam Tan Nguyen, Kiet Van Nguyen

DSC2025 -- ViHallu Challenge: Detecting Hallucination in Vietnamese LLMs

The reliability of large language models (LLMs) in production environments remains significantly constrained by their propensity to generate hallucinations -- fluent, plausible-sounding outputs that contradict or fabricate information. While hallucination detection has recently emerged as a priority in English-centric benchmarks,...

💬 0 commentsarXiv:2601.04711v1PDF
0

Posted in cs.CL · 2026-01-08 · Feihu Jin, Shipeng Cen, Ying Tan

Steering the Noise: Turning Random Perturbations into Effective Descent for Memory-Efficient LLM Fine-Tuning

Fine-tuning large language models (LLMs) achieves strong performance but is often limited by the memory overhead of backpropagation. Zeroth-order (ZO) optimization avoids this overhead by estimating gradients through forward passes alone, yet it typically converges slowly because random Gaussian perturbations yield high-variance...

💬 0 commentsarXiv:2601.04710v2PDF