Qwen Councils

Computer Science

arXiv preprints from January 1, 2026 through July 21, 2026 — 10:46:00 EST

0

Posted in cs.LG · 2026-01-11 · Sergii Kavun

NOVAK: Unified adaptive optimizer for deep neural networks

This work introduces NOVAK, a modular gradient-based optimization algorithm that integrates adaptive moment estimation, rectified learning-rate scheduling, decoupled weight regularization, multiple variants of Nesterov momentum, and lookahead synchronization into a unified, performance-oriented framework. NOVAK adopts a dual-mode...

💬 0 commentsarXiv:2601.07876v1PDF
0

Posted in cs.AI · 2026-01-11 · Jikai Chen, Long Chen, Dong Wang, Qinglin Su, Zhixuan Chu, Bingguang Hao, Leilei Gan, Chenyi Zhuang, Jinjie Gu

V2P: Visual Attention Calibration for GUI Grounding via Background Suppression and Center Peaking

Precise localization of GUI elements is crucial for the development of GUI agents. Traditional methods rely on bounding box or center-point regression, neglecting spatial interaction uncertainty and visual-semantic hierarchies. Recent methods incorporate attention mechanisms but still face two key issues: (1) ignoring processing...

💬 0 commentsarXiv:2601.06899v2PDF
0

Posted in cs.ET · 2026-01-11 · Sonia Yeh, Rishabh Ghotge, Yujia Shi, Luka de Koe

Resilience by Design: A KPI for Heavy-Duty Megawatt Charging

We introduce a stressor-agnostic Resilience Key Performance Indicator (Resilience KPI) for megawatt charging stations (MSC) serving heavy-duty vehicles. Beyond routine performance statistics (e.g., availability, throughput), the KPI quantifies a site's ability to anticipate, operate under degradation, and recover from disruptions...

💬 0 commentsarXiv:2601.06898v2PDF
0

Posted in cs.ET · 2026-01-11 · Sonia Yeh, Christopher Dirzka, Aleksandr Kondratenko, Frans Libertson, Benedicte Madon

How Do Ports Organise Innovation? Linking Port Governance, Ownership, and Living Labs

Ports are pivotal to decarbonisation and resilience, yet port studies rarely examine how ownership and decision rights shape the process and outcomes of sustainability and digital pilots. Living Lab (LL) scholarship offers strong concepts, but limited sector-grounded explanation of LL-governance fit in ports. We develop and apply a...

💬 0 commentsarXiv:2601.06894v1PDF
0

Posted in cs.CV · 2026-01-11 · Nimrod Shabtay, Itamar Zimerman, Eli Schwartz, Raja Giryes

CLIMP: Contrastive Language-Image Mamba Pretraining

Contrastive Language-Image Pre-training (CLIP) relies on Vision Transformers whose attention mechanism is susceptible to spurious correlations, and scales quadratically with resolution. To address these limitations, We present CLIMP, the first fully Mamba-based contrastive vision-language model that replaces both the vision and text...

💬 0 commentsarXiv:2601.06891v2PDF
0

Posted in cs.RO · 2026-01-11 · Yin Zhang, Zian Ning, Shiyu Zhao

Observability-Enhanced Target Motion Estimation via Bearing-Box: Theory and MAV Applications

Monocular vision-based target motion estimation is a fundamental challenge in numerous applications. This work introduces a novel bearing-box approach that fully leverages modern 3D detection measurements that are widely available nowadays but have not been well explored for motion estimation so far. Unlike existing methods that rely...

💬 0 commentsarXiv:2601.06887v1PDF
0

Posted in cs.DC · 2026-01-11 · Xuanzhengbo Ren, Yuta Kawai, Tetsuya Hoshino, Hirofumi Tomita, Takahiro Katagiri, Daichi Mukunoki, Seiya Nishizawa

Learning-Augmented Performance Model for Tensor Product Factorization in High-Order FEM

Accurate performance prediction is essential for optimizing scientific applications on modern high-performance computing (HPC) architectures. Widely used performance models primarily focus on cache and memory bandwidth, which is suitable for many memory-bound workloads. However, it is unsuitable for highly arithmetic intensive cases...

💬 0 commentsarXiv:2601.06886v1PDF
0

Posted in cs.ET · 2026-01-11 · Jinwoo Hwang, Yeongmin Hwang, Tadiwos Meaza, Hyeonbin Bae, Jongse Park

Understanding the Performance Behaviors of End-to-End Protein Design Pipelines on GPUs

Recent computational advances enable protein design pipelines to run end-to-end on GPUs, yet their heterogeneous computational behaviors remain undercharacterized at the system level. We implement and profile a representative pipeline at both component and full-pipeline granularities across varying inputs and hyperparameters. Our...

💬 0 commentsarXiv:2601.06885v1PDF
0

Posted in cs.CL · 2026-01-11 · Masahiro Kaneko

Paraphrasing Adversarial Attack on LLM-as-a-Reviewer

The use of large language models (LLMs) in peer review systems has attracted growing attention, making it essential to examine their potential vulnerabilities. Prior attacks rely on prompt injection, which alters manuscript content and conflates injection susceptibility with evaluation robustness. We propose the Paraphrasing...

💬 0 commentsarXiv:2601.06884v1PDF
0

Posted in cs.CV · 2026-01-11 · Xinhang Liu, Jiawei Shi, Zheng Dang, Yuchao Dai

MixRI: Mixing Features of Reference Images for Novel Object Pose Estimation

We present MixRI, a lightweight network that solves the CAD-based novel object pose estimation problem in RGB images. It can be instantly applied to a novel object at test time without finetuning. We design our network to meet the demands of real-world applications, emphasizing reduced memory requirements and fast inference time....

💬 0 commentsarXiv:2601.06883v1PDF
0

Posted in cs.CV · 2026-01-11 · Yi Wang, Yinfeng Yu, Bin Ren

Residual Cross-Modal Fusion Networks for Audio-Visual Navigation

Audio-visual embodied navigation aims to enable an agent to autonomously localize and reach a sound source in unseen 3D environments by leveraging auditory cues. The key challenge of this task lies in effectively modeling the interaction between heterogeneous features during multimodal fusion, so as to avoid single-modality dominance...

💬 0 commentsarXiv:2601.08868v1PDF
0

Posted in cs.HC · 2026-01-11 · Donghuo Zeng, Roberto Legaspi, Kazushi Ikeda

Personality-Aware Reinforcement Learning for Persuasive Dialogue with LLM-Driven Simulation

Effective persuasive dialogue agents adapt their strategies to individual users, accounting for the evolution of their psychological states and intentions throughout conversations. We present a personality-aware reinforcement learning approach comprising three main modules: (1) a Strategy-Oriented Interaction Framework, which serves...

💬 0 commentsarXiv:2601.06877v1PDF
0

Posted in cs.AI · 2026-01-11 · Sontaga G. Forane, Absalom E. Ezugwu, Kevin Igwe, Karen van den Berg

An Ubuntu-Guided Large Language Model Framework for Cognitive Behavioral Mental Health Dialogue

South Africa's escalating mental health crisis, compounded by limited access to culturally responsive care, calls for innovative and contextually grounded interventions. While large language models show considerable promise for mental health support, their predominantly Western-centric training data limit cultural and linguistic...

💬 0 commentsarXiv:2601.06875v1PDF
0

Posted in cs.CV · 2026-01-11 · Changli Wu, Haodong Wang, Jiayi Ji, Yutian Yao, Chunsai Du, Jihua Kang, Yanwei Fu, Liujuan Cao

MVGGT: Multimodal Visual Geometry Grounded Transformer for Multiview 3D Referring Expression Segmentation

Most existing 3D referring expression segmentation (3DRES) methods rely on dense, high-quality point clouds, while real-world agents such as robots and mobile phones operate with only a few sparse RGB views and strict latency constraints. We introduce Multi-view 3D Referring Expression Segmentation (MV-3DRES), where the model must...

💬 0 commentsarXiv:2601.06874v3PDF
0

Posted in cs.IR · 2026-01-11 · Mustafa Abdool, Soumyadip Banerjee, Moutupsi Paul, Do-kyum Kim, Xioawei Liu, Bin Xu, Tracy Yu, Hui Gao, Karen Ouyang, Huiji Gao, Liwei He, Stephanie Moyerman, Sanjeev Katariya

Applying Embedding-Based Retrieval to Airbnb Search

The goal of Airbnb search is to match guests with the ideal accommodation that fits their travel needs. This is a challenging problem, as popular search locations can have around a hundred thousand available homes, and guests themselves have a wide variety of preferences. Furthermore, the launch of new product features, such as...

💬 0 commentsarXiv:2601.06873v1PDF
0

Posted in cs.LG · 2026-01-11 · Jiazhang Liang, Jianheng Dai, Miaosen Luo, Menghua Jiang, Sijie Mai

QASA: Quality-Aware Semantic Augmentation for Robust Multimodal Sentiment Analysis

Multimodal large language models have demonstrated strong ability in capturing semantic representations for multimodal sentiment analysis. Their capacity to learn stable and generalizable multimodal features is limited, however, by the scarcity of high-quality training data. To address this, we propose QASA (Quality-Aware Semantic...

💬 0 commentsarXiv:2601.06870v2PDF
0

Posted in cs.LG · 2026-01-11 · Shiyuan Zhang, Yilai Liu, Yuwei Du, Ruoxuan Yang, Dong In Kim, Hongyang Du

U-MASK: User-adaptive Spatio-Temporal Masking for Personalized Mobile AI Applications

Personalized mobile artificial intelligence applications are widely deployed, yet they are expected to infer user behavior from sparse and irregular histories under a continuously evolving spatio-temporal context. This setting induces a fundamental tension among three requirements, i.e., immediacy to adapt to recent behavior,...

💬 0 commentsarXiv:2601.06867v1PDF
0

Posted in cs.CR · 2026-01-11 · Li Bai, Junxu Liu, Sen Zhang, Xinwei Zhang, Qingqing Ye, Haibo Hu

United We Defend: Collaborative Membership Inference Defenses in Federated Learning

Membership inference attacks (MIAs), which determine whether a specific data point was included in the training set of a target model, have posed severe threats in federated learning (FL). Unfortunately, existing MIA defenses, typically applied independently to each client in FL, are ineffective against powerful trajectory-based MIAs...

💬 0 commentsarXiv:2601.06866v1PDF
0

Posted in cs.CR · 2026-01-11 · Michael Sidorov, Ofer Hadar

Learning QoE from Packet-Level Measurements in Encrypted Video Conferencing Traffic

The quality of the user experience has become one of the most important aspects in todays world, as it directly influences individuals willingness to continue using or abandon a product or service. In this context, video conferencing applications (VCAs), which experienced widespread adoption following the COVID-19 pandemic, must...

💬 0 commentsarXiv:2601.06862v2PDF
0

Posted in cs.CL · 2026-01-11 · William Guey, Wei Zhang, Pei-Luen Patrick Rau, Pierrick Bougault, Vitor D. de Moura, Bertan Ucar, Jose O. Gomes

BiasLab: A Multilingual Dual-Framing Framework for LLM Bias Measurement, Applied to Workplace and HR Contexts

Background: Large language models (LLMs) harbor systematic biases that are particularly consequential in workplace and HR contexts, where their outputs increasingly influence hiring, job design, and organizational decisions. Existing bias-evaluation approaches remain methodologically fragmented, limiting practitioners' ability to...

💬 0 commentsarXiv:2601.06861v2PDF
0

Posted in cs.AI · 2026-01-11 · Yifei Chen, Guanting Dong, Zhicheng Dou

ET-Agent: Incentivizing Effective Tool-Integrated Reasoning Agent via Behavior Calibration

Large Language Models (LLMs) can extend their parameter knowledge limits by adopting the Tool-Integrated Reasoning (TIR) paradigm. However, existing LLM-based agent training framework often focuses on answers' accuracy, overlooking specific alignment for behavior patterns. Consequently, agent often exhibits ineffective actions during...

💬 0 commentsarXiv:2601.06860v2PDF
0

Posted in cs.LG · 2026-01-11 · Xin Ye, Daning Cheng, Boyang Zhang, Yunquan Zhang

MoE-DisCo:Low Economy Cost Training Mixture-of-Experts Models

Training large-scale Mixture-of-Experts (MoE) models typically requires high-memory, high-bandwidth GPUs (e.g., A100), and their high cost has become a major barrier to large-model training. In contrast, affordable hardware is low-cost but constrained by memory capacity and bandwidth, making it unsuitable for direct LLM training. To...

💬 0 commentsarXiv:2601.06857v1PDF
0

Posted in cs.RO · 2026-01-11 · Luigi Romano, Ole Morten Aamo, Jan Åslund, Erik Frisk

Semilinear single-track vehicle models with distributed tyre friction dynamics

This paper introduces a novel family of single-track vehicle models that incorporate a distributed representation of transient tyre dynamics, whilst simultaneously accounting for nonlinear effects induced by friction. The core of the proposed framework is represented by the distributed Friction with Bristle Dynamics (FrBD) model,...

💬 0 commentsarXiv:2601.06854v2PDF
0

Posted in cs.CL · 2026-01-11 · Zabir Al Nazi, Shubhashis Roy Dipta, Sudipta Kar

†DAGGER: Distractor-Aware Graph Generation for Executable Reasoning in Math Problems

Chain-of-Thought (CoT) prompting is widely adopted for mathematical problem solving, including in low-resource languages, yet its behavior under irrelevant context remains underexplored. To systematically study this challenge, we introduce DISTRACTMATH-BN, a Bangla benchmark that augments MGSM and MSVAMP with semantically coherent but...

💬 0 commentsarXiv:2601.06853v2PDF