Qwen Councils
arXiv is taking too long to respond. Please try again or narrow your search.
Showing downloaded papers while arXiv is unavailable.

Computer Science

arXiv preprints from January 1, 2026 through July 20, 2026 — 21:56:59 EST

0

Posted in cs.LG · 2026-01-12 · Simon Jegou, Maximilian Jeblick

KVzap: Fast, Adaptive, and Faithful KV Cache Pruning

Growing context lengths in transformer-based language models have made the key-value (KV) cache a critical inference bottleneck. While many KV cache pruning methods have been proposed, they have not yet been adopted in major inference engines due to speed--accuracy trade-offs. We introduce KVzap, a fast, input-adaptive approximation...

💬 0 commentsarXiv:2601.07891v2PDF
0

Posted in cs.RO · 2026-01-12 · Yun Chen, Bowei Huang, Fan Guo, Kang Song

Heterogeneous Multi-Expert Reinforcement Learning for Long-Horizon Multi-Goal Tasks in Autonomous Forklifts

Autonomous mobile manipulation in unstructured warehouses requires a balance between efficient large-scale navigation and high-precision object interaction. Traditional end-to-end learning approaches often struggle to handle the conflicting demands of these distinct phases. Navigation relies on robust decision-making over large...

💬 0 commentsarXiv:2601.07304v1PDF
0

Posted in cs.SD · 2026-01-12 · Xueping Zhang, Han Yin, Yang Xiao, Lin Zhang, Ting Dang, Rohan Kumar Das, Ming Li

ESDD2: Environment-Aware Speech and Sound Deepfake Detection Challenge Evaluation Plan

Audio recorded in real-world environments often contains a mixture of foreground speech and background environmental sounds. With rapid advances in text-to-speech, voice conversion, and other generation models, either component can now be modified independently. Such component-level manipulations are harder to detect, as the remaining...

💬 0 commentsarXiv:2601.07303v5PDF
0

Posted in cs.SE · 2026-01-12 · Nidhal Selmi, Jean-michel Bruel, Sébastien Mosser, Matthieu Crespo, Alain Kerbrat

Engineering Decisions in MBSE: Insights for a Decision Capture Framework Development

Decision-making is a core engineering design activity that conveys the engineer's knowledge and translates it into courses of action. Capturing this form of knowledge can reap potential benefits for the engineering teams and enhance development efficiency. Despite its clear value, traditional decision capture often requires a...

💬 0 commentsarXiv:2601.07301v1PDF
0

Posted in cs.CV · 2026-01-12 · Jianghao Yin, Qingbin Li, Kun Sun, Cheng Ding, Jie Wang, Qin Chen, Jie Zhou, Nan Wang, Changqing Li, Pei Wu, Jian Xu, Zheming Yang, Liang He

Mimic Human Cognition, Master Multi-Image Reasoning: A Meta-Action Framework for Enhanced Visual Understanding

While Multimodal Large Language Models (MLLMs) excel at single-image understanding, they exhibit significantly degraded performance in multi-image reasoning scenarios. Multi-image reasoning presents fundamental challenges including complex inter-relationships between images and scattered critical information across image sets....

💬 0 commentsarXiv:2601.07298v2PDF
0

Posted in cs.AI · 2026-01-12 · Yujin Zhou, Chuxue Cao, Jinluan Yang, Lijun Wu, Conghui He, Sirui Han, Yike Guo

LRAS: Advanced Legal Reasoning with Agentic Search

While Large Reasoning Models (LRMs) have demonstrated exceptional logical capabilities in mathematical domains, their application to the legal field remains hindered by the strict requirements for procedural rigor and adherence to legal logic. Existing legal LLMs, which rely on "closed-loop reasoning" derived solely from internal...

💬 0 commentsarXiv:2601.07296v1PDF
0

Posted in cs.CL · 2026-01-12 · Kei Saito

NRR-Phi: A Typed External Text-to-State Interface and Update Contract for Inspectable Ambiguity-State Maintenance

Ambiguity-bearing inputs reach downstream systems through interfaces that favor a single resolved response before later context arrives. Even when alternatives are externalized, their representation and relative activation depend on the update rule. We address this state-maintenance problem within Non-Resolution Reasoning (NRR) by...

💬 0 commentsarXiv:2601.19933v7PDF
0

Posted in cs.IR · 2026-01-12 · Wenhao Lai, Weike Pan, Zhong Ming

Towards Multi-Behavior Multi-Task Recommendation via Behavior-informed Graph Embedding Learning

Multi-behavior recommendation (MBR) aims to improve the performance w.r.t. the target behavior (i.e., purchase) by leveraging auxiliary behaviors (e.g., click, favourite). However, in real-world scenarios, a recommendation method often needs to process different types of behaviors and generate personalized lists for each task (i.e.,...

💬 0 commentsarXiv:2601.07294v1PDF
0

Posted in cs.CL · 2026-01-12 · Ruyuan Wan, Changye Li, Ting-Hao 'Kenneth' Huang

"Newspaper Eat" Means "Not Tasty": A Taxonomy and Benchmark for Coded Language in Real-World Chinese Online Reviews

Coded language is an important part of human communication. It refers to cases where users intentionally encode meaning so that the surface text differs from the intended meaning and must be decoded to be understood. Current language models handle coded language poorly. Progress has been limited by the lack of real-world datasets and...

💬 0 commentsarXiv:2601.19932v2PDF
0

Posted in cs.CV · 2026-01-12 · Weidong Tang, Xinyan Wan, Siyu Li, Xiumei Wang

Inference-Time Scaling for Visual AutoRegressive modeling by Searching Representative Samples

While inference-time scaling has significantly enhanced generative quality in large language and diffusion models, its application to vector-quantized (VQ) visual autoregressive modeling (VAR) remains unexplored. We introduce VAR-Scaling, the first general framework for inference-time scaling in VAR, addressing the critical challenge...

💬 0 commentsarXiv:2601.07293v1PDF
0

Posted in cs.CV · 2026-01-12 · Qi Zheng, Shuliang Liu, Yu Huang, Sihang Jia, Jungang Li, Lyuhao Chen, Junhao Chen, Hanqian Li, Aiwei Liu, Yibo Yan, Xuming Hu

A Visual Semantic Adaptive Watermark grounded by Prefix-Tuning for Large Vision-Language Model

Watermarking has emerged as a pivotal solution for content traceability and intellectual property protection in Large Vision-Language Models (LVLMs). However, vision-agnostic watermarks introduce visually irrelevant tokens and disrupt visual grounding by enforcing indiscriminate pseudo-random biases, while some semantic-aware methods...

💬 0 commentsarXiv:2601.07291v1PDF
0

Posted in cs.CV · 2026-01-12 · Jiapeng Shi, Junke Wang, Zuyao You, Bo He, Zuxuan Wu

VideoLoom: A Video Large Language Model for Joint Spatial-Temporal Understanding

This paper presents VideoLoom, a unified Video Large Language Model (Video LLM) for joint spatial-temporal understanding. To facilitate the development of fine-grained spatial and temporal localization capabilities, we curate LoomData-8.7k, a human-centric video dataset with temporally grounded and spatially localized captions. With...

💬 0 commentsarXiv:2601.07290v1PDF
0

Posted in cs.LG · 2026-01-12 · Yalan Tan, Yanyong Huang, Zongxin Shen, Dongjie Wang, Fengmao Lv, Tianrui Li

Kernel Alignment-based Multi-view Unsupervised Feature Selection with Sample-level Adaptive Graph Learning

Although multi-view unsupervised feature selection (MUFS) has demonstrated success in dimensionality reduction for unlabeled multi-view data, most existing methods reduce feature redundancy by focusing on linear correlations among features but often overlook complex nonlinear dependencies. This limits the effectiveness of feature...

💬 0 commentsarXiv:2601.07288v2PDF
0

Posted in cs.CV · 2026-01-12 · Yuanyang Yin, Yufan Deng, Shenghai Yuan, Kaipeng Zhang, Xiao Yang, Feng Zhao

Focal Guidance: Unlocking Controllability from Semantic-Weak Layers in Video Diffusion Models

The task of Image-to-Video (I2V) generation aims to synthesize a video from a reference image and a text prompt. This requires diffusion models to reconcile high-frequency visual constraints and low-frequency textual guidance during the denoising process. However, while existing I2V models prioritize visual consistency, how to...

💬 0 commentsarXiv:2601.07287v1PDF
0

Posted in cs.RO · 2026-01-12 · Haoyu Zhang, Shibo Jin, Lusong Li, Jun Li, Liang Lin, Xiaodong He, Zecui Zeng

AdaMorph: Unified Motion Retargeting via Embodiment-Aware Adaptive Transformers

Retargeting human motion to heterogeneous robots is a fundamental challenge in robotics, primarily due to the severe kinematic and dynamic discrepancies between varying embodiments. Existing solutions typically resort to training embodiment-specific models, which scales poorly and fails to exploit shared motion semantics. To address...

💬 0 commentsarXiv:2601.07284v2PDF
0

Posted in cs.CL · 2026-01-12 · Lucky Susanto, Musa Izzanardi Wijanarko, Khumaisa Nur'aini, Farid Adilazuarda, Alham Fikri Aji, Derry Tanti Wijaya

Does Visual Rendering Bypass Tokenization? Investigating Script-Tokenizer Misalignment in Pixel-Based Language Models

While pixel-based language modeling aims to bypass the sub-word tokenization bottleneck by rendering text as images, recent multimodal variants such as DualGPT reintroduce text tokenizers to improve autoregressive performance. We investigate a fundamental question, does visual rendering truly decouple a model from tokenization...

💬 0 commentsarXiv:2602.06973v1PDF
0

Posted in cs.CL · 2026-01-12 · Changzai Pan, Jie Zhang, Kaiwen Wei, Chenshuo Pan, Yu Zhao, Jingwang Huang, Jian Yang, Zhenhe Wu, Haoyang Zeng, Xiaoyan Gu, Weichao Sun, Yanbo Zhai, Yujie Mao, Zhuoru Jiang, Jiang Zhong, Shuangyong Song, Yongxiang Li, Zhongjiang He

ReasonTabQA: A Comprehensive Benchmark for Table Question Answering from Real World Industrial Scenarios

Recent advancements in Large Language Models (LLMs) have significantly catalyzed table-based question answering (TableQA). However, existing TableQA benchmarks often overlook the intricacies of industrial scenarios, which are characterized by multi-table structures, nested headers, and massive scales. These environments demand robust...

💬 0 commentsarXiv:2601.07280v1PDF
0

Posted in cs.GT · 2026-01-12 · Hodaya Barr, Eden Hartman, Yonatan Aumann, Sarit Kraus

Coalition Tactics: Bribery and Control in Parliamentary Elections

Strategic manipulation of elections is typically studied in the context of promoting individual candidates. In parliamentary elections, however, the focus shifts: voters may care more about the overall governing coalition than the individual parties' seat counts. This paper studies this new problem: manipulating parliamentary...

💬 0 commentsarXiv:2601.07279v1PDF
0

Posted in cs.CR · 2026-01-12 · Karthikeyan V. R., Premnath S., Kavinraaj S., J. Sangeetha

A High-Recall Cost-Sensitive Machine Learning Framework for Real-Time Online Banking Transaction Fraud Detection

Fraudulent activities on digital banking services are becoming more intricate by the day, challenging existing defenses. While older rule driven methods struggle to keep pace, even precision focused algorithms fall short when new scams are introduced. These tools typically overlook subtle shifts in criminal behavior, missing crucial...

💬 0 commentsarXiv:2601.07276v2PDF
0

Posted in cs.CL · 2026-01-12 · Kalvin Chang, Yiwen Shao, Jiahong Li, Dong Yu

Towards Comprehensive Semantic Speech Embeddings for Chinese Dialects

Despite having hundreds of millions of speakers, Chinese dialects lag behind Mandarin in speech and language technologies. Most varieties are primarily spoken, making dialect-to-Mandarin speech-LLMs (large language models) more practical than dialect LLMs. Building dialect-to-Mandarin speech-LLMs requires speech representations with...

💬 0 commentsarXiv:2601.07274v1PDF
0

Posted in cs.RO · 2026-01-12 · Yuxuan Hu, Kuangji Zuo, Boyu Ma, Shihao Li, Zhaoyang Xia, Feng Xu, Jianfei Yang

WaveMan: mmWave-Based Room-Scale Human Interaction Perception for Humanoid Robots

Reliable humanoid-robot interaction (HRI) in household environments is constrained by two fundamental requirements, namely robustness to unconstrained user positions and preservation of user privacy. Millimeter-wave (mmWave) sensing inherently supports privacy-preserving interaction, making it a promising modality for room-scale HRI....

💬 0 commentsarXiv:2601.07454v1PDF
0

Posted in cs.DL · 2026-01-12 · Snehasish Paul

Building Faculty Expertise Ontology using Protege: Enhancing Academic Library Research Services

Academic libraries struggle to find and access faculty expertise across disciplines. This research proposes a faculty expertise ontology with a hierarchical structure based on Protégé to enhance library services and knowledge organisation. The ontology classifies relationships between departments, subject areas, faculty members, and...

💬 0 commentsarXiv:2601.07451v1PDF
0

Posted in cs.IR · 2026-01-12 · Hao Jiang, Zhi Yang, Annan Wang, Yichi Zhang, Weisi Lin

RLPO: Residual Listwise Preference Optimization for Long-Context Review Ranking

Review ranking is pivotal in e-commerce for prioritizing diagnostic and authentic feedback from the deluge of user-generated content. While large language models have improved semantic assessment, existing ranking paradigms face a persistent trade-off in long-context settings. Pointwise scoring is efficient but often fails to account...

💬 0 commentsarXiv:2601.07449v2PDF
0

Posted in cs.CV · 2026-01-12 · Mahdi Chamseddine, Didier Stricker, Jason Rambach

PanoSAMic: Panoramic Image Segmentation from SAM Feature Encoding and Dual View Fusion

Existing image foundation models are not optimized for spherical images having been trained primarily on perspective images. PanoSAMic integrates the pre-trained Segment Anything (SAM) encoder to make use of its extensive training and integrate it into a semantic segmentation model for panoramic images using multiple modalities. We...

💬 0 commentsarXiv:2601.07447v3PDF
0

Posted in cs.LO · 2026-01-12 · Zhipeng Chen, Haolun Tang, Jingyi Zhan

Formalization of Amicable Numbers Theory

This paper presents a formalization of the theory of amicable numbers in the Lean~4 proof assistant. Two positive integers $m$ and $n$ are called an amicable pair if the sum of proper divisors of $m$ equals $n$ and the sum of proper divisors of $n$ equals $m$. Our formalization introduces the proper divisor sum function $\propersum(n)...

💬 0 commentsarXiv:2601.07444v1PDF