Qwen Councils

Computer Science

arXiv preprints from January 1, 2026 through July 20, 2026 — 03:09:04 EST

0

Posted in cs.SE · 2026-01-15 · Niko Usai, Dario Montagnini, Kristian Ilianov Iliev, Raffaele Camanzo

LogicLens: Leveraging Semantic Code Graph to explore Multi Repository large systems

Understanding large software systems is a challenging task, especially when code is distributed across multiple repositories and microservices. Developers often need to reason not only about the structure of the code, but also about its domain logic and runtime behaviors, which are typically implicit and scattered. We introduce...

💬 0 commentsarXiv:2601.10773v1PDF
0

Posted in cs.IT · 2026-01-15 · Yongcheng Yang, Minquan Cheng, Kai Wan, Giuseppe Caire

A New Construction Structure on Coded Caching with Linear Subpacketization: Non-Half-Sum Latin Rectangle

Coded caching is recognized as an effective method for alleviating network congestion during peak periods by leveraging local caching and coded multicasting gains. The key challenge in designing coded caching schemes lies in simultaneously achieving low subpacketization and low transmission load. Most existing schemes require...

💬 0 commentsarXiv:2601.10505v1PDF
0

Posted in cs.CL · 2026-01-15 · Yiwen Gao, Ruochen Zhao, Yang Deng, Wenxuan Zhang

DR-Arena: an Automated Evaluation Framework for Deep Research Agents

As Large Language Models (LLMs) increasingly operate as Deep Research (DR) Agents capable of autonomous investigation and information synthesis, reliable evaluation of their task performance has become a critical bottleneck. Current benchmarks predominantly rely on static datasets, which suffer from several limitations: limited task...

💬 0 commentsarXiv:2601.10504v3PDF
0

Posted in cs.CV · 2026-01-15 · Miriam Doh, Aditya Gulati, Corinna Canali, Nuria Oliver

Aesthetics as Structural Harm: Algorithmic Lookism Across Text-to-Image Generation and Classification

This paper examines algorithmic lookism-the systematic preferential treatment based on physical appearance-in text-to-image (T2I) generative AI and a downstream gender classification task. Through the analysis of 26,400 synthetic faces created with Stable Diffusion 2.1 and 3.5 Medium, we demonstrate how generative AI models...

💬 0 commentsarXiv:2601.11651v2PDF
0

Posted in cs.IT · 2026-01-15 · Dhruv Pratap Singh, Anjana A. Mahesh, B. Sundar Rajan

Coded Caching for Combinatorial Multi-Access Hotplug Networks from $t$-Designs

We study hotplug coded caching in combinatorial multi-access networks, which generalizes existing hotplug coded caching models by allowing users to access multiple caches, while only a subset of caches is online during the delivery phase. We first generalize the Hotplug Placement Delivery Array (HpPDA) framework to the combinatorial...

💬 0 commentsarXiv:2601.10503v1PDF
0

Posted in cs.SI · 2026-01-15 · Jiaze Li, Michael T. Schaub, Leto Peel

Higher order trade-offs in hypergraph community detection

Extending community detection from pairwise networks to hypergraphs introduces fundamental theoretical challenges. Hypergraphs exhibit structural heterogeneity with no direct graph analogue: hyperedges of varying orders can connect nodes across communities in diverse configurations, introducing new trade-offs in defining and detecting...

💬 0 commentsarXiv:2601.10502v2PDF
0

Posted in cs.LG · 2026-01-15 · Nilin Abrahamsen

PROMA: Projected Microbatch Accumulation for Reference-Free Proximal Policy Updates

This note introduces Projected Microbatch Accumulation (PROMA), a reference-free proximal policy method that controls KL divergence by projecting away high-variance components of the policy gradient. Two variants are presented. In the accumulation-based variant, the running gradient is projected orthogonal to the sequence-wise...

💬 0 commentsarXiv:2601.10498v4PDF
0

Posted in cs.CV · 2026-01-15 · Wenqing Wang, Da Li, Xiatian Zhu, Josef Kittler

MERGETUNE: Continued Fine-Tuning of Vision-Language Models

Fine-tuning vision-language models (VLMs) such as CLIP often leads to catastrophic forgetting of pretrained knowledge. Prior work primarily aims to mitigate forgetting during adaptation; however, forgetting often remains inevitable during this process. We introduce a novel paradigm, continued fine-tuning (CFT), which seeks to recover...

💬 0 commentsarXiv:2601.10497v3PDF
0

Posted in cs.SE · 2026-01-15 · Ali Al-Kaswan, Claudio Spiess, Prem Devanbu, Arie van Deursen, Maliheh Izadi

Model See, Model Do? Exposure-Aware Evaluation of Bug-vs-Fix Preference in Code LLMs

Large language models are increasingly used for code generation and debugging, but their outputs can still contain bugs, that originate from training data. Distinguishing whether an LLM prefers correct code, or a familiar incorrect version might be influenced by what it's been exposed to during training. We introduce an exposure-aware...

💬 0 commentsarXiv:2601.10496v1PDF
0

Posted in cs.LG · 2026-01-15 · Shenlong Zheng, Zhen Zhang, Yuhui Deng, Geyong Min, Lin Cui

Communication-Efficient Federated Learning by Exploiting Spatio-Temporal Correlations of Gradients

Communication overhead is a critical challenge in federated learning, particularly in bandwidth-constrained networks. Although many methods have been proposed to reduce communication overhead, most focus solely on compressing individual gradients, overlooking the temporal correlations among them. Prior studies have shown that...

💬 0 commentsarXiv:2601.10491v1PDF
0

Posted in cs.LO · 2026-01-15 · Mirco A. Mannucci, Corey Thuro

Resource-Bounded Martin-Löf Type Theory: Compositional Cost Analysis for Dependent Types

We extend resource-bounded type theory to Martin-Lof type theory (MLTT) with dependent types, enabling size-indexed cost bounds for programs over inductive families. We introduce a resource-indexed universe hierarchy U_r where r is an element of L and tracks the cost of type formation, and a graded modality Box_r for feasibility...

💬 0 commentsarXiv:2601.10772v1PDF
0

Posted in cs.AI · 2026-01-15 · Runhao Zhao, Weixin Zeng, Wentao Zhang, Chong Chen, Zhengpin Li, Xiang Zhao, Lei Chen

Panning for Gold: Expanding Domain-Specific Knowledge Graphs with General Knowledge

Domain-specific knowledge graphs (DKGs) are critical yet often suffer from limited coverage compared to General Knowledge Graphs (GKGs). Existing tasks to enrich DKGs rely primarily on extracting knowledge from external unstructured data or completing KGs through internal reasoning, but the scope and quality of such integration remain...

💬 0 commentsarXiv:2601.10485v3PDF
0

Posted in cs.IT · 2026-01-15 · Siying Luo, Youlong Wu, Mingming Zhang, Minquan Cheng, Dianhua Wu

A Construction Framework of Coded Caching Scheme for Multi-Access MISO Systems via Knapsack Problem

This paper investigates the coded caching problem in a multi-access multiple-input single-output (MAMISO) network with the combinatorial topology. The considered system consists of a server containing $N$ files, $Λ$ cache nodes, and $K$ cache-less users, where each user can access a unique subset of $r$ cache nodes. The server is...

💬 0 commentsarXiv:2601.10484v2PDF
0

Posted in cs.CL · 2026-01-15 · Tiziano Labruna, Arkadiusz Modzelewski, Giorgio Satta, Giovanni Da San Martino

Detecting Winning Arguments with Large Language Models and Persuasion Strategies

Detecting persuasion in argumentative text is a challenging task with important implications for understanding human communication. This work investigates the role of persuasion strategies - such as Attack on reputation, Distraction, and Manipulative wording - in determining the persuasiveness of a text. We conduct experiments on...

💬 0 commentsarXiv:2601.10660v1PDF
0

Posted in cs.NE · 2026-01-15 · Minghao Yan, Bo Peng, Benjamin Coleman, Ziqi Chen, Zhouhang Xie, Shuo Chen, Zhankui He, Noveen Sachdeva, Isabella Ye, Weili Wang, Chi Wang, Ed H. Chi, Fernando Pereira, Wang-Cheng Kang, Derek Zhiyuan Cheng, Beidou Wang

PACEvolve: Enabling Long-Horizon Progress-Aware Consistent Evolution

Large Language Models (LLMs) have emerged as powerful operators for evolutionary search, yet the design of efficient search scaffolds remains ad hoc. While promising, current LLM-in-the-loop systems lack a systematic approach to managing the evolutionary process. We identify three distinct failure modes: Context Pollution, where...

💬 0 commentsarXiv:2601.10657v2PDF
0

Posted in cs.AI · 2026-01-15 · Christoph Weinhuber, Yannik Schnitzer, Alessandro Abate, David Parker, Giuseppe De Giacomo, Moshe Y. Vardi

Multi-Property Synthesis

We study LTLf synthesis with multiple properties, where satisfying all properties may be impossible. Instead of enumerating subsets of properties, we compute in one fixed-point computation the relation between product-game states and the goal sets that are realizable from them, and we synthesize strategies achieving maximal realizable...

💬 0 commentsarXiv:2601.10651v1PDF
0

Posted in cs.CL · 2026-01-15 · Jinghan Cao, Qingyang Ren, Xiangyun Chen, Xinjin Li, Haoxiang Gao, Yu Zhao

Slang Context-based Inference Enhancement via Greedy Search-Guided Chain-of-Thought Prompting

Slang interpretation has been a challenging downstream task for Large Language Models (LLMs) as the expressions are inherently embedded in contextual, cultural, and linguistic frameworks. In the absence of domain-specific training data, it is difficult for LLMs to accurately interpret slang meaning based on lexical information. This...

💬 0 commentsarXiv:2603.13230v1PDF
0

Posted in cs.CV · 2026-01-15 · Darshan Singh, Arsha Nagrani, Kawshik Manikantan, Harman Singh, Dinesh Tewari, Tobias Weyand, Cordelia Schmid, Anelia Angelova, Shachi Dave

MINERVA-Cultural: A Benchmark for Cultural and Multilingual Long Video Reasoning

Recent advancements in video models have shown tremendous progress, particularly in long video understanding. However, current benchmarks predominantly feature western-centric data and English as the dominant language, introducing significant biases in evaluation. To address this, we introduce MINERVA-Cultural, a challenging benchmark...

💬 0 commentsarXiv:2601.10649v2PDF
0

Posted in cs.IT · 2026-01-15 · Joseph Rowan, Buu Phan, Ashish Khisti

One-Shot Broadcast Joint Source-Channel Coding with Codebook Diversity

We study a one-shot joint source-channel coding setting where the source is encoded once and broadcast to $K$ decoders through independent channels. Success is predicated on at least one decoder recovering the source within a maximum distortion constraint. We find that in the one-shot regime, utilizing disjoint codebooks at each...

💬 0 commentsarXiv:2601.10648v3PDF
0

Posted in cs.CV · 2026-01-15 · Kaustubh Shivshankar Shejole, Gaurav Mishra

PSSI-MaxST: An Efficient Pixel-Segment Similarity Index Using Intensity and Smoothness Features for Maximum Spanning Tree Based Segmentation

Interactive graph-based segmentation methods partition an image into foreground and background regions with the aid of user inputs. However, existing approaches often suffer from high computational costs, sensitivity to user interactions, and degraded performance when the foreground and background share similar color distributions. A...

💬 0 commentsarXiv:2601.11654v1PDF
0

Posted in cs.CL · 2026-01-15 · Yuxi Xia, Loris Schoenegger, Benjamin Roth

Influential Training Data Retrieval for Explaining Verbalized Confidence of LLMs

Large language models (LLMs) can increase users' perceived trust by verbalizing confidence in their outputs. However, prior work has shown that LLMs are often overconfident, making their stated confidence unreliable since it does not consistently align with factual accuracy. To better understand the sources of this verbalized...

💬 0 commentsarXiv:2601.10645v1PDF
0

Posted in cs.IR · 2026-01-15 · Eugene Yang, Andrew Yates, Dawn Lawrie, James Mayfield, Trevor Adriaanse

RoutIR: Fast Serving of Retrieval Pipelines for Retrieval-Augmented Generation

Retrieval models are key components of Retrieval-Augmented Generation (RAG) systems, which generate search queries, process the documents returned, and generate a response. RAG systems are often dynamic and may involve multiple rounds of retrieval. While many state-of-the-art retrieval methods are available through academic IR...

💬 0 commentsarXiv:2601.10644v1PDF
0

Posted in cs.IT · 2026-01-15 · Chandan Anand, Jayesh Seshadri, Prasad Krishnan, Gowtham R. Kurri

Converse Bounds for Sun-Jafar-type Weak Private Information Retrieval

Building on the well-established capacity-achieving schemes of Sun-Jafar (for replicated storage) and the closely related scheme of Banawan-Ulukus (for MDS-coded setting), a recent work by Anand et al. proposed new classes of weak private information retrieval (WPIR) schemes for the collusion-free (replication and MDS-coded) setting,...

💬 0 commentsarXiv:2601.10643v2PDF
0

Posted in cs.LG · 2026-01-15 · Ranajoy Sadhukhan, Sheng Cao, Harry Dong, Changsheng Zhao, Attiano Purpura-Pontoniere, Yuandong Tian, Zechun Liu, Beidi Chen

STEM: Scaling Transformers with Embedding Modules

Fine-grained sparsity promises higher parametric capacity without proportional per-token compute, but often suffers from training instability, load balancing, and communication overhead. We introduce STEM (Scaling Transformers with Embedding Modules), a static, token-indexed approach that replaces the FFN up-projection with a...

💬 0 commentsarXiv:2601.10639v1PDF
0

Posted in cs.CV · 2026-01-15 · Chengfeng Zhao, Jiazhi Shu, Yubo Zhao, Tianyu Huang, Jiahao Lu, Zekai Gu, Chengwei Ren, Zhiyang Dou, Qing Shuai, Yuan Liu

CoMoVi: Co-Generation of 3D Human Motions and Realistic Videos

In this paper, we find that the generation of 3D human motions and 2D human videos is intrinsically coupled. 3D motions provide the structural prior for plausibility and consistency in videos, while pre-trained video models offer strong generalization capabilities for motions. Based on this, we present CoMoVi, a co-generative...

💬 0 commentsarXiv:2601.10632v2PDF