Qwen Councils
arXiv is taking too long to respond. Please try again or narrow your search.
Showing downloaded papers while arXiv is unavailable.

Computer Science

arXiv preprints from January 1, 2026 through July 20, 2026 — 23:33:21 EST

0

Posted in cs.CV · 2026-01-17 · Honglin Lin, Chonghan Qin, Zheng Liu, Qizhi Pei, Yu Li, Zhanping Zhong, Xin Gao, Yanfeng Wang, Conghui He, Lijun Wu

Scientific Image Synthesis: Benchmarking, Methodologies, and Downstream Utility

While synthetic data has proven effective for improving scientific reasoning in the text domain, multimodal reasoning remains constrained by the difficulty of synthesizing scientifically rigorous images. Existing Text-to-Image (T2I) models often produce outputs that are visually plausible yet scientifically incorrect, resulting in a...

💬 0 commentsarXiv:2601.17027v1PDF
0

Posted in cs.CV · 2026-01-17 · Xiaomei Yang, Antai Liu, Xizhan Gao, Fa Zhu, Sijie Niu, Giancarlo Fortino

Learning Language-Driven Sequence-Level Modal-Invariant Representations for Video-Based Visible-Infrared Person Re-Identification

The core of video-based visible-infrared person re-identification (VVI-ReID) lies in learning sequence-level modal-invariant representations across different modalities. Recent research tends to use modality-shared language prompts generated by CLIP to guide the learning of modal-invariant representations. Despite achieving optimal...

💬 0 commentsarXiv:2601.12062v2PDF
0

Posted in cs.CL · 2026-01-17 · Jinsook Lee, Kirk Vanacore, Zhuqian Zhou, Bakhtawar Ahtisham, Jeanine Grutter, Rene F. Kizilcec

Codebook-Injected Dialogue Segmentation for Multi-Utterance Constructs Annotation: LLM-Assisted and Gold-Label-Free Evaluation

Dialogue Act (DA) annotation typically treats communicative or pedagogical intent as localized to individual utterances or turns. This leads annotators to agree on the underlying action while disagreeing on segment boundaries, reducing apparent reliability. We propose codebook-injected segmentation, which conditions boundary decisions...

💬 0 commentsarXiv:2601.12061v2PDF
0

Posted in cs.CC · 2026-01-17 · Ismael Rodriguez, David Rubio, Fernando Rubio

Complexity of adaptive testing in scenarios defined extensionally

In this paper we consider a testing setting where the set of possible definitions of the Implementation Under Test (IUT), as well as the behavior of each of these definitions in all possible interactions, are extensionally defined, i.e., on an element-by-element and case-by-case basis. Under this setting, the problem of finding the...

💬 0 commentsarXiv:2601.12056v1PDF
0

Posted in cs.CV · 2026-01-17 · Lina Meyer, Felix Wissel, Tobias Knopp, Susanne Pfefferle, Ralf Fliegert, Maximilian Sandmann, Liana Uebler, Franziska Möckl, Björn-Philipp Diercks, David Lohr, René Werner

Automating Parameter Selection in Deep Image Prior for Fluorescence Microscopy Image Denoising via Similarity-Based Parameter Transfer

Unsupervised deep image prior (DIP) addresses shortcomings of training data requirements and limited generalization associated with supervised deep learning. The performance of DIP depends on the network architecture and the stopping point of its iterative process. Optimizing these parameters for a new image requires time, restricting...

💬 0 commentsarXiv:2601.12055v1PDF
0

Posted in cs.CV · 2026-01-17 · Zaiyan Zhang, Jie Li, Shaowei Shi, Qiangqiang Yuan

Task-Driven Prompt Learning: A Joint Framework for Multi-modal Cloud Removal and Segmentation

Optical remote sensing imagery is indispensable for Earth observation, yet persistent cloud occlusion limits its downstream utility. Most cloud removal (CR) methods are optimized for low-level fidelity and can over-smooth textures and boundaries that are critical for analysis-ready data (ARD), leading to a mismatch between visually...

💬 0 commentsarXiv:2601.12052v2PDF
0

Posted in cs.CV · 2026-01-17 · Weixin Ye, Wei Wang, Yahui Liu, Yue Song, Bin Ren, Wei Bi, Rita Cucchiara, Nicu Sebe

A Unified Masked Jigsaw Puzzle Framework for Vision and Language Models

In federated learning, Transformer, as a popular architecture, faces critical challenges in defending against gradient attacks and improving model performance in both Computer Vision (CV) and Natural Language Processing (NLP) tasks. It has been revealed that the gradient of Position Embeddings (PEs) in Transformer contains sufficient...

💬 0 commentsarXiv:2601.12051v1PDF
0

Posted in cs.IT · 2026-01-17 · Saeed Razavikia, Mohammad Kazemi, Deniz Gündüz, Carlo Fischione

Function Computation Over Multiple Access Channels via Hierarchical Constellations

We study function computation over a Gaussian multiple-access channel (MAC), where multiple transmitters aim at computing a function of their values at a common receiver. To this end, we propose a novel coded-modulation framework for over-the-air computation (OAC) based on hierarchical constellation design, which supports reliable...

💬 0 commentsarXiv:2601.12050v1PDF
0

Posted in cs.CV · 2026-01-17 · Chenchen Zhao, Muxi Chen, Qiang Xu

\textit{FocaLogic}: Logic-Based Interpretation of Visual Model Decisions

Interpretability of modern visual models is crucial, particularly in high-stakes applications. However, existing interpretability methods typically suffer from either reliance on white-box model access or insufficient quantitative rigor. To address these limitations, we introduce FocaLogic, a novel model-agnostic framework designed to...

💬 0 commentsarXiv:2601.12049v1PDF
0

Posted in cs.CR · 2026-01-17 · Xiaomei Zhang, Zhaoxi Zhang, Leo Yu Zhang, Yanjun Zhang, Guanhong Tao, Shirui Pan

Less Is More -- Until It Breaks: Security Pitfalls of Vision Token Compression in Large Vision-Language Models

Visual token compression is widely adopted to improve the inference efficiency of Large Vision-Language Models (LVLMs), enabling their deployment in latency-sensitive and resource-constrained scenarios. However, existing work has mainly focused on efficiency and performance, while the security implications of visual token compression...

💬 0 commentsarXiv:2601.12042v1PDF
0

Posted in cs.AI · 2026-01-17 · Murilo da Luz, Bruno Brandão, Luana Martins, Gustavo Oliveira, Bryan de Oliveira, Luckeciano Melo, Telma Soares

Partial Reasoning in Language Models: Search and Refinement Guided by Uncertainty

The use of Large Language Models (LLMs) for reasoning and planning tasks has drawn increasing attention in Artificial Intelligence research. Despite their remarkable progress, these models still exhibit limitations in multi-step inference scenarios, particularly in mathematical and logical reasoning. We introduce PREGU (Partial...

💬 0 commentsarXiv:2601.12040v1PDF
0

Posted in cs.AI · 2026-01-17 · Beishui Liao

Subargument Argumentation Frameworks: Separating Direct Conflict from Structural Dependency

Dung's abstract argumentation frameworks model acceptability solely in terms of an attack relation, thereby conflating two conceptually distinct aspects of argumentative reasoning: direct conflict between arguments and the structural dependencies that arise from their internal composition. While this abstraction preserves...

💬 0 commentsarXiv:2601.12038v3PDF
0

Posted in cs.CE · 2026-01-17 · Gustavo Delazeri, Marcus Ritt

Wildfire Suppression: Complexity, Models, and Instances

Wildfires cause major losses worldwide, and the frequency of fire-weather conditions is likely to increase in many regions. We study the allocation of suppression resources over time on a graph-based representation of a landscape to slow down fire propagation. Our contributions are theoretical and methodological. First, we prove that...

💬 0 commentsarXiv:2603.29865v1PDF
0

Posted in cs.HC · 2026-01-17 · Yue Yang, Christoph Leuze, Brian Hargreaves, Bruce Daniel, Fred M Baik

Multimodal Feedback for Handheld Tool Guidance: Combining Wrist-Based Haptics with Augmented Reality

We investigate how vibrotactile wrist feedback can enhance spatial guidance for handheld tool movement in optical see-through augmented reality (AR). While AR overlays are widely used to support surgical tasks, visual occlusion, lighting conditions, and interface ambiguity can compromise precision and confidence. To address these...

💬 0 commentsarXiv:2601.12037v1PDF
0

Posted in cs.SI · 2026-01-17 · Qitong Liu, Hao Peng, Zuchen Li, Xihang Meng, Ziyu Yang, Jiting Li, Li Sun, Philip S. Yu

Effective and Unsupervised Social Event Detection and Evolution via RAG and Structural Entropy

With the growing scale of social media, social event detection and evolution modeling have attracted increasing attention. Graph neural networks (GNNs) and transformer-based pre-trained language models (PLMs) have become mainstream approaches in this area. However, existing methods still face three major challenges. First, the sheer...

💬 0 commentsarXiv:2601.12035v1PDF
0

Posted in cs.CL · 2026-01-17 · Ziyi Zhao, Chongming Gao, Yang Zhang, Haoyan Liu, Weinan Gan, Huifeng Guo, Yong Liu, Fuli Feng

Don't Start Over: A Cost-Effective Framework for Migrating Personalized Prompts Between LLMs

Personalization in Large Language Models (LLMs) often relies on user-specific soft prompts. However, these prompts become obsolete when the foundation model is upgraded, necessitating costly, full-scale retraining. To overcome this limitation, we propose the Prompt-level User Migration Adapter (PUMA), a lightweight framework to...

💬 0 commentsarXiv:2601.12034v1PDF
0

Posted in cs.CL · 2026-01-17 · Muhammad Alif Al Hakim, Alfan Farizki Wicaksono, Fajri Koto

Preserving Fairness and Safety in Quantized LLMs Through Critical Weight Protection

Quantization is widely adopted to reduce the computational cost of large language models (LLMs); however, its implications for fairness and safety, particularly in dynamic quantization and multilingual contexts, remain underexplored. In this work, we conduct a systematic study of how static and dynamic quantization methods impact...

💬 0 commentsarXiv:2601.12033v2PDF
0

Posted in cs.NE · 2026-01-17 · Francisco Angulo de Lafuente, Vladimir Veselov, Richard Goodman

Speaking to Silicon: Neural Communication with Bitcoin Mining ASICs

This definitive research memoria presents a comprehensive, mathematically verified paradigm for neural communication with Bitcoin mining Application-Specific Integrated Circuits (ASICs), integrating five complementary frameworks: thermodynamic reservoir computing, hierarchical number system theory, algorithmic analysis, network...

💬 0 commentsarXiv:2601.12032v1PDF
0

Posted in cs.AI · 2026-01-17 · Yilun Yao, Shan Huang, Elsie Dai, Zhewen Tan, Zhenyu Duan, Shousheng Jia, Yanbing Jiang, Tong Yang

ARC: Active and Reflection-driven Context Management for Long-Horizon Information Seeking Agents

Large language models are increasingly deployed as research agents for deep search and long-horizon information seeking, yet their performance often degrades as interaction histories grow. This degradation, known as context rot, reflects a failure to maintain coherent and task-relevant internal states over extended reasoning horizons....

💬 0 commentsarXiv:2601.12030v1PDF
0

Posted in cs.IT · 2026-01-17 · Raghav Bongole, Tobias J. Oechtering, Mikael Skoglund

Generalizing the Fano inequality further

Interactive statistical decision making (ISDM) features algorithm-dependent data generated through interaction. Existing information-theoretic lower bounds in ISDM largely target expected risk, while tail-sensitive objectives are less developed. We generalize the interactive Fano framework of Chen et al. by replacing the hard success...

💬 0 commentsarXiv:2601.12027v1PDF
0

Posted in cs.AI · 2026-01-17 · Kartikey Singh Bhandari, Tanish Jain, Archit Agrawal, Dhruv Kumar, Praveen Kumar, Pratik Narang

Beyond Sentiment: A Multi-Agent Pipeline for Actionable Business Advice from Reviews

Customer reviews contain valuable signals about service quality, but converting large-scale review corpora into actionable business recommendations remains difficult. Standard sentiment/aspect analysis is largely descriptive, while direct prompting of large language models (LLMs) often yields generic and repetitive advice that is...

💬 0 commentsarXiv:2601.12024v2PDF
0

Posted in cs.CV · 2026-01-17 · Guillermo Figueroa-Araneda, Iris Diana Jimenez, Florian Hofherr, Manny Ko, Hector Andrade-Loarca, Daniel Cremers

DIAMOND-SSS: Diffusion-Augmented Multi-View Optimization for Data-efficient SubSurface Scattering

Subsurface scattering (SSS) gives translucent materials -- such as wax, jade, marble, and skin -- their characteristic soft shadows, color bleeding, and diffuse glow. Modeling these effects in neural rendering remains challenging due to complex light transport and the need for densely captured multi-view, multi-light datasets (often...

💬 0 commentsarXiv:2601.12020v1PDF
0

Posted in cs.CL · 2026-01-17 · Chaowei Zhang, Xiansheng Luo, Zewei Zhang, Yi Zhu, Jipeng Qiang, Longwei Wang

Acting Flatterers via LLMs Sycophancy: Combating Clickbait with LLMs Opposing-Stance Reasoning

The widespread proliferation of online content has intensified concerns about clickbait, deceptive or exaggerated headlines designed to attract attention. While Large Language Models (LLMs) offer a promising avenue for addressing this issue, their effectiveness is often hindered by Sycophancy, a tendency to produce reasoning that...

💬 0 commentsarXiv:2601.12019v1PDF
0

Posted in cs.CV · 2026-01-17 · Pavan Kumar Yata, Pediredla Pradeep, Goli Himanish, Swathi M

SAR-Based Marine Oil Spill Detection Using the DeepSegFusion Architecture

Detection of oil spills from satellite images is essential for both environmental surveillance and maritime safety. Traditional threshold-based methods frequently encounter performance degradation due to very high false alarm rates caused by look-alike phenomena such as wind slicks and ship wakes. Here, a hybrid deep learning model,...

💬 0 commentsarXiv:2601.12015v1PDF
0

Posted in cs.AI · 2026-01-17 · Elio Masciari, Vincenzo Moscato, Enea Vincenzo Napolitano, Gian Marco Orlando, Marco Perillo, Diego Russo

Are LLMs Ready for TOON? Benchmarking Structural Correctness-Sustainability Trade-offs in Novel Structured Output Formats

Large Language Models (LLMs) are increasingly required to generate structured, machine-readable outputs for downstream systems. While recent benchmarks have focused on evaluating the structural correctness of such outputs, the environmental impact of inference for different output formats has largely been overlooked. In this paper, we...

💬 0 commentsarXiv:2601.12014v1PDF