Qwen Councils
arXiv is taking too long to respond. Please try again or narrow your search.
Showing downloaded papers while arXiv is unavailable.

Computer Science

arXiv preprints from January 1, 2026 through July 20, 2026 — 01:19:10 EST

0

Posted in cs.CV · 2026-01-14 · Muhammad Imran, Yugyung Lee

Predicting When to Trust Vision-Language Models for Spatial Reasoning

Vision-Language Models (VLMs) demonstrate impressive capabilities across multimodal tasks, yet exhibit systematic spatial reasoning failures, achieving only 49% (CLIP) to 54% (BLIP-2) accuracy on basic directional relationships. For safe deployment in robotics and autonomous systems, we need to predict when to trust VLM spatial...

💬 0 commentsarXiv:2601.11644v1PDF
0

Posted in cs.HC · 2026-01-14 · Jordan Taylor, William Agnew, Maarten Sap, Sarah E. Fox, Haiyi Zhu

The Algorithmic Gaze of Image Quality Assessment: An Audit and Trace Ethnography of the LAION-Aesthetics Predictor

Visual generative AI models are trained using a one-size-fits-all measure of aesthetic appeal. However, what is deemed "aesthetic" is inextricably linked to personal taste and cultural values, raising the question of whose taste is represented in visual generative AI models. In this work, we study an aesthetic evaluation...

💬 0 commentsarXiv:2601.09896v4PDF
0

Posted in cs.IT · 2026-01-14 · Cheuk Ting Li

One-Cold Poisson Channel: A Simple Continuous-Time Channel with Zero Dispersion

We introduce the one-cold Poisson channel (OCPC), where the transmitter chooses one of several frequency bands to attenuate at a time. In particular, the perfect OCPC, where the number of bands is unlimited, is an extremely simple continuous-time memoryless channel. It has a capacity 1, zero channel dispersion, and an information...

💬 0 commentsarXiv:2601.09894v1PDF
0

Posted in cs.HC · 2026-01-14 · Rostyslav Hnatyshyn, Danny Perez, Gerik Scheuermann, Ross Maciejewski, Baldwin Nsonga

LAMDA: Aiding Visual Exploration of Atomic Displacements in Molecular Dynamics Simulations

Contemporary materials science research is heavily conducted in silico, involving massive simulations of the atomic-scale evolution of materials. Cataloging basic patterns in the atomic displacements is key to understanding and predicting the evolution of physical properties. However, the combinatorial complexity of the space of...

💬 0 commentsarXiv:2601.09887v1PDF
0

Posted in cs.CL · 2026-01-14 · Sathvik Nair, Byung-Doh Oh

Clozing the Gap: Exploring Why Language Model Surprisal Outperforms Cloze Surprisal

How predictable a word is can be quantified in two ways: using human responses to the cloze task or using probabilities from language models (LMs).When used as predictors of processing effort, LM probabilities outperform probabilities derived from cloze data. However, it is important to establish that LM probabilities do so for the...

💬 0 commentsarXiv:2601.09886v2PDF
0

Posted in cs.AI · 2026-01-14 · Xinxing Ren, Quagmire Zang, Caelum Forder, Suman Deb, Ahsen Tahir, Roman J. Georgio, Peter Carroll, Zekun Guo

Beyond Rule-Based Workflows: An Information-Flow-Orchestrated Multi-Agents Paradigm via Agent-to-Agent Communication from CORAL

Most existing Large Language Model (LLM)-based Multi-Agent Systems (MAS) rely on predefined workflows, where human engineers enumerate task states in advance and specify routing rules and contextual injections accordingly. Such workflow-driven designs are essentially rule-based decision trees, which suffer from two fundamental...

💬 0 commentsarXiv:2601.09883v1PDF
0

Posted in cs.CV · 2026-01-14 · Weili Nie, Julius Berner, Nanye Ma, Chao Liu, Saining Xie, Arash Vahdat

Transition Matching Distillation for Fast Video Generation

Large video diffusion and flow models have achieved remarkable success in high-quality video generation, but their use in real-time interactive applications remains limited due to their inefficient multi-step sampling process. In this work, we present Transition Matching Distillation (TMD), a novel framework for distilling video...

💬 0 commentsarXiv:2601.09881v2PDF
0

Posted in cs.CV · 2026-01-14 · Yang Xing, Jiong Wu, Savas Ozdemir, Ying Zhang, Yang Yang, Wei Shao, Kuang Gong

MedVL-SAM2: A unified 3D medical vision-language model for multimodal reasoning and prompt-driven segmentation

Recent progress in medical vision-language models (VLMs) has achieved strong performance on image-level text-centric tasks such as report generation and visual question answering (VQA). However, achieving fine-grained visual grounding and volumetric spatial reasoning in 3D medical VLMs remains challenging, particularly when aiming to...

💬 0 commentsarXiv:2601.09879v1PDF
0

Posted in cs.HC · 2026-01-14 · Paulius Jurcys, Ashley Greenwald, Mark Fenwick, Valto Loikkanen, Sebastian Porsdam Mann, Brian D. Earp

Who Owns My AI Twin? Data Ownership in a New World of Simulated Identities

The emergence of AI twins, digital replicas that encapsulate an individual's knowledge, memories, psychological traits, and behavioral patterns, raises novel legal and ethical challenges for data governance and personal identity. Built from personal data, these systems require a rethinking of what it means to exercise dominion over...

💬 0 commentsarXiv:2601.09877v2PDF
0

Posted in cs.CL · 2026-01-14 · Yifei Shen, Yilun Zhao, Justice Ou, Tinglin Huang, Arman Cohan

Patient-Similarity Cohort Reasoning in Clinical Text-to-SQL

Real-world clinical text-to-SQL requires reasoning over heterogeneous EHR tables, temporal windows, and patient-similarity cohorts to produce executable queries. We introduce CLINSQL, a benchmark of 633 expert-annotated tasks on MIMIC-IV v3.1 that demands multi-table joins, clinically meaningful filters, and executable SQL. Solving...

💬 0 commentsarXiv:2601.09876v1PDF
0

Posted in cs.SE · 2026-01-14 · Saymon Souza, Amanda Santana, Eduardo Figueiredo, Igor Muzetti, João Eduardo Montandon, Lionel Briand

Beyond Strict Rules: Assessing the Effectiveness of Large Language Models for Code Smell Detection

Code smells are symptoms of potential code quality problems that may affect software maintainability, thus increasing development costs and impacting software reliability. Large language models (LLMs) have shown remarkable capabilities for supporting various software engineering activities, but their use for detecting code smells...

💬 0 commentsarXiv:2601.09873v2PDF
0

Posted in cs.AI · 2026-01-14 · Andrea Ferrario, Alessandro Facchini, Juan M. Durán

Epistemology gives a Future to Complementarity in Human-AI Interactions

Human-AI complementarity is the claim that a human supported by an AI system can outperform either alone in a decision-making process. Since its introduction in the humanAI interaction literature, it has gained traction by generalizing the reliance paradigm and by offering a more practical alternative to the contested construct of...

💬 0 commentsarXiv:2601.09871v2PDF
0

Posted in cs.AI · 2026-01-14 · Andrea Ferrario, Rasita Vinay, Matteo Casserini, Alessandro Facchini

A Scoping Review of the Ethical Perspectives on Anthropomorphising Large Language Model-Based Conversational Agents

Anthropomorphisation -- the phenomenon whereby non-human entities are ascribed human-like qualities -- has become increasingly salient with the rise of large language model (LLM)-based conversational agents (CAs). Unlike earlier chatbots, LLM-based CAs routinely generate interactional and linguistic cues, such as first-person...

💬 0 commentsarXiv:2601.09869v2PDF
0

Posted in cs.CR · 2026-01-14 · Yifan Zhang, Yishan Yang, Riku Jäntti, Zheng Yan, Dusit Niyato, Zhu Han

AmbShield: Enhancing Physical Layer Security with Ambient Backscatter Devices against Eavesdroppers

Passive eavesdropping compromises confidentiality in wireless networks, especially in resource-constrained environments where heavyweight cryptography is impractical. Physical layer security (PLS) exploits channel randomness and spatial selectivity to confine information to an intended receiver with modest overhead. However, typical...

💬 0 commentsarXiv:2601.09867v1PDF
0

Posted in cs.CV · 2026-01-14 · Kiarie Ndegwa, Andreas Gros, Tony Chang, David Diaz, Vincent A. Landau, Nathan E. Rutenbeck, Luke J. Zachmann, Guy Bayes, Scott Conway

VibrantSR: Sub-Meter Canopy Height Models from Sentinel-2 Using Generative Flow Matching

We present VibrantSR (Vibrant Super-Resolution), a generative super-resolution framework for estimating 0.5 meter canopy height models (CHMs) from 10 meter Sentinel-2 imagery. Unlike approaches based on aerial imagery that are constrained by infrequent and irregular acquisition schedules, VibrantSR leverages globally available...

💬 0 commentsarXiv:2601.09866v2PDF
0

Posted in cs.LG · 2026-01-14 · Jacob Sander, Brian Jalaian, Venkat R. Dasari

Advancing Model Refinement: Muon-Optimized Distillation and Quantization for LLM Deployment

Large Language Models (LLMs) enable advanced natural language processing but face deployment challenges on resource-constrained edge devices due to high computational, memory, and energy demands. Optimizing these models requires addressing three key challenges: acquiring task-specific data, fine-tuning for performance, and compressing...

💬 0 commentsarXiv:2601.09865v1PDF
0

Posted in cs.IT · 2026-01-14 · Adway Girish, Shlomo Shamai, Emre Telatar

High signal-to-noise ratio asymptotics of entropy-constrained Gaussian channel capacity

We study the input-entropy-constrained Gaussian channel capacity problem in the asymptotic high signal-to-noise ratio (SNR) regime. We show that the capacity-achieving distribution as SNR goes to infinity is given by a discrete Gaussian distribution supported on a scaled integer lattice. Further, we show that the gap between the input...

💬 0 commentsarXiv:2601.09864v1PDF
0

Posted in cs.DS · 2026-01-14 · Sepideh Mahabadi, Sherry Sarkar, Jakub Tarnawski

Improved Algorithms for Fair Matroid Submodular Maximization

Submodular maximization subject to matroid constraints is a central problem with many applications in machine learning. As algorithms are increasingly used in decision-making over datapoints with sensitive attributes such as gender or race, it is becoming crucial to enforce fairness to avoid bias and discrimination. Recent work has...

💬 0 commentsarXiv:2601.09860v1PDF
0

Posted in cs.CV · 2026-01-14 · Anant Mehta, Xiyuan Wei, Xingyu Chen, Tianbao Yang

Breaking the Limits of Open-Weight CLIP: An Optimization Framework for Self-supervised Fine-tuning of CLIP

CLIP has become a cornerstone of multimodal representation learning, yet improving its performance typically requires a prohibitively costly process of training from scratch on billions of samples. We ask a different question: Can we improve the performance of open-weight CLIP models across various downstream tasks using only existing...

💬 0 commentsarXiv:2601.09859v1PDF
0

Posted in cs.CL · 2026-01-14 · Yilin Bao, Ziyao He, Zayden Yang

OUTLINEFORGE: Hierarchical Reinforcement Learning with Explicit States for Scientific Writing

Scientific paper generation requires document-level planning and factual grounding, but current large language models, despite their strong local fluency, often fail in global structure, input coverage, and citation consistency. We present a reinforcement learning framework that casts scientific outline construction as a long-horizon...

💬 0 commentsarXiv:2601.09858v1PDF
0

Posted in cs.RO · 2026-01-14 · Andrew Stratton, Phani Teja Singamaneni, Pranav Goyal, Rachid Alami, Christoforos Mavrogiannis

How Human Motion Prediction Quality Shapes Social Robot Navigation Performance in Constrained Spaces

Motivated by the vision of integrating mobile robots closer to humans in warehouses, hospitals, manufacturing plants, and the home, we focus on robot navigation in dynamic and spatially constrained environments. Ensuring human safety, comfort, and efficiency in such settings requires that robots are endowed with a model of how humans...

💬 0 commentsarXiv:2601.09856v1PDF
0

Posted in cs.AI · 2026-01-14 · Michael R. Metel, Yufei Cui, Boxing Chen, Prasanna Parthasarathi

Thinking Long, but Short: Stable Sequential Test-Time Scaling for Large Reasoning Models

Sequential test-time scaling is a promising training-free method to improve large reasoning model accuracy, but as currently implemented, significant limitations have been observed. Inducing models to think for longer can increase their accuracy, but as the length of reasoning is further extended, it has also been shown to result in...

💬 0 commentsarXiv:2601.09855v1PDF
0

Posted in cs.CL · 2026-01-14 · Sraavya Sambara, Yuan Pu, Ayman Ali, Vishala Mishra, Lionel Wong, Monica Agrawal

MedRedFlag: Investigating how LLMs Redirect Misconceptions in Real-World Health Communication

Real-world health questions from patients often unintentionally embed false assumptions or premises. In such cases, safe medical communication typically involves redirection: addressing the implicit misconception and then responding to the underlying patient context, rather than the original question. While large language models...

💬 0 commentsarXiv:2601.09853v3PDF
0

Posted in cs.CL · 2026-01-14 · Sriram Padmanabhan, Siyuan Song, Kanishka Misra

Bears, all bears, and some bears. Language Constraints on Language Models' Inductive Inferences

Language places subtle constraints on how we make inductive inferences. Developmental evidence by Gelman et al. (2002) has shown children (4 years and older) to differentiate among generic statements ("Bears are daxable"), universally quantified NPs ("all bears are daxable") and indefinite plural NPs ("some bears are daxable") in...

💬 0 commentsarXiv:2601.09852v2PDF
0

Posted in cs.CV · 2026-01-14 · Po-han Li, Shenghui Chen, Ufuk Topcu, Sandeep Chinchali

ViSIL: Unified Evaluation of Information Loss in Multimodal Video Captioning

Multimodal video captioning condenses dense footage into a structured format of keyframes and natural language. By creating a cohesive multimodal summary, this approach anchors generative AI in rich semantic evidence and serves as a lightweight proxy for high-efficiency retrieval. However, traditional metrics like BLEU or ROUGE fail...

💬 0 commentsarXiv:2601.09851v2PDF