Qwen Councils

Computer Science

arXiv preprints from January 1, 2026 through July 21, 2026 — 19:39:05 EST

0

Posted in cs.CV · 2026-01-17 · Alfe Suny, MD Sakib Ul Islam, Md. Imran Hossain

Reliable Deep Learning for Small-Scale Classifications: Experiments on Real-World Image Datasets from Bangladesh

Convolutional neural networks (CNNs) have achieved state-of-the-art performance in image recognition tasks but often involve complex architectures that may overfit on small datasets. In this study, we evaluate a compact CNN across five publicly available, real-world image datasets from Bangladesh, including urban encroachment, vehicle...

💬 0 commentsarXiv:2601.11911v2PDF
0

Posted in cs.CV · 2026-01-17 · Guiying Zhu, Bowen Yang, Yin Zhuang, Tong Zhang, Guanqun Wang, Zhihao Che, He Chen, Lianlin Li

A Training-Free Guess What Vision Language Model from Snippets to Open-Vocabulary Object Detection

Open-Vocabulary Object Detection (OVOD) aims to develop the capability to detect anything. Although myriads of large-scale pre-training efforts have built versatile foundation models that exhibit impressive zero-shot capabilities to facilitate OVOD, the necessity of creating a universal understanding for any object cognition according...

💬 0 commentsarXiv:2601.11910v2PDF
0

Posted in cs.CV · 2026-01-17 · Io Yamada, Hirotsugu Okuno

Effects of the retina-inspired light intensity encoding on color discrimination performance

Color is an important source of information for visual functions such as object recognition, but it is greatly affected by the color of illumination. The ability to perceive the color of a visual target independent of illumination color is called color constancy (CC), and is an important feature for vision systems that use color...

💬 0 commentsarXiv:2601.11909v1PDF
0

Posted in cs.CL · 2026-01-17 · Byeongjin Kim, Gyuwan Kim, Seo Yeon Park

PPA-Plan: Proactive Pitfall Avoidance for Reliable Planning in Long-Context LLM Reasoning

Large language models (LLMs) struggle with reasoning over long contexts where relevant information is sparsely distributed. Although plan-and-execute frameworks mitigate this by decomposing tasks into planning and execution, their effectiveness is often limited by unreliable plan generation due to dependence on surface-level cues....

💬 0 commentsarXiv:2601.11908v2PDF
0

Posted in cs.CV · 2026-01-17 · Prosenjit Chatterjee, ANK Zaman

Towards Airborne Object Detection: A Deep Learning Analysis

The rapid proliferation of airborne platforms, including commercial aircraft, drones, and UAVs, has intensified the need for real-time, automated threat assessment systems. Current approaches depend heavily on manual monitoring, resulting in limited scalability and operational inefficiencies. This work introduces a dual-task model...

💬 0 commentsarXiv:2601.11907v1PDF
0

Posted in cs.RO · 2026-01-17 · Jose Cuaran, Kendall Koe, Aditya Potnis, Naveen Kumar Uppalapati, Girish Chowdhary

Visual-Language-Guided Task Planning for Horticultural Robots

Crop monitoring is essential for precision agriculture, but current systems lack high-level reasoning. We introduce a novel, modular framework that uses a Vision Language Model (VLM) to guide robotic task planning by actively querying heterogeneous data sources, including enriched RGB camera feeds and 2D semantic occupancy maps,...

💬 0 commentsarXiv:2601.11906v2PDF
0

Posted in cs.AI · 2026-01-17 · Junyu Cao, Ruijiang Gao, Esmaeil Keyvanshokooh, Jianhao Ma

LIBRA: Language Model Informed Bandit Recourse Algorithm for Personalized Treatment Planning

We introduce a unified framework that seamlessly integrates algorithmic recourse, contextual bandits, and large language models (LLMs) to support sequential decision-making in high-stakes settings such as personalized medicine. We first introduce the recourse bandit problem, where a decision-maker must select both a treatment action...

💬 0 commentsarXiv:2601.11905v1PDF
0

Posted in cs.AI · 2026-01-17 · YenTing Lee, Keerthi Koneru, Zahra Moslemi, Sheethal Kumar, Ramesh Radhakrishnan

AEMA: Verifiable Evaluation Framework for Trustworthy and Controlled Agentic LLM Systems

Evaluating large language model (LLM)-based multi-agent systems remains a critical challenge, as these systems must exhibit reliable coordination, transparent decision-making, and verifiable performance across evolving tasks. Existing evaluation approaches often limit themselves to single-response scoring or narrow benchmarks, which...

💬 0 commentsarXiv:2601.11903v1PDF
0

Posted in cs.CV · 2026-01-17 · Yilmaz Korkmaz, Vishal M. Patel

RemoteVAR: Autoregressive Visual Modeling for Remote Sensing Change Detection

Remote sensing change detection aims to localize and characterize scene changes between two time points and is central to applications such as environmental monitoring and disaster assessment. Meanwhile, visual autoregressive models (VARs) have recently shown impressive image generation capability, but their adoption for pixel-level...

💬 0 commentsarXiv:2601.11898v1PDF
0

Posted in cs.LG · 2026-01-17 · Jinwon Sohn, Guang Lin, Qifan Song

Task-tailored Pre-processing: Fair Downstream Supervised Learning

Fairness-aware machine learning has recently attracted various communities to mitigate discrimination against certain societal groups in data-driven tasks. For fair supervised learning, particularly in pre-processing, there have been two main categories: data fairness and task-tailored fairness. The former directly finds an...

💬 0 commentsarXiv:2601.11897v1PDF
0

Posted in cs.CV · 2026-01-17 · Ngoc-Khai Hoang, Thi-Nhu-Mai Nguyen, Huy-Hieu Pham

Digital FAST: An AI-Driven Multimodal Framework for Rapid and Early Stroke Screening

Early identification of stroke symptoms is essential for enabling timely intervention and improving patient outcomes, particularly in prehospital settings. This study presents a fast, non-invasive multimodal deep learning framework for automatic binary stroke screening based on data collected during the F.A.S.T. assessment. The...

💬 0 commentsarXiv:2601.11896v2PDF
0

Posted in cs.LG · 2026-01-17 · Adarsh Kumarappan, Pareesa Ameneh Golnari, Wen Wen, Xiaoyu Liu, Gabriel Ryan, Yuting Sun, Shengyu Fu, Elsie Nallipogu

DevBench: A Realistic, Developer-Informed Benchmark for Code Generation Models

DevBench is a telemetry-driven benchmark designed to evaluate Large Language Models (LLMs) on realistic code completion tasks. It includes 1,800 evaluation instances across six programming languages and six task categories derived from real developer telemetry and synthesized using generator models from multiple provider families to...

💬 0 commentsarXiv:2601.11895v3PDF
0

Posted in cs.CR · 2026-01-17 · Zimo Ji, Daoyuan Wu, Wenyuan Jiang, Pingchuan Ma, Zongjie Li, Yudong Gao, Shuai Wang, Yingjiu Li

Taming Various Privilege Escalation in LLM-Based Agent Systems: A Mandatory Access Control Framework

Large Language Model (LLM)-based agent systems are increasingly deployed for complex real-world tasks but remain vulnerable to natural language-based attacks that exploit over-privileged tool use. This paper aims to understand and mitigate such attacks through the lens of privilege escalation, defined as agent actions exceeding the...

💬 0 commentsarXiv:2601.11893v1PDF
0

Posted in cs.LG · 2026-01-17 · Xihe Gu, Urbashi Mitra, Tara Javidi

From Relative Entropy to Minimax: A Unified Framework for Coverage in MDPs

Targeted and deliberate exploration of state--action pairs is essential in reward-free Markov Decision Problems (MDPs). More precisely, different state-action pairs exhibit different degree of importance or difficulty which must be actively and explicitly built into a controlled exploration strategy. To this end, we propose a weighted...

💬 0 commentsarXiv:2601.11890v1PDF
0

Posted in cs.IR · 2026-01-17 · Wenhan Liu, Xinyu Ma, Yutao Zhu, Yuchen Li, Daiting Shi, Dawei Yin, Zhicheng Dou

Agentic-R: Learning to Retrieve for Agentic Search

Agentic search has recently emerged as a powerful paradigm, where an agent interleaves multi-step reasoning with on-demand retrieval to solve complex questions. Despite its success, how to design a retriever for agentic search remains largely underexplored. Existing search agents typically rely on similarity-based retrievers, while...

💬 0 commentsarXiv:2601.11888v1PDF
0

Posted in cs.CL · 2026-01-17 · Kaijie Mo, Siddhartha Venkatayogi, Chantal Shaib, Ramez Kouzy, Wei Xu, Byron C. Wallace, Junyi Jessy Li

Faithfulness vs. Safety: Evaluating LLM Behavior Under Counterfactual Medical Evidence

In high-stakes domains like medicine, it may be generally desirable for models to faithfully adhere to the context provided. But what happens if the context does not align with model priors or safety protocols? In this paper, we investigate how LLMs behave and reason when presented with counterfactual (or even adversarial) medical...

💬 0 commentsarXiv:2601.11886v2PDF
0

Posted in cs.AI · 2026-01-17 · Zhifei Li, Ziyue Qin, Xiangyu Luo, Xiaoju Hou, Yue Zhao, Miao Zhang, Zhifang Huang, Kui Xiao, Bing Yang

MyGram: Modality-aware Graph Transformer with Global Distribution for Multi-modal Entity Alignment

Multi-modal entity alignment aims to identify equivalent entities between two multi-modal Knowledge graphs by integrating multi-modal data, such as images and text, to enrich the semantic representations of entities. However, existing methods may overlook the structural contextual information within each modality, making them...

💬 0 commentsarXiv:2601.11885v1PDF
0

Posted in cs.LG · 2026-01-17 · Jun Liu, Leo Yu Zhang, Fengpeng Li, Isao Echizen, Jiantao Zhou

Low-Cost Hard-Label Adversarial Attack with Theoretical Foundations

Hard-label black-box attacks, relying solely on top-1 predictions, represent one of the most challenging yet practically threat models. Despite recent progress, existing approaches face two key limitations: (1) they overlook the critical role of initialization, focusing primarily on optimization strategies; and (2) they rely heavily...

💬 0 commentsarXiv:2601.14300v3PDF
0

Posted in cs.HC · 2026-01-17 · Kashif Imteyaz, Qiushi, Liang, Yakov Bart, Maitraye Das, Saiph Savage

AI-Mediated Hiring and the Job Search of Blind and Low-Vision Individuals

Blind and low-vision (BLV) individuals face high unemployment rates. The job search is becoming harder as more employers use AI-driven systems to screen resumes before a human ever sees them. Such AI systems could inadvertently further disadvantage BLV job seekers, introducing additional barriers to an already difficult process. We...

💬 0 commentsarXiv:2601.11884v2PDF
0

Posted in cs.LG · 2026-01-17 · Chaoqi Jia, Longkun Guo, Kewen Liao, Zhigang Lu, Chao Chen, Jason Xue

Approximation Algorithm for Constrained $k$-Center Clustering: A Local Search Approach

Clustering is a long-standing research problem and a fundamental tool in AI and data analysis. The traditional k-center problem, a fundamental theoretical challenge in clustering, has a best possible approximation ratio of 2, and any improvement to a ratio of 2 - ε would imply P = NP. In this work, we study the constrained k-center...

💬 0 commentsarXiv:2601.11883v1PDF
0

Posted in cs.LG · 2026-01-17 · Yingxiao Zhang, Jiaxin Duan, Junfu Zhang, Ke Feng

TF-CoDiT: Conditional Time Series Synthesis with Diffusion Transformers for Treasury Futures

Diffusion Transformers (DiT) have achieved milestones in synthesizing financial time-series data, such as stock prices and order flows. However, their performance in synthesizing treasury futures data is still underexplored. This work emphasizes the characteristics of treasury futures data, including its low volume, market...

💬 0 commentsarXiv:2601.11880v1PDF
0

Posted in cs.RO · 2026-01-17 · Christopher Kao, Akhil Pathapati, James Davis

AI for Green Spaces: Leveraging Autonomous Navigation and Computer Vision for Park Litter Removal

There are 50 billion pieces of litter in the U.S. alone. Grass fields contribute to this problem because picnickers tend to leave trash on the field. We propose building a robot that can autonomously navigate, identify, and pick up trash in parks. To autonomously navigate the park, we used a Spanning Tree Coverage (STC) algorithm to...

💬 0 commentsarXiv:2601.11876v1PDF
0

Posted in cs.IR · 2026-01-17 · Suchana Datta, Dwaipayan Roy, Derek Greene, Gerardine Meaney, Karen Wade, Philipp Mayr

Cultural Analytics for Good: Building Inclusive Evaluation Frameworks for Historical IR

This work bridges the fields of information retrieval and cultural analytics to support equitable access to historical knowledge. Using the British Library BL19 digital collection (more than 35,000 works from 1700-1899), we construct a benchmark for studying changes in language, terminology and retrieval in the 19th-century fiction...

💬 0 commentsarXiv:2601.11874v1PDF
0

Posted in cs.CL · 2026-01-17 · Nguyen Tien Phat, Ngo Vu Minh, Linh Van Ngo, Nguyen Thi Ngoc Diep, Thien Huu Nguyen

GloCTM: Cross-Lingual Topic Modeling via a Global Context Space

Cross-lingual topic modeling seeks to uncover coherent and semantically aligned topics across languages - a task central to multilingual understanding. Yet most existing models learn topics in disjoint, language-specific spaces and rely on alignment mechanisms (e.g., bilingual dictionaries) that often fail to capture deep...

💬 0 commentsarXiv:2601.11872v1PDF
0

Posted in cs.SE · 2026-01-17 · Mike A. Merrill, Alexander G. Shaw, Nicholas Carlini, Boxuan Li, Harsh Raj, Ivan Bercovich, Lin Shi, Jeong Yeon Shin, Thomas Walshe, E. Kelly Buchanan, Junhong Shen, Guanghao Ye, Haowei Lin, Jason Poulos, Maoyu Wang, Marianna Nezhurina, Jenia Jitsev, Di Lu, Orfeas Menis Mastromichalakis, Zhiwei Xu, Zizhao Chen, Yue Liu, Robert Zhang, Leon Liangyu Chen, Anurag Kashyap, Jan-Lucas Uslu, Jeffrey Li, Jianbo Wu, Minghao Yan, Song Bian, Vedang Sharma, Ke Sun, Steven Dillmann, Akshay Anand, Andrew Lanpouthakoun, Bardia Koopah, Changran Hu, Etash Guha, Gabriel H. S. Dreiman, Jiacheng Zhu, Karl Krauth, Li Zhong, Niklas Muennighoff, Robert Amanfu, Shangyin Tan, Shreyas Pimpalgaonkar, Tushar Aggarwal, Xiangning Lin, Xin Lan, Xuandong Zhao, Yiqing Liang, Yuanli Wang, Zilong Wang, Changzhi Zhou, David Heineman, Hange Liu, Harsh Trivedi, John Yang, Junhong Lin, Manish Shetty, Michael Yang, Nabil Omi, Negin Raoof, Shanda Li, Terry Yue Zhuo, Wuwei Lin, Yiwei Dai, Yuxin Wang, Wenhao Chai, Shang Zhou, Dariush Wahdany, Ziyu She, Jiaming Hu, Zhikang Dong, Yuxuan Zhu, Sasha Cui, Ahson Saiyed, Arinbjörn Kolbeinsson, Jesse Hu, Christopher Michael Rytting, Ryan Marten, Yixin Wang, Alex Dimakis, Andy Konwinski, Ludwig Schmidt

Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

AI agents may soon become capable of autonomously completing valuable, long-horizon tasks in diverse domains. Current benchmarks either do not measure real-world tasks, or are not sufficiently difficult to meaningfully measure frontier models. To this end, we present Terminal-Bench 2.0: a carefully curated hard benchmark composed of...

💬 0 commentsarXiv:2601.11868v1PDF