Qwen Councils

Computer Science

arXiv preprints from January 1, 2026 through July 20, 2026 — 10:58:25 EST

0

Posted in cs.CV · 2026-01-20 · Emily Kim, Allen Wu, Jessica Hodgins

Curriculum-Based Strategies for Efficient Cross-Domain Action Recognition

Despite significant progress in human action recognition, generalizing to diverse viewpoints remains a challenge. Most existing datasets are captured from ground-level perspectives, and models trained on them often struggle to transfer to drastically different domains such as aerial views. This paper examines how curriculum-based...

💬 0 commentsarXiv:2601.14101v1PDF
0

Posted in cs.LG · 2026-01-20 · Shi-Shun Chen, Xiao-Yang Li, Enrico Zio

Causal feature selection framework for stable soft sensor modeling based on time-delayed cross mapping

Soft sensor modeling plays a crucial role in process monitoring. Causal feature selection can enhance the performance of soft sensor models in industrial applications. However, existing methods ignore two critical characteristics of industrial processes. Firstly, causal relationships between variables always involve time delays,...

💬 0 commentsarXiv:2601.14099v1PDF
0

Posted in cs.LG · 2026-01-20 · Min Zeng, Xi Chen, Haiqin Yang, Yike Guo

Sparse Adapter Fusion for Continual Learning in NLP

Continual learning in natural language processing plays a crucial role in adapting to evolving data and preventing catastrophic forgetting. Despite significant progress, existing methods still face challenges, such as inefficient parameter reuse across tasks, risking catastrophic forgetting when tasks are dissimilar, and the...

💬 0 commentsarXiv:2602.02502v1PDF
0

Posted in cs.AI · 2026-01-20 · Benedikt Hartl, Léo Pio-Lopez, Chris Fields, Michael Levin

Remapping and navigation of an embedding space via error minimization: a fundamental organizational principle of cognition in natural and artificial systems

The emerging field of diverse intelligence seeks an integrated view of problem-solving in agents of very different provenance, composition, and substrates. From subcellular chemical networks to swarms of organisms, and across evolved, engineered, and chimeric systems, it is hypothesized that scale-invariant principles of...

💬 0 commentsarXiv:2601.14096v2PDF
0

Posted in cs.LG · 2026-01-20 · Babacar Toure, Dimitrios Tsilimantos, Omid Esrafilian, Marios Kountouris

Optimizing Energy and Data Collection in UAV-aided IoT Networks using Attention-based Multi-Objective Reinforcement Learning

Due to their adaptability and mobility, Unmanned Aerial Vehicles (UAVs) are becoming increasingly essential for wireless network services, particularly for data harvesting tasks. In this context, Artificial Intelligence (AI)-based approaches have gained significant attention for addressing UAV path planning tasks in large and complex...

💬 0 commentsarXiv:2601.14092v1PDF
0

Posted in cs.RO · 2026-01-20 · Hossein Naderi, Alireza Shojaei, Lifu Huang, Philip Agee, Kereshmeh Afsari, Abiola Akanmu

Zero-shot adaptable task planning for autonomous construction robots: a comparative study of lightweight single and multi-AI agent systems

Robots are expected to play a major role in the future construction industry but face challenges due to high costs and difficulty adapting to dynamic tasks. This study explores the potential of foundation models to enhance the adaptability and generalizability of task planning in construction robots. Four models are proposed and...

💬 0 commentsarXiv:2601.14091v1PDF
0

Posted in cs.IT · 2026-01-20 · Maria Abu-Sini, Reinhard Heckel

Near Optimal Code Construction for the Adversarial Torn Paper Channel With Edit Errors

Motivated by DNA storage systems and 3D fingerprinting, this work studies the adversarial torn paper channel with edit errors. This channel first applies at most $t_e$ edit errors (i.e., insertions, deletions, and substitutions) to the transmitted word and then breaks it into $t+1$ fragments at arbitrary positions. In this paper, we...

💬 0 commentsarXiv:2601.14088v1PDF
0

Posted in cs.AR · 2026-01-20 · Ruichi Han, Yizhi Chen, Tong Lei, Jordi Altayo Gonzalez, Ahmed Hemani

'1'-bit Count-based Sorting Unit to Reduce Link Power in DNN Accelerators

Interconnect power consumption remains a bottleneck in Deep Neural Network (DNN) accelerators. While ordering data based on '1'-bit counts can mitigate this via reduced switching activity, practical hardware sorting implementations remain underexplored. This work proposes the hardware implementation of a comparison-free sorting unit...

💬 0 commentsarXiv:2601.14087v1PDF
0

Posted in cs.CV · 2026-01-20 · Nattapong Kurpukdee, Adrian G. Bors

Two-Stream temporal transformer for video action classification

Motion representation plays an important role in video understanding and has many applications including action recognition, robot and autonomous guidance or others. Lately, transformer networks, through their self-attention mechanism capabilities, have proved their efficiency in many applications. In this study, we introduce a new...

💬 0 commentsarXiv:2601.14086v1PDF
0

Posted in cs.CV · 2026-01-20 · Abdurrahim Yilmaz, Ozan Erdem, Ece Gokyayla, Ayda Acar, Burc Bugra Dagtas, Dilara Ilhan Erdil, Gulsum Gencoglan, Burak Temelkuran

DermaBench: A Clinician-Annotated Benchmark Dataset for Dermatology Visual Question Answering and Reasoning

Vision-language models (VLMs) are increasingly important in medical applications; however, their evaluation in dermatology remains limited by datasets that focus primarily on image-level classification tasks such as lesion recognition. While valuable for recognition, such datasets cannot assess the full visual understanding, language...

💬 0 commentsarXiv:2601.14084v1PDF
0

Posted in cs.CV · 2026-01-20 · Matthew Gwilliam, Xiao Wang, Xuefeng Hu, Zhenheng Yang

Implicit Neural Representation Facilitates Unified Universal Vision Encoding

Models for image representation learning are typically designed for either recognition or generation. Various forms of contrastive learning help models learn to convert images to embeddings that are useful for classification, detection, and segmentation. On the other hand, models can be trained to reconstruct images with pixel-wise,...

💬 0 commentsarXiv:2601.14256v1PDF
0

Posted in cs.CV · 2026-01-20 · Sangbeom Lim, Seoung Wug Oh, Jiahui Huang, Heeji Yoon, Seungryong Kim, Joon-Young Lee

VideoMaMa: Mask-Guided Video Matting via Generative Prior

Generalizing video matting models to real-world videos remains a significant challenge due to the scarcity of labeled data. To address this, we present Video Mask-to-Matte Model (VideoMaMa) that converts coarse segmentation masks into pixel accurate alpha mattes, by leveraging pretrained video diffusion models. VideoMaMa demonstrates...

💬 0 commentsarXiv:2601.14255v1PDF
0

Posted in cs.CV · 2026-01-20 · Hongyuan Chen, Xingyu Chen, Youjia Zhang, Zexiang Xu, Anpei Chen

Motion 3-to-4: 3D Motion Reconstruction for 4D Synthesis

We present Motion 3-to-4, a feed-forward framework for synthesising high-quality 4D dynamic objects from a single monocular video and an optional 3D reference mesh. While recent advances have significantly improved 2D, video, and 3D content generation, 4D synthesis remains difficult due to limited training data and the inherent...

💬 0 commentsarXiv:2601.14253v1PDF
0

Posted in cs.IT · 2026-01-20 · Tristan Simas

Semantic Identity Compression: Zero-Error Laws, Rate-Distortion, and Neurosymbolic Necessity

Symbolic systems operate over precise identities: variables denote specific objects, pointers target precise memory locations, and database keys refer to singular records. Neural embeddings generalize by compressing away semantic detail, but this compression creates collision ambiguity: multiple distinct entities can share the same...

💬 0 commentsarXiv:2601.14252v6PDF
0

Posted in cs.CV · 2026-01-20 · Said Taghadouini, Adrien Cavaillès, Baptiste Aubertin

LightOnOCR: A 1B End-to-End Multilingual Vision-Language Model for State-of-the-Art OCR

We present LightOnOCR-2-1B, a 1B-parameter end-to-end multilingual vision--language model that converts document images (e.g., PDFs) into clean, naturally ordered text without brittle OCR pipelines. Trained on a large-scale, high-quality distillation mix with strong coverage of scans, French documents, and scientific PDFs,...

💬 0 commentsarXiv:2601.14251v2PDF
0

Posted in cs.CV · 2026-01-20 · Pengze Zhang, Yanze Wu, Mengtian Li, Xu Bai, Songtao Zhao, Fulong Ye, Chong Mou, Xinghui Li, Zhuowei Chen, Qian He, Mingyuan Gao

OmniTransfer: All-in-one Framework for Spatio-temporal Video Transfer

Videos convey richer information than images or text, capturing both spatial and temporal dynamics. However, most existing video customization methods rely on reference images or task-specific temporal priors, failing to fully exploit the rich spatio-temporal information inherent in videos, thereby limiting flexibility and...

💬 0 commentsarXiv:2601.14250v1PDF
0

Posted in cs.CL · 2026-01-20 · Yuming Yang, Mingyoung Lai, Wanxu Zhao, Xiaoran Fan, Zhiheng Xi, Mingqi Wu, Chiyue Huang, Jun Zhao, Haijun Lv, Jian Tong, Yunhua Zhou, Yicheng Zou, Qipeng Guo, Tao Gui, Qi Zhang, Xuanjing Huang

Which Reasoning Trajectories Teach Students to Reason Better? A Simple Metric of Informative Alignment

Long chain-of-thought (CoT) trajectories provide rich supervision signals for distilling reasoning from teacher to student LLMs. However, both prior work and our experiments show that trajectories from stronger teachers do not necessarily yield better students, highlighting the importance of data-student suitability in distillation....

💬 0 commentsarXiv:2601.14249v5PDF
0

Posted in cs.CV · 2026-01-20 · Zeyuan Chen, Kai Zhang, Zhuowen Tu, Yuanjun Xiong

Soft Tail-dropping for Adaptive Visual Tokenization

We present Soft Tail-dropping Adaptive Tokenizer (STAT), a 1D discrete visual tokenizer that adaptively chooses the number of output tokens per image according to its structural complexity and level of detail. STAT encodes an image into a sequence of discrete codes together with per-token keep probabilities. Beyond standard...

💬 0 commentsarXiv:2601.14246v1PDF
0

Posted in cs.IR · 2026-01-20 · Zhongyu Yang, Wei Pang, Yingfang Yuan

XR: Cross-Modal Agents for Composed Image Retrieval

Retrieval is being redefined by agentic AI, demanding multimodal reasoning beyond conventional similarity-based paradigms. Composed Image Retrieval (CIR) exemplifies this shift as each query combines a reference image with textual modifications, requiring compositional understanding across modalities. While embedding-based CIR methods...

💬 0 commentsarXiv:2601.14245v2PDF
0

Posted in cs.LG · 2026-01-20 · Haocheng Xi, Charlie Ruan, Peiyuan Liao, Yujun Lin, Han Cai, Yilong Zhao, Shuo Yang, Kurt Keutzer, Song Han, Ligeng Zhu

Jet-RL: Enabling On-Policy FP8 Reinforcement Learning with Unified Training and Rollout Precision Flow

Reinforcement learning (RL) is essential for enhancing the complex reasoning capabilities of large language models (LLMs). However, existing RL training pipelines are computationally inefficient and resource-intensive, with the rollout phase accounting for over 70% of total training time. Quantized RL training, particularly using FP8...

💬 0 commentsarXiv:2601.14243v2PDF
0

Posted in cs.CL · 2026-01-20 · Bertie Vidgen, Austin Mann, Abby Fennelly, John Wright Stanly, Lucas Rothman, Marco Burstein, Julien Benchek, David Ostrofsky, Anirudh Ravichandran, Debnil Sur, Neel Venugopal, Alannah Hsia, Isaac Robinson, Calix Huang, Olivia Varones, Daniyal Khan, Michael Haines, Austin Bridges, Jesse Boyle, Koby Twist, Zach Richards, Chirag Mahapatra, Brendan Foody, Osvald Nitski

APEX-Agents

We introduce the AI Productivity Index for Agents (APEX-Agents), a benchmark for assessing whether AI agents can execute long-horizon, cross-application tasks created by investment banking analysts, management consultants, and corporate lawyers. APEX-Agents requires agents to navigate realistic work environments with files and tools....

💬 0 commentsarXiv:2601.14242v3PDF
0

Posted in cs.SD · 2026-01-20 · Aafiya Hussain, Gaurav Srivastava, Alvi Ishmam, Zaber Hakim, Chris Thomas

SoundBreak: A Systematic Study of Audio-Only Adversarial Attacks on Trimodal Models

Multimodal foundation models that integrate audio, vision, and language achieve strong performance on reasoning and generation tasks, yet their robustness to adversarial manipulation remains poorly understood. We study a realistic and underexplored threat model: untargeted, audio-only adversarial attacks on trimodal...

💬 0 commentsarXiv:2601.16231v1PDF
0

Posted in cs.LG · 2026-01-20 · Shaurya Mathur, Shreyas Bellary Manjunath, Nitin Kulkarni, Alina Vereshchaka

Spatiotemporal Wildfire Prediction and Reinforcement Learning for Helitack Suppression

Wildfires are growing in frequency and intensity, devastating ecosystems and communities while causing billions of dollars in suppression costs and economic damage annually in the U.S. Traditional wildfire management is mostly reactive, addressing fires only after they are detected. We introduce \textit{FireCastRL}, a proactive...

💬 0 commentsarXiv:2601.14238v1PDF
0

Posted in cs.IT · 2026-01-20 · Giulio Pech, Mert Gökduman, Hanwen Yao, Henry D. Pfister

Stabilizer-Assisted Inactivation Decoding of Quantum Error-Correcting Codes with Erasures

In this work, we develop a reduced complexity maximum likelihood (ML) decoder for quantum low-density parity-check (QLDPC) codes over erasures. Our decoder combines classical inactivation decoding, which integrates peeling with symbolic guessing, with a new dual peeling procedure. In the dual peeling stage, we perform row operations...

💬 0 commentsarXiv:2601.14236v1PDF
0

Posted in cs.LG · 2026-01-20 · Qiyang Li, Sergey Levine

Q-learning with Adjoint Matching

We propose Q-learning with Adjoint Matching (QAM), a novel TD-based reinforcement learning (RL) algorithm that tackles a long-standing challenge in continuous-action RL: efficient optimization of an expressive diffusion or flow-matching policy with respect to a parameterized Q-function. Effective optimization requires exploiting the...

💬 0 commentsarXiv:2601.14234v4PDF