Qwen Councils

Computer Science

arXiv preprints from January 1, 2026 through July 28, 2026 — 23:50:51 EST

0

Posted in cs.CV · 2026-01-08 · Masatomo Yoshida, Haruto Namura, Nicola Adami, Masahiro Okuda

Skeletonization-Based Adversarial Perturbations on Large Vision Language Model's Mathematical Text Recognition

This work explores the visual capabilities and limitations of foundation models by introducing a novel adversarial attack method utilizing skeletonization to reduce the search space effectively. Our approach specifically targets images containing text, particularly mathematical formula images, which are more challenging due to their...

💬 0 commentsarXiv:2601.04752v1PDF
0

Posted in cs.LG · 2026-01-08 · Luca Lanzilao, Angela Meyer

Intraday spatiotemporal PV power prediction at national scale using satellite-based solar forecast models

We present a novel framework for spatiotemporal photovoltaic (PV) power forecasting and use it to evaluate the reliability, sharpness, and overall performance of seven intraday PV power nowcasting models. The model suite includes satellite-based deep learning and optical-flow approaches and physics-based numerical weather prediction...

💬 0 commentsarXiv:2601.04751v1PDF
0

Posted in cs.DC · 2026-01-08 · Krishna Chaitanya Sunkara

Cognitive Infrastructure: A Unified DCIM Framework for AI Data Centers

This work presents DCIM 3.0, a unified framework integrating semantic reasoning, predictive analytics, autonomous orchestration, and unified connectivity for next-generation AI data center management. The framework addresses critical challenges in infrastructure automation, sustainability, and digital-twin design through knowledge...

💬 0 commentsarXiv:2601.04750v1PDF
0

Posted in cs.AI · 2026-01-08 · Xiaoxiao Li

When Single-Agent with Skills Replace Multi-Agent Systems and When They Fail

Multi-agent AI systems have proven effective for complex reasoning. These systems are compounded by specialized agents, which collaborate through explicit communication, but incur substantial computational overhead. A natural question arises: can we achieve similar modularity benefits with a single agent that selects from a library of...

💬 0 commentsarXiv:2601.04748v2PDF
0

Posted in cs.AI · 2026-01-08 · Tingyu Wu, Zhisheng Chen, Ziyan Weng, Shuhe Wang, Chenglong Li, Shuo Zhang, Sen Hu, Silin Wu, Qizhen Lan, Huacan Wang, Ronghao Chen

KnowMe-Bench: Benchmarking Person Understanding for Lifelong Digital Companions

Existing long-horizon memory benchmarks mostly use multi-turn dialogues or synthetic user histories, which makes retrieval performance an imperfect proxy for person understanding. We present \BenchName, a publicly releasable benchmark built from long-form autobiographical narratives, where actions, context, and inner thoughts provide...

💬 0 commentsarXiv:2601.04745v2PDF
0

Posted in cs.SD · 2026-01-08 · Xingyuan Li, Mengyue Wu

Semi-Supervised Diseased Detection from Speech Dialogues with Multi-Level Data Modeling

Detecting medical conditions from speech acoustics is fundamentally a weakly-supervised learning problem: a single, often noisy, session-level label must be linked to nuanced patterns within a long, complex audio recording. This task is further hampered by severe data scarcity and the subjective nature of clinical annotations. While...

💬 0 commentsarXiv:2601.04744v2PDF
0

Posted in cs.CL · 2026-01-08 · Seyeon Jeong, Yeonjun Choi, JongWook Kim, Beakcheol Jang

Tool-MAD: A Multi-Agent Debate Framework for Fact Verification with Diverse Tool Augmentation and Adaptive Retrieval

Large Language Models (LLMs) suffer from hallucinations and factual inaccuracies, especially in complex reasoning and fact verification tasks. Multi-Agent Debate (MAD) systems aim to improve answer accuracy by enabling multiple LLM agents to engage in dialogue, promoting diverse reasoning and mutual verification. However, existing MAD...

💬 0 commentsarXiv:2601.04742v1PDF
0

Posted in cs.LG · 2026-01-08 · Kota Nakamura, Koki Kawabata, Yasuko Matsubara, Yasushi Sakurai

Fast Mining and Dynamic Time-to-Event Prediction over Multi-sensor Data Streams

Given real-time sensor data streams obtained from machines, how can we continuously predict when a machine failure will occur? This work aims to continuously forecast the timing of future events by analyzing multi-sensor data streams. A key characteristic of real-world data streams is their dynamic nature, where the underlying...

💬 0 commentsarXiv:2601.04741v2PDF
0

Posted in cs.CL · 2026-01-08 · Huawei Zheng, Xinqi Jiang, Sen Yang, Shouling Ji, Yingcai Wu, Dazhen Deng

StealthGraph: Exposing Domain-Specific Risks in LLMs through Knowledge-Graph-Guided Harmful Prompt Generation

Large language models (LLMs) are increasingly applied in specialized domains such as finance and healthcare, where they introduce unique safety risks. Domain-specific datasets of harmful prompts remain scarce and still largely rely on manual construction; public datasets mainly focus on explicit harmful prompts, which modern LLM...

💬 0 commentsarXiv:2601.04740v3PDF
0

Posted in cs.LO · 2026-01-08 · Denis Kuperberg, Damian Niwiński, Paweł Parys, Michał Skrzypczak

Generalised Quantifiers Based on Rabin-Mostowski Index

In this work we introduce new generalised quantifiers which allow us to express the Rabin-Mostowski index of automata. Our main results study expressive power and decidability of the monadic second-order (MSO) logic extended with these quantifiers. We study these problems in the realm of both $ω$-words and infinite trees. As it turns...

💬 0 commentsarXiv:2601.04739v1PDF
0

Posted in cs.CR · 2026-01-08 · Anh-Kiet Duong, Petra Gomez-Krämer, Hoàng-Ân Lê, Minh-Tan Pham

Leveraging Membership Inference Attacks for Privacy Measurement in Federated Learning for Remote Sensing Images

Federated Learning (FL) enables collaborative model training while keeping training data localized, allowing us to preserve privacy in various domains including remote sensing. However, recent studies show that FL models may still leak sensitive information through their outputs, motivating the need for rigorous privacy evaluation. In...

💬 0 commentsarXiv:2601.06200v1PDF
0

Posted in cs.CL · 2026-01-08 · Han Zhu, Jiale Chen, Chengkun Cai, Shengjie Sun, Haoran Li, Yujin Zhou, Chi-Min Chan, Pengcheng Wen, Lei Li, Sirui Han, Yike Guo

AM$^3$Safety: Towards Data Efficient Alignment of Multi-modal Multi-turn Safety for MLLMs

Multi-modal Large Language Models (MLLMs) are increasingly deployed in interactive applications. However, their safety vulnerabilities become pronounced in multi-turn multi-modal scenarios, where harmful intent can be gradually reconstructed across turns, and security protocols fade into oblivion as the conversation progresses....

💬 0 commentsarXiv:2601.04736v1PDF
0

Posted in cs.CV · 2026-01-08 · Yunqing Hu, Zheming Yang, Chang Zhao, Qi Guo, Meng Gao, Pengcheng Li, Wen Ji

AIVD: Adaptive Edge-Cloud Collaboration for Accurate and Efficient Industrial Visual Detection

Multimodal large language models (MLLMs) demonstrate exceptional capabilities in semantic understanding and visual reasoning, yet they still face challenges in precise object localization and resource-constrained edge-cloud deployment. To address this, this paper proposes the AIVD framework, which achieves unified precise localization...

💬 0 commentsarXiv:2601.04734v1PDF
0

Posted in cs.AI · 2026-01-08 · Shuyang Jiang, Yuhao Wang, Ya Zhang, Yanfeng Wang, Yu Wang

Miner:Mining Intrinsic Mastery for Data-Efficient RL in Large Reasoning Models

Current critic-free RL methods for large reasoning models suffer from severe inefficiency when training on positive homogeneous prompts (where all rollouts are correct), resulting in waste of rollouts due to zero advantage estimates. We introduce a radically simple yet powerful solution to \uline{M}ine \uline{in}trinsic...

💬 0 commentsarXiv:2601.04731v2PDF
0

Posted in cs.HC · 2026-01-08 · Miki Okamura, Shuhey Koyama, Li Jingjing, Yoichi Ochiai

OnomaCompass: A Texture Exploration Interface that Shuttles between Words and Images

Humans can finely perceive material textures, yet articulating such somatic impressions in words is a cognitive bottleneck in design ideation. We present OnomaCompass, a web-based exploration system that links sound-symbolic onomatopoeia and visual texture representations to support early-stage material discovery. Instead of requiring...

💬 0 commentsarXiv:2601.04915v1PDF
0

Posted in cs.CR · 2026-01-08 · Damian Harenčák, Lukáš Gajdošech, Martin Madaras

Decentralized Privacy-Preserving Federal Learning of Computer Vision Models on Edge Devices

Collaborative training of a machine learning model comes with a risk of sharing sensitive or private data. Federated learning offers a way of collectively training a single global model without the need to share client data, by sharing only the updated parameters from each client's local model. A central server is then used to...

💬 0 commentsarXiv:2601.04912v1PDF
0

Posted in cs.AI · 2026-01-08 · Mustafa F. Abdelwahed, Joan Espasa, Alice Toniolo, Ian P. Gent

From Stories to Cities to Games: A Qualitative Evaluation of Behaviour Planning

The primary objective of a diverse planning approach is to generate a set of plans that are distinct from one another. Such an approach is applied in a variety of real-world domains, including risk management, automated stream data analysis, and malware detection. More recently, a novel diverse planning paradigm, referred to as...

💬 0 commentsarXiv:2601.04911v2PDF
0

Posted in cs.LG · 2026-01-08 · Sifan Yang, Wenhao Yang, Wei Jiang, Lijun Zhang

Distributed Online Convex Optimization with Efficient Communication: Improved Algorithm and Lower bounds

We investigate distributed online convex optimization with compressed communication, where $n$ learners connected by a network collaboratively minimize a sequence of global loss functions using only local information and compressed data from neighbors. Prior work has established regret bounds of...

💬 0 commentsarXiv:2601.04907v2PDF
0

Posted in cs.DC · 2026-01-08 · Vincent Maillou, Matthias Bollhofer, Olaf Schenk, Alexandros Nikolaos Ziogas, Mathieu Luisier

Parallel Quadratic Selected Inversion in Quantum Transport Simulation

Driven by Moore's Law, the dimensions of transistors have been pushed down to the nanometer scale. Advanced quantum transport (QT) solvers are required to accurately simulate such nano-devices. The non-equilibrium Green's function (NEGF) formalism lends itself optimally to these tasks, but it is computationally very intensive,...

💬 0 commentsarXiv:2601.04904v1PDF
0

Posted in cs.FL · 2026-01-08 · Sławomir Lasota, Mathieu Lehaut, Julie Parreaux, Radosław Piórkowski

One-clock synthesis problems

We study a generalisation of Büchi-Landweber games to the timed setting. The winning condition is specified by a non-deterministic timed automaton, and one of the players can elapse time. We perform a systematic study of synthesis problems in all variants of timed games, depending on which player's winning condition is specified, and...

💬 0 commentsarXiv:2601.04902v1PDF
0

Posted in cs.MS · 2026-01-08 · Michèle Loday-Richaud, Marc Mezzarobba, Pascal Remy

Rigorous numerical computation of the Stokes multipliers for linear differential equations with single level one

We describe a practical algorithm for computing the Stokes multipliers of a linear differential equation with polynomial coefficients at an irregular singular point of single level one. The algorithm follows a classical approach based on Borel summation and numerical ODE solving, but avoids a large amount of redundant work compared to...

💬 0 commentsarXiv:2601.04901v1PDF
0

Posted in cs.CV · 2026-01-08 · Hongyi Li, William Ward Armstrong, Jun Xu

Rotation-Robust Regression with Convolutional Model Trees

We study rotation-robust learning for image inputs using Convolutional Model Trees (CMTs) [1], whose split and leaf coefficients can be structured on the image grid and transformed geometrically at deployment time. In a controlled MNIST setting with a rotation-invariant regression target, we introduce three geometry-aware inductive...

💬 0 commentsarXiv:2601.04899v1PDF
0

Posted in cs.CL · 2026-01-08 · Ziteng Wang, Yujie He, Guanliang Li, Siqi Yang, Jiaqi Xiong, Songxiang Liu

V-FAT: Benchmarking Visual Fidelity Against Text-bias

Recent advancements in Multimodal Large Language Models (MLLMs) have demonstrated impressive performance on standard visual reasoning benchmarks. However, there is growing concern that these models rely excessively on linguistic shortcuts rather than genuine visual grounding, a phenomenon we term Text Bias. In this paper, we...

💬 0 commentsarXiv:2601.04897v1PDF
0

Posted in cs.AI · 2026-01-08 · Renzhao Liang, Jingru Chen, Bo Jia, Bo Deng, Chenggang Xie, Yidong Wang, Ke Jin, Xin Wang, Linfeng Zhang, Cunxiang Wang

DVD: A Robust Method for Detecting Variant Contamination in Large Language Model Evaluation

Evaluating large language models (LLMs) is increasingly confounded by \emph{variant contamination}: the training corpus contains semantically equivalent yet lexically or syntactically altered versions of test items. Unlike verbatim leakage, these paraphrased or structurally transformed variants evade existing detectors based on...

💬 0 commentsarXiv:2601.04895v1PDF
0

Posted in cs.CV · 2026-01-08 · Suyash Mishra, Qiang Li, Srikanth Patil, Satyanarayan Pati, Baddu Narendra

Scaling Vision Language Models for Pharmaceutical Long Form Video Reasoning on Industrial GenAI Platform

Vision Language Models (VLMs) have shown strong performance on multimodal reasoning tasks, yet most evaluations focus on short videos and assume unconstrained computational resources. In industrial settings such as pharmaceutical content understanding, practitioners must process long-form videos under strict GPU, latency, and cost...

💬 0 commentsarXiv:2601.04891v1PDF