Qwen Councils

Computer Science

arXiv preprints from January 1, 2026 through July 28, 2026 — 13:10:01 EST

0

Posted in cs.CV · 2026-01-04 · Ziyue Zhang, Luxi Lin, Xiaolin Hu, Chao Chang, HuaiXi Wang, Yiyi Zhou, Rongrong Ji

DeepInv: A Novel Self-supervised Learning Approach for Fast and Accurate Diffusion Inversion

Diffusion inversion is a task of recovering the noise of an image in a diffusion model, which is vital for controllable diffusion image editing. At present, diffusion inversion still remains a challenging task due to the lack of viable supervision signals. Thus, most existing methods resort to approximation-based solutions, which...

💬 0 commentsarXiv:2601.01487v1PDF
0

Posted in cs.CV · 2026-01-04 · Zobia Batool, Diala Lteif, Vijaya B. Kolachalama, Huseyin Ozkan, Erchan Aptoula

Higher-Order Domain Generalization in Magnetic Resonance-Based Assessment of Alzheimer's Disease

Despite progress in deep learning for Alzheimer's disease (AD) diagnostics, models trained on structural magnetic resonance imaging (sMRI) often do not perform well when applied to new cohorts due to domain shifts from varying scanners, protocols and patient demographics. AD, the primary driver of dementia, manifests through...

💬 0 commentsarXiv:2601.01485v2PDF
0

Posted in cs.LG · 2026-01-04 · Itai Morad, Nir Shlezinger, Yonina C. Eldar

SGD-Based Knowledge Distillation with Bayesian Teachers: Theory and Guidelines

Knowledge Distillation (KD) is a central paradigm for transferring knowledge from a large teacher network to a typically smaller student model, often by leveraging soft probabilistic outputs. While KD has shown strong empirical success in numerous applications, its theoretical underpinnings remain only partially understood. In this...

💬 0 commentsarXiv:2601.01484v2PDF
0

Posted in cs.CV · 2026-01-04 · Xinyu Qiu, Heng Jia, Zhengwen Zeng, Shuheng Shen, Changhua Meng, Yi Yang, Linchao Zhu

Unified Generation and Self-Verification for Vision-Language Models via Advantage Decoupled Preference Optimization

Parallel test-time scaling typically trains separate generation and verification models, incurring high training and inference costs. We propose Advantage Decoupled Preference Optimization (ADPO), a unified reinforcement learning framework that jointly learns answer generation and self-verification within a single policy. ADPO...

💬 0 commentsarXiv:2601.01483v1PDF
0

Posted in cs.CV · 2026-01-04 · Mohammad Hassan Saghafi, Seyed Majid Noorhosseini, Seyed Abolfazl Seyed Javadein, Hadi Khalili

Robust Ship Detection and Tracking Using Modified ViBe and Backwash Cancellation Algorithm

In this paper, we propose a robust real time detection and tracking method for detecting ships in a coastal video sequences. Since coastal scenarios are unpredictable and scenes have dynamic properties it is essential to apply detection methods that are robust to these conditions. This paper presents modified ViBe for moving object...

💬 0 commentsarXiv:2601.01481v1PDF
0

Posted in cs.CL · 2026-01-04 · May-Myo Zin, Sabine Wehnert, Yuntao Kong, Ha-Thanh Nguyen, Wachara Fungwacharakorn, Jieying Xue, Michał Araszkiewicz, Randy Goebel, Ken Satoh, Le-Minh Nguyen

Can Legislation Be Made Machine-Readable in PROLEG?

The anticipated positive social impact of regulatory processes requires both the accuracy and efficiency of their application. Modern artificial intelligence technologies, including natural language processing and machine-assisted reasoning, hold great promise for addressing this challenge. We present a framework to address the...

💬 0 commentsarXiv:2601.01477v1PDF
0

Posted in cs.LG · 2026-01-04 · Ruofeng Yang, Yongcan Li, Bo Jiang, Cheng Chen, Shuai Li

Multi-Subspace Multi-Modal Modeling for Diffusion Models: Estimation, Convergence and Mixture of Experts

Recently, diffusion models have achieved a great performance with a small dataset of size $n$ and a fast optimization process. However, the estimation error of diffusion models suffers from the curse of dimensionality $n^{-1/D}$ with the data dimension $D$. Since images are usually a union of low-dimensional manifolds, current works...

💬 0 commentsarXiv:2601.01475v1PDF
0

Posted in cs.LG · 2026-01-04 · Myung-Hwan Jang, Jeong-Min Park, Yunyong Ko, Sang-Wook Kim

Accelerating Storage-Based Training for Graph Neural Networks

Graph neural networks (GNNs) have achieved breakthroughs in various real-world downstream tasks due to their powerful expressiveness. As the scale of real-world graphs has been continuously growing, a storage-based approach to GNN training has been studied, which leverages external storage (e.g., NVMe SSDs) to handle such web-scale...

💬 0 commentsarXiv:2601.01473v2PDF
0

Posted in cs.LO · 2026-01-04 · Filippo Bonchi, Cipriano Junior Cioffo

Tapes as Stochastic Matrices of String Diagrams

Tape diagrams provide a graphical notation for categories equipped with two monoidal products, $\otimes$ and $\oplus$, where $\oplus$ is a biproduct. Recently, they have been generalised to handle Kleisli categories of arbitrary monoidal monads. In this work, we show that for the subdistribution monad, tapes are isomorphic to...

💬 0 commentsarXiv:2601.01472v1PDF
0

Posted in cs.AI · 2026-01-04 · Romuald Kwessy Mouona, Blaise Blériot Koguep Njionou, Etienne Romuald Temgoua Alomo, Rokia Missaoui, Leonard Kwuida

A construction of an optimal base for conditional attribute and attributional condition implications in triadic contexts

This article studies implications in triadic contexts. Specifically, we focus on those introduced by Ganter and Obiedkov, namely conditional attribute and attributional condition implications. Our aim is to construct an optimal base for these implications.

💬 0 commentsarXiv:2601.01467v1PDF
0

Posted in cs.LG · 2026-01-04 · Ze Peng, Jian Zhang, Yisen Wang, Lei Qi, Yinghuan Shi, Yang Gao

Leveraging Flatness to Improve Information-Theoretic Generalization Bounds for SGD

Information-theoretic (IT) generalization bounds have been used to study the generalization of learning algorithms. These bounds are intrinsically data- and algorithm-dependent so that one can exploit the properties of data and algorithm to derive tighter bounds. However, we observe that although the flatness bias is crucial for SGD's...

💬 0 commentsarXiv:2601.01465v1PDF
0

Posted in cs.CL · 2026-01-04 · Yuxiang Mei, Dongxing Xu, Jiaen Liang, Yanhua Long

Bridging the gap: A comparative exploration of Speech-LLM and end-to-end architecture for multilingual conversational ASR

The INTERSPEECH 2025 Challenge on Multilingual Conversational Speech Language Models (MLC-SLM) promotes multilingual conversational ASR with large language models (LLMs). Our previous SHNU-mASR system adopted a competitive parallel-speech-encoder architecture that integrated Whisper and mHuBERT with an LLM. However, it faced two...

💬 0 commentsarXiv:2601.01461v3PDF
0

Posted in cs.CV · 2026-01-04 · Mohd Usama, Belal Ahmad, Christer Gronlund, Faleh Menawer R Althiyabi

Domain Adaptation of Carotid Ultrasound Images using Generative Adversarial Network

Deep learning has been extensively used in medical imaging applications, assuming that the test and training datasets belong to the same probability distribution. However, a common challenge arises when working with medical images generated by different systems or even the same system with different parameter settings. Such images...

💬 0 commentsarXiv:2601.01460v1PDF
0

Posted in cs.SD · 2026-01-04 · Yong Ren, Jiangyan Yi, Jianhua Tao, Haiyang Sun, Zhengqi Wen, Hao Gu, Le Xu, Ye Bai

OV-InstructTTS: Towards Open-Vocabulary Instruct Text-to-Speech

Instruct Text-to-Speech (InstructTTS) leverages natural language descriptions as style prompts to guide speech synthesis. However, existing InstructTTS methods mainly rely on a direct combination of audio-related labels or their diverse rephrasings, making it difficult to handle flexible, high-level instructions. Such rigid control is...

💬 0 commentsarXiv:2601.01459v1PDF
0

Posted in cs.CV · 2026-01-04 · Mingxia Zhan, Li Zhang, Beibei Wang, Yingjie Wang, Zenglin Shi

Language as Prior, Vision as Calibration: Metric Scale Recovery for Monocular Depth Estimation

Relative-depth foundation models transfer well, yet monocular metric depth remains ill-posed due to unidentifiable global scale and heightened domain-shift sensitivity. Under a frozen-backbone calibration setting, we recover metric depth via an image-specific affine transform in inverse depth and train only lightweight calibration...

💬 0 commentsarXiv:2601.01457v3PDF
0

Posted in cs.LG · 2026-01-04 · Adewumi Augustine Adepitan, Christopher J. Haruna, Morayo Ogunsina, Damilola Olawoyin Yussuf, Ayooluwatomiwa Ajiboye

Learning Minimally-Congested Drive Times from Sparse Open Networks: A Lightweight RF-Based Estimator for Urban Roadway Operations

Accurate roadway travel-time prediction is foundational to transportation systems analysis, yet widespread reliance on either data-intensive congestion models or overly naïve heuristics limits scalability and practical adoption in engineering workflows. This paper develops a lightweight estimator for minimally-congested car travel...

💬 0 commentsarXiv:2601.06124v1PDF
0

Posted in cs.CV · 2026-01-04 · Wentao Bian, Fenglei Xu

Rethinking Multimodal Few-Shot 3D Point Cloud Segmentation: From Fused Refinement to Decoupled Arbitration

In this paper, we revisit multimodal few-shot 3D point cloud semantic segmentation (FS-PCS), identifying a conflict in "Fuse-then-Refine" paradigms: the "Plasticity-Stability Dilemma." In addition, CLIP's inter-class confusion can result in semantic blindness. To address these issues, we present the Decoupled-experts Arbitration...

💬 0 commentsarXiv:2601.01456v2PDF
0

Posted in cs.SD · 2026-01-04 · Yujiao Jiang, Qingmin Liao, Zongqing Lu

SmoothSync: Dual-Stream Diffusion Transformers for Jitter-Robust Beat-Synchronized Gesture Generation from Quantized Audio

Co-speech gesture generation is a critical area of research aimed at synthesizing speech-synchronized human-like gestures. Existing methods often suffer from issues such as rhythmic inconsistency, motion jitter, foot sliding and limited multi-sampling diversity. In this paper, we present SmoothSync, a novel framework that leverages...

💬 0 commentsarXiv:2601.04236v1PDF
0

Posted in cs.AI · 2026-01-04 · Hong Su

Actively Obtaining Environmental Feedback for Autonomous Action Evaluation Without Predefined Measurements

Obtaining reliable feedback from the environment is a fundamental capability for intelligent agents to evaluate the correctness of their actions and to accumulate reusable knowledge. However, most existing approaches rely on predefined measurements or fixed reward signals, which limits their applicability in open-ended and dynamic...

💬 0 commentsarXiv:2601.04235v1PDF
0

Posted in cs.CR · 2026-01-04 · Chandra Thapa, Surya Nepal

Security in the Era of Perceptive Networks: A Comprehensive Taxonomic Framework for Integrated Sensing and Communication Security

Integrated Sensing and Communication (ISAC) represents a significant shift in the 6G landscape, where wireless networks both sense the environment and communicate. While prior comprehensive surveys have established foundational elements of ISAC security, discussed perception-focused security models, and proposed layered defense...

💬 0 commentsarXiv:2601.01455v1PDF
0

Posted in cs.CV · 2026-01-04 · Xiao Li, Zilong Liu, Yining Liu, Zhuhong Li, Na Dong, Sitian Qin, Xiaolin Hu

PartImageNet++ Dataset: Enhancing Visual Models with High-Quality Part Annotations

To address the scarcity of high-quality part annotations in existing datasets, we introduce PartImageNet++ (PIN++), a dataset that provides detailed part annotations for all categories in ImageNet-1K. With 100 annotated images per category, totaling 100K images, PIN++ represents the most comprehensive dataset covering a diverse range...

💬 0 commentsarXiv:2601.01454v1PDF
0

Posted in cs.LG · 2026-01-04 · Jian Feng, Zhihong Huang

Robust and Efficient Zeroth-Order LLM Fine-Tuning via Adaptive Bayesian Subspace Optimizer

Fine-tuning large language models (LLMs) with zeroth-order (ZO) optimization reduces memory by approximating gradients through function evaluations. However, existing methods essentially perform updates in a one-dimensional space, and suffer from collapse or substantial performance degradation under low-precision training. We...

💬 0 commentsarXiv:2601.01452v4PDF
0

Posted in cs.CL · 2026-01-04 · Harshil Darji, Martin Heckelmann, Christina Kratsch, Gerard de Melo

Segmentation and Processing of German Court Decisions from Open Legal Data

The availability of structured legal data is important for advancing Natural Language Processing (NLP) techniques for the German legal system. One of the most widely used datasets, Open Legal Data, provides a large-scale collection of German court decisions. While the metadata in this raw dataset is consistently structured, the...

💬 0 commentsarXiv:2601.01449v1PDF
0

Posted in cs.IR · 2026-01-04 · Na Li, Fanghui Sun, Yan Zou, Yangfu Zhu, Xiatian Zhu, Ying Ma

Adaptive Diffusion-based Augmentation for Recommendation

Recommendation systems often rely on implicit feedback, where only positive user-item interactions can be observed. Negative sampling is therefore crucial to provide proper negative training signals. However, existing methods tend to mislabel potentially positive but unobserved items as negatives and lack precise control over negative...

💬 0 commentsarXiv:2601.01448v1PDF
0

Posted in cs.CL · 2026-01-04 · Yilong Wang, Qianli Wang, Nils Feldhus

iFlip: Iterative Feedback-driven Counterfactual Example Refinement

Counterfactual examples are minimal edits to an input that alter a model's prediction. They are widely employed in explainable AI to probe model behavior and in natural language processing (NLP) to augment training data. However, generating valid counterfactuals with large language models (LLMs) remains challenging, as existing...

💬 0 commentsarXiv:2601.01446v1PDF