Qwen Councils

Computer Science

arXiv preprints from January 1, 2026 through July 28, 2026 — 12:08:38 EST

0

Posted in cs.LG · 2026-01-04 · Ke Xiao, Haoze Zhang, Runze Mao, Han Li, Zhi X. Chen

Towards LLM-enabled autonomous combustion research: A literature-aware agent for self-corrective modeling workflows

The rapid evolution of large language models (LLMs) is transforming artificial intelligence into autonomous research partners, yet a critical gap persists in complex scientific domains such as combustion modeling. Here, practical AI assistance requires the seamless integration of domain literature knowledge with robust execution...

💬 0 commentsarXiv:2601.01357v1PDF
0

Posted in cs.CV · 2026-01-04 · Dang H. Pham, Tu N. Nguyen, Hoa N. Nguyen

Advanced Machine Learning Approaches for Enhancing Person Re-Identification Performance

Person re-identification (ReID) plays a critical role in intelligent surveillance systems by linking identities across multiple cameras in complex environments. However, ReID faces significant challenges such as appearance variations, domain shifts, and limited labeled data. This dissertation proposes three advanced approaches to...

💬 0 commentsarXiv:2601.01356v1PDF
0

Posted in cs.CV · 2026-01-04 · Canming Xia, Peixi Peng, Guang Tan, Zhan Su, Haoran Xu, Zhenxian Liu, Luntong Li

COVR:Collaborative Optimization of VLMs and RL Agent for Visual-Based Control

Visual reinforcement learning (RL) suffers from poor sample efficiency due to high-dimensional observations in complex tasks. While existing works have shown that vision-language models (VLMs) can assist RL, they often focus on knowledge distillation from the VLM to RL, overlooking the potential of RL-generated interaction data to...

💬 0 commentsarXiv:2601.06122v1PDF
0

Posted in cs.SD · 2026-01-04 · Zhiyuan Zhao, Lijian Lin, Ye Zhu, Kai Xie, Yunfei Liu, Yu Li

LEMAS: Large A 150K-Hour Large-scale Extensible Multilingual Audio Suite with Generative Speech Models

We present the LEMAS-Dataset, which, to our knowledge, is currently the largest open-source multilingual speech corpus with word-level timestamps. Covering over 150,000 hours across 10 major languages, LEMAS-Dataset is constructed via a efficient data processing pipeline that ensures high-quality data and annotations. To validate the...

💬 0 commentsarXiv:2601.04233v1PDF
0

Posted in cs.CV · 2026-01-04 · Yixuan Lai, He Wang, Kun Zhou, Tianjia Shao

Slot-ID: Identity-Preserving Video Generation from Reference Videos via Slot-Based Temporal Identity Encoding

Producing prompt-faithful videos that preserve a user-specified identity remains challenging: models need to extrapolate facial dynamics from sparse reference while balancing the tension between identity preservation and motion naturalness. Conditioning on a single image completely ignores the temporal signature, which leads to...

💬 0 commentsarXiv:2601.01352v1PDF
0

Posted in cs.CL · 2026-01-04 · Juan Junqueras, Florian Boudin, May-Myo Zin, Ha-Thanh Nguyen, Wachara Fungwacharakorn, Damián Ariel Furman, Akiko Aizawa, Ken Satoh

FC-CONAN: An Exhaustively Paired Dataset for Robust Evaluation of Retrieval Systems

Hate speech (HS) is a critical issue in online discourse, and one promising strategy to counter it is through the use of counter-narratives (CNs). Datasets linking HS with CNs are essential for advancing counterspeech research. However, even flagship resources like CONAN (Chung et al., 2019) annotate only a sparse subset of all...

💬 0 commentsarXiv:2601.01350v1PDF
0

Posted in cs.LG · 2026-01-04 · Yuyan Pi, Min Jin, Wentao Xie, Xinhua Liu

From Classification to Generation: An Open-Ended Paradigm for Adverse Drug Reaction Prediction Based on Graph-Motif Feature Fusion

Computational biology offers immense potential for reducing the high costs and protracted cycles of new drug development through adverse drug reaction (ADR) prediction. However, current methods remain impeded by drug data scarcity-induced cold-start challenge, closed label sets, and inadequate modeling of label dependencies. Here we...

💬 0 commentsarXiv:2601.01347v1PDF
0

Posted in cs.CL · 2026-01-04 · Md Abdullah Al Kafi, Raka Moni, Sumit Kumar Banshal

Reasoning Over Recall: Evaluating the Efficacy of Generalist Architectures vs. Specialized Fine-Tunes in RAG-Based Mental Health Dialogue Systems

The deployment of Large Language Models (LLMs) in mental health counseling faces the dual challenges of hallucinations and lack of empathy. While the former may be mitigated by RAG (retrieval-augmented generation) by anchoring answers in trusted clinical sources, there remains an open question as to whether the most effective model...

💬 0 commentsarXiv:2601.01341v1PDF
0

Posted in cs.CV · 2026-01-04 · Weihang You, Hanqi Jiang, Yi Pan, Junhao Chen, Tianming Liu, Fei Dou

Achieving Fine-grained Cross-modal Understanding through Brain-inspired Hierarchical Representation Learning

Understanding neural responses to visual stimuli remains challenging due to the inherent complexity of brain representations and the modality gap between neural data and visual inputs. Existing methods, mainly based on reducing neural decoding to generation tasks or simple correlations, fail to reflect the hierarchical and temporal...

💬 0 commentsarXiv:2601.01339v1PDF
0

Posted in cs.CV · 2026-01-04 · Wenting Lu, Didi Zhu, Tao Shen, Donglin Zhu, Ayong Ye, Chao Wu

Watch Wider and Think Deeper: Collaborative Cross-modal Chain-of-Thought for Complex Visual Reasoning

Multi-modal reasoning requires the seamless integration of visual and linguistic cues, yet existing Chain-of-Thought methods suffer from two critical limitations in cross-modal scenarios: (1) over-reliance on single coarse-grained image regions, and (2) semantic fragmentation between successive reasoning steps. To address these...

💬 0 commentsarXiv:2601.02422v1PDF
0

Posted in cs.CL · 2026-01-04 · Hossam Amer, Maryam Dialameh, Hossein Rajabzadeh, Walid Ahmed, Weiwei Zhang, Yang Liu

FLOP-Efficient Training: Early Stopping Based on Test-Time Compute Awareness

Scaling training compute, measured in FLOPs, has long been shown to improve the accuracy of large language models, yet training remains resource-intensive. Prior work shows that increasing test-time compute (TTC)-for example through iterative sampling-can allow smaller models to rival or surpass much larger ones at lower overall cost....

💬 0 commentsarXiv:2601.01332v1PDF
0

Posted in cs.CY · 2026-01-04 · Hongkun Yang, Lionel Z. Wang, Wei Fan, Yiran Hu, Lixu Wang, Chenyu Liu, Yu Zeng, Shenghong Fu, Lei Gong, Zhengxin Zhang, Haoyang Li, Jiexin Zheng, Xin Xu

AppellateGen: A Benchmark for Appellate Legal Judgment Generation

Legal judgment generation is a critical task in legal intelligence. However, existing research in legal judgment generation has predominantly focused on first-instance trials, relying on static fact-to-verdict mappings while neglecting the dialectical nature of appellate (second-instance) review. To address this, we introduce...

💬 0 commentsarXiv:2601.01331v3PDF
0

Posted in cs.AI · 2026-01-04 · Shengji Tang, Weihao Lin, Peng Ye, Jingqi Ye, Hao Li, Yiqun Zhang, Xiaosong Wang, Bo Zhang, Shuyue Hu, Tao Chen, Lei Bai, Wanli Ouyang

Beyond Gemini-3-Pro: Revisiting LLM Routing and Aggregation at Scale

Large Language Models (LLMs) have rapidly advanced, with Gemini-3-Pro setting a new performance milestone. In this work, we explore collective intelligence as an alternative to monolithic scaling, and demonstrate that open-source LLMs' collaboration can surpass Gemini-3-Pro. We first revisit LLM routing and aggregation at scale and...

💬 0 commentsarXiv:2601.01330v2PDF
0

Posted in cs.CV · 2026-01-04 · Wenhui Chu, Aobo Jin, Hardik A. Gohel

A Novel Deep Learning Method for Segmenting the Left Ventricle in Cardiac Cine MRI

This research aims to develop a novel deep learning network, GBU-Net, utilizing a group-batch-normalized U-Net framework, specifically designed for the precise semantic segmentation of the left ventricle in short-axis cine MRI scans. The methodology includes a down-sampling pathway for feature extraction and an up-sampling pathway for...

💬 0 commentsarXiv:2601.01512v1PDF
0

Posted in cs.AI · 2026-01-04 · Ahmed Dawoud, Osama El-Shamy

Reading Between the Lines: Deconfounding Causal Estimates using Text Embeddings and Deep Learning

Estimating causal treatment effects in observational settings is frequently compromised by selection bias arising from unobserved confounders. While traditional econometric methods struggle when these confounders are orthogonal to structured covariates, high-dimensional unstructured text often contains rich proxies for these latent...

💬 0 commentsarXiv:2601.01511v1PDF
0

Posted in cs.CV · 2026-01-04 · Tao Li, Qing Li, Na Li, Hui Xie

DiffKD-DCIS: Predicting Upgrade of Ductal Carcinoma In Situ with Diffusion Augmentation and Knowledge Distillation

Accurately predicting the upgrade of ductal carcinoma in situ (DCIS) to invasive ductal carcinoma (IDC) is crucial for surgical planning. However, traditional deep learning methods face challenges due to limited ultrasound data and poor generalization ability. This study proposes the DiffKD-DCIS framework, integrating conditional...

💬 0 commentsarXiv:2601.01507v1PDF
0

Posted in cs.IT · 2026-01-04 · Shengcai Zhou, Luping Xiang, Yi Wang, Kun Yang, Kai Kit Wong, Chan-Byoung Chae

Extended Target Adaptive Beamforming for ISAC:A Perspective of Predictive Error Ellipse

Utilizing communication signals to extract motion parameters has emerged as a key direction in Vehicle-to- Everything (V2X) networks. Accurately modeling the relationship between communication signals and sensing performance is critical for the advancement of such systems. Unlike prior work that relies primarily on qualitative...

💬 0 commentsarXiv:2601.06125v1PDF
0

Posted in cs.LG · 2026-01-04 · Fan Xu, Wei Gong, Hao Wu, Lilan Peng, Nan Wang, Qingsong Wen, Xian Wu, Kun Wang, Xibin Zhao

Advanced Global Wildfire Activity Modeling with Hierarchical Graph ODE

Wildfires, as an integral component of the Earth system, are governed by a complex interplay of atmospheric, oceanic, and terrestrial processes spanning a vast range of spatiotemporal scales. Modeling their global activity on large timescales is therefore a critical yet challenging task. While deep learning has recently achieved...

💬 0 commentsarXiv:2601.01501v1PDF
0

Posted in cs.DC · 2026-01-04 · Jinxiao Zhang, Yunpu Xu, Xiyong Wu, Runmin Dong, Shenggan Cheng, Yi Zhao, Mengxuan Chen, Qinrui Zheng, Jianting Liu, Haohuan Fu

DiT-HC: Enabling Efficient Training of Visual Generation Model DiT on HPC-oriented CPU Cluster

Generative foundation models have become an important tool for data reconstruction and simulation in scientific computing, showing a tight integration with traditional numerical simulations. At the same time, with the development of new hardware features, such as matrix acceleration units and high-bandwidth memory, CPU-based clusters...

💬 0 commentsarXiv:2601.01500v2PDF
0

Posted in cs.CL · 2026-01-04 · Bingguang Hao, Zengzhuang Xu, Yuntao Wen, Xinyi Xu, Yang Liu, Tong Zhao, Maolin Wang, Long Chen, Dong Wang, Yicheng Chen, Cunyin Peng, Xiangyu Zhao, Chenyi Zhuang, Ji Zhang

From Failure to Mastery: Generating Hard Samples for Tool-use Agents

The advancement of LLM agents with tool-use capabilities requires diverse and complex training corpora. Existing data generation methods, which predominantly follow a paradigm of random sampling and shallow generation, often yield simple and homogeneous trajectories that fail to capture complex, implicit logical dependencies. To...

💬 0 commentsarXiv:2601.01498v1PDF
0

Posted in cs.GT · 2026-01-04 · Mikael Møller Høgsgaard

The Optimal Sample Complexity of Linear Contracts

In this paper, we settle the problem of learning optimal linear contracts from data in the offline setting, where agent types are drawn from an unknown distribution and the principal's goal is to design a contract that maximizes her expected utility. Specifically, our analysis shows that the simple Empirical Utility Maximization (EUM)...

💬 0 commentsarXiv:2601.01496v2PDF
0

Posted in cs.IR · 2026-01-04 · Annelies de Jong, Giuseppe Cascavilla, Jessica De Pascale

Breadcrumbs in the Digital Forest: Tracing Criminals through Torrent Metadata with OSINT

This work investigates the potential of torrent metadata as a source for open-source intelligence (OSINT), with a focus on user profiling and behavioral analysis. While peer-to-peer (P2P) networks such as BitTorrent are well studied with respect to privacy and performance, their metadata is rarely used for investigative purposes. This...

💬 0 commentsarXiv:2601.01492v1PDF
0

Posted in cs.CL · 2026-01-04 · Junichiro Niimi

Distortion Instead of Hallucination: The Effect of Reasoning Under Strict Constraints

With the widespread adoption of large language models (LLMs), hallucinations, which are non-factual fabrications in model outputs, have become serious concerns. Reasoning capabilities have received attention as a self-verification process to improve output reliability. However, the effect of reasoning within a closed system where LLMs...

💬 0 commentsarXiv:2601.01490v1PDF
0

Posted in cs.CL · 2026-01-04 · Vanessa Toborek, Sebastian Müller, Christian Bauckhage

Four Quadrants of Difficulty: A Simple Categorisation and its Limits

Curriculum Learning (CL) aims to improve the outcome of model training by estimating the difficulty of samples and scheduling them accordingly. In NLP, difficulty is commonly approximated using task-agnostic linguistic heuristics or human intuition, implicitly assuming that these signals correlate with what neural models find...

💬 0 commentsarXiv:2601.01488v1PDF