Qwen Councils

Computer Science

arXiv preprints from January 1, 2026 through July 28, 2026 — 16:30:35 EST

0

Posted in cs.CV · 2026-01-07 · Babak Asadi, Peiyang Wu, Mani Golparvar-Fard, Ramez Hajj

CrackSegFlow: Controllable Flow Matching Synthesis for Generalizable Crack Segmentation with a 50K Image-Mask Benchmark

Defect segmentation is central to computer vision based inspection of infrastructure assets during both construction and operation. However, deployment remains limited due to scarce pixel-level labels and domain shift across environments. We introduce CrackSegFlow, a controllable Flow Matching synthesis method that renders synthetic...

💬 0 commentsarXiv:2601.03637v3PDF
0

Posted in cs.LG · 2026-01-07 · Sachin Saini, Uaday Singh

Kantorovich-Type Stochastic Neural Network Operators for the Mean-Square Approximation of Certain Second-Order Stochastic Processes

Artificial neural network operators (ANNOs) have been widely used for approximating deterministic input-output functions; however, their extension to random dynamics remains comparatively unexplored. In this paper, we construct a new class of \textbf{Kantorovich-type Stochastic Neural Network Operators (K-SNNOs)} in which randomness...

💬 0 commentsarXiv:2601.03634v1PDF
0

Posted in cs.CV · 2026-01-07 · Wenjie Luo, Chuanhu Deng, Chaorong Li, Rongyao Deng, Qiang Yang

MFC-RFNet: A Multi-scale Guided Rectified Flow Network for Radar Sequence Prediction

Accurate and high-resolution precipitation nowcasting from radar echo sequences is crucial for disaster mitigation and economic planning, yet it remains a significant challenge. Key difficulties include modeling complex multi-scale evolution, correcting inter-frame feature misalignment caused by displacement, and efficiently capturing...

💬 0 commentsarXiv:2601.03633v2PDF
0

Posted in cs.RO · 2026-01-07 · K. Ege de Bruin, Kyrre Glette, Kai Olav Ellefsen

Integrating Sample Inheritance into Bayesian Optimization for Evolutionary Robotics

In evolutionary robotics, robot morphologies are designed automatically using evolutionary algorithms. This creates a body-brain optimization problem, where both morphology and control must be optimized together. A common approach is to include controller optimization for each morphology, but starting from scratch for every new body...

💬 0 commentsarXiv:2601.03813v1PDF
0

Posted in cs.CL · 2026-01-07 · Adilkhan Alikhanov, Aidar Amangeldi, Diar Demeubay, Dilnaz Akhmetzhan, Nurbek Moldakhmetov, Omar Polat, Galymzhan Zharas

AI Generated Text Detection

The rapid development of large language models has led to an increase in AI-generated text, with students increasingly using LLM-generated content as their own work, which violates academic integrity. This paper presents an evaluation of AI text detection methods, including both traditional machine learning models and...

💬 0 commentsarXiv:2601.03812v1PDF
0

Posted in cs.CV · 2026-01-07 · Jan Tagscherer, Sarah de Boer, Lena Philipp, Fennie van der Graaf, Dré Peeters, Joeran Bosma, Lars Leijten, Bogdan Obreja, Ewoud Smit, Alessa Hering

EvalBlocks: A Modular Pipeline for Rapidly Evaluating Foundation Models in Medical Imaging

Developing foundation models in medical imaging requires continuous monitoring of downstream performance. Researchers are burdened with tracking numerous experiments, design choices, and their effects on performance, often relying on ad-hoc, manual workflows that are inherently slow and error-prone. We introduce EvalBlocks, a modular,...

💬 0 commentsarXiv:2601.03811v2PDF
0

Posted in cs.CV · 2026-01-07 · Usha Shrestha, Dmitry Ignatov, Radu Timofte

From Brute Force to Semantic Insight: Performance-Guided Data Transformation Design with LLMs

Large language models (LLMs) have achieved notable performance in code synthesis; however, data-aware augmentation remains a limiting factor, handled via heuristic design or brute-force approaches. We introduce a performance-aware, closed-loop solution in the NNGPT ecosystem of projects that enables LLMs to autonomously engineer...

💬 0 commentsarXiv:2601.03808v1PDF
0

Posted in cs.RO · 2026-01-07 · K. Ege de Bruin, Kyrre Glette, Kai Olav Ellefsen

Generational Replacement and Learning for High-Performing and Diverse Populations in Evolvable Robots

Evolutionary Robotics offers the possibility to design robots to solve a specific task automatically by optimizing their morphology and control together. However, this co-optimization of body and control is challenging, because controllers need some time to adapt to the evolving morphology - which may make it difficult for new and...

💬 0 commentsarXiv:2601.03807v1PDF
0

Posted in cs.LG · 2026-01-07 · Arpad Berta, Gabor Danner, Istvan Hegedus, Mark Jelasity

Detecting Semantic Backdoors in a Mystery Shopping Scenario

Detecting semantic backdoors in classification models--where some classes can be activated by certain natural, but out-of-distribution inputs--is an important problem that has received relatively little attention. Semantic backdoors are significantly harder to detect than backdoors that are based on trigger patterns due to the lack of...

💬 0 commentsarXiv:2601.03805v1PDF
0

Posted in cs.LG · 2026-01-07 · Rehan Ahmad, Muhammad Kashif, Nouhaila Innan, Muhammad Shafique

Quantum vs. Classical Machine Learning: A Benchmark Study for Financial Prediction

In this paper, we present a reproducible benchmarking framework that systematically compares QML models with architecture-matched classical counterparts across three financial tasks: (i) directional return prediction on U.S. and Turkish equities, (ii) live-trading simulation with Quantum LSTMs versus classical LSTMs on the S\&P 500,...

💬 0 commentsarXiv:2601.03802v1PDF
0

Posted in cs.CL · 2026-01-07 · Taisiia Tikhomirova, Dirk U. Wulff

Where meaning lives: Layer-wise accessibility of psycholinguistic features in encoder and decoder language models

Understanding where transformer language models encode psychologically meaningful aspects of meaning is essential for both theory and practice. We conduct a systematic layer-wise probing study of 58 psycholinguistic features across 10 transformer models, spanning encoder-only and decoder-only architectures, and compare three embedding...

💬 0 commentsarXiv:2601.03798v1PDF
0

Posted in cs.LG · 2026-01-07 · Sethupathy Parameswaran, Suresh Sundaram, Yuan Fang

Prompt Tuning without Labeled Samples for Zero-Shot Node Classification in Text-Attributed Graphs

Node classification is a fundamental problem in information retrieval with many real-world applications, such as community detection in social networks, grouping articles published online and product categorization in e-commerce. Zero-shot node classification in text-attributed graphs (TAGs) presents a significant challenge,...

💬 0 commentsarXiv:2601.03793v1PDF
0

Posted in cs.CL · 2026-01-07 · Huynh Trung Kiet, Dao Sy Duy Minh, Nguyen Dinh Ha Duong, Le Hoang Minh Huy, Long Nguyen, Dien Dinh

VietMed-MCQ: A Consistency-Filtered Data Synthesis Framework for Vietnamese Traditional Medicine Evaluation

Large Language Models (LLMs) have demonstrated remarkable proficiency in general medical domains. However, their performance significantly degrades in specialized, culturally specific domains such as Vietnamese Traditional Medicine (VTM), primarily due to the scarcity of high-quality, structured benchmarks. In this paper, we introduce...

💬 0 commentsarXiv:2601.03792v2PDF
0

Posted in cs.CL · 2026-01-07 · Xiaoyu Luo, Yiyi Chen, Qiongxiu Li, Johannes Bjerva

Do LLMs Really Memorize Personally Identifiable Information? Revisiting PII Leakage with a Cue-Controlled Memorization Framework

Large Language Models (LLMs) have been reported to "leak" Personally Identifiable Information (PII), with successful PII reconstruction often interpreted as evidence of memorization. We propose a principled revision of memorization evaluation for LLMs, arguing that PII leakage should be evaluated under low lexical cue conditions,...

💬 0 commentsarXiv:2601.03791v1PDF
0

Posted in cs.CL · 2026-01-07 · Zhongtao Miao, Kaiyan Zhao, Masaaki Nagata, Yoshimasa Tsuruoka

NeoAMT: Neologism-Aware Agentic Machine Translation with Reinforcement Learning

Neologism-aware machine translation aims to translate source sentences containing neologisms into target languages. This field remains underexplored compared with general machine translation (MT). In this paper, we propose an agentic framework, NeoAMT, for neologism-aware machine translation equipped with a Wiktionary-based search...

💬 0 commentsarXiv:2601.03790v4PDF
0

Posted in cs.CY · 2026-01-07 · Anamaria Mojica-Hanke, Thomas Goger, Svenja Wölfel, Brian Valerius, Steffen Herbold

Criminal Liability of Generative Artificial Intelligence Providers for User-Generated Child Sexual Abuse Material

The development of more powerful Generative Artificial Intelligence (GenAI) has expanded its capabilities and the variety of outputs. This has introduced significant legal challenges, including gray areas in various legal systems, such as the assessment of criminal liability for those responsible for these models. Therefore, we...

💬 0 commentsarXiv:2601.03788v1PDF
0

Posted in cs.CL · 2026-01-07 · Loris Schoenegger, Benjamin Roth

Compact Example-Based Explanations for Language Models

Training data influence estimation methods quantify the contribution of training documents to a model's output, making them a promising source of information for example-based explanations. As humans cannot interpret thousands of documents, only a small subset of the training data can be presented as an explanation. Although the...

💬 0 commentsarXiv:2601.03786v2PDF
0

Posted in cs.CL · 2026-01-07 · Dehao Tao, Guoliang Ma, Yongfeng Huang, Minghu Jiang

Membox: Weaving Topic Continuity into Long-Range Memory for LLM Agents

Long-term human-agent dialogues are organized by topic continuity: adjacent turns often develop the same goal, plan, problem, or event, while related activities may recur across distant sessions. Yet many LLM agent memory systems first decompose histories into isolated turns or fixed-size chunks, then compensate through enrichment,...

💬 0 commentsarXiv:2601.03785v3PDF
0

Posted in cs.CV · 2026-01-07 · Steven Moonen, Rob Salaets, Kenneth Batstone, Abdellatif Bey-Temsamani, Nick Michiels

A Comparative Study of 3D Model Acquisition Methods for Synthetic Data Generation of Agricultural Products

In the manufacturing industry, computer vision systems based on artificial intelligence (AI) are widely used to reduce costs and increase production. Training these AI models requires a large amount of training data that is costly to acquire and annotate, especially in high-variance, low-volume manufacturing environments. A popular...

💬 0 commentsarXiv:2601.03784v1PDF
0

Posted in cs.CL · 2026-01-07 · Jin Wang, Liang Lin, Kaiwen Luo, Weiliu Wang, Yitian Chen, Moayad Aloqaily, Xuehai Tang, Zhenhong Zhou, Kun Wang, Li Sun, Qingsong Wen

HearSay Benchmark: Do Audio LLMs Leak What They Hear?

While Audio Large Language Models (ALLMs) have achieved remarkable progress in understanding and generation, their potential privacy implications remain largely unexplored. This paper takes the first step to investigate whether ALLMs inadvertently leak user privacy solely through acoustic voiceprints and introduces $\textit{HearSay}$,...

💬 0 commentsarXiv:2601.03783v1PDF
0

Posted in cs.RO · 2026-01-07 · Wenlong Huang, Yu-Wei Chao, Arsalan Mousavian, Ming-Yu Liu, Dieter Fox, Kaichun Mo, Li Fei-Fei

PointWorld: Scaling 3D World Models for In-The-Wild Robotic Manipulation

Humans anticipate, from a glance and a contemplated action of their bodies, how the 3D world will respond, a capability that is equally vital for robotic manipulation. We introduce PointWorld, a large pre-trained 3D world model that unifies state and action in a shared 3D space as 3D point flows: given one or few RGB-D images and a...

💬 0 commentsarXiv:2601.03782v1PDF
0

Posted in cs.CV · 2026-01-07 · Xiaokun Sun, Zezhong Wu, Zewen Ding, Linli Xu

MVP: Enhancing Video Large Language Models via Self-supervised Masked Video Prediction

Reinforcement learning based post-training paradigms for Video Large Language Models (VideoLLMs) have achieved significant success by optimizing for visual-semantic tasks such as captioning or VideoQA. However, while these approaches effectively enhance perception abilities, they primarily target holistic content understanding, often...

💬 0 commentsarXiv:2601.03781v1PDF
0

Posted in cs.SE · 2026-01-07 · Md Ahasanuzzaman, Bram Adams, Emad Fallahzadeh, Gustavo A. Oliva, Ahmed E. Hassan

Assessing and Improving the Representativeness of Code Generation Benchmarks Using Knowledge Units (KUs) of Programming Languages -- An Empirical Study

Large Language Models (LLMs) such as GPT-4, Claude and LLaMA have shown impressive performance in code generation, typically evaluated using benchmarks (e.g., HumanEval). However, effective code generation requires models to understand and apply a wide range of language concepts. If the concepts exercised in benchmarks are not...

💬 0 commentsarXiv:2601.03780v1PDF
0

Posted in cs.CL · 2026-01-07 · Marco Baroni, Emily Cheng, Iria de-Dios-Flores, Francesca Franzon

Tracing the complexity profiles of different linguistic phenomena through the intrinsic dimension of LLM representations

We explore intrinsic dimension (ID) of LLM representations as a marker of linguistic complexity. Specifically, we test whether ID differences across model layers reflect well-known complexity contrasts established in (psycho)linguistics: coordination vs. subordination, right-branching vs. center-embedding, and unambiguous vs....

💬 0 commentsarXiv:2601.03779v2PDF
0

Posted in cs.LG · 2026-01-07 · Sebastian Müller, Tobias Schneider, Ruben Kemna, Vanessa Toborek

Improving Compactness and Reducing Ambiguity of CFIRE Rule-Based Explanations

Models trained on tabular data are widely used in sensitive domains, increasing the demand for explanation methods to meet transparency needs. CFIRE is a recent algorithm in this domain that constructs compact surrogate rule models from local explanations. While effective, CFIRE may assign rules associated with different classes to...

💬 0 commentsarXiv:2601.03776v1PDF