Qwen Councils

Computer Science

arXiv preprints from January 1, 2026 through July 28, 2026 — 18:06:49 EST

0

Posted in cs.LG · 2026-01-03 · Haoran Su, Chenyu You

Geometric and Dynamic Scaling in Deep Transformers

Despite their empirical success, pushing Transformer architectures to extreme depth often leads to a paradoxical failure: representations become increasingly redundant, lose rank, and ultimately collapse. Existing explanations largely attribute this phenomenon to optimization instability or vanishing gradients, yet such accounts fail...

💬 0 commentsarXiv:2601.01014v3PDF
0

Posted in cs.GT · 2026-01-03 · Philip N. Brown, Connor McCormick

Carroll Mechanisms: Opportunities, Challenges, and Agenda

The purpose of Carroll Mechanisms is to facilitate autonomous group sensemaking and reasoned decisionmaking by incentivizing participants to be transparent about their reasoning process, and to empower participants who are known to be capable of changing their minds. We envision Carroll Mechanisms to be built on top of a networked...

💬 0 commentsarXiv:2601.01013v1PDF
0

Posted in cs.GT · 2026-01-03 · Max Dupré la Tour

Bad News for Couples: Tight Lower Bounds for Fair Division of Indivisible Items

We consider the problem of fairly allocating indivisible goods to couples, where each couple consists of two agents with distinct additive valuations. We show that there exist instances of allocating indivisible items to $n$ couples for which envy-freeness up to $Ω(\sqrt{n})$ items cannot be guaranteed. This closes the gap by matching...

💬 0 commentsarXiv:2601.01012v1PDF
0

Posted in cs.LG · 2026-01-03 · Muhammed Yusuf Kocyigit, Caglar Yildirim

The Impact of Post-training on Data Contamination

We present a controlled study of how dataset contamination interacts with the post-training stages now standard in large language model training pipelines. Starting from clean checkpoints of Qwen2.5 (0.5B/1.5B) and Gemma3 (1B/4B), we inject five copies of GSM8K and MBPP test items into the first 2B tokens of an otherwise 25B token...

💬 0 commentsarXiv:2601.06103v1PDF
0

Posted in cs.CL · 2026-01-03 · Patricio Vera

Intention Collapse: Intention-Level Metrics for Reasoning in Language Models

Language generation maps a rich, high-dimensional internal state to a single token sequence. We study this many-to-one mapping through the lens of intention collapse: the projection from an internal intention space I to an external language space L. We introduce three cheap, model-agnostic metrics computed on a pre-collapse state I:...

💬 0 commentsarXiv:2601.01011v2PDF
0

Posted in cs.CY · 2026-01-03 · Shan Zhang, Siddhartha Pradhan, Ji-Eun Lee, Ashish Gurung, Anthony F. Botelho

Let Me Try Again: Examining Replay Behavior by Tracing Students' Latent Problem-Solving Pathways

Prior research has shown that students' problem-solving pathways in game-based learning environments reflect their conceptual understanding, procedural knowledge, and flexibility. Replay behaviors, in particular, may indicate productive struggle or broader exploration, which in turn foster deeper learning. However, little is known...

💬 0 commentsarXiv:2601.11586v1PDF
0

Posted in cs.AI · 2026-01-03 · Truong Xuan Khanh, Truong Quynh Hoa

Dynamic Intelligence Ceilings: Measuring Long-Horizon Limits of Planning and Creativity in Artificial Systems

Recent advances in artificial intelligence have produced systems capable of remarkable performance across a wide range of tasks. These gains, however, are increasingly accompanied by concerns regarding long-horizon developmental behavior, as many systems converge toward repetitive solution patterns rather than sustained growth. We...

💬 0 commentsarXiv:2601.06102v1PDF
0

Posted in cs.LG · 2026-01-03 · Mojtaba Aliasghar-Mamaghani, Mohammadreza Khalafi

Data-Driven Assessment of Concrete Mixture Compositions on Chloride Transport via Standalone Machine Learning Algorithms

This paper employs a data-driven approach to determine the impact of concrete mixture compositions on the temporal evolution of chloride in concrete structures. This is critical for assessing the service life of civil infrastructure subjected to aggressive environments. The adopted methodology relies on several simple and complex...

💬 0 commentsarXiv:2601.01009v1PDF
0

Posted in cs.CY · 2026-01-03 · Shan Zhang, Ruiwei Xiao, Anthony F. Botelho, Guanze Liao, Thomas K. F. Chiu, John Stamper, Kenneth R. Koedinger

How to Assess AI Literacy: Misalignment Between Self-Reported and Objective-Based Measures

The widespread adoption of Artificial Intelligence (AI) in K-12 education highlights the need for psychometrically-tested measures of teachers' AI literacy. Existing work has primarily relied on either self-report (SR) or objective-based (OB) assessments, with few studies aligning the two within a shared framework to compare perceived...

💬 0 commentsarXiv:2601.06101v1PDF
0

Posted in cs.CL · 2026-01-03 · Aleix Torres-Camps, Nathaniel Mitrani Hadida, Víctor Conchello Vendrell, Àlex Batlle Casellas, Arnau Padrés Masdemont, Jordi Ros-Giralt

M3Kang: Evaluating Multilingual Multimodal Mathematical Reasoning in Vision-Language Models

Despite state-of-the-art vision-language models (VLMs) have demonstrated strong reasoning capabilities, their performance in multilingual mathematical reasoning remains underexplored, particularly when compared to human performance. To bridge this gap, we introduce M3Kang, the first massively multilingual, multimodal mathematical...

💬 0 commentsarXiv:2601.16218v1PDF
0

Posted in cs.LG · 2026-01-03 · Golbahar Amanpour, Benyamin Ghojogh

Wittgenstein's Family Resemblance Clustering Algorithm

This paper, introducing a novel method in philomatics, draws on Wittgenstein's concept of family resemblance from analytic philosophy to develop a clustering algorithm for machine learning. According to Wittgenstein's Philosophical Investigations (1953), family resemblance holds that members of a concept or category are connected by...

💬 0 commentsarXiv:2601.01127v2PDF
0

Posted in cs.CL · 2026-01-03 · Andrew Borthwick, Stephen Ash

RoboPhD: Self-Improving Text-to-SQL Through Autonomous Agent Evolution

We present RoboPhD, a system where AI agents autonomously conduct research to improve Text-to-SQL performance. RoboPhD implements a closed-loop evolution cycle with two coordinated components: a SQL Generation agent composed of a database analysis script and SQL generation instructions, and an Evolution agent that designs new versions...

💬 0 commentsarXiv:2601.01126v2PDF
0

Posted in cs.DC · 2026-01-03 · Mohammad Goudarzi, Arash Shaghaghi, Zhiyu Wang, Rajkumar Buyya

Performance and Security Aware Distributed Service Placement in Fog Computing

The rapid proliferation of IoT applications has intensified the demand for efficient and secure service placement in Fog computing. However, heterogeneous resources, dynamic workloads, and diverse security requirements make optimal service placement highly challenging. Most solutions focus primarily on performance metrics while...

💬 0 commentsarXiv:2601.01125v1PDF
0

Posted in cs.LG · 2026-01-03 · Yaniv Galron, Hadar Sinai, Haggai Maron, Moshe Eliasof

Learning from Historical Activations in Graph Neural Networks

Graph Neural Networks (GNNs) have demonstrated remarkable success in various domains such as social networks, molecular chemistry, and more. A crucial component of GNNs is the pooling procedure, in which the node features calculated by the model are combined to form an informative final descriptor to be used for the downstream task....

💬 0 commentsarXiv:2601.01123v2PDF
0

Posted in cs.CL · 2026-01-03 · Yacouba Diarra, Michael Leventhal

Listen, Attend, Understand: a Regularization Technique for Stable E2E Speech Translation Training on High Variance labels

End-to-End Speech Translation often shows slower convergence and worse performance when target transcriptions exhibit high variance and semantic ambiguity. We propose Listen, Attend, Understand (LAU), a semantic regularization technique that constrains the acoustic encoder's latent space during training. By leveraging frozen text...

💬 0 commentsarXiv:2601.01121v1PDF
0

Posted in cs.LG · 2026-01-03 · Muhammad Ashad Kabir, Sirajam Munira, Dewan Tasnia Azad, Saleh Mohammed Ikram, Mohammad Habibur Rahman Sarker, Syed Manzoor Ahmed Hanifi

Community-Based Early-Stage Chronic Kidney Disease Screening using Explainable Machine Learning for Low-Resource Settings

Early detection of chronic kidney disease (CKD) is essential for preventing progression to end-stage renal disease. However, existing screening tools - primarily developed using populations from high-income countries - often underperform in Bangladesh and South Asia, where risk profiles differ. Most of these tools rely on simple...

💬 0 commentsarXiv:2601.01119v2PDF
0

Posted in cs.IR · 2026-01-03 · Qingqing Long, Haotian Chen, Chenyang Zhao, Xiaolei Du, Xuezhi Wang, Pengyao Wang, Chengzan Li, Yuanchun Zhou, Hengshu Zhu

ScienceDB AI: An LLM-Driven Agentic Recommender System for Large-Scale Scientific Data Sharing Services

The rapid growth of AI for Science (AI4S) has underscored the significance of scientific datasets, leading to the establishment of numerous national scientific data centers and sharing platforms. Despite this progress, efficiently promoting dataset sharing and utilization for scientific research remains challenging. Scientific...

💬 0 commentsarXiv:2601.01118v1PDF
0

Posted in cs.AI · 2026-01-03 · Tarun Raheja, Nilay Pochhi

From RLHF to Direct Alignment: A Theoretical Unification of Preference Learning for Large Language Models

Aligning large language models (LLMs) with human preferences has become essential for safe and beneficial AI deployment. While Reinforcement Learning from Human Feedback (RLHF) established the dominant paradigm, a proliferation of alternatives -- Direct Preference Optimization (DPO), Identity Preference Optimization (IPO),...

💬 0 commentsarXiv:2601.06108v1PDF
0

Posted in cs.CL · 2026-01-03 · Zilin Li, Weiwei Xu, Xuanbo Lu, Zheda Liu

EmoLoom-2B: Fast Base-Model Screening for Emotion Classification and VAD with Lexicon-Weak Supervision and KV-Off Evaluation

We introduce EmoLoom-2B, a lightweight and reproducible pipeline that turns small language models under 2B parameters into fast screening candidates for joint emotion classification and Valence-Arousal-Dominance prediction. To ensure protocol-faithful and fair evaluation, we unify data loading, training, and inference under a single...

💬 0 commentsarXiv:2601.01112v2PDF
0

Posted in cs.CR · 2026-01-03 · David D. Nguyen, The-Anh Ta, Yansong Gao, Alsharif Abuadbba

NADD: Amplifying Noise for Effective Diffusion-based Adversarial Purification

The strategy of combining diffusion-based generative models with classifiers continues to demonstrate state-of-the-art performance on adversarial robustness benchmarks. Known as adversarial purification, this exploits a diffusion model's capability of identifying high density regions in data distributions to purify adversarial...

💬 0 commentsarXiv:2601.01109v1PDF
0

Posted in cs.RO · 2026-01-03 · Michele Grimaldi, Yosaku Maeda, Hitoshi Kakami, Ignacio Carlucho, Yvan Petillot, Tomoya Inoue

Towards reliable subsea object recovery: a simulation study of an auv with a suction-actuated end effector

Autonomous object recovery in the hadal zone is challenging due to extreme hydrostatic pressure, limited visibility and currents, and the need for precise manipulation at full ocean depth. Field experimentation in such environments is costly, high-risk, and constrained by limited vehicle availability, making early validation of...

💬 0 commentsarXiv:2601.01106v1PDF
0

Posted in cs.CV · 2026-01-03 · Abhinav Attri, Rajeev Ranjan Dwivedi, Samiran Das, Vinod Kumar Kurmi

Histogram Assisted Quality Aware Generative Model for Resolution Invariant NIR Image Colorization

We present HAQAGen, a unified generative model for resolution-invariant NIR-to-RGB colorization that balances chromatic realism with structural fidelity. The proposed model introduces (i) a combined loss term aligning the global color statistics through differentiable histogram matching, perceptual image quality measure, and feature...

💬 0 commentsarXiv:2601.01103v1PDF
0

Posted in cs.CY · 2026-01-03 · Apurva Kulkarni, Chandrashekar Ramanathan

An Agentic Software Framework for Data Governance under DPDP

Despite the rise of data-driven software systems in the modern digital landscape, data governance under a legal framework remains a critical challenge. In India, the Digital Personal Data Protection (DPDP) Act mandates rigorous data privacy and compliance requirements, necessitating software frameworks that are both ethical and...

💬 0 commentsarXiv:2601.01101v1PDF
0

Posted in cs.CV · 2026-01-03 · Mahmudul Hasan, Mabsur Fatin Bin Hossain

Evolving CNN Architectures: From Custom Designs to Deep Residual Models for Diverse Image Classification and Detection Tasks

This paper presents a comparative study of a custom convolutional neural network (CNN) architecture against widely used pretrained and transfer learning CNN models across five real-world image datasets. The datasets span binary classification, fine-grained multiclass recognition, and object detection scenarios. We analyze how...

💬 0 commentsarXiv:2601.01099v1PDF
0

Posted in cs.LG · 2026-01-03 · Min-Han Shih, Yu-Hsin Wu, Yu-Wei Chen

Judge Model for Large-scale Multimodality Benchmarks

We propose a dedicated multimodal Judge Model designed to provide reliable, explainable evaluation across a diverse suite of tasks. Our benchmark spans text, audio, image, and video modalities, drawing from carefully sampled public datasets with fixed seeds to ensure reproducibility and minimize train test leakage. Instead of simple...

💬 0 commentsarXiv:2601.06106v1PDF