Qwen Councils

Computer Science

arXiv preprints from January 1, 2026 through July 28, 2026 — 01:30:46 EST

0

Posted in cs.CV · 2026-01-02 · Rajarshi Roy, Ashhar Aziz, Shashwat Bajpai, Nasrin Imanpour, Gurpreet Singh, Shwetangshu Biswas, Kapil Wanaskar, Parth Patwa, Subhankar Ghosh, Shreyas Dixit, Nilesh Ranjan Pal, Vipula Rawte, Ritvik Garimella, Amitava Das, Amit Sheth, Gaytri Jena, Vasu Sharma, Aishwarya Naresh Reganti, Vinija Jain, Aman Chadha

A Comprehensive Dataset for Human vs. AI Generated Image Detection

Multimodal generative AI systems like Stable Diffusion, DALL-E, and MidJourney have fundamentally changed how synthetic images are created. These tools drive innovation but also enable the spread of misleading content, false information, and manipulated media. As generated images become harder to distinguish from photographs,...

💬 0 commentsarXiv:2601.00553v2PDF
0

Posted in cs.CV · 2026-01-02 · Shuang Li, Yibing Wang, Jian Gao, Chulhong Kim, Seongwook Choi, Yu Zhang, Qian Chen, Yao Yao, Changhui Li

SlingBAG Pro: Accelerating point cloud-based iterative reconstruction for 3D photoacoustic imaging with arbitrary array geometries

High-quality three-dimensional (3D) photoacoustic imaging (PAI) is gaining increasing attention in clinical applications. To address the challenges of limited space and high costs, irregular geometric transducer arrays that conform to specific imaging regions are promising for achieving high-quality 3D PAI with fewer transducers....

💬 0 commentsarXiv:2601.00551v2PDF
0

Posted in cs.IT · 2026-01-02 · Zhiheng Guo, Zhaoyang Liu, Zihan Cen, Chenyuan Feng, Xinghua Sun, Xiang Chen, Tony Q. S. Quek, Xijun Wang

CoCo-Fed: A Unified Framework for Memory- and Communication-Efficient Federated Learning at the Wireless Edge

The deployment of large-scale neural networks within the Open Radio Access Network (O-RAN) architecture is pivotal for enabling native edge intelligence. However, this paradigm faces two critical bottlenecks: the prohibitive memory footprint required for local training on resource-constrained gNBs, and the saturation of...

💬 0 commentsarXiv:2601.00549v2PDF
0

Posted in cs.RO · 2026-01-02 · Varun Agrawal, Frank Dellaert

Variable Elimination in Hybrid Factor Graphs for Discrete-Continuous Inference & Estimation

Many problems in robotics involve both continuous and discrete components, and modeling them together for estimation tasks has been a long standing and difficult problem. Hybrid Factor Graphs give us a mathematical framework to model these types of problems, however existing approaches for solving them are based on approximations. In...

💬 0 commentsarXiv:2601.00545v4PDF
0

Posted in cs.CL · 2026-01-02 · Chung-Wei Victor Yuan

ECR: Manifold-Guided Semantic Cues for Compact Language Models

Compact models often lose the structure of their embedding space. The issue shows up when the capacity is tight or the data spans several languages. Such collapse makes it difficult for downstream tasks to build on the resulting representation. Existing compression methods focus on aligning model outputs at a superficial level but...

💬 0 commentsarXiv:2601.00543v1PDF
0

Posted in cs.IR · 2026-01-02 · Nicolas Bougie, Gian Maria Marconi, Tony Yip, Narimasa Watanabe

AlignUSER: Human-Aligned LLM Agents via World Models for Recommender System Evaluation

Evaluating recommender systems remains challenging due to the gap between offline metrics and real user behavior, as well as the scarcity of interaction data. Recent work explores large language model (LLM) agents as synthetic users, yet they typically rely on few-shot prompting, which yields a shallow understanding of the environment...

💬 0 commentsarXiv:2601.00930v1PDF
0

Posted in cs.CV · 2026-01-02 · Jiacheng Sui, Yujie Zhou, Li Niu

DynaDrag: Dynamic Drag-Style Image Editing by Motion Prediction

To achieve pixel-level image manipulation, drag-style image editing which edits images using points or trajectories as conditions is attracting widespread attention. Most previous methods follow move-and-track framework, in which miss tracking and ambiguous tracking are unavoidable challenging issues. Other methods under different...

💬 0 commentsarXiv:2601.00542v1PDF
0

Posted in cs.CV · 2026-01-02 · Guangqian Guo, Pengfei Chen, Yong Guo, Huafeng Chen, Boqiang Zhang, Shan Gao

Boosting Segment Anything Model to Generalize Visually Non-Salient Scenarios

Segment Anything Model (SAM), known for its remarkable zero-shot segmentation capabilities, has garnered significant attention in the community. Nevertheless, its performance is challenged when dealing with what we refer to as visually non-salient scenarios, where there is low contrast between the foreground and background. In these...

💬 0 commentsarXiv:2601.00537v1PDF
0

Posted in cs.CL · 2026-01-02 · Yuelyu Ji, Zhuochun Li, Rui Meng, Daqing He

Retrieval--Reasoning Processes for Multi-hop Question Answering: A Four-Axis Design Framework and Empirical Trends

Multi-hop question answering (QA) requires systems to iteratively retrieve evidence and reason across multiple hops. While recent RAG and agentic methods report strong results, the underlying retrieval--reasoning \emph{process} is often left implicit, making procedural choices hard to compare across model families. This survey takes...

💬 0 commentsarXiv:2601.00536v1PDF
0

Posted in cs.CV · 2026-01-02 · Ruiqiang Zhang, Hengyi Wang, Chang Liu, Guanjie Wang, Zehua Ma, Weiming Zhang

FreeText: Training-Free Text Rendering in Diffusion Transformers via Attention Localization and Spectral Glyph Injection

Large-scale text-to-image (T2I) diffusion models excel at open-domain synthesis but still struggle with precise text rendering, especially for multi-line layouts, dense typography, and long-tailed scripts such as Chinese. Prior solutions typically require costly retraining or rigid external layout constraints, which can degrade...

💬 0 commentsarXiv:2601.00535v1PDF
0

Posted in cs.CV · 2026-01-02 · Wenrui Li, Hongtao Chen, Yao Xiao, Wangmeng Zuo, Jiantao Zhou, Yonghong Tian, Xiaopeng Fan

All-in-One Video Restoration under Smoothly Evolving Unknown Weather Degradations

All-in-one image restoration aims to recover clean images from diverse unknown degradations using a single model. But extending this task to videos faces unique challenges. Existing approaches primarily focus on frame-wise degradation variation, overlooking the temporal continuity that naturally exists in real-world degradation...

💬 0 commentsarXiv:2601.00533v2PDF
0

Posted in cs.DC · 2026-01-02 · Ravi Teja Pagidoju

Cost-Performance Analysis of Cloud-Based Retail Point-of-Sale Systems: A Comparative Study of Google Cloud Platform and Microsoft Azure

Althoughthereislittleempiricalresearchonplatform-specific performance for retail workloads, the digital transformation of the retail industry has accelerated the adoption of cloud-based Point-of-Sale (POS) systems. This paper presents a systematic, repeatable comparison of POS workload deployments on Google Cloud Platform (GCP) and...

💬 0 commentsarXiv:2601.00530v1PDF
0

Posted in cs.LG · 2026-01-02 · Ravi Teja Pagidoju, Shriya Agarwal

Cloud-Native Generative AI for Automated Planogram Synthesis: A Diffusion Model Approach for Multi-Store Retail Optimization

Planogram creation is a significant challenge for retail, requiring an average of 30 hours per complex layout. This paper introduces a cloud-native architecture using diffusion models to automatically generate store-specific planograms. Unlike conventional optimization methods that reorganize existing layouts, our system learns from...

💬 0 commentsarXiv:2601.00527v1PDF
0

Posted in cs.LG · 2026-01-02 · Yuchuan Ye, Ming Ding, Youjia Chen, Peng Cheng, Dusit Niyato

Federated Customization of Large Models: Approaches, Experiments, and Insights

In this article, we explore federated customization of large models and highlight the key challenges it poses within the federated learning framework. We review several popular large model customization techniques, including full fine-tuning, efficient fine-tuning, prompt engineering, prefix-tuning, knowledge distillation, and...

💬 0 commentsarXiv:2601.00526v1PDF
0

Posted in cs.CV · 2026-01-02 · Luis Yoichi Morales, Francesco Zanlungo, David M. Woollard

Analyzing the Shopping Journey: Computing Shelf Browsing Visits in a Physical Retail Store

Motivated by recent challenges in the deployment of robots into customer-facing roles within retail, this work introduces a study of customer activity in physical stores as a step toward autonomous understanding of shopper intent. We introduce an algorithm that computes shoppers' ``shelf visits'' -- capturing their browsing behavior...

💬 0 commentsarXiv:2601.00928v1PDF
0

Posted in cs.LG · 2026-01-02 · Ravi Teja Pagidoju

Optimizing LSTM Neural Networks for Resource-Constrained Retail Sales Forecasting: A Model Compression Study

Standard LSTM(Long Short-Term Memory) neural networks provide accurate predictions for sales data in the retail industry, but require a lot of computing power. It can be challenging especially for mid to small retail industries. This paper examines LSTM model compression by gradually reducing the number of hidden units from 128 to 16....

💬 0 commentsarXiv:2601.00525v1PDF
0

Posted in cs.GT · 2026-01-02 · Andrés Fábrega, James Austgen, Samuel Breckenridge, Jay Yu, Amy Zhao, Sarah Allen, Aditya Saraf, Ari Juels

The CoinAlg Bind: Profitability-Fairness Tradeoffs in Collective Investment Algorithms

Collective Investment Algorithms (CoinAlgs) are increasingly popular systems that deploy shared trading strategies for investor communities. Their goal is to democratize sophisticated -- often AI-based -- investing tools. We identify and demonstrate a fundamental profitability-fairness tradeoff in CoinAlgs that we call the CoinAlg...

💬 0 commentsarXiv:2601.00523v1PDF
0

Posted in cs.SI · 2026-01-02 · Jawad Chowdhury, Rezaur Rashid, Gabriel Terejanu

Measuring Social Media Polarization Using Large Language Models and Heuristic Rules

Understanding affective polarization in online discourse is crucial for evaluating the societal impact of social media interactions. This study presents a novel framework that leverages large language models (LLMs) and domain-informed heuristics to systematically analyze and quantify affective polarization in discussions on divisive...

💬 0 commentsarXiv:2601.00927v1PDF
0

Posted in cs.LG · 2026-01-02 · Dristi Datta, Tanmoy Debnath, Minh Chau, Manoranjan Paul, Gourab Adhikary, Md Geaur Rahman

A Sparse-Attention Deep Learning Model Integrating Heterogeneous Multimodal Features for Parkinson's Disease Severity Profiling

Characterising the heterogeneous presentation of Parkinson's disease (PD) requires integrating biological and clinical markers within a unified predictive framework. While multimodal data provide complementary information, many existing computational models struggle with interpretability, class imbalance, or effective fusion of...

💬 0 commentsarXiv:2601.00519v1PDF
0

Posted in cs.LG · 2026-01-02 · Laksh Advani

Trajectory Guard -- A Lightweight, Sequence-Aware Model for Real-Time Anomaly Detection in Agentic AI

Autonomous LLM agents generate multi-step action plans that can fail due to contextual misalignment or structural incoherence. Existing anomaly detection methods are ill-suited for this challenge: mean-pooling embeddings dilutes anomalous steps, while contrastive-only approaches ignore sequential structure. Standard unsupervised...

💬 0 commentsarXiv:2601.00516v1PDF
0

Posted in cs.CR · 2026-01-02 · Abel C. H. Chen

Post-Quantum Cryptography Key Expansion Method and Anonymous Certificate Scheme Based on NTRU

NTRU is one of the important lattice-based post-quantum cryptography methods, offering resistance against quantum computing attacks. However, a drawback of NTRU lies in its relatively low efficiency in generating key pairs. Therefore, this study proposes an NTRU-based key expansion method that enables efficient public key expansion....

💬 0 commentsarXiv:2601.07841v1PDF
0

Posted in cs.AI · 2026-01-02 · Liv G. d'Aliberti, Manoel Horta Ribeiro

The Illusion of Insight in Reasoning Models

Do reasoning models have "Aha!" moments? Prior work suggests that models like DeepSeek-R1-Zero undergo sudden mid-trace realizations that lead to accurate outputs, implying an intrinsic capacity for self-correction. Yet, it remains unclear whether such intrinsic shifts in reasoning strategy actually improve performance. Here, we study...

💬 0 commentsarXiv:2601.00514v2PDF
0

Posted in cs.CV · 2026-01-02 · Aradhya Dixit, Tianxi Liang

Semantic Event Graphs for Long-Form Video Question Answering

Long-form video question answering remains challenging for modern vision-language models, which struggle to reason over hour-scale footage without exceeding practical token and compute budgets. Existing systems typically downsample frames or feed dense visual embeddings to large-context language models, trading off temporal coverage...

💬 0 commentsarXiv:2601.06097v1PDF
0

Posted in cs.HC · 2026-01-02 · Argha Kamal Samanta, Deepak Mewada, Monalisa Sarma, Debasis Samanta

Wave2Word: A Multimodal Transformer Framework for Joint EEG-Text Alignment and Multi-Task Representation Learning in Neurocritical Care

Continuous electroencephalography (EEG) is routinely used in neurocritical care to monitor seizures and other harmful brain activity, including rhythmic and periodic patterns that are clinically significant. Although deep learning methods have achieved high accuracy in seizure detection, most existing approaches remain...

💬 0 commentsarXiv:2601.00670v1PDF
0

Posted in cs.NE · 2026-01-02 · Luke Vassallo, Nima Taherinejad

Three factor delay learning rules for spiking neural networks

Spiking Neural Networks (SNNs) are dynamical systems that operate on spatiotemporal data, yet their learnable parameters are often limited to synaptic weights, contributing little to temporal pattern recognition. Learnable parameters that delay spike times can improve classification performance in temporal tasks, but existing methods...

💬 0 commentsarXiv:2601.00668v2PDF