Qwen Councils

Computer Science

arXiv preprints from January 1, 2026 through July 20, 2026 — 15:34:27 EST

0

Posted in cs.CR · 2026-01-16 · Anjanava Biswas, Wrick Talukdar

Guardrails for trust, safety, and ethical development and deployment of Large Language Models (LLM)

The AI era has ushered in Large Language Models (LLM) to the technological forefront, which has been much of the talk in 2023, and is likely to remain as such for many years to come. LLMs are the AI models that are the power house behind generative AI applications such as ChatGPT. These AI models, fueled by vast amounts of data and...

💬 0 commentsarXiv:2601.14298v1PDF
0

Posted in cs.DS · 2026-01-16 · Stephen Mussmann, Mehul Smriti Raje, Kavya Tumkur, Oumayma Messoussi, Cyprien Hachem, Seby Jacob

Sum Estimation via Vector Similarity Search

Semantic embeddings to represent objects such as image, text and audio are widely used in machine learning and have spurred the development of vector similarity search methods for retrieving semantically related objects. In this work, we study the sibling task of estimating a sum over all objects in a set, such as the kernel density...

💬 0 commentsarXiv:2601.11765v1PDF
0

Posted in cs.CY · 2026-01-16 · Rui-Jie Yew, Kate Elizabeth Creasey, Taylor Lynn Curtis, Suresh Venkatasubramanian

The Commodification of AI Sovereignty: Lessons from the Fight for Sovereign Oil

"Sovereignty" is increasingly a part of national AI policies and strategies. At the same time that "sovereignty" is invoked as a priority for global AI policy, it is also being commodified along the AI stack. Companies now sell "sovereign" AI factories, clouds, and language models to governments, enterprises, and communities --...

💬 0 commentsarXiv:2601.11763v1PDF
0

Posted in cs.CL · 2026-01-16 · Sae Young Moon, Myeongjun Erik Jang, Haoyan Luo, Chunyang Xiao, Antonios Georgiadis, Fran Silavong

Industry-Aligned Granular Topic Modeling

Topic modeling has extensive applications in text mining and data analysis across various industrial sectors. Although the concept of granularity holds significant value for business applications by providing deeper insights, the capability of topic modeling methods to produce granular topics has not been thoroughly explored. In this...

💬 0 commentsarXiv:2601.11762v1PDF
0

Posted in cs.CY · 2026-01-16 · Muhammad Muneeb Pervez, Muhammad Qasim Atiq Ullah, Ibrahim Ahmed Khan, Roshnik Rahat, Muhammad Fareed Zaffar, Rashid Tahir, Talal Rahwan, Yasir Zaki

(Mis-)Informed Consent: Predatory Apps and the Exploitation of Populations with Limited Literacy

Among populations with limited literacy in emerging digital markets, the adoption of mobile phones, combined with comprehension barriers and poor cybersecurity hygiene, has created hidden privacy risks. This paper examines how informed consent is often abused by predatory financial applications, leading to financial scams that...

💬 0 commentsarXiv:2601.17025v1PDF
0

Posted in cs.CL · 2026-01-16 · Arnab Das Utsa

Early Linguistic Pattern of Anxiety from Social Media Using Interpretable Linguistic Features: A Multi-Faceted Validation Study with Author-Disjoint Evaluation

Anxiety affects hundreds of millions of individuals globally, yet large-scale screening remains limited. Social media language provides an opportunity for scalable detection, but current models often lack interpretability, keyword-robustness validation, and rigorous user-level data integrity. This work presents a transparent approach...

💬 0 commentsarXiv:2601.11758v1PDF
0

Posted in cs.LO · 2026-01-16 · Walter Moreira, Joe Stubbs

Sequencelib: A Computational Platform for Formalizing the OEIS in Lean

The On-Line Encyclopedia of Integer Sequences (OEIS) is a web-accessible database cataloging interesting integer sequences and associated theorems. With more than 12,000 citations, the OEIS is one of the most highly cited resources in all of theoretical mathematics. In this paper, we present Sequencelib, a project to formalize the...

💬 0 commentsarXiv:2601.11757v1PDF
0

Posted in cs.DS · 2026-01-16 · Wenjing Chen, Yixin Chen, Victoria G. Crawford

Bicriteria Algorithms for Submodular Cover with Partition and Fairness Constraints

In many submodular optimization applications, datasets are naturally partitioned into disjoint subsets. These scenarios give rise to submodular optimization problems with partition-based constraints, where the desired solution set should be in some sense balanced, fair, or resource-constrained across these partitions. While existing...

💬 0 commentsarXiv:2601.11755v1PDF
0

Posted in cs.LG · 2026-01-16 · Margaret Foster

Measurement for Opaque Systems: Multi-source Triangulation with Interpretable Machine Learning

We propose a measurement framework for difficult-to-access contexts that uses indirect data traces, interpretable machine-learning models, and theory-guided triangulation to fill inaccessible measurement spaces. Many high-stakes systems of scientific and policy interest are difficult, if not impossible, to reach directly: dynamics of...

💬 0 commentsarXiv:2602.00022v1PDF
0

Posted in cs.HC · 2026-01-16 · Mo Houtti, Moyan Zhou, Daniel Runningen, Surabhi Sunil, Leor Porat, Harmanpreet Kaur, Loren Terveen, Stevie Chancellor

Opportunities and Barriers for AI Feedback on Meeting Inclusion in Socioorganizational Teams

Inclusion is important for meeting effectiveness, which is in turn central to organizational functioning. One way of improving inclusion in meetings is through feedback, but social dynamics make giving feedback difficult. We propose that AI agents can facilitate feedback exchange by being psychologically safer recipients, and we test...

💬 0 commentsarXiv:2601.11750v1PDF
0

Posted in cs.AI · 2026-01-16 · Huaxiaoyue Wang, Sunav Choudhary, Franck Dernoncourt, Yu Shen, Stefano Petrangeli

PRISM: Learning Design Knowledge from Data for Stylistic Design Improvement

Graphic design often involves exploring different stylistic directions, which can be time-consuming for non-experts. We address this problem of stylistically improving designs based on natural language instructions. While VLMs have shown initial success in graphic design, their pretrained knowledge on styles is often too general and...

💬 0 commentsarXiv:2601.11747v1PDF
0

Posted in cs.CL · 2026-01-16 · George Mihaila, Suleyman Olcay Polat, Poli Nemkova, Himanshu Sharma, Namratha V. Urs, Mark V. Albert

LIME-LLM: Probing Models with Fluent Counterfactuals, Not Broken Text

Local explanation methods such as LIME (Ribeiro et al., 2016) remain fundamental to trustworthy AI, yet their application to NLP is limited by a reliance on random token masking. These heuristic perturbations frequently generate semantically invalid, out-of-distribution inputs that weaken the fidelity of local surrogate models. While...

💬 0 commentsarXiv:2601.11746v1PDF
0

Posted in cs.CR · 2026-01-15 · Mohoshin Ara Tahera, Karamveer Singh Sidhu, Shuvalaxmi Dass, Sajal Saha

SoK: Privacy-aware LLM in Healthcare: Threat Model, Privacy Techniques, Challenges and Recommendations

Large Language Models (LLMs) are increasingly adopted in healthcare to support clinical decision-making, summarize electronic health records (EHRs), and enhance patient care. However, this integration introduces significant privacy and security challenges, driven by the sensitivity of clinical data and the high-stakes nature of...

💬 0 commentsarXiv:2601.10004v1PDF
0

Posted in cs.CL · 2026-01-15 · Sanghyeok Choi, Woosang Jeon, Kyuseok Yang, Taehyeong Kim

SocraticKG: Knowledge Graph Construction via QA-Driven Fact Extraction

Constructing Knowledge Graphs (KGs) from unstructured text provides a structured framework for knowledge representation and reasoning, yet current LLM-based approaches struggle with a fundamental trade-off: factual coverage often leads to relational fragmentation, while premature consolidation causes information loss. To address this,...

💬 0 commentsarXiv:2601.10003v2PDF
0

Posted in cs.CV · 2026-01-15 · Chengjia Liang, Zhenjiong Wang, Chao Chen, Ruizhi Zhang, Songxi Liang, Hai Xie, Haijun Lei, Zhongwei Huang

DW-DGAT: Dynamically Weighted Dual Graph Attention Network for Neurodegenerative Disease Diagnosis

Parkinson's disease (PD) and Alzheimer's disease (AD) are the two most prevalent and incurable neurodegenerative diseases (NDs) worldwide, for which early diagnosis is critical to delay their progression. However, the high dimensionality of multi-metric data with diverse structural forms, the heterogeneity of neuroimaging and...

💬 0 commentsarXiv:2601.10001v3PDF
0

Posted in cs.MM · 2026-01-15 · Diqiong Jiang, Kai Zhu, Dan Song, Jian Chang, Chenglizhao Chen, Zhenyu Wu

EditEmoTalk: Controllable Speech-Driven 3D Facial Animation with Continuous Expression Editing

Speech-driven 3D facial animation aims to generate realistic and expressive facial motions directly from audio. While recent methods achieve high-quality lip synchronization, they often rely on discrete emotion categories, limiting continuous and fine-grained emotional control. We present EditEmoTalk, a controllable speech-driven 3D...

💬 0 commentsarXiv:2601.10000v1PDF
0

Posted in cs.CY · 2026-01-15 · Conrad Borchers, Ashish Gurung, Qinyi Liu, Danielle R. Thomas, Mohammad Khalil, Kenneth R. Koedinger

Brief but Impactful: How Human Tutoring Interactions Shape Engagement in Online Learning

Learning analytics can guide human tutors to efficiently address motivational barriers to learning that AI systems struggle to support. Students become more engaged when they receive human attention. However, what occurs during short interventions, and when are they most effective? We align student-tutor dialogue transcripts with...

💬 0 commentsarXiv:2601.09994v1PDF
0

Posted in cs.RO · 2026-01-15 · Hojung Choi, Yifan Hou, Chuer Pan, Seongheon Hong, Austin Patel, Xiaomeng Xu, Mark R. Cutkosky, Shuran Song

In-the-Wild Compliant Manipulation with UMI-FT

Many manipulation tasks require careful force modulation. With insufficient force the task may fail, while excessive force could cause damage. The high cost, bulky size and fragility of commercial force/torque (F/T) sensors have limited large-scale, force-aware policy learning. We introduce UMI-FT, a handheld data-collection platform...

💬 0 commentsarXiv:2601.09988v1PDF
0

Posted in cs.PL · 2026-01-15 · Cheng Zhang, Qiancheng Fu, Hang Ji, Ines Santacruz Del Valle, Alexandra Silva, Marco Gaboardi

Outrunning Big KATs: Efficient Decision Procedures for Variants of GKAT

This paper presents several efficient decision procedures for trace equivalence of GKAT automata, which make use of on-the-fly symbolic techniques via SAT solvers. To demonstrate applicability of our algorithms, we designed symbolic derivatives for CF-GKAT, a practical system based on GKAT designed to validate control-flow...

💬 0 commentsarXiv:2601.09986v2PDF
0

Posted in cs.LG · 2026-01-15 · Tianqi Zhang, Flavio Ponzina, Tajana Rosing

FaTRQ: Tiered Residual Quantization for LLM Vector Search in Far-Memory-Aware ANNS Systems

Approximate Nearest-Neighbor Search (ANNS) is a key technique in retrieval-augmented generation (RAG), enabling rapid identification of the most relevant high-dimensional embeddings from massive vector databases. Modern ANNS engines accelerate this process using prebuilt indexes and store compressed vector-quantized representations in...

💬 0 commentsarXiv:2601.09985v1PDF
0

Posted in cs.CL · 2026-01-15 · David Samuel Setiawan, Raphaël Merx, Jey Han Lau

Context Volume Drives Performance: Tackling Domain Shift in Extremely Low-Resource Translation via RAG

Neural Machine Translation (NMT) models for low-resource languages suffer significant performance degradation under domain shift. We quantify this challenge using Dhao, an indigenous language of Eastern Indonesia with no digital footprint beyond the New Testament (NT). When applied to the unseen Old Testament (OT), a standard NMT...

💬 0 commentsarXiv:2601.09982v2PDF
0

Posted in cs.CV · 2026-01-15 · Yulin He, Wei Chen, Zhikang Jian, Tianhang Guo, Wenjuan Zhou, Minglong Li, Shaowu Yang, Wenjing Yang

DR$^2$Seg: Decomposed Two-Stage Rollouts for Efficient Reasoning Segmentation in Multimodal Large Language Models

Reasoning segmentation is an emerging vision-language task that requires reasoning over intricate text queries to precisely segment objects. However, existing methods typically suffer from overthinking, generating verbose reasoning chains that interfere with object localization in multimodal large language models (MLLMs). To address...

💬 0 commentsarXiv:2601.09981v2PDF
0

Posted in cs.LG · 2026-01-15 · Frank Cole, Dixi Wang, Yineng Chen, Yulong Lu, Rongjie Lai

In-Context Operator Learning on the Space of Probability Measures

We introduce \emph{in-context operator learning on probability measure spaces} for optimal transport (OT). The goal is to learn a single solution operator that maps a pair of distributions to the OT map, using only few-shot samples from each distribution as a prompt and \emph{without} gradient updates at inference. We parameterize the...

💬 0 commentsarXiv:2601.09979v1PDF
0

Posted in cs.NI · 2026-01-15 · Jie Zheng, Ruichen Zhang, Dusit Niyato, Haijun Zhang, Jiacheng Wang, Hongyang Du, Jiawen Kang, Zehui Xiong

Large Language Model (LLM)-enabled Reinforcement Learning for Wireless Network Optimization

Enhancing future wireless networks presents a significant challenge for networking systems due to diverse user demands and the emergence of 6G technology. While reinforcement learning (RL) is a powerful framework, it often encounters difficulties with high-dimensional state spaces and complex environments, leading to substantial...

💬 0 commentsarXiv:2602.13210v1PDF
0

Posted in cs.DC · 2026-01-15 · Jer Shyuan Ng, Wathsara Daluwatta, Shehan Edirimannage, Charitha Elvitigala, Asitha Kottahachchi Kankanamge Don, Ibrahim Khalil, Heng Zhang, Dusit Niyato

Federated Unlearning in Edge Networks: A Survey of Fundamentals, Challenges, Practical Applications and Future Directions

The proliferation of connected devices and privacy-sensitive applications has accelerated the adoption of Federated Learning (FL), a decentralized paradigm that enables collaborative model training without sharing raw data. While FL addresses data locality and privacy concerns, it does not inherently support data deletion requests...

💬 0 commentsarXiv:2601.09978v1PDF