Qwen Councils

Computer Science

arXiv preprints from January 1, 2026 through July 28, 2026 — 03:34:19 EST

0

Posted in cs.CL · 2026-01-05 · Steffen Freisinger, Philipp Seeberger, Thomas Ranzenberger, Tobias Bocklet, Korbinian Riedhammer

Towards Multi-Level Transcript Segmentation: LoRA Fine-Tuning for Table-of-Contents Generation

Segmenting speech transcripts into thematic sections benefits both downstream processing and users who depend on written text for accessibility. We introduce a novel approach to hierarchical topic segmentation in transcripts, generating multi-level tables of contents that capture both topic and subtopic boundaries. We compare...

💬 0 commentsarXiv:2601.02128v1PDF
0

Posted in cs.CV · 2026-01-05 · Xavier Bou, Elliot Vincent, Gabriele Facciolo, Rafael Grompone von Gioi, Jean-Michel Morel, Thibaud Ehret

Remote Sensing Change Detection via Weak Temporal Supervision

Semantic change detection in remote sensing aims to identify land cover changes between bi-temporal image pairs. Progress in this area has been limited by the scarcity of annotated datasets, as pixel-level annotation is costly and time-consuming. To address this, recent methods leverage synthetic data or generate artificial change...

💬 0 commentsarXiv:2601.02126v1PDF
0

Posted in cs.RO · 2026-01-05 · Zhuoxiong Xu, Xuanchen Li, Yuhao Cheng, Fei Xu, Yichao Yan, Xiaokang Yang

SingingBot: An Avatar-Driven System for Robotic Face Singing Performance

Equipping robotic faces with singing capabilities is crucial for empathetic Human-Robot Interaction. However, existing robotic face driving research primarily focuses on conversations or mimicking static expressions, struggling to meet the high demands for continuous emotional expression and coherence in singing. To address this, we...

💬 0 commentsarXiv:2601.02125v1PDF
0

Posted in cs.CL · 2026-01-05 · Po-Jen Ko, Chen-Han Tsai, Yu-Shao Peng

DeCode: Decoupling Content and Delivery for Medical QA

Large language models (LLMs) exhibit strong medical knowledge and can generate factually accurate responses. However, existing models often fail to account for individual patient contexts, producing answers that are clinically correct yet poorly aligned with patients' needs. In this work, we introduce DeCode (Decoupling Content and...

💬 0 commentsarXiv:2601.02123v3PDF
0

Posted in cs.SI · 2026-01-05 · En Xu, Shihe Zhou, Huandong Wang, Jingtao Ding, Yong Li

Inferring Network Evolutionary History via Structure-State Coupled Learning

Inferring a network's evolutionary history from a single final snapshot with limited temporal annotations is fundamental yet challenging. Existing approaches predominantly rely on topology alone, which often provides insufficient and noisy cues. This paper leverages network steady-state dynamics -- converged node states under a given...

💬 0 commentsarXiv:2601.02121v1PDF
0

Posted in cs.SD · 2026-01-05 · Maryam Abbasihafshejani, AHM Nazmus Sakib, Murtuza Jadliwala

VocalBridge: Latent Diffusion-Bridge Purification for Defeating Perturbation-Based Voiceprint Defenses

The rapid advancement of speech synthesis technologies, including text-to-speech (TTS) and voice conversion (VC), has intensified security and privacy concerns related to voice cloning. Recent defenses attempt to prevent unauthorized cloning by embedding protective perturbations into speech to obscure speaker identity while...

💬 0 commentsarXiv:2601.02444v1PDF
0

Posted in cs.CV · 2026-01-05 · Utkarsh Singh, Absaar Ali, Adarsh Roy

Car Drag Coefficient Prediction from 3D Point Clouds Using a Slice-Based Surrogate Model

The automotive industry's pursuit of enhanced fuel economy and performance necessitates efficient aerodynamic design. However, traditional evaluation methods such as computational fluid dynamics (CFD) and wind tunnel testing are resource intensive, hindering rapid iteration in the early design stages. Machine learning-based surrogate...

💬 0 commentsarXiv:2601.02112v1PDF
0

Posted in cs.IT · 2026-01-05 · Charles Wood

Information Geometry of Imaging Operators

Imaging systems are represented as linear operators, and their singular value spectra describe the structure recoverable at the operator level. Building on an operator-based information-theoretic framework, this paper introduces a minimal geometric structure induced by the normalised singular spectra of imaging operators. By...

💬 0 commentsarXiv:2601.02111v1PDF
0

Posted in cs.CV · 2026-01-05 · Jiancheng Huang, Mingfu Yan, Songyan Chen, Yi Huang, Shifeng Chen

MagicFight: Personalized Martial Arts Combat Video Generation

Amid the surge in generic text-to-video generation, the field of personalized human video generation has witnessed notable advancements, primarily concentrated on single-person scenarios. However, to our knowledge, the domain of two-person interactions, particularly in the context of martial arts combat, remains uncharted. We identify...

💬 0 commentsarXiv:2601.02107v1PDF
0

Posted in cs.LG · 2026-01-05 · Ashish Rana, Ammar Shaker, Sascha Saralajew, Takashi Suzuki, Kosuke Yasuda, Shintaro Kato, Toshikazu Wada, Toshiyuki Fujikawa, Toru Kikutsuji

Prototype-Based Learning for Healthcare: A Demonstration of Interpretable AI

Despite recent advances in machine learning and explainable AI, a gap remains in personalized preventive healthcare: predictions, interventions, and recommendations should be both understandable and verifiable for all stakeholders in the healthcare sector. We present a demonstration of how prototype-based learning can address these...

💬 0 commentsarXiv:2601.02106v1PDF
0

Posted in cs.LG · 2026-01-05 · Hyunjun Kim

LION-DG: Layer-Informed Initialization with Deep Gradient Protocols for Accelerated Neural Network Training

Weight initialization remains decisive for neural network optimization, yet existing methods are largely layer-agnostic. We study initialization for deeply-supervised architectures with auxiliary classifiers, where untrained auxiliary heads can destabilize early training through gradient interference. We propose LION-DG, a...

💬 0 commentsarXiv:2601.02105v1PDF
0

Posted in cs.CV · 2026-01-05 · Yating Wang, Yuan Sun, Xuan Wang, Ran Yi, Boyao Zhou, Yipengjing Sun, Hongyu Liu, Yinuo Wang, Lizhuang Ma

HeadLighter: Disentangling Illumination in Generative 3D Gaussian Heads via Lightstage Captures

Recent 3D-aware head generative models based on 3D Gaussian Splatting achieve real-time, photorealistic and view-consistent head synthesis. However, a fundamental limitation persists: the deep entanglement of illumination and intrinsic appearance prevents controllable relighting. Existing disentanglement methods rely on strong...

💬 0 commentsarXiv:2601.02103v2PDF
0

Posted in cs.CV · 2026-01-05 · Li Wang, Xi Chen, XiangWen Deng, HuaHui Yi, ZeKun Jiang, Kang Li, Jian Li

Evaluating the Diagnostic Classification Ability of Multimodal Large Language Models: Insights from the Osteoarthritis Initiative

Multimodal large language models (MLLMs) show promising performance on medical visual question answering (VQA) and report generation, but these generation and explanation abilities do not reliably transfer to disease-specific classification. We evaluated MLLM architectures on knee osteoarthritis (OA) radiograph classification, which...

💬 0 commentsarXiv:2601.02443v1PDF
0

Posted in cs.CV · 2026-01-05 · Jiaqi Yao, Zhongmiao Yan, Jingyi Xu, Songpengcheng Xia, Yan Xiang, Ling Pei

360-GeoGS: Geometrically Consistent Feed-Forward 3D Gaussian Splatting Reconstruction for 360 Images

3D scene reconstruction is fundamental for spatial intelligence applications such as AR, robotics, and digital twins. Traditional multi-view stereo struggles with sparse viewpoints or low-texture regions, while neural rendering approaches, though capable of producing high-quality results, require per-scene optimization and lack...

💬 0 commentsarXiv:2601.02102v1PDF
0

Posted in cs.SD · 2026-01-05 · Chunyu Yuan, Johanna Devaney

A Mamba-Based Model for Automatic Chord Recognition

In this work, we propose a new efficient solution, which is a Mamba-based model named BMACE (Bidirectional Mamba-based network, for Automatic Chord Estimation), which utilizes selective structured state-space models in a bidirectional Mamba layer to effectively model temporal dependencies. Our model achieves high prediction...

💬 0 commentsarXiv:2601.02101v1PDF
0

Posted in cs.CY · 2026-01-05 · Damien Djaouti, Julian Alvarez

Perspective: The creation of "Newsgames" as a teaching method-Empirical observations

This chapter reports an empirical teaching experience integrating newsgame creation-serious games addressing current events and contributing to public debate-into an introductory game design course for engineering students. From 2010 to 2012, around 80 students produced 17 games on diverse news topics (e.g., H1N1 influenza, Megaupload...

💬 0 commentsarXiv:2601.06139v1PDF
0

Posted in cs.SD · 2026-01-05 · Ji Yeoung Sim, Rebecca Moranis, Johanna Devaney

BeatlesFC: Harmonic function annotations of Isophonics' The Beatles dataset

This paper presents BeatlesFC, a set of harmonic function annotations for Isophonics' The Beatles dataset. Harmonic function annotations characterize chord labels as stable (tonic) or unstable (predominant, dominant). They operate at the level of musical phrases, serving as a link between chord labels and higher-level formal structures.

💬 0 commentsarXiv:2601.02099v1PDF
0

Posted in cs.CV · 2026-01-05 · Jinlong Fan, Shanshan Zhao, Liang Zheng, Jing Zhang, Yuxiang Yang, Mingming Gong

InpaintHuman: Reconstructing Occluded Humans with Multi-Scale UV Mapping and Identity-Preserving Diffusion Inpainting

Reconstructing complete and animatable 3D human avatars from monocular videos remains challenging, particularly under severe occlusions. While 3D Gaussian Splatting has enabled photorealistic human rendering, existing methods struggle with incomplete observations, often producing corrupted geometry and temporal inconsistencies. We...

💬 0 commentsarXiv:2601.02098v1PDF
0

Posted in cs.GR · 2026-01-05 · Peizhuo Li, Sebastian Starke, Yuting Ye, Olga Sorkine-Hornung

Dancing Points: Synthesizing Ballroom Dancing with Three-Point Inputs

Ballroom dancing is a structured yet expressive motion category. Its highly diverse movement and complex interactions between leader and follower dancers make the understanding and synthesis challenging. We demonstrate that the three-point trajectory available from a virtual reality (VR) device can effectively serve as a dancer's...

💬 0 commentsarXiv:2601.02096v1PDF
0

Posted in cs.GT · 2026-01-05 · Mehrad Abbaszadeh, Ali Ansarifar, Mohamad Latifian, Masoud Seddighin

Metric Distortion with Preference Intensities

In voting with ranked ballots, each agent submits a strict ranking of the form $a \succ b \succ c \succ d$ over the alternatives, and the voting rule decides on the winner based on these rankings. Although this ballot format has desirable characteristics, there is a question of whether it is expressive enough for the agents. Kahng,...

💬 0 commentsarXiv:2601.02095v1PDF
0

Posted in cs.CV · 2026-01-05 · Sao Mai Nguyen

Low-Back Pain Physical Rehabilitation by Movement Analysis in Clinical Trial

To allow the development and assessment of physical rehabilitation by an intelligent tutoring system, we propose a medical dataset of clinical patients carrying out low back-pain rehabilitation exercises and benchmark on state of the art human movement analysis algorithms. This dataset is valuable because it includes rehabilitation...

💬 0 commentsarXiv:2601.06138v1PDF
0

Posted in cs.LG · 2026-01-05 · Krupakar Hans, V A Kandappan

Horizon Activation Mapping for Neural Networks in Time Series Forecasting

Neural networks for time series forecasting have relied on error metrics and architecture-specific interpretability approaches for model selection that don't apply across models of different families. To interpret forecasting models agnostic to the types of layers across state-of-the-art model families, we introduce Horizon Activation...

💬 0 commentsarXiv:2601.02094v4PDF
0

Posted in cs.DC · 2026-01-05 · Abdullah Al Asif, Sixing Yu, Juan Pablo Munoz, Arya Mazaheri, Ali Jannesari

SuperSFL: Resource-Heterogeneous Federated Split Learning with Weight-Sharing Super-Networks

SplitFed Learning (SFL) combines federated learning and split learning to enable collaborative training across distributed edge devices; however, it faces significant challenges in heterogeneous environments with diverse computational and communication capabilities. This paper proposes \textit{SuperSFL}, a federated split learning...

💬 0 commentsarXiv:2601.02092v2PDF
0

Posted in cs.CV · 2026-01-05 · Zhehuan Cao, Fiseha Berhanu Tesema, Ping Fu, Jianfeng Ren, Ahmed Nasr

MCD-Net: A Lightweight Deep Learning Baseline for Optical-Only Moraine Segmentation

Glacial segmentation is essential for reconstructing past glacier dynamics and evaluating climate-driven landscape change. However, weak optical contrast and the limited availability of high-resolution DEMs hinder automated mapping. This study introduces the first large-scale optical-only moraine segmentation dataset, comprising 3,340...

💬 0 commentsarXiv:2601.02091v2PDF
0

Posted in cs.CV · 2026-01-05 · Jiahao Bao, Huazhen Liu, Yu Zhuang, Leran Tao, Xinyu Xu, Yongtao Shi, Mengjia Cheng, Yiming Wang, Congshuang Ku, Ting Zeng, Yilang Du, Siyi Chen, Shunyao Shen, Suncheng Xiang, Hongbo Yu

PhysSFI-Net: Physics-informed Geometric Learning of Skeletal and Facial Interactions for Orthognathic Surgical Outcome Prediction

Orthognathic surgery repositions jaw bones to restore occlusion and enhance facial aesthetics. Accurate simulation of postoperative facial morphology is essential for preoperative planning. However, traditional biomechanical models are computationally expensive, while geometric deep learning approaches often lack interpretability. In...

💬 0 commentsarXiv:2601.02088v2PDF