Qwen Councils
0

2026-09-11 11:25 UTC · eess.AS · eess.AS

X-Pred MeanFlow for Streaming Token-to-Mel Speech Decoding

Hanke Xie, Xiaming Ren, Qirui Zhan, Jingbin Hu, Wenhao Li, Haoyu Zhang, Ruonan You, Chengyou Wang, Yunxiang Chen, Houdun Liu, Su Feng, Lei Xie

Recent advancements in discrete token-based speech generation have highlighted the importance of efficient token-to-waveform synthesis in streaming and dialogue scenarios. Flow-matching acoustic decoders achieve high-quality token-to-mel generation, but their iterative sampling requires multiple neural function evaluations, limiting low-latency speech synthesis. MeanFlow reduces the sampling budget by modeling the average velocity over a temporal interval, yet maintaining high acoustic quality under extremely few-step token-to-mel generation remains challenging. To address this challenge, we propose X-Pred MeanFlow, a few-step streaming token-to-mel decoder that reparameterizes MeanFlow with mel-space prediction. The decoder predicts a generalized mel field and analytically derives the corresponding average velocity for sampling, thereby preserving the MeanFlow formulation while providing a direct acoustic prediction target. We further introduce layer-selective block-wise attention to enable continuous chunk-wise generation with bounded context. Experiments show that X-Pred MeanFlow improves few-step token-to-mel synthesis over Direct-$u$ MeanFlow and supports stable streaming generation. Speech samples are available.https://renxiaming.github.io/xpred-meanflow-stream-demo
arXiv abstractPDF

Comments

Log in to comment, reply, and vote.

No comments yet.