MeanVoiceFlow2

Joint Optimization of Mean Flow and Content Encoder for Fast One-Step Zero-Shot Voice Conversion

Interspeech 2026

Paper

teaser
Fig. 1. Overview of MeanVoiceFlow2. The student jointly learns a computationally efficient content encoder cφ and an average velocity network uφ through (a) conversion distillation and (b) real-data reconstruction. We further incorporate diffusion-GAN training with sample mixing using a discriminator Dψ to promote realism, and teacher-guided conditioning augmentation based on xθaug to promote disentanglement.

Abstract

Flow-matching approaches to voice conversion (VC) have gained attention owing to their high speech quality and strong speaker similarity. Among them, one-step models such as MeanVoiceFlow are particularly attractive because they enable efficient inference; however, their reliance on a computationally intensive content encoder remains a bottleneck. We therefore propose MeanVoiceFlow2, a framework that jointly optimizes a flow-based conversion module and a computationally efficient content encoder. The model is trained through conversion distillation using MeanVoiceFlow and the reconstruction of real data. We further incorporate diffusion-GAN training with sample mixing and teacher-guided conditioning augmentation to enhance realism and disentanglement. Experiments on zero-shot VC showed that MeanVoiceFlow2 achieved higher perceptual quality and approximately 9 times faster inference than MeanVoiceFlow while maintaining comparable speaker similarity.


Contents


Results

Results on VCTK dataset (Table 4)
Source Target DiffVC MVF MVF2 (ours) FVG2
Speed↑ Slow (×0.04) (×1) Fast (×9) Fast (×9)
Requires pretrained vocoder
for training
No No No Yes
Female → Female
Female → Female
Male → Male
Male → Male
Female → Male
Female → Male
Male → Female
Male → Female

Results on LibriTTS dataset (Table 5)
Source Target MVF MVF2 (ours)
Speed↑ (×1) Fast (×9)
Female → Female
Female → Female
Male → Male
Male → Male
Female → Male
Female → Male
Male → Female
Male → Female

Citation

@inproceedings{kaneko2026meanvoiceflow2,
  title={MeanVoiceFlow2: Joint Optimization of Mean Flow and Content Encoder for Fast One-Step Zero-Shot Voice Conversion,
  author={Kaneko, Takuhiro and Kameoka, Hirokazu and Tanaka, Kou and Kondo, Yuto},
  booktitle={Interspeech},
  year={2026},
}