license: mit
library_name: pytorch
tags:
- speech-editing
- speech-synthesis
- text-to-speech
wavepainter
Released checkpoints for wavepainter, a masked-span speech editor.
phase2-hubert-child is the released model and the artifact of record: phase two
as trained, and the checkpoint every number in the paper comes from. On the
English full split of Ming-Freeform-Audio-Edit, against Ren et al.
(arXiv:2602.00560, the row with GRPO), scored
with oracle edit spans taken from the transcript diff:
| edit type | metric | Ren et al. | wavepainter |
|---|---|---|---|
| substitution | WER | 4.41 | 3.479 |
| insertion | WER | 4.97 | 4.400 |
| deletion | WER | 6.88 | 9.697 |
| substitution | SIM | 0.78 | 0.943 |
| insertion | SIM | 0.82 | 0.961 |
| deletion | SIM | 0.78 | 0.918 |
Five of the six compared numbers improve on theirs: substitution and insertion WER, and speaker similarity on all three edit types. We do not beat them on deletion WER -- see the repository's Limitations, which also locates most of that gap with a ground-truth waveform splice.
Signal quality is reported full-reference, since a no-reference MOS estimator rates the untouched original recordings below synthesised output on this benchmark. Over the audio the edit does not touch, PESQ is 4.090 / 4.083 / 4.079 for substitution / insertion / deletion, against 4.351 for copy-synthesis with no edit at all.
The interpolated variants
Two earlier releases interpolate the phase-two child back toward the phase-1 base, which moves the diffusion denoiser only. They are published because the paper's first version reported them and anyone who fetched them should keep working; the paper no longer reports either.
| substitution | insertion | deletion | |
|---|---|---|---|
phase2-hubert-child (released) |
3.479 | 4.400 | 9.697 |
phase2-hubert-a085 |
3.063 | 4.128 | 9.136 |
phase2-hubert-a070 |
3.381 | 3.817 | 9.525 |
WER, lower is better. Nothing is bolded because the released checkpoint is not the best row: interpolating toward the phase-1 base improves every WER column, a=0.85 most on substitution and deletion and a=0.70 on insertion. We publish the child as the artifact of record anyway, because it is the model phase two produced and the coefficient was chosen by looking at these scores. The coefficient also has no single optimum across edit types, which is why both variants are kept rather than one being presented as strictly better. Speaker similarity is 0.943 / 0.961 / 0.918 for all three except insertion at a=0.70, which is 0.960.
Files
phase1-base/-- the phase-1 editor. Start here to re-run phase 2 (~1 GPU-hour).phase2-hubert-child/-- the released model, and the artifact of record.phase2-hubert-a085/-- an earlier release, interpolated at 0.85.phase2-hubert-a070/-- the same, at 0.70.
Use
git clone https://github.com/pujariaditya/wavepainter && cd wavepainter
pip install -e . && ./setup.sh
python scripts/download_weights.py
./scripts/verify_benchmark.sh
Checkpoints are sha256-pinned in scripts/download_weights.py. Licence and
attribution, including for the components these weights derive from, are in the
repository's ATTRIBUTION.md.