wavepainter / README.md
RootAccess4Life's picture
Upload README.md with huggingface_hub
7eb601e verified
|
Raw
History Blame Contribute Delete
3.34 kB
---
license: mit
library_name: pytorch
tags:
- speech-editing
- speech-synthesis
- text-to-speech
---
# wavepainter
Released checkpoints for [wavepainter](https://github.com/pujariaditya/wavepainter),
a masked-span speech editor.
`phase2-hubert-child` is the released model and the artifact of record: phase two
as trained, and the checkpoint every number in the paper comes from. On the
English `full` split of Ming-Freeform-Audio-Edit, against Ren et al.
([arXiv:2602.00560](https://arxiv.org/abs/2602.00560), the row with GRPO), scored
with oracle edit spans taken from the transcript diff:
| edit type | metric | Ren et al. | wavepainter |
|---|---|---|---|
| substitution | WER | 4.41 | **3.479** |
| insertion | WER | 4.97 | **4.400** |
| deletion | WER | **6.88** | 9.697 |
| substitution | SIM | 0.78 | **0.943** |
| insertion | SIM | 0.82 | **0.961** |
| deletion | SIM | 0.78 | **0.918** |
Five of the six compared numbers improve on theirs: substitution and insertion
WER, and speaker similarity on all three edit types. We do not beat them on
deletion WER -- see the repository's Limitations, which also locates most of that
gap with a ground-truth waveform splice.
Signal quality is reported full-reference, since a no-reference MOS estimator
rates the untouched original recordings below synthesised output on this
benchmark. Over the audio the edit does not touch, PESQ is 4.090 / 4.083 / 4.079
for substitution / insertion / deletion, against 4.351 for copy-synthesis with no
edit at all.
### The interpolated variants
Two earlier releases interpolate the phase-two child back toward the phase-1 base,
which moves the diffusion denoiser only. They are published because the paper's
first version reported them and anyone who fetched them should keep working; the
paper no longer reports either.
| | substitution | insertion | deletion |
|---|---|---|---|
| `phase2-hubert-child` (released) | 3.479 | 4.400 | 9.697 |
| `phase2-hubert-a085` | 3.063 | 4.128 | 9.136 |
| `phase2-hubert-a070` | 3.381 | 3.817 | 9.525 |
WER, lower is better. Nothing is bolded because the released checkpoint is not
the best row: interpolating toward the phase-1 base improves every WER column,
a=0.85 most on substitution and deletion and a=0.70 on insertion. We publish the
child as the artifact of record anyway, because it is the model phase two
produced and the coefficient was chosen by looking at these scores. The
coefficient also has no single optimum across edit types, which is why both
variants are kept rather than one being presented as strictly better. Speaker
similarity is 0.943 / 0.961 / 0.918 for all three except insertion at a=0.70,
which is 0.960.
## Files
- `phase1-base/` -- the phase-1 editor. Start here to re-run phase 2 (~1 GPU-hour).
- `phase2-hubert-child/` -- the released model, and the artifact of record.
- `phase2-hubert-a085/` -- an earlier release, interpolated at 0.85.
- `phase2-hubert-a070/` -- the same, at 0.70.
## Use
```bash
git clone https://github.com/pujariaditya/wavepainter && cd wavepainter
pip install -e . && ./setup.sh
python scripts/download_weights.py
./scripts/verify_benchmark.sh
```
Checkpoints are sha256-pinned in `scripts/download_weights.py`. Licence and
attribution, including for the components these weights derive from, are in the
repository's `ATTRIBUTION.md`.