malarinv/tacotron2

mirror of https://github.com/malarinv/tacotron2 synced 2026-06-11 10:02:07 +00:00

Go to file

Rafael Valle 151eef9466 README.md: typo

2018-05-03 15:13:06 -07:00

LICENSE

adding readme and license

2018-05-03 15:10:51 -07:00

README.md

README.md: typo

2018-05-03 15:13:06 -07:00

tensorboard.png

tensorboard.png: adding tensorboard image

2018-05-03 15:11:54 -07:00

README.md

Tacotron 2 (without wavenet)

Tacotron 2 PyTorch implementation of Natural TTS Synthesis By Conditioning Wavenet On Mel Spectrogram Predictions.

This implementation includes distributed and fp16 support and uses the LJSpeech dataset.

Distributed and FP16 support relies on work by Christian Sarofeen and NVIDIA's frameworks team.

Pre-requisites

NVIDIA GPU + CUDA cuDNN

Setup

Download and extract the LJ Speech dataset
Clone this repo: git clone https://github.com/NVIDIA/tacotron2.git
CD into this repo: cd tacotron2
Update .wav paths: sed -i -- 's,DUMMY,ljs_dataset_folder/wavs,g' *.txt
Install pytorch 0.4
Install python requirements or use docker container (tbd)
- Install python requirements: pip install requirements.txt
- OR
- Docker container (tbd)

Training

python train.py --output_directory=outdir --log_directory=logdir
(OPTIONAL) tensorboard --logdir=outdir/logdir

Multi-GPU (distributed) and FP16 Training

python -m multiproc train.py --output_directory=/outdir --log_directory=/logdir --hparams=distributed_run=True

Inference

jupyter notebook --ip=127.0.0.1 --port=31337
load inference.ipynb

nv-wavenet: Faster than real-time wavenet inference

Acknowledgements

This implementation is inspired or uses code from the following repos: Ryuchi Yamamoto, Keith Ito, Prem Seetharaman.

We are thankful to the Tacotron 2 paper authors, specially Jonathan Shen, Yuxuan Wang and Zongheng Yang.