Torchaudio Documentation

Torchaudio is a library for audio and signal processing with PyTorch. It provides I/O, signal and data processing functions, datasets, model implementations and application components.

Tutorials

AM inference with CUDA CTC Beam Seach Decoder

Topics: Pipelines, ASR, CTC Decoder, CUDA CTC-Decoder

Learn how to perform ASR beam search decoding with GPU, using torchaudio.models.decoder.cuda_ctc_decoder.

Loading waveform Tensors from files and saving them

Topics: I/O

Learn how to query/load audio files and save waveform tensors to files, using torchaudio.info, torchaudio.load and torchaudio.save functions.

CTC Forced Alignment API

Topics: CTC, Forced Alignment

Learn how to use TorchAudio's CTC forced alignment API (torchaudio.functional.forced_align).

Forced alignment for multilingual data

Topics: Forced Alignment

Learn how to use align multiligual data using TorchAudio's CTC forced alignment API (torchaudio.functional.forced_align) and a multiligual Wav2Vec2 model.

Streaming media decoding with StreamReader

Topics: I/O, StreamReader

Learn how to load audio/video to Tensors using torchaudio.io.StreamReader class.

Device input, synthetic audio/video, and filtering with StreamReader

Topics: I/O, StreamReader

Learn how to load media from hardware devices, generate synthetic audio/video, and apply filters to them with torchaudio.io.StreamReader.

Citing torchaudio

If you find torchaudio useful, please cite the following paper:

Yang, Y.-Y., Hira, M., Ni, Z., Chourdia, A., Astafurov, A., Chen, C., Yeh, C.-F., Puhrsch, C., Pollack, D., Genzel, D., Greenberg, D., Yang, E. Z., Lian, J., Mahadeokar, J., Hwang, J., Chen, J., Goldsborough, P., Roy, P., Narenthiran, S., Watanabe, S., Chintala, S., Quenneville-Bélair, V, & Shi, Y. (2021). TorchAudio: Building Blocks for Audio and Speech Processing. arXiv preprint arXiv:2110.15018.

In BibTeX format:

@article{yang2021torchaudio,
  title={TorchAudio: Building Blocks for Audio and Speech Processing},
  author={Yao-Yuan Yang and Moto Hira and Zhaoheng Ni and
          Anjali Chourdia and Artyom Astafurov and Caroline Chen and
          Ching-Feng Yeh and Christian Puhrsch and David Pollack and
          Dmitriy Genzel and Donny Greenberg and Edward Z. Yang and
          Jason Lian and Jay Mahadeokar and Jeff Hwang and Ji Chen and
          Peter Goldsborough and Prabhat Roy and Sean Narenthiran and
          Shinji Watanabe and Soumith Chintala and
          Vincent Quenneville-Bélair and Yangyang Shi},
  journal={arXiv preprint arXiv:2110.15018},
  year={2021}
}

@misc{hwang2023torchaudio,
   title={TorchAudio 2.1: Advancing speech recognition, self-supervised learning, and audio processing components for PyTorch},
   author={Jeff Hwang and Moto Hira and Caroline Chen and Xiaohui Zhang and Zhaoheng Ni and Guangzhi Sun and Pingchuan Ma and Ruizhe Huang and Vineel Pratap and Yuekai Zhang and Anurag Kumar and Chin-Yun Yu and Chuang Zhu and Chunxi Liu and Jacob Kahn and Mirco Ravanelli and Peng Sun and Shinji Watanabe and Yangyang Shi and Yumeng Tao and Robin Scheibler and Samuele Cornell and Sean Kim and Stavros Petridis},
   year={2023},
   eprint={2310.17864},
   archivePrefix={arXiv},
   primaryClass={eess.AS}
}

Torchaudio Documentation

Tutorials

AM inference with CUDA CTC Beam Seach Decoder

Loading waveform Tensors from files and saving them

CTC Forced Alignment API

Forced alignment for multilingual data

Streaming media decoding with StreamReader

Device input, synthetic audio/video, and filtering with StreamReader

Citing torchaudio

Docs

Tutorials

Resources