Automatic Speech Recognition (ASR), Speaker Verification, Speech Synthesis, Text-to-Speech (TTS), Language Modelling, Singing Voice Synthesis (SVS), Voice Conversion (VC)
-
Updated
Oct 19, 2023
Automatic Speech Recognition (ASR), Speaker Verification, Speech Synthesis, Text-to-Speech (TTS), Language Modelling, Singing Voice Synthesis (SVS), Voice Conversion (VC)
End-to-end Automatic Speech Recognition for Madarian and English in Tensorflow
The DARPA TIMIT Acoustic-Phonetic Continuous Speech Corpus.
End-to-End speech recognition implementation base on TensorFlow (CTC, Attention, and MTL training)
Python implementation of pre-processing for End-to-End speech recognition
End-to-end Automatic Speech Recognition for Madarian and English in Tensorflow
Extract mfcc vectors and phones from TIMIT dataset
Sum-Product Networks (SPNs) for Robust Automatic Speaker Identification.
Keyword spotting using RNNs + Edit distance
End-to-end ASR system on TIMIT
Speaker verification using Gaussian Mixture Model (GMM)
My bachelor thesis on Phoneme recognition and alignment on the TIMIT dataset
A simple CRDNN based ASR model for my own understanding of how ASR works and are trained. (Work in progress) If anyone finds any error or have any suggestion please do let me know.
Kaldi recipes for frozen SSL features (HuBERT, mHuBERT, AV-HuBERT) in audio and audio-visual ASR. Includes PCA reduction, TDNN-F chain models, and a GRID lipreading recipe at 6.36% WER on unseen speakers.
Main objective of this model is to develop Automatic Speech Recognition using Deep Neural Network.
Variational Decomposition Autoencoders - Codebase repository for paper: "Variational decomposition autoencoding improves disentanglement of latent representations"
The initial CNN experiments of my bachelor thesis
To associate your repository with the timit-dataset topic, visit your repo's landing page and select "manage topics."