Real time voice cloning

Loading...
Thumbnail Image
Files
Date
2024
Journal Title
Journal ISSN
Volume Title
Publisher
NHCE
Abstract
This Project describe a text-to-speech (TTS) synthesis this is able to generate speech audio inside facet the voice of diverse audio gadget, consisting of those unseen withinside the route of education. Our device consists of three independently knowledgeable components: speaker encoder network, knowledgeable on a speaker verification assignment the usage of an impartial dataset of noisy speech without transcripts from masses of audio gadget, to generate a fixed-dimensional embedding vector from great seconds of reference speech from a purpose speaker; a series-to-series synthesis network based totally mostly on Tacotron 2 that generates a mel spectrogram from text, conditioned at the speaker embedding;
Description
Keywords
Citation
Collections