Real time voice cloning
Loading...
Files
Date
2024
Authors
Journal Title
Journal ISSN
Volume Title
Publisher
NHCE
Abstract
This Project describe a text-to-speech (TTS) synthesis this is able to generate speech audio inside facet
the voice of diverse audio gadget, consisting of those unseen withinside the route of education. Our
device consists of three independently knowledgeable components: speaker encoder network,
knowledgeable on a speaker verification assignment the usage of an impartial dataset of noisy speech
without transcripts from masses of audio gadget, to generate a fixed-dimensional embedding vector
from great seconds of reference speech from a purpose speaker; a series-to-series synthesis network
based totally mostly on Tacotron 2 that generates a mel spectrogram from text, conditioned at the
speaker embedding;