VoiceSmith

VoiceSmith makes it possible to train and infer on both single and multispeaker models without any coding experience. It fine-tunes a pretty solid text to speech pipeline based on a modified version of DelightfulTTS and UnivNet on your dataset. Both models were pretrained on a proprietary 5000 speaker dataset. It also provides some tools for dataset preprocessing like automatic text normalization. Windows (only CPU supported currently) or any Linux based operating system. If you want to run this on macOS you have to follow the steps in build from source in order to create the installer. This is untested since I don't currently own a Mac. NVIDIA GPU with CUDA support is highly recommended, you can train on CPU otherwise but it will take days if not weeks. VoiceSmith currently uses a two-stage modified DelightfulTTS and UnivNet pipeline.

Features

Windows (only CPU supported currently)
Train and infer on both single and multispeaker models
No coding experience needed
It fine-tunes a pretty solid text to speech pipeline based on a modified version of DelightfulTTS
Models were pretrained on a proprietary 5000 speaker dataset
Provides some tools for dataset preprocessing like automatic text normalization

Project Samples

Project Activity

See All Activity >

License

Apache License V2.0

Follow VoiceSmith

VoiceSmith Web Site

Other Useful Business Software

Data management solutions for confident marketing

For companies wanting a complete Data Management solution that is native to Salesforce

Verify, deduplicate, manipulate, and assign records automatically to keep your CRM data accurate, complete, and ready for business.

Learn More

Rate This Project

User Reviews

Be the first to post a review of VoiceSmith!

Additional Project Details

Operating Systems

Mac, Windows

Programming Language

Python

Related Categories

Python Voice Cloning Software

Registered

2023-03-23

Similar Business Software

Murf AI

Murf API is an advanced text-to-speech (TTS) solution that transforms written text into natural, lifelike voiceovers with remarkable accuracy and ease. It empowers developers and businesses with a suite of sophisticated features, including pitch and speed modulation, audio duration adjustments,...

See Software
ElevenLabs

The most realistic and versatile AI speech software, ever. Eleven brings the most compelling, rich and lifelike voices to creators and publishers seeking the ultimate tools for storytelling. Generate top-quality spoken audio in any voice and style with the most advanced and multipurpose AI...

See Software
Zyphra Zonos

Zyphra is excited to announce the release of Zonos-v0.1 beta, featuring two expressive and real-time text-to-speech models with high-fidelity voice cloning. We are releasing our 1.6B transformer and 1.6B hybrid under an Apache 2.0 license. It is difficult to quantitatively measure quality in the...

See Software
Play.ht

AI Powered Text to Voice Generation. Play.ht offers uncanny, high-fidelity AI Voices for any project where you need human-sounding voice overs and performances. Hollywood studios, auto manufacturers, and other large enterprises use Play.ht to create realistic and engaging voiceovers...

See Software
Chatterbox

Chatterbox is a free, open source voice cloning AI model developed by Resemble AI, licensed under MIT. It enables zero-shot voice cloning using just 5 seconds of reference audio, eliminating the need for training. The model offers expressive speech synthesis with unique emotion control, allowing...

See Software
Kukarella

Kukarella is an AI-powered audio and voice-content platform that enables users to create professional voice-overs, multi-speaker dialogues, transcriptions, and visual content all within one integrated environment. The platform features a text-to-speech tool with access to hundreds of...

See Software

Report inappropriate content

VoiceSmith

[WIP] VoiceSmith makes training text to speech models easy

Get an email when there's a new version of VoiceSmith

Features

Project Samples

Project Activity

Categories

License

Follow VoiceSmith

User Reviews

Additional Project Details

Operating Systems

Programming Language

Related Categories

Registered