Starting material
Text to speech
Written words
Speech to text
Recorded or live speech
Voice guide
The novita voice models listing is a starting point, not a substitute for listening to a result. First identify whether you need speech generated from text or words recovered from audio; then test a short, representative sample.
These are two different directions for a speech workflow. Use the input and output columns to narrow your search before comparing individual models.
Text to speech
Written words
Speech to text
Recorded or live speech
Text to speech
Audible speech
Speech to text
Written transcript
Text to speech
A suitable voice and delivery
Speech to text
A transcription model suited to the recording
Text to speech
The spoken result preserves the intended wording and pronunciation
Speech to text
The transcript captures the words that were spoken
Text to speech
Ambiguous names, abbreviations or punctuation
Speech to text
Background noise, overlapping speakers or unclear audio
Text to speech
Listen while reading the source text
Speech to text
Read the transcript while replaying the source
Neither direction carries every cue from its source. These use cases show what to check rather than assuming a clear-looking result is complete.
Turns a short script into narration. Written emphasis may not translate into the intended pause or tone.
Listen for names, timing and emphasis against the edited scene. For visual generation rather than narration, novita image models covers image-focused options.
novita image modelsTurns recorded answers into text. A transcript can omit hesitation, speaker changes or emotion.
Keep the recording available and replay passages where attribution or exact wording matters. For a separate visual-output workflow, novita image models describes image options.
novita image modelsPrepares spoken versions of written instructions. Formatting alone may not tell a listener where one step ends.
Test the audio without looking at the page and revise sentences that are hard to follow. If the project also needs generated visuals, novita image models addresses that different output type.
novita image modelsCompares voice samples using the same short passage. A favorable result on one sentence may not carry over to longer material.
Include a name, a number and a complete instruction in each test. For comparisons involving visual assets, novita image models focuses on image outputs.
novita image modelsMake the test small enough to repeat. A consistent input makes differences easier to hear or read than an improvised demonstration.
Decide whether your source is text or audio and write down the output you need. Check each candidate's stated input and output before testing it.
For generated speech, use a passage with the names and punctuation your project uses. For transcription, use a short recording with the same noise and speaking style as the real source.
Hold the sample constant across candidates. Note omitted words, unexpected pronunciation, speaker confusion and places that require a second listen.
For generated voice, listen once without the script and again while following each word. For a transcript, replay uncertain passages and confirm names and numbers manually. Keep the original input so you can repeat the check after changing a model or revising the material.
The voice-model listing is the relevant starting point for exploring voice-related options. Check each listing's stated task, accepted input and output before preparing a sample; the category name alone does not establish that every model handles the same job.
Test the same passage with each candidate voice so the comparison stays consistent. Listen for intelligibility, pronunciation and whether the delivery suits the intended audience, rather than judging from a voice name alone.
Do not assume language support is identical across a voice-model listing. Check the information for the individual model and test the language, accent and names that appear in your actual material.
For generated audio, compare the recording with its source text and listen for skipped words, awkward pauses and mispronounced names. For a transcription result, replay the original audio while checking uncertain words, speakers and numbers.