A singing voice synthesizer. Write a melody in the piano roll, type a word on each note, and Roboto sings it. The voice is synthesized with no samples, from a glottal pulse source, five formant resonators and shaped noise for consonants. Roboto is free and runs as a plugin or standalone on macOS, Windows, Linux, iPhone and iPad.
Piano roll where every note carries a word. A note named + sings the next syllable of the previous word.
Piano roll editing: box select, Shift + click, group drag, Cmd + drag to duplicate, and cut, copy and paste at the marker. A note named - keeps singing the vowel before it at its own pitch.
Text tab that holds the whole song as text (word:note:beats), so songs can be typed or pasted. Lines such as @tempo 120, @transpose -12, @chorus on, @talkbox on, @language de and @voice Tracy set the song, and further down they change it from that line on. Cmd + Return (Ctrl + Enter on Windows and Linux) sings from the note at the cursor, and the word being sung is highlighted.
Automation curves along the song for volume, pitch bend, vibrato depth, breath, formant shift and chorus. Points snap to the grid and to semitones for pitch bend.
Built-in songs in every language to try.
Import MIDI, karaoke MIDI (.kar), UltraStar, LRC lyrics, SRT subtitles and song text, from the menu or by dropping the file on the window. MIDI notes without lyrics are sung on "ah".
Musicalize button: plain lyric lines get a pentatonic melody, one note per syllable.
Voice
Synthesized voice with no samples: a glottal pulse source, five formant resonators and shaped noise for consonants.
Sings in English, German, French, Spanish, Japanese, Russian, Finnish and Swedish, with spelling rules for each language and the vowels and consonants of its speakers. Spanish follows the pronunciation of Spain. Japanese kanji are read on macOS and iOS.
English words are turned into phonemes by built-in rules and a word list. Phonemes can also be typed directly, for example S-T-R-IY-T.
Ten voice presets: Default, Robot, Soprano, Bass, Child, Whisper, Choir, Opera, Giant and Alien.
Classic voices with the sound of 1990s singing software: Robert, Sarah, Tracy, Andy, Abe, Webster, Richard, Desmond, Male Breath and Female Breath. A second formant engine sings them, and the same sliders and curves shape them.
Talk box mode: a plucked string, driven through a speaker driver and a tube, takes the place of the voice and is shaped by the mouth, like a guitar or synth talk box. It works with every voice, Classic voices included, and stays as loud as the singing voice.
Controls
Transpose, tempo with optional host sync, glide between notes, vibrato depth and rate, jitter and breath.
Consonant level and length, formant shift (vocal tract size) and formant bandwidth.
Reverb and double track.
Display
Animated robot face with LED eyes and a mouth that follow the voice.
Live waveform, and lyrics that light up in time with the song.
MIDI and export
MIDI keys from C1 sing from the matching bar while held.
Play the melody live: keys on MIDI channel 2 replace the written notes while the words keep going, one note at a time, sliding between keys at the Glide speed.
MIDI controllers drive volume, vibrato, breath, formant shift and chorus.
Export to WAV, AAC (.m4a, on macOS, Windows and iOS), MIDI, karaoke MIDI (.kar), UltraStar, LRC lyrics and SRT subtitles. The MIDI file holds notes, lyrics, tempo and the curves as controllers, and the times in the lyric files match the WAV.
Video export (MP4, 1280x720 at 60 fps) of the robot, waveform and lyrics with the song, on macOS, Windows and iOS.
Plugin
VST3, AU and standalone on macOS, VST3 and standalone on Windows and Linux, AUv3 and standalone on iPhone and iPad.
All voice parameters can be automated by the host.
System Requirements
macOS 12 or later (Apple Silicon and Intel): VST3, AU and standalone.
Windows and Linux (64-bit): VST3 and standalone.
iOS 15 or later (iPhone and iPad): standalone and AUv3.