Another HN user suggested embedding sound effects. The lead author of MOSS-TTS emailed me that the next version will have a SOTA voice designer. Poke around the voice design space, you'll find text-to-effects here and there. I certainly have my eyes on this space.
ninininino 1 days ago [-]
KeenLore is really cool.
soundworlds 1 days ago [-]
You should look into using AI to generate code that synthesizes sounds. Not exactly what you asked for, but I do think it is an approach worth considering: https://m.youtube.com/watch?v=1-i45X5aj94
SyneRyder 1 days ago [-]
I've had some luck with that approach - giving a frontier LLM a reference sound, asking it to replicate the sound via code (basically asking it to create a physical model, I guess) and being able to steer the output by talking with the LLM.
I was working with very simple elements, but I was surprised by some of the outputs.
I would have also suggested Waves Illugen, but it turns out that is text-only, you can't give it an audio reference.
dmos62 12 hours ago [-]
Great video!
moonu 1 days ago [-]
Even though they're technically trained for music, it might be worth testing Suno/Lyria to see if they're able to do this. You might have to isolate it afterwards, but seems like it could be viable
xg15 1 days ago [-]
Maybe a dumb idea, but how good are image models with manipulating spectrogram images? Then a workaround could be: convert input audio to spectrogram -> pass spectrogram + prompt to an image model -> convert modified spectrogram back to audio waveform.
I wouldn't call it a dumb idea, but there's soooo much subtlety to sound that wont be visible in any reasonably sized image.
narrationbox 1 days ago [-]
Could be interesting, though spectrogram voice gen is a step back in terms of advancement. Vocoder models that take in Mel Specs suffer from a lot of issues like hissing due to Griffin Lim issues. Would be curious to see if diffusion models can work on neural codecs directly.
I think Google had one called riffusion (the first version was designed for specs)
The DEMON realtime engine is open source. The hosted service has a number of integrations with popular audio tools. (I'm on the team)
jajazheng 20 hours ago [-]
Hi! I’m unsure if this is what you were looking for, QWEN3 TTS has a “clone” feature. You can use the local version, give it a voice reference and text and it will read the text with your instructions with the sample voice you selected. I used this combo to make a customised “audiobook” for my mother. It worked quite well.
chr15m 1 days ago [-]
Apparently the AudioX and AudioLDM(2) models do this but I think you've found a genuine gap.
ruiqingcn 7 hours ago [-]
INDEX TTS2
narrationbox 1 days ago [-]
Plenty of models take in text + audio and spits out audio. It's the format of most newer generation accent conversion/voice cloning models.
What's your exact use case?
chr15m 1 days ago [-]
Sound effects are completely different to voice, which those models are trained to output.
KeenLore is my locally hosted full-cast emotive audiobook generator I'm developing for my novel:
https://www.youtube.com/watch?v=WAeHgE94rVo
Another HN user suggested embedding sound effects. The lead author of MOSS-TTS emailed me that the next version will have a SOTA voice designer. Poke around the voice design space, you'll find text-to-effects here and there. I certainly have my eyes on this space.
I was working with very simple elements, but I was surprised by some of the outputs.
I would have also suggested Waves Illugen, but it turns out that is text-only, you can't give it an audio reference.
I think Google had one called riffusion (the first version was designed for specs)
This generates audio embeddings - much like CLIP does for visual inputs.
The DEMON realtime engine is open source. The hosted service has a number of integrations with popular audio tools. (I'm on the team)
What's your exact use case?