Suno推出Speech公开测试,可同时生成语音和背景音乐
AI music maker Suno now generates spoken words
Suno已在网页端和移动端推出Speech公开测试,可依据脚本或提示描述生成语音,并同时生成配套背景音乐。在Create页选择Speech后,Simple模式通过提示描述需求,Advanced模式可写入自定义脚本并调节声音性别、说话风格和变化程度;背景音乐可以关闭,单条最长约八分钟。
Suno is branching out from the world of AI music, launching a new feature that generates spoken voices based on scripts or prompted descriptions. Speech is now available in public beta across Suno’s web and mobile platforms, and allows you to simultaneously generate voiceovers and background music to accompany them.
“Music will always be at the heart of Suno and what we build. At the same time, our vision has always extended to other forms of human expression,” Suno chief product officer, Jack Brody, said in the announcement. “Today, we’re expanding what’s possible in Suno with Speech: the first audio model that generates voice and music together as one cohesive track.”
AI-generated speech is hardly new — DeepMind has been experimenting with deep learning speech synthesis for a decade, Adobe has a text-to-speech tool, and ElevenLabs has become one of the most recognizable platforms for it since launching in 2023. Suno is just throwing its hat into the ring — likely in an attempt to diversify the platform, given its music generator has attracted so many lawsuits.
Pairing AI music with generated voices is Suno’s spin on text-to-speech tools. It’s optional, meaning you can easily turn off the background music with a toggle if you just want clean speech, but the idea is that it’ll compliment certain use cases for generative spoken word — such as having a calming soundtrack for poems, or something more energetic for dramatic voiceovers and encouraging speeches.
To use the feature, select the “Create” tab, and navigate to the Speech option. There are two modes: Simple, which allows you to describe what you want to create via the provided prompt box (such as “a pirate captain rallying his crew”), or the Advanced mode that lets you add a custom script if you already know exactly what you want it to say. Advanced settings also let you adjust the gender of the AI voice, speech style, and how much variety each voice generation will have. Speech has a maximum duration of around eight minutes.
Suno admits that the feature is far from perfect, but says it’ll keep improving Speech around user feedback. “Beta really does mean beta,” said Brody. “Occasionally, British accents can wander off to Australia and back. Dramatic pauses may be very dramatic. You will almost certainly discover uses for this that never occurred to us.”
Follow topics and authors from this story to see more like this in your personalized homepage feed and to receive email updates.
- Jess Weatherbed
来源:The Verge · AI · theverge.com