October 3, 2026 — Suno has opened a public beta of Speech, a feature that generates spoken audio and original background music together as a single track, on web, iOS, and Android.
What happened
Suno announced Speech (beta) on October 1 in a company blog post by chief product officer Jack Brody. Users enter an idea, a poem, or existing writing, then describe the voice and musical style they want. The result is spoken audio set to original background music, generated together in one take.
The feature opened to everyone after about a month of testing with a small group. Suno describes it as the first model to generate voice and music together as one cohesive track, and frames it as a new canvas for creative entertainment — personal creations like the birthday songs and wedding tributes its community already makes.
How it works in practice
Suno’s release notes give examples including bedtime stories over soft piano, hype speeches over stadium drums, guided meditations, and pep talks. A toggle switches the background score on or off, so the tool doubles as a bare text-to-speech generator when music would get in the way.
You can type out exact dialogue to be spoken, or describe a tone and let the model improvise. Each generation costs 10 credits for two takes — the same as a song generation — and Suno has not announced separate pricing or usage limits for Speech.
Why it matters
Suno built its name on turning text prompts into full songs. Now it wants to own the whole audio workflow — and it lands in a crowded field. Dedicated voice AI from ElevenLabs has been a real business since 2023, and Adobe and DeepMind have their own speech tooling. Suno’s differentiator is the combination: one model, one take, voice and score together, with a pause in the reading able to land on a change in the music.
Be clear-eyed about the fine print: Suno itself warns the beta can produce inconsistent accents and unusually dramatic pauses, and independent reviewers note there’s no published length cap, no confirmed commercial-use terms for Speech, and no API. As one reviewer put it: fine to play with if you already use Suno, but don’t build a podcast or audiobook workflow on it yet.
The subscriber’s take: this is the generative-media convergence in miniature. Music generators, voice models, and video tools all started as separate categories; they’re being assembled into integrated creative suites. For Suno specifically, a second product line is also a hedge — its core music business remains tangled in copyright lawsuits from major labels.
Frequently asked questions
What is Suno Speech?
A new public-beta feature in Suno that turns a script or prompted description into spoken audio with AI-generated background music, produced together as one track.
Which platforms support it?
Speech is available to all users on Suno’s web platform and its iOS and Android apps.
Can I use it for voice without music?
Yes — a toggle switches the background score on or off, making it usable as a straightforward text-to-speech tool. Suno hasn’t promised consistently isolated voice output, so test before relying on it.
What does it cost?
Suno hasn’t announced Speech-specific pricing. Reviewers report each generation costs 10 credits for two takes, the same as a song generation.
What are the limitations?
Suno warns of accent inconsistencies and exaggerated dramatic pauses. There is no published length cap, no confirmed separate commercial-use terms, and no API yet.
Sources: Unite.AI, Okay News, Tech Startups, The Verge, Songtested.

