Amazon Polly - Text-to-Speech with Impressive Results
Polly is an Amazon cloud service that converts text into speech (text to speech). While text-to-speech isn’t novel—Google Translate can easily do it too—Polly can produce voices that sound as natural as possible based on the text provided, which is a huge boon for language learners. Beyond that, it has a wide range of applications, such as turning subtitles into audio, scripts, narration, dialogues, or even directly recording podcasts. Readers who want to try it out can visit Amazon Polly.
Testing Common Languages
In terms of practical use, I wanted to test Chinese, English, Japanese, and Korean. Here are a few audio samples I generated using Polly:
- Japanese
- Chinese (Mandarin)
- English (US)
- English (UK)
- Korean
English goes without saying—the support is remarkably rich. In addition to choosing between American and British accents, there are multiple voices to pick from. It sounds quite natural, and without listening closely, you could easily mistake it for a real person. Chinese sounds a bit unnatural; it doesn’t sound like Taiwanese Mandarin, nor does it quite sound like standard Mainland Mandarin. The level of Japanese support exceeded my expectations: not only are the sentences read smoothly, but when English words are mixed in, Polly will even pronounce the English with a Japanese accent before reading it aloud. For example, for “この件についてはbug ticket必要でしょうか?” (Do we need to file a bug ticket for this issue?), Polly reads it as:
Although it can’t match the nuanced vocal performance of a voice actor across different contexts, it is already a very practical tool for my needs.
Polly offers two options. One is “Neural Voices,” which aims to generate the most natural, human-like sound possible. The other is “Standard,” which already sounds reasonably natural for a synthesized voice, though you can still detect the robotic undertone. Currently, only some languages support “Neural Voices.” Among Chinese, English, Japanese, and Korean, English, Japanese, and Korean all support Neural voices.
Polly also supports SSML (Speech Synthesis Markup Language)1, which lets you add tags to insert pauses into specific sentences or adjust the tone of the voice depending on the scenario, adding a greater sense of immersion.
Pricing
You can refer to the official website for pricing details. Standard voices cost 16.00 per 1 million characters. Unless your product requires high-volume text-to-speech conversion, it is extremely affordable for general or auxiliary use, making it easily accessible for indie developers as well.
You are billed monthly for the number of characters of text that you processed. Amazon Polly’s Standard voices are priced at 16.00 per 1 million characters for speech or Speech Marks requests (outside the free tier).
Integration (Using Node.js as an Example)
Integrating Polly is straightforward using the aws-sdk. Here is a sample code snippet:
polly.synthesizeSpeech(
{
Text: "おはようございます",
TextType: "text",
VoiceId: "Takumi",
LanguageCode: "ja-JP",
OutputFormat: "mp3",
},
(err, data) => {
if (err) {
console.log(err);
}
fs.writeFileSync("./result.mp3", data.AudioStream);
}
);
polly.startSpeechSynthesisTask()
Writing it like this will save the converted audio into result.mp3.
Conclusion
Amazon Polly is both easy to use and inexpensive. It feels like something that could be integrated into many applications to enrich content. Personally, I would use it for language learning—being able to input text and immediately hear a near-realistic pronunciation is incredibly convenient.
For Chinese speakers, while the voice is acceptable, it isn’t an accent that people in Taiwan are used to, which can be somewhat off-putting. It’s a bit of a pity, and I hope a localized Taiwanese accent will be available in the future.
Footnotes
Related Posts
- When a Measure Becomes a Target: From the Window Tax to Pull Request Counts I once wrote a script to tally how many PRs I contributed in a quarter, how many reviews I left, and how many tickets I closed, hoping to use numbers to prove my output to my manager. My manager simply remarked that performance isn't just about output. Years later, I finally understood—when a measure becomes a target, it ceases to be a good measure. From the British window tax and the Hanoi rat bounty to evaluating developers by PR counts today, the underlying mechanism is exactly the same.
- Using Cloudflare Images for Image Storage and Transformation Putting an image on a webpage is the simplest task in frontend development. But doing it properly—including resizing, generating multiple formats, and withstanding heavy traffic—is actually an entire end-to-end solution. Eventually, I offloaded everything to Cloudflare Images, keeping only a single original image.
- Stop Using AWS Access Keys Access Keys are an easily overlooked security risk in AWS. By pairing OIDC with IAM Roles, GitHub Actions can securely operate AWS resources without storing any secrets.
- Database Primary Keys: AUTO_INCREMENT, UUID, and UUIDv7 Backend developers often face the choice of primary keys: should you use auto-increment or UUID? What about collisions? How does UUIDv7 compare to created_at + index in performance? Here are the design decisions and benchmark results from testing 20 million rows.