· 4 min read

Amazon Polly - Text-to-Speech with Impressive Results

This article was auto-translated from Chinese. Some nuances may be lost in translation.

Polly is an Amazon cloud service that converts text into speech (text to speech). While text-to-speech isn’t novel—Google Translate can easily do it too—Polly can produce voices that sound as natural as possible based on the text provided, which is a huge boon for language learners. Beyond that, it has a wide range of applications, such as turning subtitles into audio, scripts, narration, dialogues, or even directly recording podcasts. Readers who want to try it out can visit Amazon Polly.

Testing Common Languages

In terms of practical use, I wanted to test Chinese, English, Japanese, and Korean. Here are a few audio samples I generated using Polly:

  • Japanese
  • Chinese (Mandarin)
  • English (US)
  • English (UK)
  • Korean

English goes without saying—the support is remarkably rich. In addition to choosing between American and British accents, there are multiple voices to pick from. It sounds quite natural, and without listening closely, you could easily mistake it for a real person. Chinese sounds a bit unnatural; it doesn’t sound like Taiwanese Mandarin, nor does it quite sound like standard Mainland Mandarin. The level of Japanese support exceeded my expectations: not only are the sentences read smoothly, but when English words are mixed in, Polly will even pronounce the English with a Japanese accent before reading it aloud. For example, for “この件についてはbug ticket必要でしょうか?” (Do we need to file a bug ticket for this issue?), Polly reads it as:

Although it can’t match the nuanced vocal performance of a voice actor across different contexts, it is already a very practical tool for my needs.

Polly offers two options. One is “Neural Voices,” which aims to generate the most natural, human-like sound possible. The other is “Standard,” which already sounds reasonably natural for a synthesized voice, though you can still detect the robotic undertone. Currently, only some languages support “Neural Voices.” Among Chinese, English, Japanese, and Korean, English, Japanese, and Korean all support Neural voices.

Polly also supports SSML (Speech Synthesis Markup Language)1, which lets you add tags to insert pauses into specific sentences or adjust the tone of the voice depending on the scenario, adding a greater sense of immersion.

Pricing

You can refer to the official website for pricing details. Standard voices cost 4.00per1millioncharacters,whileNeuralvoicescost4.00 per 1 million characters, while Neural voices cost 16.00 per 1 million characters. Unless your product requires high-volume text-to-speech conversion, it is extremely affordable for general or auxiliary use, making it easily accessible for indie developers as well.

You are billed monthly for the number of characters of text that you processed. Amazon Polly’s Standard voices are priced at 4.00per1millioncharactersforspeechorSpeechMarksrequests(outsidethefreetier).AmazonPolly’sNeuralvoicesarepricedat4.00 per 1 million characters for speech or Speech Marks requests (outside the free tier). Amazon Polly’s Neural voices are priced at 16.00 per 1 million characters for speech or Speech Marks requests (outside the free tier).

Integration (Using Node.js as an Example)

Integrating Polly is straightforward using the aws-sdk. Here is a sample code snippet:

polly.synthesizeSpeech(
  {
    Text: "おはようございます",
    TextType: "text",
    VoiceId: "Takumi",
    LanguageCode: "ja-JP",
    OutputFormat: "mp3",
  },
  (err, data) => {
    if (err) {
      console.log(err);
    }
    fs.writeFileSync("./result.mp3", data.AudioStream);
  }
);

polly.startSpeechSynthesisTask()

Writing it like this will save the converted audio into result.mp3.

Conclusion

Amazon Polly is both easy to use and inexpensive. It feels like something that could be integrated into many applications to enrich content. Personally, I would use it for language learning—being able to input text and immediately hear a near-realistic pronunciation is incredibly convenient.

For Chinese speakers, while the voice is acceptable, it isn’t an accent that people in Taiwan are used to, which can be somewhat off-putting. It’s a bit of a pity, and I hope a localized Taiwanese accent will be available in the future.

Footnotes

  1. https://docs.aws.amazon.com/polly/latest/dg/ssml.html ↩

Related Posts

Explore Other Topics