Skip to main content

Overview

Tool identifier text_to_speech. text_to_speech synthesizes text into speech and saves the audio locally. Import path: from veadk.tools.builtin_tools.tts import text_to_speech

Environment & prerequisites

Requirements:
  1. Configure the API key for the agent’s reasoning model.
  2. Configure the VeSpeech service App ID and API key.
Environment variables:
  • MODEL_AGENT_API_KEY: API key for the agent’s reasoning model
  • TOOL_VESPEECH_APP_ID: VeSpeech service App ID
  • TOOL_VESPEECH_API_KEY: VeSpeech service API key
  • TOOL_VESPEECH_SPEAKER: voice, defaults to zh_female_vv_uranus_bigtts
  • TOOL_VESPEECH_AUDIO_OUTPUT_PATH: audio output directory, defaults to the system temp directory

Usage

examples/tools/tts/agent.py

Parameters and output format

Success returns {"saved_audio_path": "..."}; failures may return error. Audio is saved locally as 24 kHz PCM in a .pcm file, not a WAV or MP3 file. Use matching PCM parameters when converting it. Missing output directories are created, and the process needs write permission. The tool uses the seed-tts-2.0 resource and the Volcengine speech endpoint. Enable the resource and selected voice for the App ID. CLOUD_PROVIDER=byteplus does not switch the speech endpoint. The tool may also attempt playback on the host machine; check the output file even if no audio device is available. Check error, then verify that saved_audio_path exists and is nonempty. A server-local path is not a user-downloadable link; publish the file through your application when needed.
Last modified on September 19, 2026