> ## Documentation Index
> Fetch the complete documentation index at: https://docs.veadk.xyz/llms.txt
> Use this file to discover all available pages before exploring further.

# Text-to-speech

## Overview

Tool identifier `text_to_speech`.

`text_to_speech` synthesizes text into speech and saves the audio locally.

Import path: `from veadk.tools.builtin_tools.tts import text_to_speech`

## Environment & prerequisites

<Warning>
  Requirements:

  1. Configure the API key for the agent's reasoning model.
  2. Configure the VeSpeech service App ID and API key.
</Warning>

Environment variables:

* `MODEL_AGENT_API_KEY`: API key for the agent's reasoning model
* `TOOL_VESPEECH_APP_ID`: VeSpeech service App ID
* `TOOL_VESPEECH_API_KEY`: VeSpeech service API key
* `TOOL_VESPEECH_SPEAKER`: voice, defaults to `zh_female_vv_uranus_bigtts`
* `TOOL_VESPEECH_AUDIO_OUTPUT_PATH`: audio output directory, defaults to the system temp directory

## Usage

```python title="examples/tools/tts/agent.py" lines theme={null}
import asyncio

from veadk import Agent, Runner
from veadk.memory.short_term_memory import ShortTermMemory
from veadk.tools.builtin_tools.tts import text_to_speech

agent = Agent(
    name="tts_agent",
    model_name="doubao-seed-2-1-pro-260628",
    description="An agent that speaks text aloud.",
    instruction="Use the text_to_speech tool to synthesize the user's text into speech.",
    tools=[text_to_speech],
)

runner = Runner(agent=agent, short_term_memory=ShortTermMemory())


async def main():
    response = await runner.run("Read this aloud: Hello, welcome to VeADK")
    print(response)


if __name__ == "__main__":
    asyncio.run(main())
```

## Parameters and output format

| Parameter | Type | Default | Description |
| :- | :- | :- | :- |
| `text` | `str` | Required | Text to synthesize |
| `tool_context` | `ToolContext` | Injected | Invocation context |

Success returns `{"saved_audio_path": "..."}`; failures may return `error`. Audio is saved locally as 24 kHz PCM in a `.pcm` file, not a WAV or MP3 file. Use matching PCM parameters when converting it. Missing output directories are created, and the process needs write permission.

The tool uses the `seed-tts-2.0` resource and the Volcengine speech endpoint. Enable the resource and selected voice for the App ID. `CLOUD_PROVIDER=byteplus` does not switch the speech endpoint. The tool may also attempt playback on the host machine; check the output file even if no audio device is available.

Check `error`, then verify that `saved_audio_path` exists and is nonempty. A server-local path is not a user-downloadable link; publish the file through your application when needed.
