> ## Documentation Index
> Fetch the complete documentation index at: https://docs.veadk.xyz/llms.txt
> Use this file to discover all available pages before exploring further.

# Model

By default an agent uses the globally configured model — the one set via
environment variables or `config.yaml` (see the [Quickstart](/productions/veadk/preview/en/get-started/quickstart)).
VeADK uses `doubao-seed-2-1-pro-260628` as the default inference model.
You can also set a model per agent when you create it.

## Set the model for a single agent

Override the global default with `model_name` and `model_provider`:

```python lines theme={null}
from veadk import Agent

agent = Agent(
    model_name="doubao-seed-2-1-pro-260628",
    model_provider="openai",
)
```

When omitted, `model_provider`, `model_api_base`, and `model_api_key` fall back
to the global configuration.

When a non-native model interface is used (i.e. `model_provider` is set), VeADK
automatically retries once on an HTTP 429 rate-limit error before the model has
produced any output. The retry delay honours the response's `Retry-After` header,
capped at 2 seconds, and defaults to 0.5 seconds when not provided. Once the
model has started producing output, no retry is attempted, to avoid duplicate
content.

### Resolve an API key by name

In addition to setting `MODEL_AGENT_API_KEY` directly, set
`MODEL_AGENT_API_KEY_NAME` to resolve an Ark API key by name. An explicit key
always takes precedence:

```bash lines theme={null}
export MODEL_AGENT_API_KEY_NAME="production-key"
```

You can also pass `model_api_key_name` to `Agent`. This capability is available
in VeADK 1.0.2 and later.

## Configure fallback models

`model_name` also accepts a list: the first entry is the primary model and the
rest are fallbacks, tried in order when the primary model is unavailable.

```python lines theme={null}
agent = Agent(
    model_name=["doubao-seed-2-1-pro-260628", "deepseek-r1-250528"],
)
```

Fallbacks apply to both the default model interface and the Responses API when
`enable_responses=True`.

## Configure cross-provider fallback models

The `model_name` list only supports fallbacks within the same provider. When a fallback model uses a different provider, API base, or API key, use the `model_fallbacks` parameter. It accepts a list of strings or `ModelFallbackEndpoint` objects; VeADK merges them with any `model_name` list fallbacks and passes the combined chain to LiteLLM.

String entries are automatically prefixed with the primary model's `model_provider`, equivalent to appending same-provider model names to the `model_name` list:

```python lines theme={null}
from veadk import Agent

agent = Agent(
    model_name="doubao-seed-2-1-pro-260628",
    model_provider="ark",
    model_fallbacks=["deepseek-r1-250528"],
)
```

For cross-provider fallbacks, use `ModelFallbackEndpoint` to specify a separate provider, API base, and credentials for each fallback endpoint. `ModelFallbackEndpoint` is exported from the top-level `veadk` package:

```python lines theme={null}
from veadk import Agent, ModelFallbackEndpoint

agent = Agent(
    model_name="doubao-seed-2-1-pro-260628",
    model_provider="ark",
    model_api_key="primary-key",
    model_api_base="https://ark.example.com/api/v3",
    model_fallbacks=[
        ModelFallbackEndpoint(
            model_provider="openai",
            model_name="gpt-4o-mini",
            model_api_base="https://api.openai.com/v1",
            model_api_key_env="BACKUP_MODEL_API_KEY",
        ),
    ],
)
```

You can also pass dicts matching `ModelFallbackEndpoint` fields. VeADK accepts both full field names and LiteLLM-style aliases:

```python lines theme={null}
agent = Agent(
    model_name="doubao-seed-2-1-pro-260628",
    model_provider="ark",
    model_api_key="primary-key",
    model_api_base="https://ark.example.com/api/v3",
    model_fallbacks=[
        {
            "model": "openai/gpt-4o-mini",
            "api_key": "openai-key",
            "api_base": "https://api.openai.com/v1",
        }
    ],
)
```

### `ModelFallbackEndpoint` parameters

| Parameter | Type | Default | Description |
| :- | :- | :- | :- |
| `model_name` | `str` | — | Fallback model name. When it does not include a provider prefix and `model_provider` is not set, the primary model's `model_provider` is used. |
| `model_provider` | `Optional[str]` | `None` | Fallback model provider. When it differs from the primary model, the fallback is treated as cross-provider and requires its own credentials and API base. |
| `model_api_base` | `Optional[str]` | `None` | Fallback model API base URL. Required for cross-provider fallbacks. |
| `model_api_key` | `Optional[str]` | `None` | Fallback model API key, passed as a literal value. |
| `model_api_key_env` | `Optional[str]` | `None` | Name of the environment variable holding the fallback model API key; VeADK reads its value at runtime. |
| `model_extra_config` | `dict` | `{}` | Extra configuration passed to LiteLLM. The `extra_headers` and `extra_body` fields are deep-merged with the corresponding fields in the primary model's `model_extra_config`; all other fields use the fallback endpoint's values. |

<Note>
  `ModelFallbackEndpoint` also accepts LiteLLM-style aliases: `model`, `provider`, `api_base` (or `base_url`), `api_key`, `api_key_env`, `extra_config`.
</Note>

<Warning>
  When `enable_responses=True` (Responses API), `model_fallbacks` only supports string entries (same-provider model names). Passing `ModelFallbackEndpoint` or dict-form endpoints raises an error.
</Warning>

<Note>
  When `model_fallbacks` is configured in the `codex` or `piagent` runtime, the fallback chain is ignored because external runtimes do not build a LiteLLM client. Use the default ADK runtime if you need fallbacks. See [Runtime](/productions/veadk/preview/en/components/agent/runtime#switch-execution-backend).
</Note>

<Note>
  When both `model` (a custom LiteLLM client) and `model_fallbacks` are provided, `model_fallbacks` has no effect; configure fallbacks on the custom model object instead.
</Note>

## Responses API

The Responses API is a Volcengine Ark interface with native, efficient context
management, a simpler I/O format, and stronger tool-calling and multimodal
capabilities. Once enabled in VeADK, every turn of the agent's conversation goes
through this interface, giving it native context caching and image, video, and
document understanding.

### Enable

Set `enable_responses=True` when creating the agent:

```python lines theme={null}
from veadk import Agent

agent = Agent(enable_responses=True)
```

Enabling the Responses API requires `google-adk>=1.34.0`, and the model must
support the interface (doubao models after version 0615 support it by default).

### Multimodal input

Beyond text, the Responses API understands images, video, and documents. Pass
multimodal data with `google.genai.types.FileData`; `file_uri` accepts three
sources:

* **Local file path**: `file://{local_path}` — uploaded automatically via the Files API.
* **Files API resource**: `file_id://{file_id}` — for already-uploaded files.
* **Web URL**: a plain `https://` link, typed by its `mime_type`.

For a local image:

```python lines theme={null}
import os
from google.genai import types
from google.genai.types import FileData

local_path = os.path.abspath("example.png")
message = types.UserContent(
    parts=[
        types.Part(text="Describe this image."),
        types.Part(
            file_data=FileData(
                file_uri=f"file://{local_path}",
                mime_type="image/png",
            )
        ),
    ],
)
```

For video, `FileData` may include `video_metadata` with `fps` to control the
frame-sampling rate (default 1, adjustable between 0.2 and 5).

### Configure Ark context management

When the Responses API is enabled, pass Ark-supported `context_management`
settings through `model_extra_config`. This example clears older thinking
content and keeps the most recent thinking turn:

```python lines theme={null}
from veadk import Agent

agent = Agent(
    enable_responses=True,
    model_extra_config={
        "context_management": {
            "edits": [
                {
                    "type": "clear_thinking",
                    "keep": {"type": "thinking_turns", "value": 1},
                }
            ]
        }
    },
)
```

### Context caching

In Responses API mode, session caching is on by default: the initial context is
stored and updated each turn, and later requests merge the cached content with
the new input before calling the model. This significantly reduces repeated-token
cost in long-context scenarios such as multi-turn conversations and complex tool
calls.

Cache hits are visible in the returned event's `usage_metadata`, where
`cached_content_token_count` is the number of tokens served from cache and
`prompt_token_count` is the total input tokens; the hit rate is their ratio.

<Warning>
  When the agent sets `output_schema`, that field conflicts with the caching
  mechanism, so VeADK automatically disables context caching.
</Warning>
