config.yaml (see the Quickstart).
VeADK uses doubao-seed-2-1-pro-260628 as the default inference model.
You can also set a model per agent when you create it.
Set the model for a single agent
Override the global default withmodel_name and model_provider:
model_provider, model_api_base, and model_api_key fall back
to the global configuration.
When a non-native model interface is used (i.e. model_provider is set), VeADK
automatically retries once on an HTTP 429 rate-limit error before the model has
produced any output. The retry delay honours the response’s Retry-After header,
capped at 2 seconds, and defaults to 0.5 seconds when not provided. Once the
model has started producing output, no retry is attempted, to avoid duplicate
content.
Resolve an API key by name
In addition to settingMODEL_AGENT_API_KEY directly, set
MODEL_AGENT_API_KEY_NAME to resolve an Ark API key by name. An explicit key
always takes precedence:
model_api_key_name to Agent. This capability is available
in VeADK 1.0.2 and later.
Configure fallback models
model_name also accepts a list: the first entry is the primary model and the
rest are fallbacks, tried in order when the primary model is unavailable.
enable_responses=True.
Configure cross-provider fallback models
Themodel_name list only supports fallbacks within the same provider. When a fallback model uses a different provider, API base, or API key, use the model_fallbacks parameter. It accepts a list of strings or ModelFallbackEndpoint objects; VeADK merges them with any model_name list fallbacks and passes the combined chain to LiteLLM.
String entries are automatically prefixed with the primary model’s model_provider, equivalent to appending same-provider model names to the model_name list:
ModelFallbackEndpoint to specify a separate provider, API base, and credentials for each fallback endpoint. ModelFallbackEndpoint is exported from the top-level veadk package:
ModelFallbackEndpoint fields. VeADK accepts both full field names and LiteLLM-style aliases:
ModelFallbackEndpoint parameters
ModelFallbackEndpoint also accepts LiteLLM-style aliases: model, provider, api_base (or base_url), api_key, api_key_env, extra_config.When
model_fallbacks is configured in the codex or piagent runtime, the fallback chain is ignored because external runtimes do not build a LiteLLM client. Use the default ADK runtime if you need fallbacks. See Runtime.When both
model (a custom LiteLLM client) and model_fallbacks are provided, model_fallbacks has no effect; configure fallbacks on the custom model object instead.Responses API
The Responses API is a Volcengine Ark interface with native, efficient context management, a simpler I/O format, and stronger tool-calling and multimodal capabilities. Once enabled in VeADK, every turn of the agent’s conversation goes through this interface, giving it native context caching and image, video, and document understanding.Enable
Setenable_responses=True when creating the agent:
google-adk>=1.34.0, and the model must
support the interface (doubao models after version 0615 support it by default).
Multimodal input
Beyond text, the Responses API understands images, video, and documents. Pass multimodal data withgoogle.genai.types.FileData; file_uri accepts three
sources:
- Local file path:
file://{local_path}— uploaded automatically via the Files API. - Files API resource:
file_id://{file_id}— for already-uploaded files. - Web URL: a plain
https://link, typed by itsmime_type.
FileData may include video_metadata with fps to control the
frame-sampling rate (default 1, adjustable between 0.2 and 5).
Configure Ark context management
When the Responses API is enabled, pass Ark-supportedcontext_management
settings through model_extra_config. This example clears older thinking
content and keeps the most recent thinking turn:
Context caching
In Responses API mode, session caching is on by default: the initial context is stored and updated each turn, and later requests merge the cached content with the new input before calling the model. This significantly reduces repeated-token cost in long-context scenarios such as multi-turn conversations and complex tool calls. Cache hits are visible in the returned event’susage_metadata, where
cached_content_token_count is the number of tokens served from cache and
prompt_token_count is the total input tokens; the hit rate is their ratio.