config.yaml (see the Quickstart).
You can also set a model per agent when you create it.
In a Volcengine environment, version 1.0.0 defaults to doubao-seed-1-8-251228 with provider openai and endpoint https://ark.cn-beijing.volces.com/api/v3/. When CLOUD_PROVIDER=byteplus, it defaults to seed-1-6-250915 at https://ark.ap-southeast.bytepluses.com/api/v3.
Set the model for a single agent
Override the global default withmodel_name and model_provider:
model_provider, model_api_base, and model_api_key fall back
to the global configuration.
Configure fallback models
model_name also accepts a list: the first entry is the primary model and the
rest are fallbacks, tried in order when the primary model is unavailable.
In VeADK 1.0.0, model fallback applies only to the default model API. When the Responses API is enabled, pass a single string to
model_name.Responses API
The Responses API is a Volcengine Ark interface with native, efficient context management, a simpler I/O format, and stronger tool-calling and multimodal capabilities. Once enabled in VeADK, every turn of the agent’s conversation goes through this interface, giving it native context caching and image, video, and document understanding.Enable
Setenable_responses=True when creating the agent:
google-adk>=1.21.0, and the model must
support the interface (doubao models after version 0615 support it by default).
Multimodal input
Beyond text, the Responses API understands images, video, and documents. Pass multimodal data withgoogle.genai.types.FileData; file_uri accepts three
sources:
- Local file path:
file://{local_path}— uploaded automatically via the Files API. - Files API resource:
file_id://{file_id}— for already-uploaded files. - Web URL: a plain
https://link, typed by itsmime_type.
FileData may include video_metadata with fps to control the
frame-sampling rate (default 1, adjustable between 0.2 and 5).
Context caching
In Responses API mode, session caching is on by default: the initial context is stored and updated each turn, and later requests merge the cached content with the new input before calling the model. This significantly reduces repeated-token cost in long-context scenarios such as multi-turn conversations and complex tool calls. Cache hits are visible in the returned event’susage_metadata, where
cached_content_token_count is the number of tokens served from cache and
prompt_token_count is the total input tokens; the hit rate is their ratio.