Skip to main content
The model interprets input, selects tools, and generates responses. Agents read global model settings from environment variables or config.yaml by default. Use Agent parameters to give individual agents their own models, endpoints, and credentials

Set the model for a single agent

Complete installation, then select one platform configuration below. Enable the model in the corresponding account, or replace its name with an accessible model or inference endpoint
Run this script from the same terminal:
main.py
Run python main.py to print a response. If all agents share the settings above, Agent(name="assistant") is sufficient

Model parameters and global settings

Explicit constructor arguments override their corresponding global settings. Unspecified fields retain global values. VeADK searches for config.yaml from the current directory upward, and existing environment variables take precedence over matching file settings. See quickstart for YAML configuration Without global overrides, Volcengine uses doubao-seed-2-1-pro-260628 at https://ark.cn-beijing.volces.com/api/v3/. With CLOUD_PROVIDER=byteplus, the defaults are seed-2-0-lite-260228 and https://ark.ap-southeast.bytepluses.com/api/v3. Check existing MODEL_AGENT_* variables when switching platforms so settings from the previous platform are not retained

Resolve an API key by name

Key precedence is: nonempty model_api_key, MODEL_AGENT_API_KEY, a lookup using model_api_key_name or MODEL_AGENT_API_KEY_NAME, then the account’s default key lookup
Name-based lookup requires account credentials permitted to read Ark API keys. The name is not the secret value. An existing key value takes precedence over lookup. This feature is available from VeADK 1.0.2

Rate-limit retries

With the default adk runtime, Responses disabled, and no custom model, VeADK retries an HTTP 429 once if no model response has been emitted. A valid numeric Retry-After controls the delay, capped at 2 seconds; missing or invalid values use 0.5 seconds. This retry layer does not replay a request after output has started

Configure fallback models

Set BACKUP_MODEL_NAME in addition to the earlier settings, then replace the agent definition with:
The first entry is primary; subsequent entries are fallback candidates sharing the provider, endpoint, and credentials. You can also append same-provider candidates with model_fallbacks=["backup-model-name"]. When both are configured, candidates from model_name precede those from model_fallbacks Same-provider name fallbacks are available for the default model API and Responses API. Every candidate must support the request’s tools, input modalities, and output format. Fallbacks handle request failures; low-quality answers do not automatically trigger a switch

Configure cross-provider fallback models

Use ModelFallbackEndpoint when a backup has a different provider, endpoint, or credentials. In addition to the primary settings, set BACKUP_MODEL_NAME, BACKUP_MODEL_PROVIDER, BACKUP_MODEL_API_BASE, and BACKUP_MODEL_API_KEY to the backup service’s actual values
fallback.py

ModelFallbackEndpoint parameters

Dictionaries can replace endpoint objects. Supported aliases are model, provider, api_base or base_url, api_key, api_key_env, and extra_config. Explicitly specify the provider, endpoint, and credentials across providers to avoid inheriting unsuitable settings

Fallback limitations

  • enable_responses=True accepts only same-provider string fallbacks; endpoint objects or dictionaries fail at initialization
  • codex and piagent ignore fallback chains; use adk when fallbacks are required. See runtime
  • With a custom model object in the default runtime, model_fallbacks is ignored; configure fallbacks on that object instead

Responses API

The Responses API supports conversation continuation, multimodal input, and structured output. Confirm that the model and endpoint support the Ark Responses protocol. An OpenAI-compatible endpoint does not automatically support every feature described here

Enable

Requires google-adk>=1.34.0. After configuring the model, set Agent(enable_responses=True); the feature is disabled by default. For BytePlus, select an available model and endpoint that explicitly support the required capabilities; the flag alone does not add support

Multimodal input

FileData.file_uri accepts the following sources. Set mime_type to match the content:
Local files are sent to the model service. Confirm that they are suitable for upload and check supported formats and size limits before running the example
Save a PNG as example.png in the current directory, then run python image_reader.py:
image_reader.py
Set video metadata on types.Part, alongside file_data. This fragment can replace the image Part; the service determines the supported frame-rate range:

Configure Ark context management

For models supporting context_management, pass settings through model_extra_config. This fragment requests removal of older thinking content while retaining the most recent thinking turn:

Context caching

Responses caching is enabled by default. Cache hits depend on the service and request content; reduced usage is not guaranteed. Set enable_responses_cache=False to disable it In event usage_metadata, cached_content_token_count counts cached input tokens and prompt_token_count counts total input tokens. Calculate their ratio only when the latter is greater than zero. Read events with Runner.run_async to inspect usage When output_schema is configured, VeADK removes cache settings that conflict with structured output. See structured output
Last modified on September 19, 2026