Skip to main content
The Harness server is a standalone VeADK agent runtime that you create, configure, and deploy with the veadk harness command line. Once deployed, it runs as an AgentKit Runtime and provides conversation endpoints, session management, and runtime config override capabilities. It suits scenarios where you need to host an agent independently and call it through an HTTP API.

When to use

Prerequisites

  • veadk-python[harness] is installed;
  • Volcengine credentials are configured (VOLCENGINE_ACCESS_KEY and VOLCENGINE_SECRET_KEY);
  • Do not commit .env, API keys, or access credentials to the repository.

Command overview

veadk harness provides the following subcommands:

Create a project

This command generates the following files in the my-harness directory:

Configure the agent

Use veadk harness add to write parameters into harness.yaml:

Configuration parameters

veadk harness add accepts the following parameters (unset fields keep their existing value): Each backend’s connection parameters also has its own flag, for example --knowledgebase-project, --knowledgebase-region, --long-term-memory-host, --short-term-memory-host, written under the matching component section.

harness.yaml configuration structure

harness.yaml organizes configuration in sections. On deploy, top-level fields and the model section are flattened into runtime environment variables (e.g. model.name maps to MODEL_AGENT_NAME); each component’s type selects a backend, and its other parameters map to the environment variables that backend reads.
harness.yaml
You can also use a harness: wrapper section in harness.yaml to nest all Harness parameters in a single object. Fields inside the harness: section are extracted and merged with top-level fields.

Structured resource configuration

Beyond the basic fields written by veadk harness add, harness.yaml supports the following structured fields for configuring resources dispatched by the AgentKit control plane. These fields are converted to corresponding JSON environment variables on deploy: Structured resource configuration example:
harness.yaml

View configuration

This command prints the parameters configured in harness.yaml, followed by the list of parameters that can be overridden per invocation through veadk harness invoke.
Knowledge base, long-term memory, sampling parameters (temperature, top_p, etc.), and registry fields are overridden only through the HTTP API and are not exposed as CLI flags.

Deploy

This command reads harness.yaml, converts it into runtime environment variables, and performs an AgentKit cloud build and Runtime creation. After deployment, the Runtime endpoint, Runtime id, and API key are recorded in harness.json.
Deployment creates cloud runtime resources. Confirm the project, region, and Runtime name before executing. Destroying a Runtime with veadk agentkit destroy deletes cloud runtime resources; preserve any needed logs and data beforehand.

Authentication

The default authentication is API key (key_auth). Add an auth section in harness.yaml or use --discovery-url and --allowed-id to enable OAuth2/JWT authentication (custom_jwt):
harness.yaml
With OAuth2/JWT authentication, calls must include an Authorization: Bearer <user-pool JWT> header. The CLI does not mint this token.

Invoke the Harness server

--name specifies the Harness name; its URL and API key are read from harness.json. You can also pass --url and --key directly.

Invoke parameters

At invoke time you can use override flags to override the deployed agent’s configuration for this single call, such as --tools, --skills, --system-prompt, --model-name. The override applies only to this call and does not modify harness.yaml.

HTTP API

The deployed Harness server exposes the following HTTP endpoints.

Conversation endpoints

Session and config endpoints

Runtime config overrides

The Harness server supports overriding the deployed agent’s configuration on each request. Overrides are passed in the harness field of the request body and apply only to that call. /harness/invoke request body structure:
Request body

harness_merge behavior

harness_merge controls how the request config is combined with the default config:

Overridable fields

The following fields can be overridden per request through harness:
Request-level knowledge base and long-term memory overrides are resolved from AgentKit control-plane resource ids into runtime configuration. When an id is provided, the Harness server fetches the corresponding resource’s connection info from the AgentKit control plane.

Create a session

The request body can include the following fields:

Query default configuration

Returns the current Harness default configuration, including model name, runtime backend, and max LLM calls. Supports camelCase query parameters (appName, userId, sessionId) and also accepts a POST request body.

Environment variables

The Harness server configures runtime behavior through environment variables. harness.yaml is automatically converted to the corresponding environment variables on deploy; you can also set them directly in the runtime environment.

Model and tools

Resource configuration

Session and memory backends

The default value of max_llm_calls is 10. When not explicitly set, a single run makes at most 10 LLM calls. Adjust it in harness.yaml or through a request-level override.
Last modified on September 19, 2026