Skip to main content
The Harness server is a standalone VeADK agent runtime that you create, configure, and deploy with the veadk harness command line. Once deployed, it runs as an AgentKit Runtime and provides conversation endpoints, session management, and runtime config override capabilities. It suits scenarios where you need to host an agent independently and call it through an HTTP API.

When to use

Prerequisites

  • Python 3.10–3.13 with veadk-python[harness] installed; this extra provides Headroom compression support
  • Volcengine VOLCENGINE_ACCESS_KEY and VOLCENGINE_SECRET_KEY with permission to build images, access TOS, create Runtimes, and configure IAM roles in the target region
  • An enabled model or endpoint accessible to the Runtime IAM role, plus separate connection settings and permissions for knowledge bases, databases, and tools
veadk harness deploy uses the Volcengine deployment flow and has no BytePlus options. For no-code deployment on BytePlus, use Harness in the standalone AgentKit CLI. The command groups and configuration files are not interchangeable The generated Dockerfile builds the service from VeADK main by default. To reproduce a released version, set its VEADK_REF to the required tag before building. Commands on this page were checked against VeADK 1.1.13

Command overview

veadk harness provides the following subcommands: create, add, and show operate on local files. deploy uses cloud resources, and invoke calls the deployed service and may consume model usage

Create a project

If the target directory is non-empty, confirmation allows existing project files to be overwritten. Back up required configuration or use a new directory first
This command generates the following files in the my-harness directory: A non-empty directory triggers overwrite confirmation, so retain required files first. create only generates local files and does not deploy cloud resources Inside the project, copy .env.example to .env and fill in deployment AK/SK values. Exclude .env, harness.json, and configurations containing real secrets from version control. The API key in invocation metadata is also a credential

Configure the agent

Use veadk harness add to write parameters into harness.yaml. Replace your-model-name with an available model or endpoint ID. The knowledge-base example also requires usable VikingDB resources in the selected project and region; omit its three knowledge-base options when not needed:

Configuration parameters

add updates only explicitly supplied fields. Defaults below describe omitted options, not initial server defaults. --path defaults to the current directory, and every subcommand supports --help Set the component backend type before its connection settings. Connection options accept text, such as --knowledgebase-use-ssl true; this differs from the boolean switch in the standalone AgentKit CLI. Passwords and API keys are written to local harness.yaml; do not commit files containing real credentials
Although help lists --builtin-tools, --mcp-router-id, --selected-skills, --mcp, and --registry, this version of add does not persist them. Edit YAML directly for structured resources as shown below. OIDC options belong to deploy, not add
A connection field accepted by the CLI is not necessarily used by every backend. Redis usernames work for knowledge bases but are ignored by long-term memory. OpenSearch secret_token is not used by current knowledge-base or memory connections. See Redis memory and OpenSearch knowledge bases

harness.yaml configuration structure

harness.yaml groups model, tool, skill, knowledge-base, and memory settings. Deployment passes them to the Runtime. Each component selects its backend with type and then supplies that backend’s connection settings. The initial session backend is local, and the server defaults to at most 10 model calls per run
harness.yaml
Agent fields also accept a harness: wrapper, whose values take precedence over matching top-level fields. Keep deployment harness_name and gateway auth at the top level. add edits top-level fields, so avoid duplicate nested settings. This command does not expand ${VAR} in YAML; do not apply the standalone AgentKit CLI interpolation rules

Structured resource configuration

Beyond the basic fields written by veadk harness add, harness.yaml supports the following structured fields for configuring resources dispatched by the AgentKit control plane. These fields are converted to corresponding JSON environment variables on deploy: Structured resource configuration example:
harness.yaml

Execution enhancement configuration

Set harness_enhance in harness.yaml to enable context preparation, tool-result compression, and answer verification. This example uses built-in compression and does not require Headroom
harness.yaml
HTTP callers can override these settings for one request through the top-level harness_enhance field, with components supplied as a comma-separated string. Compression can lose details and answer verification does not guarantee factual accuracy. See Harness extension for component behavior and limits

View configuration

This command prints the parameters configured in harness.yaml, followed by the list of parameters that can be overridden per invocation through veadk harness invoke.
Override knowledge bases, long-term memory, and sampling settings through HTTP. Structured options such as --registry appear in help, but the CLI passes text without parsing JSON objects; use the corresponding HTTP request fields
show prints configuration values without automatically redacting passwords or API keys; remove secrets before sharing output

Deploy

Deployment builds an image and creates or updates cloud resources, which may incur charges. Check the region, Runtime name, model access, and runtime environment settings first. Retain the current configuration before changing a production service. On failure, inspect resources already created before retrying
This command reads harness.yaml, converts it into runtime environment variables, and performs an AgentKit cloud build and Runtime creation. After deployment, the Runtime endpoint, Runtime id, and API key are recorded in harness.json. harness.json is written only when deployment returns an endpoint. API-key mode records the key; OAuth mode records discovery and client settings without storing a user JWT. Default in-memory sessions disappear when the process stops; configure MySQL or PostgreSQL for persistence

Authentication

The default authentication is API key (key_auth). Add an auth section in harness.yaml or use --discovery-url and --allowed-id to enable OAuth2/JWT authentication (custom_jwt):
harness.yaml
With OAuth2/JWT authentication, calls must include an Authorization: Bearer <user-pool JWT> header. The CLI does not mint this token.

Invoke the Harness server

--name specifies the Harness name; its URL and API key are read from harness.json. You can also pass --url and --key directly.

Invoke parameters

Overrides affect only the current request and do not modify harness.yaml. Use distinct --session-id values for independent conversations; the default cli-user and cli-session reuse the same identifiers After receiving a non-empty reply, verify tools, knowledge, and memory as needed. If CLI output is empty, inspect the HTTP response error field. HTTP 200 alone does not prove successful agent execution

HTTP API

The deployed Harness server exposes the following HTTP endpoints.

Conversation endpoints

Session and config endpoints

Runtime config overrides

The Harness server supports overriding the deployed agent’s configuration on each request. Overrides are passed in the harness field of the request body and apply only to that call. /harness/invoke request body structure:
Request body

harness_merge behavior

harness_merge controls how the request config is combined with the default config:

Overridable fields

The following fields can be overridden per request through harness:
Request-level knowledge base and long-term memory overrides are resolved from AgentKit control-plane resource ids into runtime configuration. When an id is provided, the Harness server fetches the corresponding resource’s connection info from the AgentKit control plane.
HTTP examples require the deployed endpoint and a real Bearer credential. Read the API key from local harness.json for API-key mode; for OAuth, use a valid JWT issued by the user pool

Create a session

The request body can include the following fields:

Query default configuration

Returns the current Harness default configuration, including model name, runtime backend, and max LLM calls. Supports camelCase query parameters (appName, userId, sessionId) and also accepts a POST request body.

Environment variables

The Harness server configures runtime behavior through environment variables. harness.yaml is automatically converted to the corresponding environment variables on deploy; you can also set them directly in the runtime environment. CLI invocation also accepts HARNESS_URL, HARNESS_KEY, and HARNESS_TIMEOUT, with a 600-second default timeout. The first two supply fallback values for --url and --key; they are not automatically written to deployment configuration

Model and tools

Resource configuration

Session and memory backends

The default value of max_llm_calls is 10. When not explicitly set, a single run makes at most 10 LLM calls. Adjust it in harness.yaml or through a request-level override.
Last modified on September 19, 2026