Skip to main content
content_safety sends text to a moderation service around model and tool calls. A blocking decision can replace content or stop a tool call. Use it to connect an agent to existing content policies
This capability sends user text, model responses, and tool data to the configured moderation endpoint. In VeADK 1.1.13, non-200 HTTP responses, timeouts, and some request errors are logged and execution continues. These failures do not mean moderation approved the content. Workflows that must block on moderation failure need additional enforcement and error handling

Overview

Tool identifier content_safety. content_safety is a content-safety guardrail tool that VeADK provides through the agent plugin mechanism: it hooks into the agent’s execution callbacks and uses the Volcengine LLM Application Firewall to review content at each stage and block unsafe input and output. Audit points that are active today:
  • Before Model Callback — review user input before it reaches the model
  • After Model Callback — review the model’s output
  • Before Tool Callback — review tool arguments before a tool call
  • After Tool Callback — review the tool’s return value
content_safety relies on the content-safety policies of the Volcengine LLM Application Firewall to detect and block different types of risky content. The top-level categories are: Notes:
  • The General Topic Control policy is not configured by default when you add firewall assets; after adding them, configure the topic-control policy yourself.
  • The Computational Resource Consumption policy does not trigger on a single request: it blocks requests only once the system detects a behavior pattern with similar attack vectors accompanied by high compute output over a time window.

Environment & prerequisites

Before using content_safety, purchase an instance, add assets, and obtain its AppID, then configure the access information through environment variables or config.yaml. content_safety reads the following settings at initialization. Configure AppID before importing it; initialization fails when AppID is missing. Configure in config.yaml:
config.yaml
VeADK loads the region from config.yaml at startup. Alternatively, set the environment variable before starting Python; it takes precedence:
Values set through environment variables take precedence over the matching keys in config.yaml.

Endpoints & authentication

content_safety supports two content-moderation endpoints, selected automatically based on TOOL_LLM_SHIELD_URL.

Default firewall endpoint

Used when the TOOL_LLM_SHIELD_URL path does not contain /OpenTOP/V1/Lumen/Moderate (the default behavior when url is not set). Requests are sent to <url>/v2/moderate, with authentication chosen as follows:
  • When TOOL_LLM_SHIELD_API_KEY is configured, API key authentication is used (the x-api-key request header is sent).
  • When no API key is configured, Volcengine AK/SK authentication is used: the VOLCENGINE_ACCESS_KEY and VOLCENGINE_SECRET_KEY environment variables are read first; if both are missing, temporary credentials are obtained from the IAM role of the runtime environment. The request headers carry the service region and service identifier.
With CLOUD_PROVIDER=byteplus, global configuration can map BytePlus AK/SK to the credential settings above. Confirm that the moderation service accepts those account credentials. Using a BytePlus model does not configure the moderation endpoint, region, or permissions

Lumen Moderate endpoint

Used when the TOOL_LLM_SHIELD_URL path contains /OpenTOP/V1/Lumen/Moderate. Requests are sent directly to the configured full address, authenticating with the AppID as the endpoint identifier and the API key as the endpoint secret.
You must configure TOOL_LLM_SHIELD_API_KEY before using the Lumen Moderate endpoint; otherwise the request fails because no valid key is present.
Compared with the default endpoint, the Lumen Moderate endpoint attaches the current session identifier and invocation identifier during tool-call reviews, so that multiple reviews within the same session can be correlated on the audit side. Model-input review, model-output review, tool-argument review, and tool-return review each issue a request at the corresponding callback point, with blocking behavior matching the default endpoint.

Usage

Hook the content_safety callbacks onto the agent to audit its execution:
To switch to the Lumen Moderate endpoint, specify the corresponding access address and API key in environment variables or config.yaml; the calling code stays the same:
config.yaml

Checking moderation

Configure settings before importing content_safety, and provide the key through TOOL_LLM_SHIELD_API_KEY. Save the example and run python app.py. Ordinary input should receive a response; verify blocking with test text covered by your configured policy and check the service audit records. A fixed sample sentence does not guarantee a particular policy match Model-input moderation checks the first text part of the last user content in the current request; model-output moderation checks its first text part. It does not automatically cover all history, images, audio, or multipart content. Tool callbacks moderate serialized arguments and results. The default request timeout is 50 seconds

Adjusting the request timeout

In the example above, replace the content_safety import with this configuration fragment and retain the same four callbacks:
Last modified on September 19, 2026