content_safety sends text to a moderation service around model and tool calls. A blocking decision can replace content or stop a tool call. Use it to connect an agent to existing content policies
Overview
Tool identifiercontent_safety.
content_safety is a content-safety guardrail tool that VeADK provides through the agent plugin mechanism: it hooks into the agent’s execution callbacks and uses the Volcengine LLM Application Firewall to review content at each stage and block unsafe input and output.
Audit points that are active today:
- Before Model Callback — review user input before it reaches the model
- After Model Callback — review the model’s output
- Before Tool Callback — review tool arguments before a tool call
- After Tool Callback — review the tool’s return value
content_safety relies on the content-safety policies of the Volcengine LLM Application Firewall to detect and block different types of risky content. The top-level categories are:
Notes:
- The General Topic Control policy is not configured by default when you add firewall assets; after adding them, configure the topic-control policy yourself.
- The Computational Resource Consumption policy does not trigger on a single request: it blocks requests only once the system detects a behavior pattern with similar attack vectors accompanied by high compute output over a time window.
Environment & prerequisites
Before usingcontent_safety, purchase an instance, add assets, and obtain its AppID, then configure the access information through environment variables or config.yaml. content_safety reads the following settings at initialization. Configure AppID before importing it; initialization fails when AppID is missing.
Configure in
config.yaml:
config.yaml
config.yaml at startup. Alternatively, set the environment variable before starting Python; it takes precedence:
Values set through environment variables take precedence over the matching keys in
config.yaml.Endpoints & authentication
content_safety supports two content-moderation endpoints, selected automatically based on TOOL_LLM_SHIELD_URL.
Default firewall endpoint
Used when theTOOL_LLM_SHIELD_URL path does not contain /OpenTOP/V1/Lumen/Moderate (the default behavior when url is not set). Requests are sent to <url>/v2/moderate, with authentication chosen as follows:
- When
TOOL_LLM_SHIELD_API_KEYis configured, API key authentication is used (thex-api-keyrequest header is sent). - When no API key is configured, Volcengine AK/SK authentication is used: the
VOLCENGINE_ACCESS_KEYandVOLCENGINE_SECRET_KEYenvironment variables are read first; if both are missing, temporary credentials are obtained from the IAM role of the runtime environment. The request headers carry the service region and service identifier.
CLOUD_PROVIDER=byteplus, global configuration can map BytePlus AK/SK to the credential settings above. Confirm that the moderation service accepts those account credentials. Using a BytePlus model does not configure the moderation endpoint, region, or permissions
Lumen Moderate endpoint
Used when theTOOL_LLM_SHIELD_URL path contains /OpenTOP/V1/Lumen/Moderate. Requests are sent directly to the configured full address, authenticating with the AppID as the endpoint identifier and the API key as the endpoint secret.
Compared with the default endpoint, the Lumen Moderate endpoint attaches the current session identifier and invocation identifier during tool-call reviews, so that multiple reviews within the same session can be correlated on the audit side. Model-input review, model-output review, tool-argument review, and tool-return review each issue a request at the corresponding callback point, with blocking behavior matching the default endpoint.
Usage
Hook thecontent_safety callbacks onto the agent to audit its execution:
config.yaml; the calling code stays the same:
config.yaml
Checking moderation
Configure settings before importingcontent_safety, and provide the key through TOOL_LLM_SHIELD_API_KEY. Save the example and run python app.py. Ordinary input should receive a response; verify blocking with test text covered by your configured policy and check the service audit records. A fixed sample sentence does not guarantee a particular policy match
Model-input moderation checks the first text part of the last user content in the current request; model-output moderation checks its first text part. It does not automatically cover all history, images, audio, or multipart content. Tool callbacks moderate serialized arguments and results. The default request timeout is 50 seconds
Adjusting the request timeout
In the example above, replace thecontent_safety import with this configuration fragment and retain the same four callbacks: