Overview
Tool identifiercontent_safety.
content_safety is a content-safety guardrail tool that VeADK provides through the agent plugin mechanism: it hooks into the agent’s execution callbacks and uses the Volcengine LLM Application Firewall to review content at each stage and block unsafe input and output.
Audit points that are active today:
- Before Model Callback — review user input before it reaches the model
- After Model Callback — review the model’s output
- Before Tool Callback — review tool arguments before a tool call
- After Tool Callback — review the tool’s return value
content_safety relies on the content-safety policies of the Volcengine LLM Application Firewall to detect and block different types of risky content. The top-level categories are:
Notes:
- The General Topic Control policy is not configured by default when you add firewall assets; after adding them, configure the topic-control policy yourself.
- The Computational Resource Consumption policy does not trigger on a single request: it blocks requests only once the system detects a behavior pattern with similar attack vectors accompanied by high compute output over a time window.
Environment & prerequisites
Before usingcontent_safety, purchase an instance, add assets, and obtain its AppID. Set the environment variable TOOL_LLM_SHIELD_APP_ID, or configure it in config.yaml:
config.yaml
Usage
Hook thecontent_safety callbacks onto the agent to audit its execution: