> ## Documentation Index
> Fetch the complete documentation index at: https://docs.veadk.xyz/llms.txt
> Use this file to discover all available pages before exploring further.

# 网页抓取

## 功能说明

对应工具标识 `web_fetch`。

`web_fetch` 对给定 URL 发起一次普通 HTTP GET 并抽取正文：HTML 转 Markdown 或纯文本，PDF 用 `pypdf` 抽取文字。它**不执行 JavaScript**——纯前端渲染或需要登录的页面可能抽取不全。与 `link_reader` 不同，该工具**自身无需任何凭证**（纯 HTTP 抓取），适合让智能体阅读用户给出的文章、文档或任意公开链接。

导入路径：`from veadk.tools.builtin_tools.web_fetch import web_fetch`

## 环境变量与前提

环境变量：

* `MODEL_AGENT_API_KEY`：智能体推理模型的 API Key（`web_fetch` 工具本身无需额外凭证）

## 使用方法

参数：

* `url`：要抓取的 `http(s)` 链接；
* `extract_mode`：`markdown`（默认，保留标题 / 链接 / 列表）或 `text`（纯文本）；
* `max_chars`：抽取内容的最大字符数（默认 `50000`）。

返回 `{"url", "title", "content", "truncated"}`，失败时返回 `{"error": ...}`。

```python title="examples/tools/web_fetch/agent.py" lines theme={null}
import asyncio

from veadk import Agent, Runner
from veadk.memory.short_term_memory import ShortTermMemory
from veadk.tools.builtin_tools.web_fetch import web_fetch

agent = Agent(
    name="web_fetch_agent",
    model_name="doubao-seed-2-1-pro-260628",
    description="An agent that reads web pages and PDFs.",
    instruction="Use the web_fetch tool to fetch the given URL, then answer based on its content.",
    tools=[web_fetch],
)

runner = Runner(agent=agent, short_term_memory=ShortTermMemory())


async def main():
    response = await runner.run(
        "抓取 https://arxiv.org/pdf/1706.03762 并总结这篇论文的核心思想"
    )
    print(response)


if __name__ == "__main__":
    asyncio.run(main())
```

## 额外说明

<Note>
  安全与限制：

  * **SSRF 防护**：解析域名后拦截私网 / 环回 / 链路本地 / 保留地址，并对每一跳重定向（含 `<meta refresh>`）重新校验，最多跟随 3 跳。
  * **上限**：HTML 下载上限 2MB、PDF 10MB；请求超时 30 秒；结果在进程内缓存 15 分钟。
  * 不渲染 JavaScript；域名检查不能替代部署环境的网络访问限制
</Note>

## 独立调用与参数边界

| 参数 | 类型 | 默认值 | 说明 |
| :- | :- | :- | :- |
| `url` | `str` | 必填 | 可公开访问的 HTTP 或 HTTPS URL |
| `extract_mode` | `str` | `"markdown"` | `markdown` 或 `text`，其他值按 `markdown` 处理 |
| `max_chars` | `int` | `50000` | 结果字符上限，限制在 1–200000；无法转换的值使用默认值 |
| `tool_context` | `ToolContext \| None` | `None` | 智能体调用时自动注入，独立调用可省略 |

直接调用不需要模型 API Key；读取 PDF 需要安装 `pypdf`。扫描版 PDF 不会自动进行 OCR。`truncated` 只表示抽取后的正文超过字符上限，不保证下载的页面完整或覆盖所有内容

```python lines theme={null}
from veadk.tools.builtin_tools.web_fetch import web_fetch

result = web_fetch("https://example.com", extract_mode="text", max_chars=2000)
if "error" in result:
    print("Fetch failed:", result["error"])
else:
    print(result["url"], result["title"])
    print(result["content"])
    print("Truncated:", result["truncated"])
```
