> ## Documentation Index
> Fetch the complete documentation index at: https://docs.veadk.xyz/llms.txt
> Use this file to discover all available pages before exploring further.

# Web fetch

## Overview

Tool identifier `web_fetch`.

`web_fetch` does a plain HTTP GET on a given URL and extracts its readable content: HTML is converted to markdown or plain text, and PDFs are extracted to text via `pypdf`. It does **not** execute JavaScript, so pages that render entirely client-side or require login may come back incomplete. Unlike `link_reader`, this tool needs **no credentials of its own** (it is a plain HTTP fetch) — use it to let the agent read articles, docs, or any public URL the user references.

Import path: `from veadk.tools.builtin_tools.web_fetch import web_fetch`

## Environment & prerequisites

Environment variables:

* `MODEL_AGENT_API_KEY`: API key for the agent's reasoning model (the `web_fetch` tool itself needs no extra credentials)

## Usage

Parameters:

* `url`: the `http(s)` URL to fetch;
* `extract_mode`: `markdown` (default, keeps headings / links / lists) or `text` (plain text);
* `max_chars`: maximum characters of extracted content (default `50000`).

Returns `{"url", "title", "content", "truncated"}`, or `{"error": ...}` on failure.

```python title="examples/tools/web_fetch/agent.py" lines theme={null}
import asyncio

from veadk import Agent, Runner
from veadk.memory.short_term_memory import ShortTermMemory
from veadk.tools.builtin_tools.web_fetch import web_fetch

agent = Agent(
    name="web_fetch_agent",
    model_name="doubao-seed-2-1-pro-260628",
    description="An agent that reads web pages and PDFs.",
    instruction="Use the web_fetch tool to fetch the given URL, then answer based on its content.",
    tools=[web_fetch],
)

runner = Runner(agent=agent, short_term_memory=ShortTermMemory())


async def main():
    response = await runner.run(
        "Fetch https://arxiv.org/pdf/1706.03762 and summarize the paper's core idea"
    )
    print(response)


if __name__ == "__main__":
    asyncio.run(main())
```

## Notes

<Note>
  Security & limits:

  * **SSRF protection**: after DNS resolution it blocks private / loopback / link-local / reserved addresses, and re-validates every redirect hop (including `<meta refresh>`), following at most 3 hops.
  * **Limits**: 2 MB download cap for HTML (10 MB for PDFs); 30 s request timeout; results cached in-process for 15 minutes.
  * No JavaScript rendering; no socket-level DNS pinning (resolve-then-revalidate only).
</Note>
