Skip to main content

Overview

Tool identifier web_fetch. web_fetch does a plain HTTP GET on a given URL and extracts its readable content: HTML is converted to markdown or plain text, and PDFs are extracted to text via pypdf. It does not execute JavaScript, so pages that render entirely client-side or require login may come back incomplete. Unlike link_reader, this tool needs no credentials of its own (it is a plain HTTP fetch) — use it to let the agent read articles, docs, or any public URL the user references. Import path: from veadk.tools.builtin_tools.web_fetch import web_fetch

Environment & prerequisites

Environment variables:
  • MODEL_AGENT_API_KEY: API key for the agent’s reasoning model (the web_fetch tool itself needs no extra credentials)

Usage

Parameters:
  • url: the http(s) URL to fetch;
  • extract_mode: markdown (default, keeps headings / links / lists) or text (plain text);
  • max_chars: maximum characters of extracted content (default 50000).
Returns {"url", "title", "content", "truncated"}, or {"error": ...} on failure.
examples/tools/web_fetch/agent.py

Notes

Security & limits:
  • SSRF protection: after DNS resolution it blocks private / loopback / link-local / reserved addresses, and re-validates every redirect hop (including <meta refresh>), following at most 3 hops.
  • Limits: 2 MB download cap for HTML (10 MB for PDFs); 30 s request timeout; results cached in-process for 15 minutes.
  • No JavaScript rendering; hostname checks do not replace network access restrictions in the deployment environment.

Direct calls and parameter limits

Direct calls need no model API key. PDF extraction requires pypdf and does not perform OCR on scanned pages. truncated only indicates that extracted text exceeded the character limit; it does not guarantee a complete download or full content coverage.
Last modified on September 19, 2026