One web tool for Pi: fetch, search, and site-map discovery.
Install / load
Loaded through the root pi-ext package.
See ../../README.md.
Commands / tools / settings
Commands:
none.
Tool: web, with required action: fetch | map | search and flat action-specific arguments.
Invalid or irrelevant arguments are rejected before I/O.
The former separate tools and MiniMax image analysis are removed, without aliases.
Update tool allowlists to web.
Actions:
fetch
url:
URL to fetch.
format:
optional; markdown default, or text, html, json, raw.
maxChars:
optional; default 20000, max 100000, min 100.
timeout:
optional; default 120000 ms, min 5000, max 600000.
Passed to fetchUrlStream alongside the Pi/tool abort signal.
linesMatching:
optional; returns matching line excerpts instead of the full visible output.
contextLines:
optional; line context around linesMatching hits, max 100.
map
url:
site URL; discovers site maps and related links without fetching page bodies.
maxUrls:
optional; default 1000, min 1, max 10000.
maxSitemaps:
optional; default 20, min 1, max 100.
maxChars:
optional; default 8000, max 100000, min 100.
search
query:
search query.
count:
optional; default 5, max 20.
format:
optional; text default, or json.
timeout:
optional; default 120000 ms, min 5000, max 600000.
Hooks:
session_shutdown:
removes temp dirs created for fetched binary/image/media/PDF/office content.
Settings:
none registered.
External binaries used when needed:
pandoc:
HTML conversion; DOCX/ODT/EPUB/RTF extraction.
Missing binary emits a warning notification on session start.
pdftotext:
PDF extraction.
Missing binary emits a warning notification on session start.
identify:
image metadata.
Missing binary emits an info notification on session start.
Missing/failing binary returns fallback metadata text.
ffprobe:
audio/video metadata.
Missing binary emits an info notification on session start.
Missing/failing binary returns fallback metadata text.
Provider credentials, env only:
Z.AI search key:
Pi model registry providers zai, z-ai, z.ai, bigmodel; fallback env ZAI_API_KEY, Z_AI_API_KEY.
MiniMax key:
Pi model registry provider minimax; fallback env MINIMAX_API_KEY.
OpenAI Codex search auth:
Pi model registry provider openai-codex OAuth.
Keys never go into settings.json.
Settings in Pi settings.json under web, with env overrides:
There are no CLI flags.
Precedence per key: env, trusted project .pi/settings.json, user ~/.pi/agent/settings.json, default.
When project and user values are both objects they deep-merge with project winning; scalars replace.
Project settings are read only when Pi trusts the project.
web.zai.freshness accepts day, week, month, year, or none, nolimit, no_limit, empty for no limit.
All other keys take non-empty strings; surrounding whitespace is trimmed, and an empty env var counts as set.
web.zai.mcpUrl is normalized to the /api/mcp/web_search_prime/mcp endpoint when it lacks a /mcp path.
MiniMax prefers the baseUrl of a registered minimax or minimax-cn model over web.minimax.apiHost.
An invalid value never falls back to a lower source.
Session start shows a warning notification per invalid key, and the affected provider fails its search with the same message.
Run mise run bench here, or mise //extensions/web:bench from the repository root, to benchmark web processing hot paths.
Live search comparison
Run mise run //extensions/web:evaluate-search to compare all configured providers on five fixed English/German technical and consumer-information queries.
It uses existing Pi credentials through ModelRegistry, makes no model-catalog refresh, applies a 45-second deadline per provider, rotates provider order, and records target rank, elapsed time, and full results in the Pi state directory.
Providers with authentication, quota, or CAPTCHA failures are not retried during that run.
These small-sample target-rank checks are not a general search-quality benchmark.
Behavior
web with action: fetch uses native fetch with browser-like headers and a per-host
referer.
Fetch accepts only HTTP and HTTPS URLs.
Private/reserved IP targets, blocked DNS answers, and unsafe redirects are
rejected.
Fetch output is stored in the Pi state directory with a responseId and
fullOutputPath.
Capped output, matching excerpts, and no-match results include the path in
tool-visible text so the complete saved output can be read directly.
linesMatching returns numbered excerpts with optional surrounding context.
Content routing is based on Content-Type:
HTML/XHTML:
strips common prompt-injection delivery vectors, extracts main content,
converts with pandoc and llm-extract.lua.
Text/XML/YAML/CSV/TSV:
returns UTF-8 text.
JSON:
pretty-prints when parseable, otherwise returns raw text.
PDF:
writes temp file, runs pdftotext -layout.
DOCX/ODT/EPUB/RTF:
writes temp file, runs pandoc to plain text.
Images:
writes temp file, returns metadata from identify, plus inline image
content.
Audio/video:
writes temp file, returns summary and ffprobe JSON.
Other binary:
writes temp file and returns path plus content type.
Fetch raw format skips conversion, decodes the body as UTF-8, and
truncates to maxChars.
HTML cleanup removes comments, scripts, styles, iframes, objects, templates,
noscript blocks, embeds, hidden inputs, long meta content,
hidden/visually-hidden elements, and non-image aria-label/alt/title
attrs.
Main-content extraction prefers <article>, #mw-content-text,
#docContent, #content, <main>, then role="main"; fallback is full
page.
Converted fetch text, raw fetch text, and search output are wrapped as
untrusted external data.
Image, media, and binary early-return paths are not wrapped by the current
implementation.
web with action: map discovers site maps from robots.txt, sitemap XML, and llms links.
Its URL listing shows at most 50 entries and obeys maxChars.
If either limit omits URLs, the result identifies the listing as partial and
includes the saved map's fullOutputPath.
web with action: search provider order:
Google WML, anonymous Parallel, OpenAI Codex native search when authenticated, MiniMax search when configured, DuckDuckGo HTML, then Z.AI Coding Plan MCP when configured.
Cancellation stops fallback.
Parallel uses the anonymous hosted MCP at https://search.parallel.ai/mcp without API keys, model metadata, or conversation identifiers.
All excerpts from each selected result are preserved as its snippet, without character caps.
The existing result-count limit still applies.
Google uses /wml/search with a fixed Nokia user agent and native HTTP, without JavaScript or browser automation.
Requests reserve slots at least two seconds apart.
CAPTCHA, consent pages, and unrecognized layouts fail into the next provider; no CAPTCHA solving or proxy rotation.
This is an undocumented endpoint and can stop working.
Parallel and Z.AI share the MCP transport in mcp.ts: JSON/SSE decoding, response-ID matching, initialization, session headers, protocol negotiation, and redirect rejection.
Authentication, tool arguments, result decoding, and provider-specific errors remain separate.
Z.AI calls MCP tool web_search_prime with web.zai.freshness, web.zai.contentSize, and web.zai.location.
DuckDuckGo submits the HTML no-JavaScript form using POST, query/region fields, and origin/referer headers, instead of searching with GET.
It throttles requests to at least two seconds apart, includes the throttle wait in its deadline, and reports CAPTCHA as a provider failure.
A local two-query comparison returned ten results per query for HTML/Lite POST, while both GET interfaces returned CAPTCHA with no results.
A follow-up five-query evaluation through the production HTML provider found the intended target first for every query.
This fixes the observed request pattern, not all possible DDG blocks; no CAPTCHA solver, proxy rotation, or browser dependency is added.
See DDG's non-JavaScript interfaces and SearXNG's request and bot-blocker notes.
MiniMax search posts to /v1/coding_plan/search.
# web
One web tool for Pi: fetch, search, and site-map discovery.
## Install / load
Loaded through the root `pi-ext` package.
See [../../README.md](../../README.md).
## Commands / tools / settings
Commands:
none.
Tool: `web`, with required `action: fetch | map | search` and flat action-specific arguments.
Invalid or irrelevant arguments are rejected before I/O.
The former separate tools and MiniMax image analysis are removed, without aliases.
Update tool allowlists to `web`.
Actions:
- `fetch`
- `url`:
URL to fetch.
- `format`:
optional; `markdown` default, or `text`, `html`, `json`, `raw`.
- `maxChars`:
optional; default `20000`, max `100000`, min `100`.
- `timeout`:
optional; default `120000` ms, min `5000`, max `600000`.
Passed to `fetchUrlStream` alongside the Pi/tool abort signal.
- `linesMatching`:
optional; returns matching line excerpts instead of the full visible output.
- `contextLines`:
optional; line context around `linesMatching` hits, max `100`.
- `map`
- `url`:
site URL; discovers site maps and related links without fetching page bodies.
- `maxUrls`:
optional; default `1000`, min `1`, max `10000`.
- `maxSitemaps`:
optional; default `20`, min `1`, max `100`.
- `maxChars`:
optional; default `8000`, max `100000`, min `100`.
- `search`
- `query`:
search query.
- `count`:
optional; default `5`, max `20`.
- `format`:
optional; `text` default, or `json`.
- `timeout`:
optional; default `120000` ms, min `5000`, max `600000`.
Hooks:
- `session_shutdown`:
removes temp dirs created for fetched binary/image/media/PDF/office content.
Settings:
none registered.
External binaries used when needed:
- `pandoc`:
HTML conversion; DOCX/ODT/EPUB/RTF extraction.
Missing binary emits a warning notification on session start.
- `pdftotext`:
PDF extraction.
Missing binary emits a warning notification on session start.
- `identify`:
image metadata.
Missing binary emits an info notification on session start.
Missing/failing binary returns fallback metadata text.
- `ffprobe`:
audio/video metadata.
Missing binary emits an info notification on session start.
Missing/failing binary returns fallback metadata text.
Provider credentials, env only:
- Z.AI search key:
Pi model registry providers `zai`, `z-ai`, `z.ai`, `bigmodel`; fallback env `ZAI_API_KEY`, `Z_AI_API_KEY`.
- MiniMax key:
Pi model registry provider `minimax`; fallback env `MINIMAX_API_KEY`.
- OpenAI Codex search auth:
Pi model registry provider `openai-codex` OAuth.
Keys never go into `settings.json`.
Settings in Pi `settings.json` under `web`, with env overrides:
| Key | Env | Default |
| --- | --- | --- |
| `web.zai.freshness` | `ZAI_SEARCH_FRESHNESS` | none |
| `web.zai.mcpUrl` | `ZAI_SEARCH_MCP_URL`, then `ZAI_SEARCH_BASE_URL` | `https://api.z.ai/api/mcp/web_search_prime/mcp` |
| `web.zai.contentSize` | `ZAI_SEARCH_CONTENT_SIZE` | `medium` |
| `web.zai.location` | `ZAI_SEARCH_LOCATION` | `us` |
| `web.minimax.apiHost` | `MINIMAX_API_HOST` | `https://api.minimax.io` |
Example:
```json
{ "web": { "zai": { "freshness": "week", "location": "eu" } } }
```
There are no CLI flags.
Precedence per key: env, trusted project `.pi/settings.json`, user `~/.pi/agent/settings.json`, default.
When project and user values are both objects they deep-merge with project winning; scalars replace.
Project settings are read only when Pi trusts the project.
`web.zai.freshness` accepts `day`, `week`, `month`, `year`, or `none`, `nolimit`, `no_limit`, empty for no limit.
All other keys take non-empty strings; surrounding whitespace is trimmed, and an empty env var counts as set.
`web.zai.mcpUrl` is normalized to the `/api/mcp/web_search_prime/mcp` endpoint when it lacks a `/mcp` path.
MiniMax prefers the `baseUrl` of a registered `minimax` or `minimax-cn` model over `web.minimax.apiHost`.
An invalid value never falls back to a lower source.
Session start shows a warning notification per invalid key, and the affected provider fails its search with the same message.
## Debug
Opt-in metadata diagnostics: [debug contract](../DEBUG.md).
Safe events: `session.start`, `session.shutdown`, `action.execute.start`, `action.execute.finish`, `action.execute.error`.
## Bench
Run `mise run bench` here, or `mise //extensions/web:bench` from the repository root, to benchmark web processing hot paths.
## Live search comparison
Run `mise run //extensions/web:evaluate-search` to compare all configured providers on five fixed English/German technical and consumer-information queries.
It uses existing Pi credentials through ModelRegistry, makes no model-catalog refresh, applies a 45-second deadline per provider, rotates provider order, and records target rank, elapsed time, and full results in the Pi state directory.
Providers with authentication, quota, or CAPTCHA failures are not retried during that run.
These small-sample target-rank checks are not a general search-quality benchmark.
## Behavior
- `web` with `action: fetch` uses native `fetch` with browser-like headers and a per-host
referer.
- Fetch accepts only HTTP and HTTPS URLs.
Private/reserved IP targets, blocked DNS answers, and unsafe redirects are
rejected.
- Fetch output is stored in the Pi state directory with a `responseId` and
`fullOutputPath`.
Capped output, matching excerpts, and no-match results include the path in
tool-visible text so the complete saved output can be read directly.
- `linesMatching` returns numbered excerpts with optional surrounding context.
- Content routing is based on `Content-Type`:
- HTML/XHTML:
strips common prompt-injection delivery vectors, extracts main content,
converts with `pandoc` and `llm-extract.lua`.
- Text/XML/YAML/CSV/TSV:
returns UTF-8 text.
- JSON:
pretty-prints when parseable, otherwise returns raw text.
- PDF:
writes temp file, runs `pdftotext -layout`.
- DOCX/ODT/EPUB/RTF:
writes temp file, runs `pandoc` to plain text.
- Images:
writes temp file, returns metadata from `identify`, plus inline image
content.
- Audio/video:
writes temp file, returns summary and `ffprobe` JSON.
- Other binary:
writes temp file and returns path plus content type.
- Fetch `raw` format skips conversion, decodes the body as UTF-8, and
truncates to `maxChars`.
- HTML cleanup removes comments, scripts, styles, iframes, objects, templates,
noscript blocks, embeds, hidden inputs, long meta content,
hidden/visually-hidden elements, and non-image `aria-label`/`alt`/`title`
attrs.
- Main-content extraction prefers `<article>`, `#mw-content-text`,
`#docContent`, `#content`, `<main>`, then `role="main"`; fallback is full
page.
- Converted fetch text, raw fetch text, and search output are wrapped as
untrusted external data.
Image, media, and binary early-return paths are not wrapped by the current
implementation.
- `web` with `action: map` discovers site maps from robots.txt, sitemap XML, and llms links.
Its URL listing shows at most 50 entries and obeys `maxChars`.
If either limit omits URLs, the result identifies the listing as partial and
includes the saved map's `fullOutputPath`.
- `web` with `action: search` provider order:
Google WML, anonymous Parallel, OpenAI Codex native search when authenticated, MiniMax search when configured, DuckDuckGo HTML, then Z.AI Coding Plan MCP when configured.
Cancellation stops fallback.
- Parallel uses the anonymous hosted MCP at `https://search.parallel.ai/mcp` without API keys, model metadata, or conversation identifiers.
All excerpts from each selected result are preserved as its snippet, without character caps.
The existing result-count limit still applies.
- Google uses `/wml/search` with a fixed Nokia user agent and native HTTP, without JavaScript or browser automation.
Requests reserve slots at least two seconds apart.
CAPTCHA, consent pages, and unrecognized layouts fail into the next provider; no CAPTCHA solving or proxy rotation.
This is an undocumented endpoint and can stop working.
- Parallel and Z.AI share the MCP transport in `mcp.ts`: JSON/SSE decoding, response-ID matching, initialization, session headers, protocol negotiation, and redirect rejection.
Authentication, tool arguments, result decoding, and provider-specific errors remain separate.
- Z.AI calls MCP tool `web_search_prime` with `web.zai.freshness`, `web.zai.contentSize`, and `web.zai.location`.
- DuckDuckGo submits the HTML no-JavaScript form using POST, query/region fields, and origin/referer headers, instead of searching with GET.
It throttles requests to at least two seconds apart, includes the throttle wait in its deadline, and reports CAPTCHA as a provider failure.
A local two-query comparison returned ten results per query for HTML/Lite POST, while both GET interfaces returned CAPTCHA with no results.
A follow-up five-query evaluation through the production HTML provider found the intended target first for every query.
This fixes the observed request pattern, not all possible DDG blocks; no CAPTCHA solver, proxy rotation, or browser dependency is added.
See [DDG's non-JavaScript interfaces](https://duckduckgo.com/duckduckgo-help-pages/features/non-javascript) and [SearXNG's request and bot-blocker notes](https://docs.searxng.org/dev/engines/online/duckduckgo.html).
- MiniMax search posts to `/v1/coding_plan/search`.