@bacnh85/pi-web

extensionmaintained

Pi extension for web search, page extraction, Firecrawl scraping/crawling, and Crawl4AI headless browser crawling.

by · v0.6.0 · published 2d ago

$ pi install npm:@bacnh85/pi-web
downloads/mo
1.9K
stars
8
last push
12h ago
open issues
1

Signals

license: MITtestspi manifest: missinginstall size: —deps: 0peer deps: 0

Download trend

3.7K downloads · last 12 weeks (weekly)

README

@bacnh85/pi-web

Pi extension for unified web search, content extraction, site crawling, and page capture.

Auto-selects the best backend from SearXNG (self-hosted), Brave Search, Firecrawl, Crawl4AI, and agy (Gemini/Claude, when installed) — so agents don't have to know which backend to use. Search selection is adaptive: broad discovery prefers self-hosted SearXNG, while precision-sensitive searches and inline content prefer Brave.

Install

pi install npm:@bacnh85/pi-web

Configuration

Environment lookup order:

  1. Process environment
  2. Current working directory .env.local
  3. Current working directory .env
  4. Pi global config ~/.pi/agent/.env.local
  5. Pi global config ~/.pi/agent/.env

Variables:

VariableRequiredDefaultNotes
BRAVE_API_KEYNo (1)Brave Search API key
SEARXNG_BASE_URLNohttp://172.30.55.22:8888Self-hosted SearXNG
FIRECRAWL_API_URLNohttps://api.firecrawl.dev/v2Self-hosted or hosted
FIRECRAWL_API_KEYNo (2)Required for hosted Firecrawl
CRAWL4AI_API_URLNohttp://172.30.55.22:11235Self-hosted Crawl4AI
CRAWL4AI_API_TOKENNo (3)Required if Crawl4AI auth enabled

(1) At least one search backend (SearXNG, Brave, or Firecrawl) must be configured for web_search. (2) Required for hosted Firecrawl; optional for self-hosted instances without auth. (3) Required for Crawl4AI v0.9+ default config.

Secrets are never printed; web_status reports only presence/source.

Always-on routing guidance

When any web_* tool is active, pi-web injects a condensed backend-selection protocol (SearXNG → Brave → Firecrawl ordering, Firecrawl precision/scrape caveats, source-citation rule) into the system prompt via a before_agent_start hook. This travels with the package — no edits to ~/.pi/agent/AGENTS.md are required — and carries zero overhead when pi-web is not loaded.

Tools

web_search — Unified search

Searches the web. Auto-selects backends adaptively: SearXNG for broad self-hosted discovery, Brave for precision-sensitive queries and include_content, Firecrawl as last resort.

web_search query="ansible podman quadlet" count=5
web_search query="ansible documentation" backend=brave count=10
web_search query="latest python release" engines="google,github"
web_search query="riven media" include_content=true

Parameters:

ParameterTypeDefaultDescription
querystringSearch query
countnumber5Number of results (max 20)
freshnessstringTime filter: pw, pm, py, or YYYY-MM-DDtoYYYY-MM-DD
countrystringUSTwo-letter country code
backendstringautoForce backend: auto, searxng, brave, firecrawl
enginesstringSearXNG engine override, e.g. google,github
include_contentbooleanfalseFetch page content alongside results
content_charsnumber5000Max content chars per result

Auto-selection behavior:

  1. SearXNG — first for broad/general discovery, especially when engines is supplied.
  2. Brave — first for precision-sensitive queries (site:, quoted phrases, docs/API/source lookups, short proper-name queries) and whenever include_content is true. Requires BRAVE_API_KEY.
  3. Firecrawl Search — last resort. ⚠️ Poor semantic accuracy on domain-specific/ambiguous queries (e.g., "Riven" returns League of Legends results). Prefer SearXNG or Brave for precision.

Tool output includes search diagnostics showing attempted backends and the selected backend.

Use backend parameter to force a specific backend when needed.

web_extract — Unified content extraction

Extracts readable content from a URL. Auto-selects backend: static (JSDOM) → dynamic (Firecrawl) → full (Crawl4AI) → agy (model-backed), with extraction diagnostics showing fallback attempts.

web_extract url="https://docs.ansible.com/..."
web_extract url="https://riven.tv/" mode=static
web_extract url="https://example.com" mode=dynamic prompt="Extract pricing plans"
web_extract url="https://blocked.example.com" mode=agy

Parameters:

ParameterTypeDefaultDescription
urlstringURL to extract
modestringautoauto, static, dynamic, full, or agy
promptstringPrompt for JSON extraction (dynamic/agy modes)
schemaanyJSON schema for structured extraction (dynamic/agy modes)
content_charsnumber20000Max content chars
wait_fornumberMilliseconds to wait for Firecrawl dynamic rendering. Crawl4AI /md full mode may ignore this.
mobilebooleanfalseEmulate mobile viewport (dynamic mode)

Mode behavior:

ModeBackendBest forAPI key needed
staticJSDOM+ReadabilitySimple static pages, blog posts, docsNo
dynamicFirecrawl ScrapeJS-rendered pages, dynamic contentMaybe
fullCrawl4AIJS-heavy SPA, complex renderingMaybe
agyagy (Gemini/Claude)Bot-protected / anti-AI-scraping pagesagy CLI installed
auto (default)static → dynamic → full → agyUnknown page typeMaybe

In auto mode, fallbacks are noted in the output (e.g., [Extraction fell back to Firecrawl Scrape (dynamic mode)]). If static extraction fails, the tool gracefully escalates to heavier backends.

⚠️ Note on Firecrawl Scrape: Fails on bot-protected sites (Ansible docs, many CDN-backed doc sites). Falls back to full mode (Crawl4AI) in auto mode, and to agy mode as a last resort.

agy mode (optional): Uses the Antigravity CLI with Gemini/Claude — its native read_url browser tool can fetch pages that block Firecrawl/Crawl4AI. Install with curl -fsSL https://antigravity.google/cli/install.sh | bash, authenticate once with agy, then auto mode falls back to it automatically. If agy is not installed, auto mode skips it silently; web_status reports agy.installed.

web_map — Site URL discovery

Discovers URLs from a site using Firecrawl Map. Best on base domains; may return fewer results on sub-paths.

web_map url="https://riven.tv"
web_map url="https://docs.example.com" sitemap=only

Parameters: url, limit (default 100), include_subdomains, search, sitemap, use_index, ignore_cache.

web_crawl — Site crawl

Crawls pages from a site. Two modes:

  • light (default): Firecrawl Crawl — conservative, docs-focused, single URL.
  • full: Crawl4AI Crawl — headless browser, rendered data, media, links, up to 100 URLs.
web_crawl url="https://docs.example.com" limit=10          # Firecrawl light mode
web_crawl urls=["https://a.com","https://b.com"] mode=full  # Crawl4AI full mode
web_crawl url="https://example.com" mode=light poll=true    # Poll for completion

web_screenshot — Page screenshot

Captures a full-page PNG screenshot using Crawl4AI. Returns base64-encoded PNG.

web_screenshot url="https://example.com"
web_screenshot url="https://example.com" wait_for=5 wait_for_images=true

web_pdf — Page PDF

Generates a PDF document using Crawl4AI. Returns base64-encoded PDF.

web_pdf url="https://example.com/article"

web_status — Provider status

Shows all provider configuration status and Crawl4AI server health.

web_status

Typical output:

{
  "brave": { "apiKeyFound": true, "apiKeySource": "process.env" },
  "searxng": { "baseUrl": "http://172.30.55.22:8888", ... },
  "firecrawl": { "baseUrl": "http://172.30.55.22:3002/v2", ... },
  "crawl4ai": {
    "baseUrl": "http://172.30.55.22:11235",
    ...
    "health": { "status": "healthy", "version": "0.5.0", ... }
  },
  "agy": { "installed": true }
}

Library structure

ModuleContents
lib/config.tsEnvironment loading, config helpers for all providers
lib/format.tsText sanitization, truncation, crawl/scrape result formatting
lib/content.tsReadable content extraction (JSDOM + Readability + Turndown)
lib/retry.tsRetry with exponential backoff for transient HTTP failures
lib/brave.tsBrave Search API fetch client (internal)
lib/searxng.tsSearXNG metasearch fetch client (internal)
lib/firecrawl.tsFirecrawl API fetch client with v2→v1 fallback (internal)
lib/crawl4ai.tsCrawl4AI Docker API fetch client (internal)
lib/agy.tsagy (Antigravity CLI) spawn helper — read_url extraction via Gemini/Claude
lib/search.tsUnified search orchestrator — probes backends, fallback chain
lib/extract.tsUnified extraction orchestrator — mode-based backend selection

Migration from 0.3.x

v0.4 replaces the 14 individual backend-specific tools with 7 unified tools:

v0.3 toolv0.4 replacement
brave_searchweb_search with backend: "brave"
searxng_searchweb_search with backend: "searxng"
firecrawl_searchweb_search with backend: "firecrawl"
web_contentweb_extract with mode: "static"
firecrawl_scrapeweb_extract with mode: "dynamic"
crawl4ai_scrapeweb_extract with mode: "full"
firecrawl_mapweb_map (same behavior)
firecrawl_crawlweb_crawl with mode: "light"
crawl4ai_crawlweb_crawl with mode: "full"
crawl4ai_stream(removed — use web_crawl with mode: "full")
crawl4ai_screenshotweb_screenshot (same behavior)
crawl4ai_pdfweb_pdf (same behavior)
crawl4ai_statusMerged into web_status
web_statusweb_status (enhanced with Crawl4AI health)

All v0.3 tool names were removed in v0.4. Update any agent instructions or skills that reference the old names.

Changelog

See CHANGELOG.md for release history.

Development

# Run all tests
npm test

# Run only unit tests
npm run test:unit