Vision
Image understanding, OCR, and screenshot tooling.
162 packages · auto-seeded from the “vision” category.
A persistent, policy-guarded Playwright browser for AI agents with network controls, trusted credential filling, proof screenshots, and CAPTCHA helpers.
browservision$ pi install npm:betterwrightMCP server that gives AI agents the web: fetch, search, crawl and screenshot from one local Rust binary. Zero API keys, zero accounts, Chrome-true TLS.
mcpwebbrowservision$ pi install npm:donsetchPi extension for image generation via OpenAI gpt-image, Google Nano Banana (Gemini), Alibaba Qwen-Image, OpenRouter, and custom providers.
vision$ pi install npm:@amaster.ai/pi-image-genRun AI coding agents in pi TUI overlays with interactive, hands-free, and dispatch supervision
vision$ pi install npm:pi-interactive-shellPi package that adds document_parse, document_search, document_screenshot, and a companion skill for local document understanding with LiteParse v2.
visionskill$ pi install npm:pi-docparserImprove OpenAI in pi with fast mode, usage stats, realtime voice, image generation, and footer polish.
vision$ pi install npm:@monotykamary/pi-better-openaiTurn an unfamiliar codebase into a validated reimplementation spec, then synthesize confirmed specs and a product vision into a traceable plan.
visionplan$ pi install npm:codecartographer-piPi extension for AI video generation plus local video composition: lossless clip concat and mixed image/video timelines with overlays, TTS, soft or burned subtitles, source audio, BGM, and bundled LGPL/GPL FFmpeg runtimes.
vision$ pi install npm:@amaster.ai/pi-video-genImage generation and image recognition tools for the Pi coding agent
vision$ pi install npm:@pi-unipi/imageGive text-only pi models vision — describe images with a vision model you pick via an interactive picker, then hand off the text description to non-vision models
vision$ pi install npm:pi-vision-handoffClipboard image paste for pi-archimedes
vision$ pi install npm:@pi-archimedes/image-pasteLinux desktop-control MCP server: AT-SPI accessibility trees, Wayland/X11 input, screenshots, and compositor window targeting.
mcpvision$ pi install npm:@agent-sh/computer-use-linuxAutomatic image, video and audio description for any model in Pi. Routes media to a multimodal model and injects descriptions into context.
vision$ pi install npm:pi-multimodal-proxyEnhanced read for pi – office docs (docx/pdf/pptx/xlsx etc.) via anydoc, header-only, OCR for scanned pages (Firecrawl or local rapidocr, in the order you configure).
webvision$ pi install npm:@everyx/pi-read-docCapability-aware vision + paste extension for the pi coding agent. Delegates image analysis to a vision model only when the active primary model is text-only; passes images through natively for multimodal models (zero delegation).
subagentvision$ pi install npm:@getpipher/visionPaste screenshots into the Pi coding agent without pasting file paths — stable [Image #N] references, clickable history, and compact 480px model thumbnails
vision$ pi install npm:pi-image-viewImage generation and editing for Pi using your ChatGPT Codex login.
vision$ pi install npm:pi-codex-image-genImage and file attachments for Pi — converts pasted/dropped file paths into real image attachments, collapses large text pastes into readable file references, inlines text files, and pastes clipboard file references.
vision$ pi install npm:@bacnh85/pi-attachmentsBrowser control extension for pi — navigate, click, type, screenshot, and extract data from real Chrome via CDP
browservision$ pi install npm:pi-browser-harnessRender LaTeX as terminal images in Pi's TUI.
vision$ pi install npm:@monotykamary/pi-mathPi extension for AI multimodal generation and media processing via the multix CLI: images, video, speech, music, 3D, document conversion, and ffmpeg/ImageMagick optimization.
vision$ pi install npm:pi-multix- @lamplitisles/pi-imagegenextension
Kepos bridge image generation extension for Pi
vision$ pi install npm:@lamplitisles/pi-imagegen Bridge a live pi coding session to Discord: share current work as a digest and keep going from chat. Text and image prompts, one progress card per turn, no bot SDK dependencies.
vision$ pi install npm:pi-bot-connectPreview pasted images in PI Agent before sending them, with inline thumbnails and a compact attachment gallery.
vision$ pi install npm:@prjct.app/pi-clipboardImage attachment and rendering extension for Pi TUI
vision$ pi install npm:pi-image-tools- pi-provider-freellmapiextension
Register the FreeLLM API gateway (freeapi.n.cofire.cn) as an OpenAI-compatible provider in pi, with automatic model discovery and tools for embeddings, image/video generation, speech, and transcription
vision$ pi install npm:pi-provider-freellmapi Registers a describe_image tool that routes image analysis to a configured vision model through pi's official pipeline.
vision$ pi install npm:pi-aux-vision- pi-vision-fallbackextension
Pi extension: describes images with a configured vision model when the active model cannot see
vision$ pi install npm:pi-vision-fallback Unpublished pi-codex-imagegen workspace for the staged Codex split
vision$ pi install npm:@oai404iao/pi-codex-imagegenImage and optional sprite-sheet generation for Pi via ChatGPT Codex subscriptions, OpenAI, Gemini, Qwen, Seedream, Meta Muse Image, OpenRouter, and custom providers.
vision$ pi install npm:@abhishek944/pi-image-genGive text-only pi models vision — describe images with a vision model you pick via an interactive picker, then hand off the text description to non-vision models
vision$ pi install npm:@bismawy/pi-vision-watcherAsk LLM review, comparison, brainstorming, image, verification, and pairing workflows for Claude Code, Cursor Agent, and Pi
vision$ pi install npm:@ask-llm/pluginGenerate, edit, and analyze images in pi using Google Nano Banana (image gen) and Gemini Vision (analysis). Inline terminal preview, reference-image editing, auto-save.
vision$ pi install npm:pi-bananapi extension: /shake tools|images|thinking — rebuilds session history in place to free context space
vision$ pi install npm:pi-context-shakeExpose Codex subscription text, vision, image, search, Fast, and usage APIs to Pi
vision$ pi install npm:@99percentpeople/pi-codex-apiMinimal Codex/OpenAI native tools for Pi: Codex image_generation, view_image, apply_patch
vision$ pi install npm:@vanillagreen/pi-codex-minimal-toolsPi extension that turns pasted image paths into first-class image attachments.
vision$ pi install npm:pi-pasterPi quality-of-life extension: compact statusline/π prompt, reliable multiline input, styled pasted-image chips, session naming/search/context import, scheduled prompts, handoff, permission prompts, notifications, custom compaction, and a collapsed-thinkin
visionpromptsecurity$ pi install npm:@vanillagreen/pi-qolPi Agent extension that adds a describe_image tool, letting non-multimodal models delegate image analysis to a vision-capable model (like Qwen VL)
subagentvision$ pi install npm:pi-vision-toolCodex-compatible native image attachment tool for Pi
vision$ pi install npm:pi-view-image- pi-aia-workspaceextension
One-command installer for the Ai Applied pi stack (pi-vigilant, pi-aia-asf, pi-aia-browser, pi-smart-web-search, pi-smart-fetch, pi-vision-handoff, pi-intercom, pi-safe-compact, compaction-fix).
webbrowservision$ pi install npm:pi-aia-workspace Pi extension exposing a gpt-5.5+ image generation tool backed by gpt-image-2.5.
vision$ pi install npm:pi-codex-image-tool- @earendil-works/pi-radius-workextension
Radius Google Workspace and image generation extensions for Pi
vision$ pi install npm:@earendil-works/pi-radius-work TypeScript port of the nano-banana-imagegen pi skill. Generate and edit images with Google Gemini image models via the @the-focus-ai/nano-banana CLI, with GEMINI_API_KEY resolution, output-path handling and batch generation. Exposed as a pi skill and the
visionskill$ pi install npm:@blackbelt-technology/pi-dashboard-nano-bananaConnect pi to 9Router — multi-provider chat models plus image, speech, search, and fetch tools
vision$ pi install npm:@qmahyar/pi-9routerGive DeepSeek (text-only) models vision in Claude Code, Codex, and Agent Plugins clients: describe images via any OpenAI-compatible vision endpoint.
vision$ pi install npm:@limccn/deepseek-vl-supportPaste an image in Pi TUI; relay it with recent chat context and your prompt to a vision model, and inject the analysis into the main model's context
visionprompt$ pi install npm:pi-paste-image-to-modelAgnes AI for pi: /model text model catalog (intl + CN) plus image/video generation as callable tools + auto-loaded skill. No model switch needed for media.
visionskill$ pi install npm:pi-agnes-toolsA Jupyter-shaped Python notebook for the pi agent: an ordered list of cells over one persistent namespace, with staleness hints, percent-format persistence and plots the model can see.
visionlinter$ pi install npm:@ocramz/pi-notebook-py- deepseek-vl-supportextension
Give DeepSeek (text-only) models vision in Claude Code, Codex, and Agent Plugins clients: describe images via any OpenAI-compatible vision endpoint.
vision$ pi install npm:deepseek-vl-support Pi TUI configuration and an on-demand Skill/CLI for image generation.
visionskill$ pi install npm:@bytetrue/pi-image-gen- pi-localterm-kitty-imagesextension
Enable Kitty image protocol for localterm in pi's TUI
vision$ pi install npm:pi-localterm-kitty-images Vega-Lite chart extension for pi coding agent - render data visualizations as inline images
vision$ pi install npm:@walterra/pi-chartsReadable LaTeX in Pi: terminal-native inline math and MathJax display images.
vision$ pi install npm:pi-formulaGenerate and edit images using Codex CLI
vision$ pi install npm:@pi-lab/codex-imageCodex-compatible native image attachment tool for Pi
vision$ pi install npm:@luan.sh/pi-view-image- @speclip/pi-subvisionextension
Workspace-safe hard-subtitle OCR through the local SubVision macOS service for Pi
vision$ pi install npm:@speclip/pi-subvision Pi Coding Agent extension for OmniRoute — view combos, browse providers, and sync models with enriched metadata (context windows, max tokens, reasoning, and vision) to the Ctrl+P picker
contextvision$ pi install npm:omniroute-pi-ext-integrationPi extension package for direct HTTP GET plus web_search, web_fetch, and web_image via Firecrawl, Exa, Tavily, and Brave.
webvision$ pi install npm:@xl0/pi-lovely-webCynos universal search, vision, and browser tools for the pi coding agent.
browservision$ pi install npm:@cynos-ai/toolsFill any standard Word .docx template — inject AI-authored sections, images, and tables from a template bundle. Generalizes the former manual-creator; controls manuals are one template.
vision$ pi install npm:sylo-template-docx-writerSynthetic (synthetic.new) model provider for pi - Dynamic model fetching with reasoning, vision, and tools support
vision$ pi install npm:@benvargas/pi-synthetic-providerLocal hybrid RAG pipeline for the Pi coding agent. SQLite FTS5 + sqlite-vec, ONNX embeddings via Transformers.js, PDF/DOCX/HTML extraction (with OCR fallback), per-project storage, auto-injection. Zero cloud dependency.
vision$ pi install npm:pi-local-ragPi TUI clipboard image attachment bridge
vision$ pi install npm:@jingoz/pi-image-inputpi coding agent extension for OpenCodeReview (ocr): /ocr-review commands and ocr_review / ocr_delegate tools for AI-powered code review of Git changes.
subagentvisiongit$ pi install npm:open-code-review-piSocratic planning and shared-understanding sessions for pi.
vision$ pi install npm:@majorgilles/pi-grill-meImage generation and reference-image editing for pi.
vision$ pi install npm:oira666_pi-image-generationPi extension that resizes oversize images at Read-time so they fit model byte and pixel ceilings.
vision$ pi install npm:@blackbelt-technology/pi-image-fit-extensionWindows image picker for pi running in WSL or Windows
vision$ pi install npm:@leokon3/pi-image-pasteHound web research for Pi agent - keyless search, anti-bot fetch, deep crawl, screenshot, Internet Archive dead-link recovery. Free, MIT, no API key.
mcpwebvision$ pi install npm:@houndmcp/hound-mcp-piUnofficial pi package that exposes Z.ai MCP server tools for web search, URL reading, repository reading, and vision workflows.
mcpwebvision$ pi install npm:pi-zai-mcpEncrypted Bark notifications and image paste placeholders for Pi
vision$ pi install npm:@aerok/pi-toolkitPi extension + skill for a ground→contract→mockup→test→fix→learn frontend design loop. Ships a live mockup server tool, a Playwright breakpoint-screenshot tool, and a design-contract scaffolder. Works in any React/Tailwind/shadcn project.
browservisionskill$ pi install npm:@blackbelt-technology/frontend-mockup-loop- @joemccann/pi-pdfextension
PDF manipulation, processing, and management toolkit for Pi coding agent — extract text/tables, merge/split, fill forms, create PDFs, OCR, watermark, encrypt/decrypt, and more
vision$ pi install npm:@joemccann/pi-pdf Agnes AI provider for pi — registers agnes (token billing) and agnes-plan (subscription) providers with text and image input models
visionplan$ pi install npm:@d3ara1n/pi-provider-agnesProgressive exact-revision project graph reviews for Pi and Claude Code
vision$ pi install npm:@alexjercan/quick-reviewMCP vision bridge for text-only coding agents
mcpvision$ pi install npm:atlas-vision-mcp- @xinizai/pi-vision-toolextension
A Pi extension that delegates current-turn image understanding to a separately configured OpenAI-compatible Vision Provider.
subagentvision$ pi install npm:@xinizai/pi-vision-tool Pi extension that gives non-vision GLM models (z.ai) image understanding via GLM-4.6V
vision$ pi install npm:glm-visionScreenshot picker extension for pi coding agent - quickly select and attach screenshots to your prompts
vision$ pi install npm:pi-screenshots-pickerThin Pi subagent primitive: one synchronous call for isolated context, live progress, timeout supervision, resumable sessions, and per-task model/tool control.
subagentvisionplan$ pi install npm:@eggmasonvalue/pi-subagentAutomatic jj revision management — guards file edits to keep Jujutsu revisions focused
vision$ pi install npm:pi-jj-autoPi extension: Zero-setup multi-backend OCR — MinerU (free cloud), Ollama (local GPU, LaTeX formulas), Pix2Text (local Python). Extract text, formulas, and tables from images and PDFs. Default: zero config, works out of the box.
vision$ pi install npm:pi-ocrPi extension: relay tool-result images to Gemini-family user attachments across Responses proxies.
vision$ pi install npm:pi-gemini-image-bridgeGenerate or edit PNG images in Pi using a ChatGPT Plus/Pro Codex subscription.
vision$ pi install npm:@crazygit/pi-codex-image-genDeepSeek vision image offload for pi: Files API file_id references on the official gateway (deepseek-v4-flash-vision-exp), offload trimming on third-party gateways — kills the 50 MB 413 ceiling.
vision$ pi install npm:pi-imagefilesBuild knowledge artifacts that survive revision and session boundaries — plans, specs, repo or data analyses, reports — with declarative seams, a checkpoint, provenance, and a domain-agnostic patch tool.
vision$ pi install npm:knowledge-artifactspi-web plugin that browses and previews screenshots saved by pi-web into .pi-web/paste/ — gallery panel, lightbox, and chat-inline previews for folder-mode attachments.
webvision$ pi install npm:@yieldcraft/screenshot-pasteComfyUI image/video generation extension for pi coding agent
vision$ pi install npm:pi-comfyui-paintPi extension: let a text-only model read images by asking a vision-capable model from your own models.json.
vision$ pi install npm:@bytetrue/pi-visionLog in to the airpx LLM proxy (https://airpx.cc) with an sk-proxy API key and auto-import its model catalog (prices, context windows, reasoning, vision) into the pi coding agent.
webcontextvision$ pi install npm:pi-airpxpi extension + skill: visual understanding via a configurable vision model (pi-registry reuse or custom responses API; default gpt-5.6-luna) when the main model cannot see images — pure TypeScript, single runtime
visionskill$ pi install npm:@yceachan/pi-vision-helperTransparent vision fallback for text-only pi models — overrides read so image files are described by a vision model when the active model cannot see images. Supports any OpenAI-compatible endpoint.
vision$ pi install npm:@arhen/pi-core-vision- @lvrged/lvrged-factoryextension
GPU infrastructure control for Pi: provision, deploy, run, monitor, pause, and destroy GPU video workloads on RunPod (the one first-class adapter; more plug in via the adapter pattern) — MiniMax H3 on RTX PRO 6000 with the Turbo 8-step / SageAttention2 /
vision$ pi install npm:@lvrged/lvrged-factory - @feniix/pi-code-reasoningextension
Archived and no longer maintained: Code Reasoning tools for pi and MCP — reflective sequential thinking with branching and revision support
mcpvisiongit$ pi install npm:@feniix/pi-code-reasoning Give text-only pi models vision — a native vision tool with a bundled vision script and a multi-backend model chain (auto-fallback, timeout, retry). Configure models via .env, JSON, or a custom script.
vision$ pi install npm:pi-multivisionPi agent extension that parses image/audio/video inputs and forwards them to an OpenAI-compatible multimodal proxy (e.g. dots.ai).
vision$ pi install npm:multi-content-proxyLocal MLX vision for text-only models on Apple Silicon Macs: a describe_image tool backed by LiquidAI LFM2.5-VL-3B-MLX-8bit, so non-vision models (DeepSeek, etc.) can see images.
vision$ pi install npm:pi-mlx-vision- @lvrged/video-factoryextension
GPU infrastructure control for Pi: provision, deploy, run, monitor, and destroy GPU workloads on RunPod, Vast.ai, SaladCloud, Modal, Lambda, and Prime — with H3/Wan/Hunyuan video generation as the first-class workload. First-run onboarding, spend policy,
vision$ pi install npm:@lvrged/video-factory Browser automation and web research tools for pi — search, visit, screenshot, interact, and inspect console output
webbrowservision$ pi install npm:@dreki-gg/pi-browser-tools- pi-ui-hephaestusextension
Muted thinking blocks, framed editor, animated header, response time, rich footer, and clipboard image paste for pi
vision$ pi install npm:pi-ui-hephaestus Pi extension for image generation: Codex (ChatGPT OAuth via pi, no Codex CLI) plus any OpenAI-compatible provider (llm-center, OpenAI, proxies).
vision$ pi install npm:@mics8128/pi-imagegenAn issue-tracker extension for pi that turns goals into linked, tracked user stories stored in a project-local SQLite database.
vision$ pi install npm:@ocramz/pi-issue-tracker- pi-agent-pi-markitdownextension
Convert any document to Markdown with Microsoft's markitdown CLI — PDF, DOCX, PPTX, XLSX, images (OCR), audio (transcription), YouTube, HTML, CSV, JSON, XML, EPUB, ZIP
vision$ pi install npm:pi-agent-pi-markitdown Smart clipboard paste for pi on Windows — copy files and paste their paths; falls back to images.
vision$ pi install npm:pi-smart-paste- @kassing/pi-visionextension
pi 视觉理解桥接扩展:当前模型不支持图片时,自动调用外部视觉模型分析图片并注入结果
vision$ pi install npm:@kassing/pi-vision Paste images from the Mac clipboard into pi running on a remote machine over ssh. A launchd socket service on the Mac serves the clipboard image; an ssh RemoteForward carries it; the extension fetches it on ctrl+v and attaches it like a native paste.
vision$ pi install npm:@guygrigsby/pi-image-pastePaste Windows clipboard text and images into Pi from WSL with Ctrl+V.
vision$ pi install npm:pi-wsl-clipboardFind skipped Mermaid diagrams and show them in a terminal image viewer.
vision$ pi install npm:@herbertgao/pi-mermaid-openCodex image generation and editing for Pi, Code Mode and Notebook Mode.
vision$ pi install npm:@howaboua/pi-codex-imagegenPi skill — a mechanical, countable anti-slop checklist for AI-generated frontend. Catches the specific tells an undirected model defaults to (AI-purple, Inter-everywhere, em-dashes, div-based fake screenshots, eyebrow-per-section, Jane Doe / Acme data). A
visionskill$ pi install npm:@blackbelt-technology/anti-slop-frontendPi extension for web search, URL content extraction, and image search (Tavily, Kimi, DeepSeek, Mimo, Z.AI, DashScope, Unsplash, and more)
webvision$ pi install npm:@amaster.ai/pi-web-accessAsynchronous Pi subagents and agent fleet orchestration for parallel coding agents with nested delegation, background work, and supervision in herdr.
subagentvision$ pi install npm:pi-herdsmanTransparent image-to-text bridge for text-only Pi models using a configurable vision model
vision$ pi install npm:@fradser/pi-visionOpenAI toolkit for Pi: Codex Remote Context windows, remote compaction v2, routed Web Search, image gen, auto mode.
webcontextvision$ pi install npm:pi-openai-toolkitGraphviz chart extension for pi coding agent - render DOT diagrams as inline images
vision$ pi install npm:@walterra/pi-graphviz- pi-agent-browserextension
Browser automation tool for pi — interactive browsing, screenshots with inline vision, and session cleanup via agent-browser CLI
browservision$ pi install npm:pi-agent-browser Pi extension for web search, page extraction, Firecrawl scraping/crawling, Crawl4AI headless browser crawling, real-browser interaction (trusted click/type/evaluate via CDP), Gemini web-tier research, free upstream image generation (Gemini/ChatGPT web/Z.a
webbrowservision$ pi install npm:@bacnh85/pi-webFind skipped Mermaid diagrams and show them in a terminal image viewer.
vision$ pi install npm:@tifan/pi-mermaid-openEssential extensions for pi — screenshots, image context pruning, a daily log tool, and a markdown viewer.
vision$ pi install npm:@samfp/pi-essentialsTemporary pasted-image cache with compact placeholders for the pi coding agent.
vision$ pi install npm:@tian.zuo/pi-image-cacheQwenCloud provider for pi — access Qwen3.8, DeepSeek V4, GLM-5.2, Wan image generation, and HappyHorse video generation through QwenCloud's OpenAI-compatible API
vision$ pi install npm:pi-qwencloud-providerPi-native ralph loop — autonomous coding iterations with mid-turn supervision
vision$ pi install npm:@lnilluv/pi-ralph-loopPi package for OpenAI/Codex image generation with a local browser studio.
browservision$ pi install npm:pi-imagegenA cute, colorful pet companion for the pi coding agent. Use classic ASCII pets or Petdex image pets that live below your editor and react to tools, tests, commits, PRs, reviews, subagents, model swaps and more.
subagentvisiongit$ pi install npm:pi-pokepetImage preview extension for pi coding agent — renders inline image thumbnails above the editor using kitty graphics protocol with tmux support
vision$ pi install npm:pi-image-previewPi extension for literature search, Zotero, PDF parsing, citations, materials data, and scientific image generation
vision$ pi install npm:@luffysolution/pi-scholarEnhanced Qwen Token Plan CN provider for pi. Speaks the OpenAI Responses API (not chat-completions) to activate the platform's server-side Harness tools — web search, code interpreter, web extraction, and image search — that the built-in qwen-token-plan-c
webvisionplan$ pi install npm:pi-extension-qwen-token-plan-cn-exTranslate EPUB books in Pi while preserving structure, styling, images, footnotes, and packaging.
vision$ pi install npm:pi-epub-translatorOCR image-to-text tool for Pi — extracts text from screenshots, terminal output, and code images using Tesseract + ImageMagick
vision$ pi install npm:@k3_2o/pi-read-imageOpenRouter multimodal tools for Pi — search, fetch, image gen, vision, video, PDF, TTS, STT
vision$ pi install npm:@dtmirizzi/pi-openrouter-multimodalTyped Pi tools for browser automation, rendered-page context extraction, screenshots, and saved session state via agent-browser.
browservision$ pi install npm:@53able/pi-agent-browserOpenCode-grade clipboard paste for pi on Windows: instant Ctrl+V image attachments, Explorer file drops as collapsible blocks, large-text collapsing, persistent clipboard helper
vision$ pi install npm:pi-power-pasteStandalone Pi extension for OpenAI fast mode, GPT-5.6 Pro request injection, and image generation/editing.
vision$ pi install npm:@tryinget/pi-better-openaiPi extension for NVIDIA NIM — chat, vision (VLM), embeddings, and image generation
vision$ pi install npm:pi-nvidiaInteractive Playwright web browsing for Pi. Chromium and Firefox built in; author stealth backends on either engine. Snapshots and screenshots cache to disk, saving context. Configurable profiles/cookies. Agents write site guides, auto-matched by domain.
webbrowservision$ pi install npm:pi-lean-portalDeepInfra provider for pi: dynamic model catalog, reasoning-effort thinking levels, vision, session usage + monthly billing footer
vision$ pi install npm:pi-deepinfraThe terminal previewer for pi — tabbed side drawer for markdown, syntax-highlighted code, images, and PDFs (text + visual pages).
vision$ pi install npm:@iamps6/pi-lensAtomic image placeholders for pi 0.85.0 on macOS
vision$ pi install npm:pi-image-placeholderMultimodal perception for Pi — image/audio/video/document understanding, transcription & image generation via Gemini (gemini_api or local antigravity CLI).
vision$ pi install npm:pi-gemini-multimodalOllama (local + Ollama Cloud) and LM Studio providers for the pi coding agent - live model discovery with real capabilities (tools/vision/thinking) and context windows.
contextvision$ pi install npm:pi-lm-providersPi extension for Cloudflare Workers AI — chat, image generation, TTS, and speech-to-text over the free tier
vision$ pi install npm:pi-cloudflare-workers-ai- @d-ai/piextension
Pi extension for D.AI image, video, and music generation
vision$ pi install npm:@d-ai/pi Generate and edit images in Pi using existing Codex and Grok subscription account quotas.
vision$ pi install npm:@specode/pi-subscription-imageModel-profiled Codex Responses tools for Pi: WebSocket, Lite, web search, images, compaction, and apply_patch
webvision$ pi install npm:@oai404iao/pi-codex-minimal-toolsDisplay images inline in the pi coding agent: native protocol (iTerm2/Kitty/WezTerm/Ghostty), Sixel in Windows Terminal, ANSI half-blocks everywhere else (works under Herdr/tmux).
vision$ pi install npm:pi-imgcat- pi-image-pasteextension
Pi extension that turns pasted image paths into first-class image attachments with automatic optimization.
vision$ pi install npm:pi-image-paste Lean Z.ai tools for pi: URL reader, GitHub zread lookup, and image/video vision. No search tool (use pi-web-access). Minimal token footprint.
webvisiongit$ pi install npm:zai-tools-lite- pi-warp-kitty-imagesextension
Enable Kitty image protocol for Warp terminal in pi's TUI
vision$ pi install npm:pi-warp-kitty-images Pi extension for GMI Cloud Inference Engine — OpenAI-compatible chat, vision (VLM), and embeddings
vision$ pi install npm:pi-gmicloud-aiIgnition 8.3 gateway + project assistant — file-based resource authoring, REST scan/hot-reload, tag scaffolding, screenshot-verified Perspective UI
vision$ pi install npm:sylo-ignition- pi-lemonade-linkextension
Unified pi extension for a self-hosted Lemonade server: dynamic chat-model discovery, multimodal agent tools (transcription, image gen/edit/upscale, TTS, audio, 3D mesh), a live /lemonade-setup management TUI, and a lemonade-only below-editor status bar (
vision$ pi install npm:pi-lemonade-link Native pi agent image generation and editing through OpenAI-compatible Images APIs.
vision$ pi install npm:@deqiying/pi-image-genPi skill + subagent: three parallel code-review channels (OCR / Standards / Spec) reviewing the same diff bound under one scope, then an oracle deduplicates, resolves conflicts and re-grades them into a single adjudicated report (00-adjudication.md). Pure
subagentvisionskill$ pi install npm:pi-multi-code-reviewBrowser automation skill — drive your real Chrome via the browser-use CDP CLI (navigate, screenshot, coordinate-click, extract, tabs). Penetrates closed shadow-DOM & cross-origin iframes via compositor-level coordinate clicks where js selectors fail.
browservisionskill$ pi install npm:@getpipher/browser- @maheidem/pi-auto-image-attachextension
Pi extension: auto-attaches pasted/dragged image file paths as ImageContent blocks so the model actually sees the image.
vision$ pi install npm:@maheidem/pi-auto-image-attach Image gallery, remote URL caching, multimodal delivery, and macOS Quick Look for Pi
vision$ pi install npm:@cqh6666/pi-image- @ganziliang/zhizh-pi-ai2imageextension
Image generation and editing for Pi Agent through the Zhizhengroup LLM Gateway (generate_image tool)
vision$ pi install npm:@ganziliang/zhizh-pi-ai2image pi coding agent extension: zero-switch image routing for text-only models — transcribes image blocks through a configurable VLM chain before every LLM call (parallel, deduped, cached, with fallback)
vision$ pi install npm:pi-vision-route- @ghoulm370/pi-zai-visionextension
Pi extension: Z.AI GLM-4.6V vision tools — image analysis, OCR, error diagnosis, diagram reading, UI diff, UI-to-code, video analysis
vision$ pi install npm:@ghoulm370/pi-zai-vision Tmux-safe, TUI-only local image previews for Pi
vision$ pi install npm:pi-tmux-imagesPi extension: review and test JS frontends headlessly — open dev-server pages in Playwright Chromium, interact with them, capture screenshots the model can see, and collect console/network errors.
browservision$ pi install npm:pi-frontend-check