-
Vision bridge for text-only models: paste an image, get structured JSON evidence (OCR, layout, semantics).
Vision & MultimodalInstall ▾
-
Free vision for text-only agents: built-in keyless vision chain plus pixel tools (Q&A, grounding, crop, pixel diff, colors, OCR, SVG trace, cutout, screenshots); paste an image to use it.
Vision & MultimodalInstall ▾
-
Vision for text-only models: paste an image and the model switches to a Vision Toolkit variant for image Q&A, multi-image comparison, long-screenshot OCR, screenshot-to-UI reproduction, element grounding, and pixel diff. No API key by default — images are processed by the author-hosted free service, 100 per machine per day; configurable to your own provider.
Vision & MultimodalInstall ▾
-
AI image generation for the DSH Web GUI: text-to-image and image-to-image through a configurable OpenAI-compatible endpoint (gpt-image-2 / gpt-image-1 / dall-e-3), with an api_url/api_key settings card and a sidebar split-pane generation studio.
Vision & MultimodalInstall ▾
-
Routes images to a vision model of your choice - auto-rewrite, explicit tools, or hybrid - so a text-only chat model does not fail a turn that contains a picture.
Vision & MultimodalInstall ▾
-
Upgrades DSH image viewing with pointer-centered zoom, pan, galleries, downloads, keyboard support, and inline region notes.
Vision & MultimodalInstall ▾
-
Drive a local or remote ComfyUI server from DeepSeek Harness: comfyui_run / comfyui_object_info / comfyui_workflow tools generate and edit images and videos, with a workflow library (graph extraction: per component / main flow / all), a load area with resolution auto-match, a live queue, SDXL and Wan 2.1 templates, a companion skill, and a same-origin media proxy.
Vision & MultimodalInstall ▾
-
Multi-engine text-to-image generation (OpenAI Images and Zhipu CogView presets) with per-session quota tracking, engine failover, credential-safe config, and a result card with regenerate.
Vision & MultimodalInstall ▾
-
Send local images or HTTP(S) image URLs to a configured OpenAI-compatible vision endpoint with the inspect_image tool and return its text response to the conversation.
Vision & MultimodalInstall ▾
-
Local OCR for attached images via the built-in Windows engine (Windows.Media.Ocr): only the recognized text is sent to the model, never the image bytes; vision passthrough is opt-in.
Vision & MultimodalInstall ▾
-
Lets text-only models handle pasted chat images, with a native vision experience, batch image viewing, and a built-in OpenAI-compatible analyze_image tool; vision-capable models are unaffected.
Vision & MultimodalInstall ▾
-
Free vision bridge for text-only models: image understanding, OCR, UI and debug analysis via free-tier providers (Qwen3-VL-Flash, Doubao, DeepSeek-OCR) with a settings GUI.
Vision & MultimodalInstall ▾
-
Local OCR for attached images via Tesseract: only the recognized text is sent to the model, never the image bytes; vision passthrough is opt-in.
Vision & MultimodalInstall ▾
-
DeepSeek brain + automatic image transcription: attach images in the GUI and each one is transcribed via the official deepseek-v4-flash-vision-exp by default (a pure-text V4-Pro brain can see images), with any OpenAI-compatible VLM or local Ollama as alternatives.
Vision & MultimodalInstall ▾
-
A vision-language gateway provider route: pasted images are described by a configurable VL model (Qwen-VL by default) before the DeepSeek wire.
Vision & MultimodalInstall ▾
-
Configurable image recognition for text-only DSH models: image messages are first transcribed by an OpenAI-compatible vision model (Base URL, model ID and API key set in a Settings section) and then passed to the main model as text, while image-input support is advertised. The API key auth scheme is selectable — OpenAI, Anthropic, Gemini or Azure style request headers — and a missing key is reported before any request is sent.
Vision & MultimodalInstall ▾
-
Image "reading" for text-only models: downscale + reduce color depth + structure/color fingerprints into text grids fed back to the conversation, letting the model zoom, sample and OCR autonomously like a multimodal model; fully local with zero external model dependency, ships an image-reading methodology skill and optional PaddleOCR.
Vision & MultimodalInstall ▾
-
Native image attachments for text-only DeepSeek in the Web GUI: pasted or dropped images appear as thumbnails in the session, and before dispatch the host reads them with the free Zhipu GLM-4V-Flash vision API (glm-4v-flash fallback chain) and substitutes the description, so DeepSeek answers about the image while the original is kept in history.
Vision & MultimodalInstall ▾
-
Codex Appshots for DSH: capture the frontmost active window via global shortcut and seamlessly mount it into the composer for agent queries.
Vision & MultimodalInstall ▾
-
Codex-style window capture for DSH Desktop on macOS and Windows: press both Command keys (macOS) or both Ctrl keys (Windows), or the camera button, to grab the frontmost window, attach it to the current chat, and inject a VoiceOver-style accessibility tree as hidden context.
Vision & MultimodalInstall ▾
-
Turn a phone camera into a live viewfinder and photo input for your dsh session, over LAN or USB.
Vision & MultimodalInstall ▾
-
Local OCR fallback for text-only routes: when the session model declares it cannot accept images, the attached image is cached locally and its path injected so the model can call ocr_image — PP-OCRv5 + ONNX Runtime on CPU, no API key, images never leave the machine. Silent when the model can see images.
Vision & MultimodalInstall ▾
-
Analyzes conversation images through mode-specific prompts (caption, UI, document, grounding, topology, etc.) and injects structured evidence with coordinate primitives (boxes, points, refs) as text, with session-level caching for reuse across replay and compaction.
Vision & MultimodalInstall ▾
-
Vision for text-only DeepSeek via Doubao Web by default (zero-cost, no API key — drives your logged-in Chrome through a Windows CDP bridge), with Antigravity IDE quota (flash/pro) or Gemini fallback; auto detail escalation, vision evidence memory with compaction rehydration, content-hash cache, and a bilingual client panel.
Vision & MultimodalInstall ▾
-
Zero-dependency screen capture for DSH: Lightweight — zero deps, zero binaries; Stage & shoot — one-click full screen, window layout, hover-snap capture of occluded windows; Agent self-service — path-only delivery; paths are universal, pair with modlens (optional) for one-call structured evidence.
Vision & MultimodalInstall ▾
-
Combine text, vision, and image-generation APIs into one Mix model with automatic routing: text-only requests go to the chat model, user images and agent screenshots go to the vision model, follow-ups keep using the same session image, and agents can generate or edit images with session-scoped call history.
Vision & MultimodalInstall ▾
-
Configure multiple image providers in Settings and call image_generate with the one selected model; images save under generate/image and show inline in the conversation.
Vision & MultimodalInstall ▾
-
Lets a text-only DeepSeek agent read images in the same session by delegating to a vision-capable subagent, with send-time image-to-path conversion.
Vision & MultimodalInstall ▾
-
Vision toolkit for text-only DeepSeek: model-invokable `vision` tool, wrapper adapters for deepseek/opencode-go (v4 flash/pro), Antigravity IDE quota (default, flash/pro) / any OpenAI-compatible VLM / Gemini / local Ollama channels, evidence memory with compaction rehydration, content-hash cache, and a bilingual client panel.
Vision & MultimodalInstall ▾
-
Digests the images in a prompt with a vision model before admission, so any model — a text-only one included — can read them without the session ever switching models, and adds a describe_image tool for image paths.
Vision & MultimodalInstall ▾
-
Labnana image generation for DeepSeek Harness: text-to-image / image-to-image / precise editing with credits estimation, subscription balance and web settings UI.
Vision & MultimodalInstall ▾
-
Plug-in vision for text-only models on DSH, with native interaction for image understanding and generation, GUI automation, through layered evidence memory and cache.
Vision & MultimodalInstall ▾
-
Design-fidelity QA for text-only models: a `deepseek_vision` tool borrows an eye from any OpenAI-compatible vision route, so the model can judge whether an implementation matches its mock — shipped with the benchmark behind that judgement (four fixtures, 23 injected defects, raw transcripts) and the questioning discipline it depends on.
Vision & MultimodalInstall ▾
-
Local image, voice, music and SFX generation plus transcription, with pinned identity: characters, animals, objects and actor voices are defined once and reused on every later call, degenerate output (a near-flat image, silent audio) is rejected instead of returned as success, and a generated line can be read back as text so a clone that swallowed its ending becomes visible. The engines unload when idle, and image generation and transcription can each be pointed at an OpenAI-shaped API instead of the local Vulkan backend.
Vision & MultimodalInstall ▾
-
Hands repetitive text and vision labor (OCR, image analysis, comparison) to a local KoboldCpp (llama.cpp) server through koboldcpp_run and koboldcpp_vision tools, with on-demand server lifecycle management.
Vision & MultimodalInstall ▾
-
Vision provider route that transcribes attached images to text through a configurable model (15+ OpenAI-compatible and Anthropic vendors) while DeepSeek keeps answering.
Vision & MultimodalInstall ▾
-
Model-facing `vision` tool for DeepSeek Harness: describe and OCR image files by calling the free Zhipu GLM vision API directly (glm-4v-flash fallback chain), no external CLI required.
Vision & MultimodalInstall ▾
-
Registers a `deepseek-vision` provider route: the Web GUI accepts pasted images and transcribes them to text via the free Zhipu GLM vision API before delegating to the DeepSeek adapter.
Vision & MultimodalInstall ▾
-
Full vision-capability bundle for DeepSeek Harness: a vision_understand tool (OpenAI-compatible vision APIs, free Zhipu GLM-4V-Flash by default) plus paste/drag-and-drop/button entry points for image recognition.
Vision & MultimodalInstall ▾
-
Hands repetitive text and vision labor (OCR, image analysis, comparison) to a locally running Unsloth Desktop (Unsloth Studio) server through unsloth_run and unsloth_vision tools; pure HTTP client, never spawns or owns processes.
Vision & MultimodalInstall ▾
-
Auto-discovery vision bridge for text-only DeepSeek Harness agents: automatically finds an image-capable model from your configured providers and returns picture descriptions as plain text via a vision tool.
Vision & MultimodalInstall ▾
-
Vision for any DSH route: paste images in the Web composer with intent-aware auto-analysis, delegate workspace image reads to a Kimi/MiniMax vision subagent, and materialize pasted originals for editing.
Vision & MultimodalInstall ▾
-
Routes chat images to a fixed OpenAI-compatible vision model, returns factual observations to the selected main model, and reuses session-scoped observations across replay, compaction, and restarts.
Vision & MultimodalInstall ▾
-
Image-to-text input for the Web UI: paste or drag an image and it is transcribed into structured text and sent, giving text-only LLMs image-input takeover (OpenAI-compatible vision API).
Vision & MultimodalInstall ▾
-
Renders pasted/uploaded images inline in the DSH Web chat and gives text-only models vision: the model-invokable visual_review tool calls any OpenAI-compatible multimodal API first, falling back to a local Qwen3-VL worker.
Vision & MultimodalInstall ▾
-
Transparent image guard for text-only routes: paste images without the 400 session deadlock, plus a vision_analyze tool for OCR/PDF/docx/pptx/video.
Vision & MultimodalInstall ▾
-
Vision MCP and DSH bundle for text-only DeepSeek: analyze_image, analyze_clipboard, compare_images and vision_status tools, a visual settings page, free GLM-4.6V-Flash by default, result caching and rate-limit tolerance; keys stay out of logs.
Vision & MultimodalInstall ▾
-
Two-tier image reading for text-only models: fast local OCR (RapidOCR, offline) first, vision-model fallback (modlens).
Vision & MultimodalInstall ▾
-
`describe_image` tool: a vision bridge that sends images to mimo-v2.5 through the opencode Zen API (credential `OPENCODE_GO_API_KEY`, free route first with paid fallback) and returns text descriptions for text-only models, with native passthrough and ImageMagick transcoding of SVG/TIFF/HEIC formats.
Vision & MultimodalInstall ▾
-
Adds a take screenshot tool plus two optional automatic capture points, putting the screen into the conversation as an image block; requires an image-capable model.
Vision & MultimodalInstall ▾
-
Native LLM-provider vision bridge: images pasted in the chat are described by a vision model (Qwen3-VL via pi-ai/llama.cpp) and the text description is fed to text-only DeepSeek for the reply — image admission, routing and compaction all run through harness-native mechanisms, with an LRU description cache and 503 retry.
Vision & MultimodalInstall ▾
-
Free vision OCR with adaptive tile recognition for long documents and Markdown/Word/PNG/Excel export.
Vision & MultimodalInstall ▾
-
Automatically generates and displays images in the DSH chat via API channels or local CLIs (mmx / codex / agy), and can also recognize images using the corresponding CLI.
Vision & MultimodalInstall ▾
-
Vision bridge for text-only DeepSeek models: transcribes attached images with deepseek-v4-flash-vision-exp before they reach deepseek-v4-pro, with no third-party dependencies.
Vision & MultimodalInstall ▾
-
Gives dsh the ability to generate images and videos through the grok2api API.
Vision & MultimodalInstall ▾
-
Omni-modal workstation for DSH: analyze_image over an ordered multi-card VLM failover chain, a six-tool vision toolkit (zoom, colour sampling, pixel diff, OCR, element detection, inline display) sharing the same image resolver, generate_image over OpenAI/DashScope/ComfyUI protocols, multi-card async generate_video with a /build-video-tool builder, and speak/clone_voice TTS over seven providers, all driven by one auto-saving settings page.
Vision & MultimodalInstall ▾
-
Vision for text-only dsh models: paste an image and a configured multimodal model transcribes it to text automatically — transparent twin routing, an agent-callable read-image tool, no built-in keys or relay.
Vision & MultimodalInstall ▾
-
Trims historical images in outgoing chat requests down to a recent-image count, learns the provider per-prompt image cap from HTTP 400 responses, and retries with fewer images so image-heavy sessions keep working.
Vision & MultimodalInstall ▾
-
Scene-awareness probe for DSH on macOS: a privacy-first, pull-model `look` tool that reads the frontmost non-self window (app info, AX title/selection, gated local OCR), plus a summon-hotkey snapshot captured the instant you press.
Vision & MultimodalInstall ▾
-
Gives a text-only model image understanding via a vision-language model, exposed as a `describe_image` tool.
Vision & MultimodalInstall ▾
-
Free vision bridge and image generation for text-only models: paste-image reading, GLM-4V-Flash and Gemini engine failover, ModLens-style structured evidence, and a seeded free vision model route.
Vision & MultimodalInstall ▾
-
Gives text-only models vision: forwards user images to an OpenAI-compatible vision model and shows the descriptions in a Web UI right panel.
Vision & MultimodalInstall ▾
-
Adds a configurable vision model to text-only main models: a vision_read_image tool, a composer-bar vision-model selector, and automatic image-to-text conversion for text-only routes.
Vision & MultimodalInstall ▾
-
MCP bridge to gemini.google.com: vision analysis of images and videos, Imagen image and Veo video generation, and conversation management using the logged-in browser session with no API key.
Vision & MultimodalInstall ▾
-
AI image and video studio as a floating DSH panel, so creation runs alongside the conversation instead of replacing it: text-to-image, image-to-image and multi-image composition at 1K-4K across eight aspect ratios; text-to-video and image-to-video with first-frame control at 4-12 seconds; short-drama mode that imports a script (.txt/.md/.json), breaks it into storyboard shots for preview and batch generation; and a prompt-expert workspace. Zero runtime dependencies.
Vision & MultimodalInstall ▾
-
External vision plugin for DeepSeek Harness: whale-button config panel, image recognition with auto-reply, and agent screenshot/recognize tools.
Vision & MultimodalInstall ▾
-
For the DeepSeek Harness native vision model deepseek-v4-flash-vision-exp, raises image admission limits to 32 MiB / 8192 px / 600 images and adds a highres_read tool that tiles large images, then returns the whole image plus 800x800 tiles through the host read_image tool.
Vision & MultimodalInstall ▾
-
MiniMax-powered multimodal plugin: real-time voice call mode (streaming conversation, floating dock UI), voice mode and mic voice input, plus image/video/music/speech generation and vision inspection tools.
Vision & MultimodalInstall ▾
-
Let text-only models see images: when a picture arrives with a placeholder like \[image omitted because this model accepts text only], the model calls the vision_describe tool and the plugin forwards the image reference plus the question to a multimodal model, retrying with a fallback model on failure. Configured in the settings page and persisted via the official settings API.
Vision & MultimodalInstall ▾
-
A content-aware PDF reading plugin for vision models: profiles each page for figures (vector and raster), tables, formula risk and double-column layout, then applies content-aware hybrid extraction, rendering figure/table/formula pages as high-DPI region crops. Provides a low-resolution preview to understand the page layout, and renders a specified region at high resolution. Packaged as multiple tools for agents.
Vision & MultimodalInstall ▾
-
Bridges session images to configurable vision providers and returns text-only analysis to eligible DeepSeek Harness model routes.
Vision & MultimodalInstall ▾
-
Fully-local image understanding & OCR via macOS Vision Framework: `ocr_image` (text, table layout + coordinates) and `view_image` (scene, faces, QR) — paste multiple images into the web input box or pass path/URL/base64; images never leave your Mac.
Vision & MultimodalInstall ▾
-
Unified image generation router for DeepSeek Harness (DSH): auto-discovers image models from any OpenAI-compatible endpoint, provides draw_image and draw_list_sources tools, supports SenseNova, StepFun, Agnes, Qwen, Flux, SD, Imagen and more.
Vision & MultimodalInstall ▾
-
Enhanced vision toolbox: 14 pixel-level vision tools (describe, ground, detect, crop, pixel-diff, OCR, long-screenshot OCR, vectorize, colors, cutout, screenshot, present, materialize, html-screenshot) driven by one OpenAI-compatible endpoint, with clean \[图片: path] bridge markers, content-safety classification and rate-limit auto-retry.
Vision & MultimodalInstall ▾
-
Let your AI agent see and operate a real Android phone: phone_look (vision + UI-tree fusion), tap/swipe/type, screenshot — over adb, for any MCP client.
Vision & MultimodalInstall ▾
-
GPT Image 2 `image_gen` with Codex subscription OAuth by default or explicit API-key mode: developing card, up to three live API partials, durable attachment replay/lightbox/download, text-only model output, and bounded credential-safe requests.
Vision & MultimodalInstall ▾
-
Control a HarmonyOS phone from the DSH web UI with AI: live H.264 screen mirroring, mouse touch and system keys, hilog streaming, and agentic tools that let the model read the screen, locate UI controls, then tap, long-press, press keys, or type.
Vision & MultimodalInstall ▾
-
Advertise image paste on text-only DeepSeek routes, describe attachments with a vision sidecar, and leave native vision models untouched.
Vision & MultimodalInstall ▾
-
Agent-callable vision tool that describes local images via any OpenAI-compatible vision endpoint you configure, with an optional multi-model cross-check and no built-in keys.
Vision & MultimodalInstall ▾
-
Image understanding for any DSH model: vision, OCR, grounding, and crop tools with domain presets for histopathology, cell biology, anatomy, clinical images and scientific figures.
Vision & MultimodalInstall ▾
-
Vision bridge for text-only DeepSeek: send images (alone or mixed with text) and a fast tiny vision model describes them behind the scenes — the chat keeps the picture, DeepSeek sees only text. Pure plugin: uninstall restores everything.
Vision & MultimodalInstall ▾
-
Generate images through a logged-in web ChatGPT session and download them to a local directory.
Vision & MultimodalInstall ▾
-
Vision bridge for text-only DeepSeek routes that analyzes attached and local images through configurable OpenAI Responses, Chat Completions, or Anthropic Messages endpoints while leaving image-capable routes native.
Vision & MultimodalInstall ▾
-
Structured vision evidence (OCR/layout/semantics) plus a USB camera capture tool for a "shoot-look-adjust" debug loop; backends: Ollama, DeepSeek, Xiaomi.
Vision & MultimodalInstall ▾
-
Dual-engine screen capture for DSH: inside the SSiD desktop shell, a global hotkey or tray opens a fullscreen box-select overlay (all monitors, per-display pixel-perfect frames) with in-place red-box annotation and a WeChat-style toolbar; in plain DSH (browser), the composer camera button captures the display via getDisplayMedia and reuses the same single-phase box-select + annotation overlay in-page. The cropped image lands in the current conversation composer through the official attachment intake.
Vision & MultimodalInstall ▾
-
Native vision capability extension, using either Zhipu (free) or Qwen-VL (local).
Vision & MultimodalInstall ▾
-
Composer-attached images are transcribed to text by an OpenAI-compatible vision model before reaching text-only DeepSeek models.
Vision & MultimodalInstall ▾
-
Screen capture and external vision recognition: take_screenshot, list_windows, analyze_image and view_image tools with a configurable GPT vision channel (gpt-5.5 / gpt-5.6-sol / gpt-5.6-terra), API key via the credentials service, and a settings card; view_image shows the screenshot in the Web UI while the model context keeps text only.
Vision & MultimodalInstall ▾
-
Reuses DeepSeek web's built-in vision mode for text-only models: the deepseek_vision tool drives the local deepseek-vision-cli browser automation (manual login helper, deep-think enabled, auto-closes browser) and returns image descriptions as text.
Vision & MultimodalInstall ▾
-
Synesthesia Encoder for DSH: a vision model translates images into compact structured spatial text (canvas/elements/percentage coordinates), giving text-only LLMs pixel-level image understanding via the `mm_vision` tool.
Vision & MultimodalInstall ▾
-
Local-first structured vision for text-only agents: images go to a local OpenAI-compatible VLM and come back as JSON evidence (summary, verbatim OCR, layout regions, entities/relations, colors, explicit uncertainty), with anti-hallucination fallback and an optional paste/upload bridge; zero cloud cost, images never leave the machine.
Vision & MultimodalInstall ▾
-
DeepSeek Harness vision plugin: 8 analysis modes (describe, OCR, chart data, UI review, object detection, compare, code-gen, debug), any OpenAI- or Anthropic-compatible vision API, with a built-in free vision model and automatic rate-limit failover.
Vision & MultimodalInstall ▾
-
Transparent multimodal routing for text-only models: every image in every model call is fully transcribed (verbatim OCR, data, uncertainty zones, injection-hardened) by your own multimodal understander, with focused re-look via vision_relook and automatic retry on image-related failures. No bundled endpoints, no borrowed logins.
Vision & MultimodalInstall ▾
-
On-demand vision for text-only DeepSeek models: upload images, and the model calls a view_image tool backed by any OpenAI-compatible vision endpoint (Qwen/DashScope by default).
Vision & MultimodalInstall ▾
-
Vision-augmented DeepSeek adapter: a vision-capable model describes image input, then a text-only DeepSeek model reasons over the description.
Vision & MultimodalInstall ▾
-
Integrated visual creation toolchain for DSH: prompt reverse-engineering and audit optimization plus multi-provider image generation with async background rendering.
Vision & MultimodalInstall ▾
-
Model-facing image_describe (识图) tool over the DashScope OpenAI-compatible API (qwen3.7-flash), plus a paste bridge: on text-only sessions, pasted images auto-convert to file paths at send time and render back in the transcript, so they never trip image admission. Bring your own DASHSCOPE_API_KEY; endpoint/model/budgets configurable, redirect-proof HTTP client, works in every agent preset.
Vision & MultimodalInstall ▾
-
Chat image-attachment bridge with a `view_image` tool for any OpenAI-compatible VLM (local Ollama or cloud): pasted/dropped images become `view_image` path markers before reaching text-only DeepSeek models.
Vision & MultimodalInstall ▾
-
Adds image input and recognition through configured DSH providers or an OpenAI-compatible endpoint.
Vision & MultimodalInstall ▾
-
Model-facing image-generation tool with a configurable channel and normalized image parameters.
Vision & MultimodalInstall ▾
-
Give DSH text-only models vision: an image/OCR/document recognition skill (race pool → custom channels → local) plus an idempotent host patch so image messages reach the model.
Vision & MultimodalInstall ▾
-
Agent screen capture on macOS and Windows: one tool captures the screen and returns the image itself, so the model sees it without a second call. On Windows a whole burst runs in one engine call; on macOS a resident helper drops a region capture to 13ms and a change check to 23ms. A second tool, macOS only, reports whether Screen Recording is granted and opens the settings pane that fixes it.
Vision & MultimodalInstall ▾
-
Apple on-device Vision tools for text-only dsh models: local OCR (zh-Hans + 30 langs), image classification, face detection, document layout, and a combined describe — 100% offline, no API key, tall-screenshot slicing.
Vision & MultimodalInstall ▾
-
Local WeChat OCR tool for DSH: `wechat_ocr_recognize` returns recognized text and the engine structured result for a local image path.
Vision & MultimodalInstall ▾
-
Local Windows 11 OneOCR tool for DSH: `oneocr_recognize` returns OCR text and structured line/word polygons, confidence, rotation, and handwriting style.
Vision & MultimodalInstall ▾
-
Auto-switch the DeepSeek route to the vision model on demand: flash main session switches (A), pro keeps deep reasoning and delegates image reading to a vision subagent (B), subagents always switch, with fatal-failure fallback. No manual model switching.
Vision & MultimodalInstall ▾
-
Turn every idea into an image or video with Kling AI. In DeepSeek Harness, use natural language for text-to-image, image-to-image, text-to-video, image-to-video, reference-image creation, task tracking, and result previews.
Vision & MultimodalInstall ▾
-
Give your text-only model eyes - chat image attachments are auto-described via a vision model (default prompt), with iterative re-parsing through model-generated prompts when details are missing; system/custom model modes + GUI config panel, key-safe secret handling, and a small host patch for DSH 0.1.0-rc.6 (see repo README).
Vision & MultimodalInstall ▾
-
Local eyes for text-only models: eyes_render draws text/shapes/Mermaid onto a canvas in the Web GUI, eyes_paste captures pasted images, eyes_ocr reads text via the built-in Windows OCR (offline), and eyes_analyze inspects pixels as structured data — no vision model required.
Vision & MultimodalInstall ▾
-
Give text-only models vision: a describe_image tool that delegates images to a vision model from your dsh model list over the harness's own LLM runtime.
Vision & MultimodalInstall ▾
Installing
# from npm (prebuilt) dsh plugin --profile web add <npm-package> # from GitHub (first run asks for allowBuilds approval — follow the hint, retry) dsh plugin --profile web add github:owner/repo
Any plugin you install runs third-party code with your own permissions — it can read your files, use your credentials, and reach the network, and tool approvals don’t sandbox it. GitHub-sourced plugins also run build scripts at install time — pnpm blocks those until you allow them, so an install can stop with ERR_PNPM_GIT_DEP_PREPARE_NOT_ALLOWED or ERR_PNPM_IGNORED_BUILDS; dsh prints the exact key to add under allowBuilds in your profile’s pnpm-workspace.yaml, and the install works on the next run. Allowing a build is a trust decision: only install sources you trust, and pin a commit (github:owner/repo#sha).
Get your plugin listed
Open a PR against awesome-dsh-plugin — one YAML file under data/plugins/ is the whole submission; the READMEs and this site regenerate automatically. Add the dsh-plugin topic to your repo too.