Let your AI agent see and operate a real Android phone: phone_look (vision + UI-tree fusion), tap/swipe/type, screenshot — over adb, for any MCP client.
Install
# from GitHub (first run asks for allowBuilds approval — follow the hint, retry)
dsh plugin --profile web add github:boheastill/phone-eye
Any plugin you install runs third-party code with your own permissions — it can read your files, use your credentials, and reach the network, and tool approvals don’t sandbox it. GitHub-sourced plugins also run build scripts at install time — pnpm blocks those until you allow them, so an install can stop with ERR_PNPM_GIT_DEP_PREPARE_NOT_ALLOWED or ERR_PNPM_IGNORED_BUILDS; dsh prints the exact key to add under allowBuilds in your profile’s pnpm-workspace.yaml, and the install works on the next run. Allowing a build is a trust decision: only install sources you trust, and pin a commit (github:owner/repo#sha).
README
English | 简体中文
Your AI agent can finally see and operate a real Android phone.
phone_look("what's on screen, where is the login button?")
→ "Login button at (540, 1830) — a green 'Sign in' …"
phone_tap(540, 1830)
phone_look("did the next page load?")
Works with any MCP client — Claude Code, Codex, Cursor, dsh, and friends.
What you need (plain words)
| You provide | One-time or every time? | How hard? |
|---|---|---|
| An Android phone with USB debugging on | one-time, ~60 seconds (tap "Build number" 7× → enable USB debugging) | easy, step-by-step below |
| Plug the USB cable once | one-time — after that the tool switches the phone to Wi-Fi and you never need the cable again (unless the phone factory-resets) | trivial |
| A computer with Python 3.10+ and git (or just Docker — see step 3) | — | n/a |
| A vision model | one-time setup — bring any one of these:• an OpenAI API key (or any OpenAI-compatible: GLM, DeepSeek, local llama.cpp/Ollama…)• an existing MCP vision server | one env var, most people already have a key |
That's everything. No app to install on the phone. No root. No extra server.
What it can and can't do (honest table)
| ✅ Stable | ⚠️ Works but with caveats | ❌ Not possible (any tool, not just us) |
|---|---|---|
| See the screen (screenshot + read it) | Locked screens can be read but not operated | Fully control a brand-new phone before you enable USB debugging yourself |
| Tap / swipe / type | Typing is ASCII (Chinese input needs a clipboard trick — known adb limit) | The very first "allow USB debugging?" popup on a new computer key — that one tap is yours |
| Survive Wi-Fi adb drops (auto-reconnect) | Some vendor ROMs restrict input on lock screens (e.g. MIUI) | iOS — different universe |
| Run 24/7 unattended; unexpected popups get read & handled by your agent | Vision quality depends on the model you bring | |
Multiple phones (one phone-eye process per phone, set ANDROID_SERIAL for each) |
The 60-second rule: every Android requires one human moment — enable debugging + authorize once. After that, the phone belongs to your agent, even over Wi-Fi, even after reboots of the computer.
Setup (3 steps)
1. Enable USB debugging (once per phone)
Settings → About phone → tap Build number 7× (unlocks Developer options) → Developer options → USB debugging ON. (Got stuck? Tell us your phone model in Discussions — we'll walk you through it. MIUI/HyperOS may also ask you to sign into a Xiaomi account first.)
2. Install adb (if you don't have it), plug USB once, then go wireless
# macOS: brew install android-platform-tools · Ubuntu/Debian: sudo apt install adb
# Windows: scoop install adb (or download Android platform-tools)
adb devices # phone shows up? tap "Allow" on its popup — check "always allow"
adb shell ip route # ← note the phone's Wi-Fi IP (e.g. 192.168.1.23) while still plugged
adb tcpip 5555 # switch to Wi-Fi mode (adb restarts; the USB entry disappears — normal)
adb connect 192.168.1.23:5555 # use the IP from above; then unplug the cable, forever
Settings → Wi-Fi → your network → details shows the IP; or adb shell ip addr show wlan0 | grep inet.
3. Start phone-eye
git clone https://github.com/boheastill/phone-eye && cd phone-eye
pip install -r requirements.txt
# your eyes — pick ONE:
export PHONE_EYE_VISION_API_KEY=<key> # OpenAI / GLM / any compatible
# (optional: PHONE_EYE_VISION_BASE_URL, PHONE_EYE_VISION_MODEL)
# local & offline: ..._API_KEY=sk-noauth ..._BASE_URL=http://<host>:8080/v1 ..._MODEL=<your qwen-vl>
python server.py # stdio MCP server — wire into your client:
Wire it into your client — pick yours:
# Claude Code (easiest):
claude mcp add phone-eye -- python /path/to/phone-eye/server.py
# Codex:
codex mcp add phone-eye --url stdio://python /path/to/phone-eye/server.py # or see codex docs
// any MCP client (generic stdio shape):
{ "mcpServers": { "phone-eye": { "command": "python", "args": ["/path/to/phone-eye/server.py"] } } }
Docker (optional — no Python needed on the host)
The repo ships a Dockerfile (Python 3.12 + adb):
podman build -t phone-eye . # or: docker build -t phone-eye .
# smoke: a JSON-RPC initialize reply on stdout means it boots:
printf '{"jsonrpc":"2.0","id":1,"method":"initialize","params":{"protocolVersion":"2025-03-26","capabilities":{},"clientInfo":{"name":"smoke","version":"0"}}}\n' \
| podman run --rm -i phone-eye
Use --network host so adb reaches a Wi-Fi phone and your vision endpoint:
"phone-eye": { "command": "podman", "args": ["run","--rm","-i","--network","host",
"-e","ANDROID_SERIAL=192.168.1.23:5555","-e","PHONE_EYE_VISION_API_KEY=<key>","phone-eye"] }
Recommended vision models (any vision-capable chat model works): gpt-4o-mini (default),
GLM glm-4.6v-flash (cheap), or run a local Qwen-VL via llama.cpp for fully-offline —
screenshots then never leave your LAN.
What happens when something breaks
- "No Android device reachable" → the tool already tried reconnecting; run
adb connect <ip>:5555, or replug USB. - "No vision server reachable" → you haven't set a key; the error message tells you the exact two fixes.
- Phone rebooted → Wi-Fi adb survives phone reboots on most ROMs; if not, one
adb connectagain. - Still stuck? Open a discussion — we answer, and we'll debug your setup with you. Bug reports and "it works on my X" notes are equally welcome.
Tools
| Tool | What it does |
|---|---|
phone_look(question?) |
Ask a vision model about the live screen; fuses a UI-tree dump for exact text/button bounds |
phone_tap(x, y) |
Tap |
phone_swipe(x1, y1, x2, y2, ms?) |
Swipe |
phone_type(text) |
Type ASCII text |
phone_key(key) |
Press a hardware key — wake revives a sleeping phone (the unattended essential), back/home/recents navigate |
phone_intent(action, uri?, component?) |
Open any screen by Android intent (deep settings pages, app pages) without coordinates |
phone_screenshot() |
Save screenshot to disk, return path |
Examples
- Mobile web QA loop — the agent verifies its own work on a real screen
- Surviving an OEM setup wizard — vision handles whatever pops up
- Form regression check
- Unattended sentinel — your agent on night watch: wake → unlock → intent → look → screenshot
Curious how it works — or want to modify it? Read the whitepaper (architecture, failure-classification decision tree, security model, how to add a verb). Running it as an always-on HTTP service behind your own fleet? docs/fleet.md.
Why "see", isn't this just adb?
UI-tree-only tools are blind to game canvases, images, and anything the accessibility tree can't show. Vision-only tools drift on coordinates. phone_look fuses both: the model answers what is this, the UI tree supplies exactly where. The day we dogfooded it, it discovered a USB-debugging popup on its own screen, read the buttons, and tapped "Allow" by itself — the story.
License
MIT. Verified on Redmi K40 Gaming / Android 13 — add your device to the table via PR.
Author: Bohea — independent industrial software engineer (Shenzhen). Operator HMIs · device integration · machine data into your customer's ERP. More runnable demos & engineering notes on the site.
Links
More in this category
liustack/modlens★ 4192
Vision bridge for text-only models: paste an image, get structured JSON evidence (OCR, layout, semantics).
ysr666/dsh-vision-router★ 1141
Free vision for text-only agents: built-in keyless vision chain plus pixel tools (Q&A, grounding, crop, pixel diff, colors, OCR, SVG trace, cutout, screenshots); paste an image to use it.
Anionex/dsh-vision-toolkit★ 887
Vision for text-only models: paste an image and the model switches to a Vision Toolkit variant for image Q&A, multi-image comparison, long-screenshot OCR, screenshot-to-UI reproduction, element grounding, and pixel diff. No API key by default — images are processed by the author-hosted free service, 100 per machine per day; configurable to your own provider.
fandc520/dsh-comfyui★ 105
Drive a local or remote ComfyUI server from DeepSeek Harness: comfyui_run / comfyui_object_info / comfyui_workflow tools generate and edit images and videos, with a workflow library (graph extraction: per component / main flow / all), a load area with resolution auto-match, a live queue, SDXL and Wan 2.1 templates, a companion skill, and a same-origin media proxy.
dickpy/dsh-imagegen★ 104
AI image generation for the DSH Web GUI: text-to-image and image-to-image through a configurable OpenAI-compatible endpoint (gpt-image-2 / gpt-image-1 / dall-e-3), with an api_url/api_key settings card and a sidebar split-pane generation studio.
sunxin-ai/dsh-design-qa★ 44
Design-fidelity QA for text-only models: a `deepseek_vision` tool borrows an eye from any OpenAI-compatible vision route, so the model can judge whether an implementation matches its mock — shipped with the benchmark behind that judgement (four fixtures, 23 injected defects, raw transcripts) and the questioning discipline it depends on.
Community comments
Comments are public GitHub Discussions. Loading them connects to GitHub and Giscus; a GitHub account is required to post.