Search built for agents: multilingual coverage across web, academic, code, shopping, finance, news, and encyclopedias.
Install
# from npm (prebuilt)
dsh plugin --profile web add argo-search
# from GitHub (first run asks for allowBuilds approval — follow the hint, retry)
dsh plugin --profile web add github:taxueseek/argo
GitHub-sourced plugins run build scripts on your machine at install time. Only install sources you trust, and pin a commit (github:owner/repo#sha).
README
What it is
Argo is multilingual search infrastructure for AI agents.
Real-world retrieval is never “one language + one search box”: someone asks for A-share quotes, someone else asks about the World Cup, someone searches anime in Japanese, someone wants a film director from IMDb. Argo’s premise is simple—route by domain, language, and intent to the right sources, instead of always scraping generic web titles. Web search and local file search work together.
Output is not a “list of links”, but evidence candidates + credibility breakdown. Good routing is what makes evidence stand up.
vs. wrapping yet another search API
| Common approach | Argo |
|---|---|
| Hard-wired to one engine and one key | Multi-engine auto-routing; free first, budget-aware |
| Every query is generic web search | Vertical sources first: markets, film, sports, macro, chemistry… answer-shaped results |
| Optimized only for Chinese/English | Language detection + engine locale params + cross-language fallback |
| Summarize snippets and ship it | Selection × evidence density × freshness × multi-source consensus |
| One dead engine kills the chain | Circuit breakers, negative cache, staged recovery (no vertical cross-contamination) |
| Hit the network every time | Two-layer cache (memory + SQLite); hot queries ~10ms |
| Same slow path for daily and research | Fewer engines day-to-day; open up for deep research |
| Long JSON blows agent context | Compact MCP responses; controllable snippets |
Query-shaped routing
| You ask | What tends to happen |
|---|---|
| 贵州茅台股价 | A-share market domain; snapshot sources first; early-stop when enough |
| AAPL / US pre-market | US equities domain, split from A-shares |
| 肖申克的救赎 主演 / Inception director | Film domain → IMDb etc. |
| 梅西 俱乐部 / 库里 球队 | Sports domain → TheSportsDB etc. |
| 埃菲尔铁塔在哪 / where is Eiffel Tower | Geo entity → OpenStreetMap etc. |
| NASA founding year / 国务院职能 | Org entity → Wikidata etc. |
| 周杰伦 专辑 / Taylor Swift album | Media domain → iTunes etc. |
| アニメ おすすめ / 한국 영화 추천 | Detect JA/KO → language-friendly sources; avoid Chinese-only sites |
| US CPI, China GDP | Macro domain; country split |
| 阿司匹林 分子式 | Chemistry → PubChem-style answers |
| TSMC valuation debate (deep research) | Sub-questions + parallel sources; verticals boosted |
How it works
query
├─ intent clarify (optional)
├─ query rewrite (optional; routing still sees original intent)
├─ language detect + language preference
├─ route (domain rules + TF-IDF + budget + lang supplements + hot-path cache)
├─ multi-engine recall (circuit breaker / negative cache / parallel)
├─ staged empty-result recovery (widen → same family/general → cross-lang; anti-pollution)
├─ RRF fusion + optional re-rank
├─ evidence skim (authority · density · freshness · consensus)
└─ unified JSON (incl. engine_outcomes / recovery)
Evidence scoring (short)
selection ≈ domain authority; SERP / redirect shells ranked very low
absorption ≈ density of numbers / definitions / comparisons / disclosures
freshness ≈ publish time (ignores historical comparison years like “since 2015”)
composite ≈ 0.40·selection + 0.35·absorption + 0.15·freshness + 0.10·engine score
Results include selection, absorption, credibility_fast, evidence_flags, etc. so agents can sort directly.
Agent discipline (recommended)
- High-stakes questions (positions, safety, “is this true?”): search → read fast scores →
fetchtop hits → then conclude - Numbers: state the口径 (definition/scope); when sources conflict, list them—don’t force a merge
- SERP / redirect pages: never treat as primary sources
- Social posts: sentiment and narrative, not ground truth
- Fact-check: prefer a few stratified queries (source / comparison / subject)
Quick start
Pick any path. You do not need the npm registry package for the latest build (since v2.5.1 GitHub is the install source of truth; current recommendation v2.7.2).
Zero-config works: without API keys, free engines + local local_* engines run; keyed engines are skipped when missing (and usually better when present).
Option 1: Install script (best for long-term local use)
curl -fsSL https://raw.githubusercontent.com/taxueseek/argo/main/scripts/install.sh | bash
Custom home + Skill link:
curl -fsSL https://raw.githubusercontent.com/taxueseek/argo/main/scripts/install.sh \
| bash -s -- --home "$HOME/.local/share/argo" --link "$HOME/.claude/skills/argo"
Verify:
python3 ~/.local/share/argo/scripts/search.py "贵州茅台股价" --json
python3 ~/.local/share/argo/scripts/search.py --list-engines
Option 2: MCP from GitHub (fast agent attach)
Needs Node.js 18+ and Python 3.10+. Once:
pip3 install pyyaml
npx -y github:taxueseek/argo
Client config (Claude Code / Cursor / Kimi, etc.):
{
"mcpServers": {
"argo": {
"command": "npx",
"args": ["-y", "github:taxueseek/argo"]
}
}
}
More stable, no Node: install via Option 1, point at local Python:
{
"mcpServers": {
"argo": {
"command": "python3",
"args": ["/path/to/argo/scripts/mcp_server.py"]
}
}
}
Unusual Python path: export ARGO_PYTHON=/path/to/python3 (read by the npx entry only).
Option 3: git clone (dev / patch source)
git clone https://github.com/taxueseek/argo.git
cd argo
pip3 install pyyaml
bash scripts/install.sh --link ~/.claude/skills/argo # optional
python3 scripts/search.py --list-engines
Platforms
| Platform | Integration | Notes |
|---|---|---|
| Claude Code | MCP / Skill link | npx or mcp_server.py; link_source.py ok |
| Kimi / Grok Build | MCP Server | same |
| Cursor / Cline / Continue | MCP | any MCP-capable IDE plugin |
| CLI | search.py / bin/argo |
scripts, cron, manual debug |
| Python projects | from search import super_search |
library call |
Post-install check
python3 --version # 3.10+
python3 -c "import yaml; print('PyYAML OK')"
python3 -m pytest tests/test_unit.py -q # optional
python3 scripts/search.py --list-engines
Capabilities
| Capability | What it does | Entry |
|---|---|---|
| Unified search | route → recall → fuse → skim score | search.py / argo_search |
| Local file search | on-disk code/notes/memory (offline) | argo_local_search |
| Deep research | sub-questions, multi-source, gap hints | research.py / argo_research |
| Credibility | authority / density / freshness / cross-check | evidence.py / argo_evidence |
| Intent clarify | polysemy, brand collisions, strategy hints | clarify.py / argo_clarify |
| Page fetch | HTTP first, browser fallback when needed | argo_fetch (mode=extract for structure) |
| Screenshot / PDF | page shots, structured PDF extract | argo_screenshot / argo_pdf |
| Site crawl | list-page batch crawl | argo_crawl |
| Social / sentiment | Weibo / Xiaohongshu / Bilibili / Reddit / X … | argo_social_search |
Budget modes
| Mode | Best for | Behavior |
|---|---|---|
fast |
simple Q, need speed | free engines first; skip paid re-rank |
auto |
daily default | cost-aware quality/spend tradeoff |
deep |
research, surveys | quality first; more engines allowed |
budget |
tight quota | quota control; degrade when exhausted |
Rough capability set (v2.6.0)
- ~120+ sources, 60+ domains: general web + finance / macro / film / sports / geo / orgs / media / chemistry / academic / code (source of truth:
config.yaml) - 10 MCP tools: search, research, evidence, clarify, fetch, screenshot, PDF, social, local files, crawl
- Multilingual search: Chinese, English, Japanese, Korean, Cyrillic, Thai, Arabic, Hebrew, Greek, Devanagari, …; routing and engine params follow language; non-Chinese queries avoid Chinese-only sources (Zhihu / Sogou WeChat / A-share snapshots, etc.)
- Vertical recovery gates: empty-result recovery will not “leak” pypi / npm / flash news into film or sports
- Faster daily, fuller research:
engine_policytiers—tight daily combo, open long-tail for deep / research
Engines & routing
Config currently has about 120+ sources and 60+ domains (see config.yaml and --list-engines).
Direct & vertical (excerpt)
| Engine | Scenario | Cost bias |
|---|---|---|
| anysearch / duckduckgo | general / tech | free |
| sina_quote / tencent_quote / eastmoney | A-share quotes / flows | free |
| finviz / seeking_alpha | US & overseas finance | depends |
| imdb / itunes / thesportsdb | film / music / sports | mostly free |
| local_openstreetmap / wikidata / wikipedia | geo / org / encyclopedia | free |
| arxiv / semantic_scholar / openalex | academic | mostly free |
| pubchem / gbif / rfc_editor | chemistry / species / standards | free |
| github / stackoverflow / pypi / npm | code & packages | depends |
| byted / bocha / metaso / octen | Chinese web / AI search | API / low cost |
| zhihu / wechat_sogou | Chinese opinion / WeChat | API / free |
| tavily / felo / exa | international / semantic | paid or quota |
| twitter / reddit / xiaohongshu / bilibili / weibo | social UGC | free (some need login) |
Local zero-cost layer (local_*)
No separate SearXNG service. Main path uses in-process HTML / RSS / JSON parsing (local_bing, local_sogou, local_google, local_arxiv, …). For multilingual queries, routing rewrites engine language params (e.g. Bing setlang) and fuses with RRF.
Examples
Finance
python3 scripts/search.py "贵州茅台股价" --explain
# typical: stock_query → quote snapshot sources
Academic
python3 scripts/search.py "transformer attention mechanism paper" --json
# domain often academic; combo includes arxiv etc.
Research & verify
python3 scripts/research.py "2026 mutual fund Q2 holdings structure" --depth deep --json
python3 scripts/search.py "same query" --json | \
python3 scripts/evidence.py "same query" --stdin --json
MCP tools (10)
| Tool | Purpose |
|---|---|
argo_search |
unified search |
argo_local_search |
local files (offline) |
argo_research |
deep research (incl. social-sentiment mode) |
argo_evidence |
credibility scoring |
argo_clarify |
intent disambiguation |
argo_fetch |
smart fetch (mode=extract structured extract) |
argo_crawl |
site crawl |
argo_screenshot |
page screenshot |
argo_pdf |
PDF extract |
argo_social_search |
multi-platform social (mode=sentiment) |
Install & config
Requirements
| Item | Requirement |
|---|---|
| Python | 3.10+ (CLI + MCP core) |
| Deps | pip install pyyaml (only hard dependency) |
| Node.js | only for npx entry, 18+ |
| SearXNG | not required (built-in local engines) |
API keys (all optional)
Missing keys skip that engine; free engines backstop. Use env vars—never commit real keys or paste them into issues.
# recommended (better quality)
export TAVILY_API_KEY="your_key"
export BOCHA_API_KEY="your_key"
export METASO_API_KEY="your_key"
export ZHIHU_ACCESS_SECRET="your_key"
# optional
export BRAVE_API_KEY="your_key"
export FELO_API_KEY="your_key"
export GITHUB_TOKEN="your_key"
export WEB_SEARCH_API_KEY="your_key"
export ANYSEARCH_API_KEY="your_key"
export OCTEN_API_KEY="your_key"
config.yaml only stores {ENV_NAME} placeholders—no plaintext secrets in git.
Cache
Default SQLite path is cache.db_path in config.yaml (usually ~/.cache/unified-search/cache.db).
| Type | Approx. TTL |
|---|---|
| Finance | ~5 min |
| News / realtime | ~10–15 min |
| General | ~1 hour |
| Research / evergreen | ~2–24 hours |
| Empty results | very short (avoid freezing “no hits”) |
FAQ
Works without API keys?
Yes. Many local free engines and free APIs; unkeyed path is automatic.
Install script vs npx?
Script: fixed local install, config, Skill link. npx: attach MCP fast. Same Python core.
How to check engines?python3 scripts/search.py --list-engines, or add --explain.
Multiple code copies in the repo?
No. Prefer one source + symlinks via link_source.py, not rsync clones.
CLI flags
python3 scripts/search.py [options] query
--engine, -e engine, default auto
--max-results, -n count, default 5
--depth, -d fast | balanced | deep
--mode fast | auto | deep | budget
--no-cache skip cache
--explain print routing explanation
--json JSON output
--timeout, -t timeout seconds
--list-engines list engines
Design trade-offs
- Agent absorption first, link count second.
- Free and local first; paid is optional lift.
- Failures are observable: empty / timeout / breaker are labeled—no silent swallow.
- Config-driven engines;
config.yamlis the single source of truth. - Single-source install: link entries, don’t rsync copies.
- Social is not a truth library; good for expansion and sentiment, not sole ground truth.
Good fits
- Search backend for Claude Code / Grok Build / Codex / Kimi agents
- Multilingual, multi-domain Q&A: CJK + EN + finance / film / sports / academic / code
- Scripts and pipelines that need reproducible, cacheable retrieval
- Fact-check and multi-source comparison of public finance / entity data
Not a great sole solution for: platform-native engagement ranking, or long-lived max-recall aggregators (embedded local engines replace external SearXNG as the main path).
Tree (short)
argo/
├── README.md # Chinese (default)
├── README.en.md # English
├── README.ja.md # Japanese
├── README.ko.md # Korean
├── README.es.md # Spanish
├── SKILL.md
├── package.json # npx entry
├── bin/argo.js # Node MCP launcher
├── bin/argo # Python CLI
├── config.yaml # engines & domains (source of truth)
├── assets/readme/ # README visuals
├── backends/
├── scripts/ # search / research / mcp / install …
├── sub-skills/local-search/
├── tests/
└── docs/
Changelog
| Version | Notes |
|---|---|
| v2.6.0 | Multilingual search (detect / engine params / cross-lang fallback); film·sports·geo·org·media verticals; recovery anti-pollution; capability families + matrix regression; ~120+ sources. See release notes |
| v2.5.1 | Thicker finance/macro/chemistry answer sources; engine tiers + combo budget; v2.5.1 notes |
| v2.5.0 | Install script + npx; rewrite decoupled from routing; hot-path cache; compact MCP |
| v2.4.0 | Low-score route fallback + social mis-route filters; cache depth / soft hits; breakers & negative cache; engine_outcomes |
| v2.2–v2.3 | Two-stage evidence, Chinese source table, content_signals, fetch stack, more engines |
| v2.1 | Social engine layer (multi-platform UGC) |
| v1.x | Unified name Argo; multi-engine routing + two-layer cache |
Contributing
Issues and PRs welcome. When you change routing or evidence logic, please add tests:
python3 -m pytest tests/test_unit.py tests/test_multilingual.py -q
python3 scripts/regression_p0p1.py --offline
python3 scripts/matrix_search_eval.py --offline
python3 scripts/ab_eval_p0p1.py # optional, online
Before commit: no real API keys, absolute machine paths, or account cookies. Local Skill paths belong in installs.local.yaml (gitignored).
License
MIT License © 2026 taxueseek
Good search is not about seeing more—it is about concluding with confidence, and knowing when you still should not.
Links
More in this category
liustack/modlens★ 1199
Vision bridge for text-only models: paste an image, get structured JSON evidence (OCR, layout, semantics).
Anionex/dsh-vision-toolkit★ 308
Vision tasks for text-only models: intent-aware image Q&A, long-screenshot OCR, UI reproduction, grounding, and pixel diff.
zhaoolee/notes★ 138
Export DSH conversations as Smartisan Notes-style PNGs, or create and update Markdown notes in a configured account-scoped workspace.
liustack/modsearch★ 85
Web search bridge for text-only agents: ask the web or X, get structured JSON evidence (search, fetch, citations).
Lum1104/dsh-browser★ 80
Chrome sidebar extension that lets DSH operate your browser directly, no vision capabilities required.
ysr666/dsh-vision-router★ 36
Free vision for text-only agents: built-in keyless vision chain plus pixel tools (Q&A, grounding, crop, pixel diff, colors, OCR, SVG trace, cutout, screenshots); paste an image to use it.