CN A-share / HK stock data CLI — built to be piped into AI agents
sift is a single-binary Rust CLI that pulls CN A-share + HK stock data (the US financials adapter is stubbed and not wired end-to-end yet) — listings, financial statements, announcements, quote snapshots, OHLC bars, and PDF/OCR text extracts — from public endpoints (cninfo / 东方财富 / sina / tencent / 巨潮) and emits Unix-friendly TSV / NDJSON / aligned tables.
sift exists for one reason: let LLM agents (Claude, Kimi, Hermes, OpenClaw, …) read Chinese-market financials and OHLC bars as fluently as a human analyst.
Most investment-data tools target humans (GUIs) or assume an SDK consumer (deeply-nested JSON, stateful sessions). Neither works well for tool-use loops: GUIs can't be scripted, and nested JSON wastes tokens and breaks LLM parsing. sift goes the other way — a pure-stdout CLI where every command can emit #header\tcol\tcol\n TSV (awk/pandas/Polars friendly) or NDJSON (one object per line; the default is a human-aligned table). Any tool-using model (Claude Code, OpenCode, Kimi K2, …) can pipe results straight back to the LLM at zero parser cost.
In one line: sift is plumbing — turn the LLM tap and stock data comes out.
| Feature | What it gives you |
|---|---|
| 🤖 Agent-first output | Human-aligned table by default; --format tsv emits a #header row (doubles as column names and a comment marker, VCF/MAF convention) and --format json emits NDJSON. Any model parses it reliably. |
| 🔗 Pipe-friendly | sift search 银行 --format json | jq -r .code | xargs sift quote — subcommands compose naturally. |
| 📊 Multi-source financials, first-success-wins | Eastmoney + Sina (+ future cninfo). The default auto mode is served by Eastmoney (Sina uses Chinese item labels, so it stays out of the race to keep field names stable); pin with --source sina for reproducible runs. |
| 🗂 Full announcement pipeline | announce list → browse, show → metadata, download → PDF, extract → Markdown with OCR escalation for scanned pages. |
| 📈 Quotes + OHLC bars | quote for live snapshots, bars for daily/weekly/monthly with pre/none/post adjustment. |
| 🏷 Verbatim item labels | Financial line items keep their raw upstream names — no lossy remapping. A-share (东方财富) uses EM's English column codes (TOTAL_OPERATE_INCOME), HK/sina use their native Chinese labels. --items filters against those exact labels. |
| 💾 Local cache | Listings (24h file cache), financials (DuckDB + 3-tier TTL, fresher periods get shorter TTL), announcement metadata (DuckDB, no TTL), PDFs (forever). Re-runs return in ms. |
| 🧮 Local fact store + SQL | report / market automatically ingest fetched rows into ~/.sift/facts.duckdb; query it with sift sql (read-only by default, --write as the escape hatch), curate with sift fact / metric / map — hand-written facts, standard-metric vocabulary, and raw_key → std_key mappings. |
| 🔒 Offline-first, graceful degradation | $HOME unresolvable → caches disable but commands keep working. Per-symbol failure → [warn] on stderr, other symbols still print to stdout. Never crashes the whole run. |
| 🚀 Single static binary | DuckDB bundled. No Python / Node / system libs. One ~/.local/bin/sift, done. |
Add sift to your agent's allowed-tool list (Claude Code, OpenCode, Kimi K2, etc.) and it can do this:
User: Show me Moutai's net-profit trend over the last 4 quarters vs. peers.
Agent calls:
1. sift search 茅台 --limit 1 --format json → resolves 600519
2. sift report income 600519 --last 4 --unit yi --format tsv
3. sift search 白酒 --limit 5 --format json | ... → peer codes
4. sift report income <codes> --last 4 --unit yi --format tsv
5. (LLM synthesizes the two TSVs into prose + chart)
User: Summarize Moutai's 2024 annual report, pages 1-30.
Agent calls:
1. sift announce list 600519 --type 年报 --start 2024-01-01 --format json
2. sift announce download <id> -o /tmp
3. sift extract <id> --mode auto --pages 1-30 > /tmp/report.md # scanned pages auto-OCR
4. (LLM reads the markdown and writes the summary)
User: Which A-shares are rallying on heavy volume today?
Agent calls:
1. sift search 600 --limit 50 --format json | jq -r .code | xargs sift quote --format tsv
2. (LLM ranks by pct_change and compares against 5-day avg volume)
User: Which 20 A-shares had the highest ROE in the 2024 annual reports?
Agent calls:
1. sift market --period 2024A --sort WEIGHTAVG_ROE --desc --limit 20 --format tsv # whole-market snapshot (also ingested)
2. sift metric add roe --label ROE --unit-kind ratio
3. sift map set --source eastmoney WEIGHTAVG_ROE roe
4. sift sql "SELECT symbol,name,value FROM v_facts WHERE key='roe' AND period='2024A' ORDER BY value DESC LIMIT 20"
5. (LLM synthesizes the results)
One-liner — auto-detects OS + arch, downloads from GitHub Releases, and installs a single binary into the first writable directory on your PATH (falling back to ~/.local/bin):
curl -fsSL https://raw.githubusercontent.com/eavae/sift-cli/main/scripts/install.sh | bashOptional env overrides:
SIFT_VERSION=v0.2.0 \
SIFT_INSTALL_DIR=/usr/local/bin \
curl -fsSL https://raw.githubusercontent.com/eavae/sift-cli/main/scripts/install.sh | bashSupported targets: x86_64-unknown-linux-gnu, aarch64-unknown-linux-gnu, aarch64-apple-darwin, x86_64-pc-windows-msvc. The script verifies the SHA-256 checksum that ships next to each archive.
Where github.com is blocked (e.g. mainland China): the script auto-detects an unreachable GitHub and pulls the binary through a public mirror. Since fetching the script itself also goes through GitHub, route that through the mirror too:
curl -fsSL https://cdn.gh-proxy.org/https://raw.githubusercontent.com/eavae/sift-cli/main/scripts/install.sh | bashForce the behavior with SIFT_MIRROR: auto (default), off (direct only), or a mirror URL (e.g. SIFT_MIRROR=https://cdn.gh-proxy.org).
Windows (x86_64-pc-windows-msvc) ships as a .zip. Run the same one-liner under Git Bash / MSYS2 and it detects Windows, downloads the .zip, and installs sift.exe for you. (Prefer to do it by hand? Grab the archive from the Releases page, extract sift.exe, and put it on your PATH.) The store lives at %USERPROFILE%\.sift (e.g. C:\Users\you\.sift); set HOME to override the whole tree.
If your platform isn't covered by the releases, or you want bleeding-edge main:
git clone https://github.com/eavae/sift-cli.git
cd sift-cli
./install.sh # builds release + copies to ~/.local/bin
# or manually:
cargo build --release && cp target/release/sift ~/.local/bin/Requires Rust ≥ 1.95.
On Windows, use the MSVC toolchain (rustup default stable-x86_64-pc-windows-msvc) plus the Visual Studio C++ Build Tools — the bundled DuckDB is C++ and needs a C++ compiler. Then cargo build --release produces target\release\sift.exe. (install.sh is bash-only; copy the .exe onto your PATH yourself.)
Make sure ~/.local/bin is on PATH, then:
sift --help
sift search 茅台
sift report income 600519 --last 4 --unit yiThis repo ships a ready-made sift-facts skill for AI agents (how to query and curate the local fact store). Codex users can install it with one command — no manual copying:
npx skills add eavae/sift-cli --skill sift-facts -g -a codex -y-ginstalls to your user directory (omit it to install into the current project, which can then be shared with your team)-a codextargets Codex; drop the flag to pick another agent interactively (Claude Code, Cursor, …)-yskips confirmation prompts (CI-friendly)- Update later with
npx skills update sift-facts
The skill is an instruction guide for the agent — you still need the sift binary itself, installed as above. Once installed, just ask Codex something like "query the fact store for X with sift" or "hand-write a fact for X with sift" and it kicks in automatically.
sift search <kw> [--limit N] [--no-cache] # fuzzy search (code / name / pinyin initials)
sift report {income|balance|cashflow|indicator|periods} <code...> [--last N | --years N | --period 2024Q3 | --start --end] [--qmode cumulative|single] [--source eastmoney|sina] [--no-cache] [--no-ingest]
sift announce {types|list|show|download} ... # announcement types / browsing / details / downloading
sift extract <id|path> [--pages 1-30] [--mode fast|fine|auto] [--image-dir DIR] # PDF → Markdown
sift quote <code...> # live snapshot
sift bars <code...> [--period daily|weekly|monthly] [--limit N] [--adjust pre|none|post] [--source tencent|eastmoney]
sift market [--period 2024A] [--where 'WEIGHTAVG_ROE>15'] [--sort COL] [--desc] [--limit N] [--market main|star|gem|bj] [--no-cache] [--no-ingest] # whole-market snapshot (auto-ingested)
sift sql "SELECT ..." [--write] # query the local fact store (read-only by default)
sift fact {set|rm} ... # write / delete facts by hand
sift metric {add|ls|rm} ... # manage the standard-metric vocabulary
sift map {set|ls|rm} ... # manage raw_key → std_key mappingsEvery command accepts --format tsv|json (omit for human-aligned table). Multi-symbol calls are best-effort: one failure doesn't sink the rest. extract outputs Markdown, so --format is meaningless for it.
Symbol forms: bare 6 digits = A-share (600519), bare 5 digits = HK (00700); 600519.SH / 00700.HK / 430718.BJ suffixes and sh600519 / bj430718 prefixes all work. Indices need an explicit exchange prefix and are served by quote / bars only: sh000001 = 上证指数, sz399001 = 深证成指 (output keeps the lowercase prefixed form; report / announce reject indices). Note 000001 without a prefix is always 平安银行 the stock, never the index.
extract runs purely locally in fast mode (zero API calls). When you hit scanned pages and need real OCR, switch to fine or auto — at which point sift calls a PaddleOCR cloud backend. Two configuration modes are supported; pick the one that matches your account.
For the PaddleOCR official async jobs API (AI Studio). One-shot token auth — simplest path.
export PADDLEOCR_API_BASE="https://paddleocr.aistudio-app.com" # API host, no /api/... suffix
export PADDLEOCR_API_TOKEN="<your-access-token>"
# optional: override the model (defaults to PaddleOCR-VL-1.6)
# export PADDLEOCR_MODEL="PP-StructureV3"- Async pipeline: submit file → poll job → download per-page Markdown (JSONL).
- Defaults to PaddleOCR-VL-1.6;
PADDLEOCR_MODELaccepts the document-parsing modelsPP-StructureV3,PaddleOCR-VL,PaddleOCR-VL-1.5,PaddleOCR-VL-1.6. OCR-only models (PP-OCRv5,PP-OCRv5-latin,PP-OCRv6) return plain OCR boxes instead of Markdown, soextractrejects them at startup. - Good for: individual developers, AI Studio tokens, free-tier trials.
For Baidu AI Open Platform enterprise accounts. API Key/Secret auth via OAuth, async task endpoint.
export PADDLEOCR_API_KEY="<your-baidu-api-key>"
export PADDLEOCR_SECRET_KEY="<your-baidu-secret-key>"
# optional override (rarely needed):
export SIFT_BAIDU_HOST="https://aip.baidubce.com"- Auto OAuth → access token (30-day TTL, in-process cache).
- Async pipeline: submit task → poll → download structured result (layout / tables / images).
- Good for: enterprise volume, per-account billing, Baidu-cloud compliance.
At OCR startup sift checks Token → OAuth:
- Both
PADDLEOCR_API_BASE+PADDLEOCR_API_TOKENnon-empty → Token mode wins. - Otherwise both
PADDLEOCR_API_KEY+PADDLEOCR_SECRET_KEYnon-empty → OAuth mode. - Neither configured →
--mode fine|autoerrors with a hint.
Configure only one. If both are set, Token mode takes precedence.
(Add a LICENSE file — MIT / Apache-2.0 / proprietary as appropriate.)