An autonomous, agentic knowledge compiler designed to maintain and synthesize a "Second Brain" within an Obsidian Vault. Built natively with the Google Antigravity SDK and Gemini 3.1 Pro, this system automates the ingestion of technical articles, YouTube videos, and research papers, filtering content via an automated Technical Density Grader, indexing structural relationships via Graphify.net, and maintaining a decoupled skill-data architecture.
What this is. A personal automation system that turns a large, ever-growing pile of technical reading (articles, videos, podcasts, books) into a structured, searchable knowledge base β automatically. Instead of bookmarking content and never revisiting it, the system reads it, judges whether it's actually worth keeping, extracts the key ideas, cross-links them to everything already known, and flags contradictions as understanding evolves.
Who it's for. Built and maintained by one engineer for personal use β a working example of applying real software-engineering discipline (automated tests, architecture decision records, staged rollouts, health monitoring) to a personal productivity problem, not a commercial product or team tool.
Why it matters. It demonstrates hands-on experience with skills relevant well beyond this project: designing multi-stage automation pipelines, orchestrating multiple AI models efficiently (grading content before spending compute on it, avoiding unnecessary API costs), building systems that degrade gracefully when a data source fails or changes instead of breaking outright, and documenting the reasoning behind architectural decisions so the system stays legible over time (see docs/adr/).
Maturity. Actively developed and in daily personal use, managing a growing multi-thousand-note vault. Not intended for external users or production deployment β no support contract, no SLA, no multi-tenant design.
What's under the hood, in plain terms:
- Reads and judges content before saving it β automatically scores incoming articles and videos for how substantive they are, so low-value content doesn't clutter the knowledge base.
- Builds a live map of how ideas connect β every note is linked to related concepts automatically, without paying an AI model to work out the connections every time.
- Catches contradictions β flags when a newer note disagrees with something recorded earlier, instead of silently letting outdated conclusions persist.
- Runs a full health check on demand β one command audits the entire knowledge base (broken links, duplicate content, orphaned files, structural integrity) and can auto-repair common issues.
- Keeps working when a data source misbehaves β web scraping has three fallback tiers, so a paywall or a blocked scraper doesn't just fail silently.
The rest of this document is the technical reference β architecture, setup, and workflow detail for engineers evaluating or extending the system. Start with Key Architectural Innovations for the plain-terms list above translated into implementation detail.
- For Non-Technical Readers
- Project Overview
- Problem Statement
- Key Architectural Innovations
- System Architecture
- Getting Started & Agent Installation
- The "Zone" Knowledge Architecture
- End-to-End Workflow
- AI & Agent Components
- Security & Sandboxing
- Technology Stack
- Future Improvements & Lessons Learned
This project serves as an advanced production implementation of the LLM Wiki Paradigm (inspired by Andrej Karpathy's concept of LLMs as knowledge compilers rather than simple chatbots).
Instead of treating the AI as a search engine over raw documents (like traditional RAG), the Obsidian Knowledge Curator acts as an active maintainer of a local filesystem database. It reads raw inputs, grades source density, extracts core concepts, updates existing Wiki pages, flags contradictions natively, and builds a chronological trace of evolving ideas across thousands of notes.
The Engineering Challenge: Knowledge workers and AI Engineers consume vast amounts of technical content (papers, documentation, videos). Traditional PKM (Personal Knowledge Management) systems rely on manual synthesis, while modern LLM chatbots (chat-with-PDF) fail to compound knowledge over time because they lack persistent state and cross-document reasoning. Furthermore, naive LLM graph compilation introduces massive API token costs and high query latencies.
The Solution: A headless, zero-token-overhead automation pipeline that ingests content, parses transcripts, evaluates technical density, and uses an Offline AST Graphify Indexer to map vault relationships without external LLM API costs.
Before any source (YouTube video, Web article, Tweet) is written to the vault, a 3,000-character preview is evaluated across three dimensions:
- Information Density (lack of fluff, factual saturation).
- Provenance & References (citations, data points, verified authors).
- Technical Level (code architecture relevance, concrete implementations).
If the composite score falls below MIN_TECHNICAL_SCORE (configured in .env, default: 60), ingestion halts, presenting a detailed scorecard and summary to the user for explicit override confirmation.
To enable graph-aware context retrieval across 13,000+ notes without incurring API costs or latency penalties:
- AST & Wikilink Parser: Uses local Python regex and Markdown AST parsing to extract document headings, parent-child nesting, and
[[wikilinks]]. - Hybrid Local Context Engine (
GraphifyMapper): Pre-predicts targetraw/categories and matches existingwiki/concepts (<50ms latency, $0.00 token cost) directly fromgraphify-out/graph.json. It falls back to lightweight LLM queries only if local graph matching confidence drops below 70%. - Injected Ingestion Context: Automatically injects
"graphify_context"objects intotemp/fetched_data.jsonacross all ingestion scripts (fetch_article_data.py,fetch_youtube_data.py,fetch_twitter_data.py,fetch_book_data.py). - dswok Integration: Indexes protected external knowledge directories (
dataScienceKnowledgeBase/dswok) as a read-only information graph without modifying any files within them.
- Multi-Platform Support: Ingests podcasts from Siemens.FM, Spotify, Apple Podcasts, RSS feeds, YouTube audio, and direct
.mp3/.m4afiles usingyt-dlpandcurl. - Offline Whisper Transcription: Uses local Buzz CLI (
/Applications/Buzz.app/Contents/MacOS/Buzz) to transcribe audio tracks offline without external API costs.
- Native FastMCP Tool: Exposed as
@mcp.tool()fetch_article_data(url, download_all_images)on theokcFastMCP server (scripts/mcp_server.py) for zero-permission, in-process agent execution. - Full DOM Extraction Without Limits: Targets semantic containers (
article,[itemprop='articleBody'],.post-content,.post__content,.body.markup,div.available-content, etc.) usingBeautifulSoupandmarkdownify, extracting 100% of the body and eliminating the 5,000-character default truncation limit inherited frommcp-server-fetch. - Integrated Technical Quality Scoring: Automatically scores candidate articles (0-100 scale) across information density, provenance, code architecture depth, and educational infographics before ingestion.
- Tolerant 3-Tier Fallback Chain: Tier 1 uses direct requests with realistic browser headers and SSL verification fallbacks (
ssl.CERT_NONE); Tier 2 usesmcp-server-fetchconfigured withmax_length: 10,000,000for Cloudflare/Captcha bypass; Tier 3 usessearch_webfor locked paywalls. - Audio Redirection: Automatically detects podcast/audio URLs or audio media files and delegates execution to
fetch_podcast_data.py.
5. Concurrent High-Density Image Preservation & Avatar Filtering (fetch_article_data.py & fetch_book_data.py)
- Concurrent Asset Downloads: Uses
ThreadPoolExecutor(max_workers=6)to concurrently fetch all article infographics and diagrams directly into<VAULT_ROOT>/assets/images/<slug>-<idx>-<alt_slug>.<ext>. - Modern CDN &
<picture>/<noscript>Support: Preserves images hidden inside<picture><source srcset="...">and<noscript><img ...>elements (e.g., Medium, Substack), using regex pattern matching on high-resolution CDN URLs (miro.medium.com) so no architectural diagrams are dropped. - Intelligent Avatar & Tracking Pixel Filter: Discards author profile pictures, emojis, badges, and tracking pixels (<100px) through DOM attribute inspection and real pixel dimension verification via
PIL.Image. - Resilient CDN URL Sanitization: Cleans control characters (
\x00-\x1f), whitespace, alt fragments, and comma-delimited Substack/Medium transform parameters before downloading. - Exact Obsidian Wikilink Embeddings: Replaces remote image links in generated Markdown with native Obsidian wikilinks
![[assets/images/<file>.png]]at their exact locations, with automatic C2PA/EXIF metadata sanitization.
To prevent system prompt inflation and context degradation:
- Behavior Prompt (
SKILL.md): Contains pure agent execution rules, wikilink mandates, and non-hallucination constraints. - Compiled Static Database (
KNOWLEDGE.md): An automatically regenerated index containing ~600+ concept cards with absolute file links across the vault.
A unified 7-stage health check, integrity audit, and synchronization suite:
- SQLite Differential Index: Rapid scan and synchronization across 3,000+ files.
- Multi-Category Master Plans: Dynamic regeneration of all navigation maps.
- Wikilink & Contradiction Linter: Deep scan for broken links, orphan notes, and explicit
[!contradiction]tags. - Invisible Unicode (ZWSP) Hygiene: Detects and sanitizes zero-width space characters with
--fix. - Visual Asset Inspector: Audits
assets/images/total vs unreferenced visual assets. - Protected Zones Immutability Audit: Ensures zero unauthorized modifications across protected engineering zones.
- Graphify &
KNOWLEDGE.mdRebuild: Updatesgraph.json,graph_cache.json, and the concept index in one pass.
An automated audit, deduplication, and header standardization engine (ADR 0006):
- Content Hash Deduplication: Computes SHA-256 digests of markdown content to identify identical notes across folders, safely archiving duplicates into a local
_archive/directory. - Empty & Low-Value Stub Elimination: Detects and isolates notes lacking substantial content or structure, preserving a high signal-to-noise ratio in the knowledge graph.
- Canonical Blockquote Injection: Detects missing metadata headers and non-destructively prepends canonical Obsidian blockquote schema (
Author,Source,Type,Processed, andTags: #no-read-yet). - Atomic Database & Master Plan Sync: Includes
--syncto automatically update the SQLite differential index and Category Master Plans in a single operation.
An automated spaced repetition, active recall, and pedagogical flashcard synthesis engine (ADR 0007):
- SuperMemo 20-Rules Compliance: Generates atomic
Basic(Q/A) andCloze({{c1::...}}) flashcards grounded in verifiableSourceSpanevidence slices. - AST Parsing & Math/Table Fidelity: Uses
markdown-it-pyto parse source notes while maintaining LaTeX formulas ($...$,$$...$$), fenced code blocks, and markdown tables. - Direct Anki MCP & AnkiConnect Push: Automatically creates decks, uploads media assets, and synchronizes cards via the Anki MCP server or local AnkiConnect HTTP API (
http://127.0.0.1:8765). - Three-Way Merge with Human Preservation: Retains user modifications to card fronts and backs in Markdown while updating source notes; records deleted cards in
study_suppressionsto prevent resurrection. - Deep Visual Architecture: Clones diagram and figure assets by SHA-256 hash into
<VAULT_ROOT>/assets/images/study/, generating dedicated architectural visual recall cards. - Two-Phase Commit (2PC) Journal & Crash Recovery: Logs write transactions (
PREPARED->COMMITTED), performs atomic staging swaps, and providesstudy_deck.py recoverto prevent partial state corruption.
10. High-Density Active Recall (HDAR) Book Flashcard Engine (scripts/book_flashcards_engine.py, /okc-bookFlashcards)
An end-to-end active recall flashcard generation and synchronization engine for technical books (ADR 0008):
- HDAR 7-Rule Pedagogical Rubric: Enforces zero-hallucination factual grounding, atomic information retrieval, causal mechanisms over superficial facts, and balanced bilingual technical terminology.
- Whole-Book Decomposition & Extraction: Parses EPUB and PDF non-fiction books into structured semantic chapters, section hierarchies, and embedded visual media.
- Resilient Chapter Checkpointing: Maintains incremental state in
checkpoint.jsonto allow resume-on-failure across 10+ chapters without reprocessing or data loss. - Automated Deduplication & Global Audit: Evaluates candidate cards across chapters, eliminates semantic duplicates, and outputs comprehensive audit reports (
audit_report.md). - Multi-Platform Convergence: Simultaneously generates 1-click Anki import files (
.tsv), structured JSON databases, native Obsidian study notes (<VAULT_ROOT>/.../study/<Deck Name>.md), and executes direct AnkiConnect synchronization with embedded diagrams.
11. Atomic FastMCP Note Curation & High-Speed Incremental Sync (scripts/mcp_server.py, scripts/sync_vault.py --fast)
An end-to-end atomic curation tool and sub-second incremental vault synchronizer (ADR 0009):
- Atomic Curation MCP Tool (
commit_curated_note): Writes the raw source note and multiple synthesized wiki concept notes directly toVAULT_ROOTin a single tool call. Eliminates permission boundaries associated with agent artifact tool restrictions (write_to_file) and avoids terminal prompt interrupts. - Fast Incremental Vault Synchronization (
sync_vault.py --fast): Performs lightweight SQLite differential updates and dynamically rebuilds Master Plans while surgically updating modified notes in the Graphify AST cache, bypassing the full 27,000-node crawl and dropping sync latency from ~20s to ~2.5s. - Self-Contained Automated Ingestion: Combines article DOM fetching, diagram downloads, quality grading, note synthesis, and vault synchronization into a clean 2-step agent interaction loop.
12. Universal Zero-Permission FastMCP Ingestion & Multi-Note Batch Commit (scripts/mcp_server.py, scripts/fetch_playlist_data.py)
An architectural paradigm shift establishing a strict 2-step FastMCP workflow across all content modalities, eliminating 100% of terminal permission prompts and context truncation (ADR 0010):
- Native FastMCP Suite (22 Tools): Expanded
scripts/mcp_server.pyto expose headless tools for all ingestion pipelines (fetch_playlist_data,fetch_youtube_data,fetch_article_data,fetch_twitter_data,fetch_podcast_data,fetch_spotify_data,fetch_doc_data,fetch_book_data,clean_staging_temp,inspect_flashcards_deck). - YouTube Playlist Batch Extractor (
fetch_playlist_data): Parses whole playlist structures viayt-dlp --flat-playlist, extracting video metadata, titles, and full transcripts in a single tool call or CLI command (scripts/fetch_playlist_data.py), staging structured data intemp/fetched_playlist_data.jsonand disk-persisted transcript files intemp/playlist_transcripts/. - Atomic Multi-Note Batch Commit (
commit_curated_batch): High-throughput atomic macro-tool that writes multiple raw source notes (raw_notes), an optional Series Master Plan (master_plan_path&master_plan_content), and all compiled wiki concepts (wiki_notes) directly toVAULT_ROOTin a single tool call, followed by instant--fastincremental synchronization. - Universal 2-Step Agent Protocol: Standardizes all ingestion skills (
okc-urlPlaylist,okc-urlYoutube,okc-urlTwitter,okc-urlPodcast,okc-urlSpotify,okc-doc,okc-bookSummary) on a deterministic 2-turn cycle: Turn 1 (Fetch & Stage via FastMCP) -> Turn 2 (In-Memory Synthesis & Atomic Batch/Single Commit via FastMCP), completely eliminating subshell escaping vulnerabilities, Cortex artifact boundaries, and permission fatigue.
13. Spotify Chrome DOM Hydration Bridge (scripts/fetch_spotify_data.py, /okc-urlSpotify & FastMCP fetch_spotify_data)
An automated browser-driven extraction engine overcoming Widevine DRM encryption (MP4_128_CBCS) without paid third-party APIs (ADR 0011):
- AppleScript Session Integration: Interacts directly with an active, user-authenticated Google Chrome session on macOS.
- Automated Transcript View: Navigates to the Spotify episode, targets the native "Transcript" UI tab, and triggers live cue hydration.
- Virtualized DOM Smooth-Scroll Harvesting: Executes a 35ms iterative smooth-scroll JavaScript loop over virtualized cue rows (
[data-testid="transcript-cue-row"]), capturing 100% of timestamped spoken content while avoiding missing off-screen elements. - Zero-Loss Markdown Staging: Normalizes speaker cues, combines sequential sentences, and stages output to
temp/fetched_data.jsonandtemp/fetched_data.txtfor atomic curation viacommit_curated_note.
14. Instagram Saved Collections Batch Ingestion Engine (scripts/process_instagram_saved_batch.py, /okc-instagram)
A headless, high-throughput social video ingestion pipeline for curated Instagram collections (ADR 0011):
- DOM Link Harvesting: Extracts saved reel URLs from the active browser session, bypassing anti-scraping blocks and login walls.
- Audio Extraction & Offline Whisper Transcription: Uses
yt-dlpto download reel audio streams and transcribes them offline using Whisper (basemodel viaBuzz CLIor local Whisper) at $0.00 token cost. - Batch Atomic Curation & Master Plan Builder: Generates individual Obsidian notes conforming to vault standards in
Leadership and Coach/raw/Instagram/, automatically computes creator statistics, and constructs a dedicated collection Master Plan (Master Plan β Santo Trabajo.md) with cross-links.
15. Modular Flashcards Subsystem & Declarative DeckRegistry (src/agent_tools/flashcards/, study_multi_choice.py)
A unified, AST-driven spaced repetition engine adhering to the "One Tool + N Inputs > N Scripts" standard (ADR 0011):
- Unified AST & Regex Parsers (
parsers.py): Robustly parses Basic TSV and 6-dimension Markdown multiple-choice cards (header, prompt, 4 options, isolated?delimiter, answer key, and distractor analyses across options AβD). - Declarative DeckRegistry (
registry.py): Centralizes canonical mappings for major technical courses and authors (Soledad Galli, Chip Huyen, ByteByteGo), infers categories dynamically, and resolves destination vault paths viaresolve_study_path. - FastMCP Deck Inspection (
inspect_flashcards_deck): Exposes live card counting, model validation, and subdeck inspection directly to the AI agent with zero terminal prompts.
flowchart TD
A["External Content (YouTube / Web / PDF)"] -->|Fetch Raw Data| B["Stage 1: Ingestion & Extraction"]
B -->|3000-char Preview| C{"Technical Density Grader"}
C -->|"Below Threshold (< MIN_TECHNICAL_SCORE)"| D["User Override Prompt (y/n)"]
C -->|"Pass (>= MIN_TECHNICAL_SCORE)"| E["Antigravity Agent Context"]
D -->|Approved| E
E -->|Write Source Note| F["Zone 1: raw/"]
E -->|Compile Concepts & Contradictions| G["Zone 2: wiki/"]
F -->|Offline AST & Wikilink Extraction| H["Graphify Indexer (graphify_helper.py)"]
G -->|Offline AST & Wikilink Extraction| H
H -->|Update Structural Graph| I["graphify-out/graph.json (legacy)"]
H -->|Regenerate Concept Cards| J["KNOWLEDGE.md Index Card"]
F -.->|Live vault watch| M["graphify-daemon (resident process, MCP)"]
M -.->|"query_graph / get_node -- always current"| N["Agents (Claude Code, Antigravity)"]
M -.->|Periodic snapshot flush| Q["graphify-daemon/out/graph.json"]
I -->|"GRAPHIFY_BACKEND=local/auto"| P["run_explore() / okc_doctor.py"]
Q -->|"GRAPHIFY_BACKEND=remote/auto"| P
K["Vault Linter / Health Check"] -.->|Scan Links & Orphans| G
L["Sync Vault Pipeline"] -.->|Auto-Rebuild Master Plans| J
This project runs locally and relies on Python 3.12+ and external command-line utilities.
- Python Package Manager: uv (required for high-speed, isolated environment management).
- Browser Scraping: Google Chrome or Chrome Canary (installed locally, required for dynamic CDP JS rendering).
- Media Processing: ffmpeg (required by
yt-dlpto extract audio streams). - Local Transcription Fallback: Buzz CLI (required for offline Whisper transcription fallback).
macOS (via Homebrew):
# Install uv and ffmpeg
brew install uv ffmpeg
# Install Buzz (GUI + CLI)
brew install --cask buzz
# Install Graphify CLI
uv tool install "graphifyy[gemini]"Linux (Ubuntu/Debian):
# Install uv
curl -LsSf https://astral.sh/uv/install.sh | sh
# Install ffmpeg
sudo apt update && sudo apt install -y ffmpeg# Clone the repository
git clone https://github.com/c-ibarra/obsidianKnowledgeCurator.git
cd obsidianKnowledgeCurator
# Install dependencies and setup virtual environment
uv syncRun the automated setup wizard to configure your local .env securely from .env.template, auto-detect your Obsidian Vault (~/Documents/Obsidian), and validate system dependencies:
# Interactive setup in terminal
uv run python scripts/setup_project.py
# Or within Antigravity chat:
/okc-setupAlternatively, manually copy the template and edit your .env:
cp .env.template .envBuild the initial structural graph across your vault and compile the KNOWLEDGE.md concept cards index:
# Verify vault health, rebuild Master Plans, and compile Graphify index
uv run python scripts/sync_vault.py --target-kb allTo enable bi-directional spaced repetition flashcard generation and automatic synchronization (/okc-study):
-
Install and Open Anki:
- Download and install Anki.
- Ensure Anki is running in the background during synchronization.
-
Install the AnkiConnect Add-on:
- In Anki, navigate to:
ToolsβAdd-onsβGet Add-ons... - Enter the AnkiConnect code:
2055492159and click OK. - Restart Anki.
- In Anki, navigate to:
-
Configure AnkiConnect CORS / Allowed Origins:
- Go to
ToolsβAdd-ons, select AnkiConnect, and click Config. - Ensure
webCorsOriginListincludes local loopbacks andapiKeyis empty (default) or matches your configuration:{ "apiKey": null, "apiHost": "127.0.0.1", "apiPort": 8765, "webCorsOriginList": [ "http://localhost", "http://127.0.0.1", "*" ] } - Restart Anki. Test connection by running
curl http://127.0.0.1:8765(should return"AnkiConnect").
- Go to
-
Configuring / Updating the Anki MCP Server in Antigravity / Claude Code:
- The Anki MCP server is registered in your agent client configuration (
mcp_config.jsonor Antigravity's settings). - If using
npxor standard node MCP runner:{ "mcpServers": { "anki": { "command": "npx", "args": ["-y", "@modelcontextprotocol/server-anki"], "env": { "ANKI_CONNECT_URL": "http://127.0.0.1:8765" } } } } - If using python-based
anki-connect-mcp:{ "mcpServers": { "anki": { "command": "uvx", "args": ["anki-connect-mcp"] } } } - To update the Anki MCP server, force cache refresh via
npx -y @modelcontextprotocol/server-anki@latestoruvx --upgrade anki-connect-mcp. - Note: Even if the MCP server is idle or offline,
scripts/study_deck.pyautomatically falls back to native direct HTTP calls againsthttp://127.0.0.1:8765to guarantee reliable synchronization without interruptions.
- The Anki MCP server is registered in your agent client configuration (
The vault is strictly divided into four zones to separate immutable sources from synthesized concepts:
| Zone | Purpose | Agent Permissions |
|---|---|---|
raw/ |
Immutable sources (Video transcripts, Web clippings). | Append-Only. The agent saves summaries here but never modifies historical sources. |
wiki/ |
Synthesized concepts and entities. | Read-Write. Fully maintained by the LLM. The agent creates pages, injects wikilinks, and merges updates. |
dev/ |
Architecture Decision Records (ADRs) and project files. | Collaborative. The agent acts as a co-pilot but requires explicit human approval to modify. |
dswok/ |
Protected external personal knowledge base. | Read-Only / Indexed. The agent scans and indexes relationships into graph.json but never writes or modifies files. |
1. Multimedia, Video & Podcast Ingestion (fetch_youtube_data.py, fetch_podcast_data.py, fetch_twitter_data.py, fetch_spotify_data.py & FastMCP)
- YouTube (
/okc-urlYoutube): FastMCP toolfetch_youtube_dataextracts audio streams viayt-dlp(--live-from-start), transcribing viayoutube-transcript-apiwith fallback to local Buzz CLI Whisper. Curates in memory and commits viacommit_curated_note. - Twitter/X (
/okc-urlTwitter): FastMCP toolfetch_twitter_dataextracts tweet text and attached video audio, transcribing via Buzz Whisper and committing atomically. - Podcasts & Audio (
/okc-urlPodcast): FastMCP toolfetch_podcast_dataingests Siemens.FM, Apple Podcasts, RSS feeds, or direct.mp3/.m4afiles viayt-dlp/curland Buzz Whisper. Automatically delegates Spotify URLs to the Chrome DOM bridge. - Spotify Podcasts (
/okc-urlSpotify): FastMCP toolfetch_spotify_dataleverages AppleScript and Chrome DOM hydration scrolling to extract 100% of timestamped transcripts, bypassing Widevine DRM (MP4_128_CBCS) without audio decoding fees. - Pre-calculates
"graphify_context"viaGraphifyMapperto select raw target categories and link existing wiki concept notes. - Atomically writes curated notes to
raw/and compiled concepts towiki/viacommit_curated_note.
- Native FastMCP Batch Extractor (
fetch_playlist_data): Extracts the full playlist architecture viayt-dlp --flat-playlistand pulls transcripts for all items intotemp/fetched_playlist_data.jsonandtemp/playlist_transcripts/<id>.txtin a single tool call. - In-Memory Batch Curation: Synthesizes all sequential raw source notes, compiles cross-video wiki concepts, and generates the Series Master Plan without context thrashing.
- Atomic Multi-Note Commit (
commit_curated_batch): Commits all raw notes, the Series Master Plan, and wiki concepts in a single zero-permission transaction, followed by instant--fastincremental vault synchronization.
- Native FastMCP Tool (
fetch_article_data): Headless, zero-terminal-popup execution via stdio FastMCP server (okc). - Full DOM Extraction: Targets semantic content containers (
article,[itemprop='articleBody'],.post-content,.body.markup, etc.) withBeautifulSoupandmarkdownify, extracting 100% of article body without 5,000-character truncation limits. - Concurrent Image Preservation: Concurrently downloads infographics via
ThreadPoolExecutor(max_workers=6), preserves<picture>/<noscript>images from Medium and Substack, filters out author avatars (<100px), sanitizes CDN URLs, and injects exact Obsidian wikilinks![[assets/images/...]]. - Integrated Quality Grading: Evaluates technical density and architectural relevance (0-100 score) before note generation.
- Streamlined 2-Step Curation Flow:
fetch_article_data(url="...", download_all_images=True)stages content and downloads media.commit_curated_note(raw_note_path="...", raw_content="...", wiki_notes={...}, fast_sync=True)atomically writes raw and wiki notes, updates SQLite index, and runs incremental synchronization.
- Audio Auto-Detection: Automatically detects audio/podcast URLs and delegates to
fetch_podcast_data.py. - Graphify Context Enrichment: Pre-calculates
"graphify_context"viaGraphifyMapperfor raw target categories and wiki concept linking.
- High-Density Actionable Synthesis (HDAS Standard): Produces rich, 7-section modular chapter notes combining dense narrative prose with active learning tools (Central Thesis & 1-Sentence Insight, Inquiry Questions, Enriched Summary Development with Mental Models, Visual Metaphor / Key Quote / Common Pitfall Callouts, Smart Cross-Domain Commentary, Practical Application Guide with 15-minute challenges, and Executive 1-Sentence Takeaway).
- Automated Structure Parsing: Ingests PDF, EPUB, DOCX, and TXT files, segmenting chapters and sanitizing text.
- Default Obsidian Vault Output: Writes main book summaries and individual chapter notes (
Chapter XX β <Title>.md) directly toVAULT_ROOT/dataScienceKnowledgeBase/<Category>/raw/books/. - Enforced Chapter Depth (1,600β2,650 words total): Requires dense explanatory narrative for Section 3 (Enriched Summary Development) to ensure thorough technical and conceptual depth.
- Visual Content & Mermaid.js Diagrams: Reconstructs mindmaps, architecture flows, sequence diagrams, and embeds extracted figures (
assets/images/). - Executive Master Note Hub: Compiles comprehensive master notes with full-book architecture mindmaps, mental models index,
#flashcardspaced repetition cards, and specialized glossaries. - Automatic Temporary File Cleanup: Cleans up all working files in
temp/via--cleanupon completion.
- Multi-Format Extraction: Ingests
.docx,.pptx,.xlsx,.epub,.pdf,.odt, and.csvusing the unified AnyDoc engine (src/agent_tools/anydoc_engine.py). - Embedded Asset Extraction: Extracts embedded figures and charts directly to
<VAULT_ROOT>/assets/images/and links them with native Obsidian wikilinks.
- Incrementally updates
graphify-out/graph.jsonafter every note edit (legacy pipeline). - Supports
--fast/--incrementalflag (sync_vault.py --fast) to synchronize SQLite indexes and dynamic Master Plans while performing surgical graph updates without crawling all 27,000+ nodes. - Automatically regenerates
.agents/skills/obsidian-knowledge-curator/KNOWLEDGE.mdwith updated concept links. - Rebuilds Category Master Plans and audits wikilink health via
vault_linter.py. graphify-daemon(a separately-maintained resident process,~/projects/graphify-daemon) serves the same concept graph over MCP from an always-current in-RAM snapshot β no rebuild step, republishes on every vault batch. Agents (Claude Code, Antigravity) query it directly via its MCP tools (query_graph,get_node,shortest_path, etc.) whenever it's running.scripts/knowledge_commands.py'srun_explore()andscripts/okc_doctor.py's dashboard node/edge counts read from the daemon by default too, controlled byGRAPHIFY_BACKENDin.env(auto/local/remoteβ see ADR 0005 and.agents/rules/graphify.mdfor the full architecture, including the scope/freshness trade-offs between the legacy pipeline and the daemon).vault_linter.py/update_master_plan.pyremain local-only β the daemon doesn't index vault metadata (titles, links, contradictions), only the concept graph.
- Comprehensive 7-stage health check across the entire vault.
- Runs SQLite index differential sync, multi-category Master Plan updates, dead wikilink & contradiction scans, invisible unicode (ZWSP) sanitation (
--fix), visual assets inspection, and Graphify rebuild.
- Recursively audits target vault folders for duplicate files, empty stubs, and missing canonical headers.
- Runs in safe dry-run mode by default, or with
--fixto archive duplicates/stubs into_archive/and inject canonical blockquotes. - Seamlessly integrates with
--syncto trigger an atomic SQLite index rebuild and Master Plan update.
A multi-modal spaced repetition and active recall engine:
create: Ingests notes, parses AST structures (markdown-it-py), extracts atomicKnowledgeUnits, applies SuperMemo 20-rules validation, writes<RootFolder>/study/<Deck Name>.md, and syncs directly to Anki via MCP.uv run python scripts/study_deck.py create --source "<folder_or_note>" --deck "<DeckName>" [--anki-deck "<AnkiName>"]
add: Dynamically binds new notes or source folders to an existing deck without duplicate card creation.uv run python scripts/study_deck.py add --source "<new_source>" --deck "<DeckName>"
update: Runs a Three-Way Merge across Markdown edits, SQLite persistence, and Anki. Automatically preserves human card modifications, suppresses user-deleted cards from reappearing, and pushes edits to Anki viaupdateNoteFieldswithout resetting SRS review intervals.uv run python scripts/study_deck.py update --deck "<DeckName>"sync-anki: Pushes pending or unsynced flashcards and frozen image media to Anki.uv run python scripts/study_deck.py sync-anki --deck "<DeckName>"recover: Scans the Two-Phase Commit (PREPARED->COMMITTED) journal to clean up or complete dangling transactions after system halts or unexpected crashes.uv run python scripts/study_deck.py recover [--clean-only]
10. High-Density Active Recall (HDAR) Book Flashcard Ingestion (scripts/book_flashcards_engine.py, /okc-bookFlashcards)
A specialized engine for processing entire technical books into active recall decks:
- Full Book Ingestion: Extracts chapters, sections, and figures directly from EPUB or PDF non-fiction books:
uv run python scripts/book_flashcards_engine.py --input "<path_to_epub_or_pdf>" --deck "<DeckName>"
- Direct Anki Sync & Overwrite Protection:
# Ingest and automatically push cards and media to Anki: uv run python scripts/book_flashcards_engine.py --input "<path>" --deck "<DeckName>" --sync-anki # Replace an existing deck cleanly: uv run python scripts/book_flashcards_engine.py --input "<path>" --deck "<DeckName>" --sync-anki --replace-deck
- Resumed Execution & Incremental Processing:
# Check current checkpoint status: uv run python scripts/book_flashcards_engine.py --deck "<DeckName>" --status # Re-compile multi-format outputs from existing checkpoint cards: uv run python scripts/book_flashcards_engine.py --deck "<DeckName>" --compile
11. Multiple-Choice Flashcard Engine & Validation Linter (scripts/study_multi_choice.py, scripts/validate_flashcards.py, /okc-study-multichoice)
A parameterized, idempotent active recall engine built on the DeckRegistry (ADR 0011):
validate: Enforces the 6-dimension rubric across multiple-choice decks (validating headers, isolated?delimiters, exactly 4 options AβD, answer key parity, and comprehensive distractor analysis):uv run python scripts/study_multi_choice.py validate --file "<path_to_deck.md>"sync-anki: Synchronizes multiple-choice questions to Anki with responsive styling:uv run python scripts/study_multi_choice.py sync-anki --file "<path_to_deck.md>"status: Reports card counts, model types, and target subdeck mapping viaDeckRegistry:uv run python scripts/study_multi_choice.py status --file "<path_to_deck.md>"
| Category | Technology | Purpose |
|---|---|---|
| Core AI Model | Gemini 3.1 Pro | Primary reasoning, text generation, and synthesis engine. |
| Agent Framework | Google Antigravity SDK | Tool calling, subagent orchestration, and skill workflows. |
| MCP Automation Server | FastMCP (mcp SDK) |
Native stdio server exposing 22 zero-permission ingestion, study, and vault commit tools. |
| Knowledge Graph | Graphify.net (graphifyy) |
Structural AST/wikilink graph indexing ($0.00 token cost). |
| Backend & Scripting | Python 3.12+, uv |
High-speed, deterministic local execution environment. |
| Media & Scraping | Chrome CDP/AppleScript, yt-dlp, Buzz CLI (Whisper), BeautifulSoup4 |
Dynamic CSR web scraping, Spotify/Instagram DOM bridges, audio extraction, and local speech-to-text. |
| Knowledge Base | Obsidian | Markdown-based local filesystem database. |
| Spaced Repetition & Study | Anki, AnkiConnect, Anki MCP, DeckRegistry |
AST parsers, declarative deck routing, active recall flashcards, and SRS scheduling. |
| AST Parsing | markdown-it-py |
Strict markdown tokenization and structure extraction. |
Future Improvements:
- Granular File-Level AST Caching: While
--fastprovides sub-second differential Master Plan and SQLite indexing, expand per-file hash caching across the full 27,000+ node AST graph parser to avoid any full graph rebuilding even on full syncs. - Local Hybrid Search (AST Graph + Offline Embeddings): Incorporate a local vector database using a lightweight model (e.g.,
sentence-transformersviauvin Python) to enable semantic search alongside structural AST link-graphs at zero API token cost. - Platform-Agnostic Whisper Fallback: Transition the macOS-only desktop
Buzz CLIWhisper dependency to a standalone python library (e.g.faster-whisperoropenai-whisperviauv) to enable headless/remote environment compatibility. - Structured Note Schema Validation: Use a Python schema validator (e.g., Pydantic) to ensure generated note structures strictly conform to vault conventions before writing, eliminating any potential markdown structure formatting drifts.
- Autoshared Session Scraping via CDP: Allow the headless Chrome browser to load user Chrome profiles or cookies in order to fetch paywalled or subscriber-only technical publications (e.g. paid Medium or Substack newsletters).
- Rejected Source Audit Logs: Log low technical density scorecards to
temp/rejected_sources.jsonfor batch human audits instead of halting ingestion processes on immediate interactive prompts.
Lessons Learned:
- AST vs. LLM Indexing: Using local Python AST parsers (
graphify_helper.py) to build graph relationships saves 95%+ in query latency and 100% in graph index cost compared to LLM-based graph extraction. - Quality Gates Matter: Implementing the
MIN_TECHNICAL_SCOREpre-fetch grader prevents low-quality web fluff from polluting thewiki/concept graph. - Decoupled Skill Architecture: Separating agent rules (
SKILL.md) from static compiled knowledge (KNOWLEDGE.md) maintains lightweight system prompts while giving agents instant access to 500+ concept cards. - Sandbox Boundaries Require Native FastMCP Handlers: Cortex agent environments intentionally restrict file modification tools (
write_to_file) to brain artifact directories (.gemini/antigravity/brain/). Implementing atomic FastMCP tools (commit_curated_note) bypasses permission roadblocks and delivers reliable 1-shot vault writes without terminal interruptions or subshell escaping issues. - Incremental vs. Full Rebuilds: Bypassing the 27,000-node AST graph crawl on single-note ingestions via
--fastcuts synchronization latency by 87% (from ~20s to ~2.5s) while preserving complete SQLite and Master Plan integrity. - Zero-Permission FastMCP Macro-Tools Eliminate Permission Fatigue: In high-governance agentic environments (Google Antigravity SDK, Claude Code), executing multi-command shell pipelines (
python3 -c, bash heredocs, or multiple sequential scripts) triggers excessive security permission prompts and risks context fragmentation. Bundling batch ingestion (fetch_playlist_data) and atomic multi-file writing (commit_curated_batch) into dedicated FastMCP tools cuts agent workflow latency, eliminates 100% of terminal popups, and enforces transactional vault consistency.
