Skip to content

About

Autonomous, modular knowledge lifecycle agent built on the Antigravity 2.0 SDK and Model Context Protocol (MCP) for deterministic Obsidian graph curation.

Topics

Resources

Contributing

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

Β 

History

61 Commits

Folders and files

Repository files navigation

Obsidian Knowledge Curator β€” Antigravity 2.4.0

Curator Project Banner

Python Version Gemini 3.1 Pro Obsidian Anki Antigravity Graphify

An autonomous, agentic knowledge compiler designed to maintain and synthesize a "Second Brain" within an Obsidian Vault. Built natively with the Google Antigravity SDK and Gemini 3.1 Pro, this system automates the ingestion of technical articles, YouTube videos, and research papers, filtering content via an automated Technical Density Grader, indexing structural relationships via Graphify.net, and maintaining a decoupled skill-data architecture.


πŸ‘‹ For Non-Technical Readers

What this is. A personal automation system that turns a large, ever-growing pile of technical reading (articles, videos, podcasts, books) into a structured, searchable knowledge base β€” automatically. Instead of bookmarking content and never revisiting it, the system reads it, judges whether it's actually worth keeping, extracts the key ideas, cross-links them to everything already known, and flags contradictions as understanding evolves.

Who it's for. Built and maintained by one engineer for personal use β€” a working example of applying real software-engineering discipline (automated tests, architecture decision records, staged rollouts, health monitoring) to a personal productivity problem, not a commercial product or team tool.

Why it matters. It demonstrates hands-on experience with skills relevant well beyond this project: designing multi-stage automation pipelines, orchestrating multiple AI models efficiently (grading content before spending compute on it, avoiding unnecessary API costs), building systems that degrade gracefully when a data source fails or changes instead of breaking outright, and documenting the reasoning behind architectural decisions so the system stays legible over time (see docs/adr/).

Maturity. Actively developed and in daily personal use, managing a growing multi-thousand-note vault. Not intended for external users or production deployment β€” no support contract, no SLA, no multi-tenant design.

What's under the hood, in plain terms:

  • Reads and judges content before saving it β€” automatically scores incoming articles and videos for how substantive they are, so low-value content doesn't clutter the knowledge base.
  • Builds a live map of how ideas connect β€” every note is linked to related concepts automatically, without paying an AI model to work out the connections every time.
  • Catches contradictions β€” flags when a newer note disagrees with something recorded earlier, instead of silently letting outdated conclusions persist.
  • Runs a full health check on demand β€” one command audits the entire knowledge base (broken links, duplicate content, orphaned files, structural integrity) and can auto-repair common issues.
  • Keeps working when a data source misbehaves β€” web scraping has three fallback tiers, so a paywall or a blocked scraper doesn't just fail silently.

The rest of this document is the technical reference β€” architecture, setup, and workflow detail for engineers evaluating or extending the system. Start with Key Architectural Innovations for the plain-terms list above translated into implementation detail.


πŸ“– Table of Contents


🎯 Project Overview

This project serves as an advanced production implementation of the LLM Wiki Paradigm (inspired by Andrej Karpathy's concept of LLMs as knowledge compilers rather than simple chatbots).

Instead of treating the AI as a search engine over raw documents (like traditional RAG), the Obsidian Knowledge Curator acts as an active maintainer of a local filesystem database. It reads raw inputs, grades source density, extracts core concepts, updates existing Wiki pages, flags contradictions natively, and builds a chronological trace of evolving ideas across thousands of notes.


🧠 Problem Statement

The Engineering Challenge: Knowledge workers and AI Engineers consume vast amounts of technical content (papers, documentation, videos). Traditional PKM (Personal Knowledge Management) systems rely on manual synthesis, while modern LLM chatbots (chat-with-PDF) fail to compound knowledge over time because they lack persistent state and cross-document reasoning. Furthermore, naive LLM graph compilation introduces massive API token costs and high query latencies.

The Solution: A headless, zero-token-overhead automation pipeline that ingests content, parses transcripts, evaluates technical density, and uses an Offline AST Graphify Indexer to map vault relationships without external LLM API costs.


πŸ’‘ Key Architectural Innovations

1. Technical Density Ingestion Grader (MIN_TECHNICAL_SCORE)

Before any source (YouTube video, Web article, Tweet) is written to the vault, a 3,000-character preview is evaluated across three dimensions:

  • Information Density (lack of fluff, factual saturation).
  • Provenance & References (citations, data points, verified authors).
  • Technical Level (code architecture relevance, concrete implementations).

If the composite score falls below MIN_TECHNICAL_SCORE (configured in .env, default: 60), ingestion halts, presenting a detailed scorecard and summary to the user for explicit override confirmation.

2. Zero-Cost Offline Graphify Indexing & Hybrid Mapper (graphify_mapper.py)

To enable graph-aware context retrieval across 13,000+ notes without incurring API costs or latency penalties:

  • AST & Wikilink Parser: Uses local Python regex and Markdown AST parsing to extract document headings, parent-child nesting, and [[wikilinks]].
  • Hybrid Local Context Engine (GraphifyMapper): Pre-predicts target raw/ categories and matches existing wiki/ concepts (<50ms latency, $0.00 token cost) directly from graphify-out/graph.json. It falls back to lightweight LLM queries only if local graph matching confidence drops below 70%.
  • Injected Ingestion Context: Automatically injects "graphify_context" objects into temp/fetched_data.json across all ingestion scripts (fetch_article_data.py, fetch_youtube_data.py, fetch_twitter_data.py, fetch_book_data.py).
  • dswok Integration: Indexes protected external knowledge directories (dataScienceKnowledgeBase/dswok) as a read-only information graph without modifying any files within them.

3. Podcast & Audio Ingestion Pipeline (/okc-urlPodcast & fetch_podcast_data.py)

  • Multi-Platform Support: Ingests podcasts from Siemens.FM, Spotify, Apple Podcasts, RSS feeds, YouTube audio, and direct .mp3/.m4a files using yt-dlp and curl.
  • Offline Whisper Transcription: Uses local Buzz CLI (/Applications/Buzz.app/Contents/MacOS/Buzz) to transcribe audio tracks offline without external API costs.

4. Full DOM Article Extraction & FastMCP Native Tool (fetch_article_data.py & okc FastMCP)

  • Native FastMCP Tool: Exposed as @mcp.tool() fetch_article_data(url, download_all_images) on the okc FastMCP server (scripts/mcp_server.py) for zero-permission, in-process agent execution.
  • Full DOM Extraction Without Limits: Targets semantic containers (article, [itemprop='articleBody'], .post-content, .post__content, .body.markup, div.available-content, etc.) using BeautifulSoup and markdownify, extracting 100% of the body and eliminating the 5,000-character default truncation limit inherited from mcp-server-fetch.
  • Integrated Technical Quality Scoring: Automatically scores candidate articles (0-100 scale) across information density, provenance, code architecture depth, and educational infographics before ingestion.
  • Tolerant 3-Tier Fallback Chain: Tier 1 uses direct requests with realistic browser headers and SSL verification fallbacks (ssl.CERT_NONE); Tier 2 uses mcp-server-fetch configured with max_length: 10,000,000 for Cloudflare/Captcha bypass; Tier 3 uses search_web for locked paywalls.
  • Audio Redirection: Automatically detects podcast/audio URLs or audio media files and delegates execution to fetch_podcast_data.py.

5. Concurrent High-Density Image Preservation & Avatar Filtering (fetch_article_data.py & fetch_book_data.py)

  • Concurrent Asset Downloads: Uses ThreadPoolExecutor(max_workers=6) to concurrently fetch all article infographics and diagrams directly into <VAULT_ROOT>/assets/images/<slug>-<idx>-<alt_slug>.<ext>.
  • Modern CDN & <picture>/<noscript> Support: Preserves images hidden inside <picture><source srcset="..."> and <noscript><img ...> elements (e.g., Medium, Substack), using regex pattern matching on high-resolution CDN URLs (miro.medium.com) so no architectural diagrams are dropped.
  • Intelligent Avatar & Tracking Pixel Filter: Discards author profile pictures, emojis, badges, and tracking pixels (<100px) through DOM attribute inspection and real pixel dimension verification via PIL.Image.
  • Resilient CDN URL Sanitization: Cleans control characters (\x00-\x1f), whitespace, alt fragments, and comma-delimited Substack/Medium transform parameters before downloading.
  • Exact Obsidian Wikilink Embeddings: Replaces remote image links in generated Markdown with native Obsidian wikilinks ![[assets/images/<file>.png]] at their exact locations, with automatic C2PA/EXIF metadata sanitization.

6. Decoupled Skill Factory (SKILL.md vs KNOWLEDGE.md)

To prevent system prompt inflation and context degradation:

  • Behavior Prompt (SKILL.md): Contains pure agent execution rules, wikilink mandates, and non-hallucination constraints.
  • Compiled Static Database (KNOWLEDGE.md): An automatically regenerated index containing ~600+ concept cards with absolute file links across the vault.

7. OKC Doctor & Full Diagnostics Suite (/okc-doctor & scripts/okc_doctor.py)

A unified 7-stage health check, integrity audit, and synchronization suite:

  • SQLite Differential Index: Rapid scan and synchronization across 3,000+ files.
  • Multi-Category Master Plans: Dynamic regeneration of all navigation maps.
  • Wikilink & Contradiction Linter: Deep scan for broken links, orphan notes, and explicit [!contradiction] tags.
  • Invisible Unicode (ZWSP) Hygiene: Detects and sanitizes zero-width space characters with --fix.
  • Visual Asset Inspector: Audits assets/images/ total vs unreferenced visual assets.
  • Protected Zones Immutability Audit: Ensures zero unauthorized modifications across protected engineering zones.
  • Graphify & KNOWLEDGE.md Rebuild: Updates graph.json, graph_cache.json, and the concept index in one pass.

8. Note Normalizer & Header Sanitizer (scripts/normalize_notes.py, /okc-normalize)

An automated audit, deduplication, and header standardization engine (ADR 0006):

  • Content Hash Deduplication: Computes SHA-256 digests of markdown content to identify identical notes across folders, safely archiving duplicates into a local _archive/ directory.
  • Empty & Low-Value Stub Elimination: Detects and isolates notes lacking substantial content or structure, preserving a high signal-to-noise ratio in the knowledge graph.
  • Canonical Blockquote Injection: Detects missing metadata headers and non-destructively prepends canonical Obsidian blockquote schema (Author, Source, Type, Processed, and Tags: #no-read-yet).
  • Atomic Database & Master Plan Sync: Includes --sync to automatically update the SQLite differential index and Category Master Plans in a single operation.

9. Managed Study Decks & Anki MCP Synchronization Engine (scripts/study_deck.py, /okc-study)

An automated spaced repetition, active recall, and pedagogical flashcard synthesis engine (ADR 0007):

  • SuperMemo 20-Rules Compliance: Generates atomic Basic (Q/A) and Cloze ({{c1::...}}) flashcards grounded in verifiable SourceSpan evidence slices.
  • AST Parsing & Math/Table Fidelity: Uses markdown-it-py to parse source notes while maintaining LaTeX formulas ($...$, $$...$$), fenced code blocks, and markdown tables.
  • Direct Anki MCP & AnkiConnect Push: Automatically creates decks, uploads media assets, and synchronizes cards via the Anki MCP server or local AnkiConnect HTTP API (http://127.0.0.1:8765).
  • Three-Way Merge with Human Preservation: Retains user modifications to card fronts and backs in Markdown while updating source notes; records deleted cards in study_suppressions to prevent resurrection.
  • Deep Visual Architecture: Clones diagram and figure assets by SHA-256 hash into <VAULT_ROOT>/assets/images/study/, generating dedicated architectural visual recall cards.
  • Two-Phase Commit (2PC) Journal & Crash Recovery: Logs write transactions (PREPARED -> COMMITTED), performs atomic staging swaps, and provides study_deck.py recover to prevent partial state corruption.

10. High-Density Active Recall (HDAR) Book Flashcard Engine (scripts/book_flashcards_engine.py, /okc-bookFlashcards)

An end-to-end active recall flashcard generation and synchronization engine for technical books (ADR 0008):

  • HDAR 7-Rule Pedagogical Rubric: Enforces zero-hallucination factual grounding, atomic information retrieval, causal mechanisms over superficial facts, and balanced bilingual technical terminology.
  • Whole-Book Decomposition & Extraction: Parses EPUB and PDF non-fiction books into structured semantic chapters, section hierarchies, and embedded visual media.
  • Resilient Chapter Checkpointing: Maintains incremental state in checkpoint.json to allow resume-on-failure across 10+ chapters without reprocessing or data loss.
  • Automated Deduplication & Global Audit: Evaluates candidate cards across chapters, eliminates semantic duplicates, and outputs comprehensive audit reports (audit_report.md).
  • Multi-Platform Convergence: Simultaneously generates 1-click Anki import files (.tsv), structured JSON databases, native Obsidian study notes (<VAULT_ROOT>/.../study/<Deck Name>.md), and executes direct AnkiConnect synchronization with embedded diagrams.

11. Atomic FastMCP Note Curation & High-Speed Incremental Sync (scripts/mcp_server.py, scripts/sync_vault.py --fast)

An end-to-end atomic curation tool and sub-second incremental vault synchronizer (ADR 0009):

  • Atomic Curation MCP Tool (commit_curated_note): Writes the raw source note and multiple synthesized wiki concept notes directly to VAULT_ROOT in a single tool call. Eliminates permission boundaries associated with agent artifact tool restrictions (write_to_file) and avoids terminal prompt interrupts.
  • Fast Incremental Vault Synchronization (sync_vault.py --fast): Performs lightweight SQLite differential updates and dynamically rebuilds Master Plans while surgically updating modified notes in the Graphify AST cache, bypassing the full 27,000-node crawl and dropping sync latency from ~20s to ~2.5s.
  • Self-Contained Automated Ingestion: Combines article DOM fetching, diagram downloads, quality grading, note synthesis, and vault synchronization into a clean 2-step agent interaction loop.

12. Universal Zero-Permission FastMCP Ingestion & Multi-Note Batch Commit (scripts/mcp_server.py, scripts/fetch_playlist_data.py)

An architectural paradigm shift establishing a strict 2-step FastMCP workflow across all content modalities, eliminating 100% of terminal permission prompts and context truncation (ADR 0010):

  • Native FastMCP Suite (22 Tools): Expanded scripts/mcp_server.py to expose headless tools for all ingestion pipelines (fetch_playlist_data, fetch_youtube_data, fetch_article_data, fetch_twitter_data, fetch_podcast_data, fetch_spotify_data, fetch_doc_data, fetch_book_data, clean_staging_temp, inspect_flashcards_deck).
  • YouTube Playlist Batch Extractor (fetch_playlist_data): Parses whole playlist structures via yt-dlp --flat-playlist, extracting video metadata, titles, and full transcripts in a single tool call or CLI command (scripts/fetch_playlist_data.py), staging structured data in temp/fetched_playlist_data.json and disk-persisted transcript files in temp/playlist_transcripts/.
  • Atomic Multi-Note Batch Commit (commit_curated_batch): High-throughput atomic macro-tool that writes multiple raw source notes (raw_notes), an optional Series Master Plan (master_plan_path & master_plan_content), and all compiled wiki concepts (wiki_notes) directly to VAULT_ROOT in a single tool call, followed by instant --fast incremental synchronization.
  • Universal 2-Step Agent Protocol: Standardizes all ingestion skills (okc-urlPlaylist, okc-urlYoutube, okc-urlTwitter, okc-urlPodcast, okc-urlSpotify, okc-doc, okc-bookSummary) on a deterministic 2-turn cycle: Turn 1 (Fetch & Stage via FastMCP) -> Turn 2 (In-Memory Synthesis & Atomic Batch/Single Commit via FastMCP), completely eliminating subshell escaping vulnerabilities, Cortex artifact boundaries, and permission fatigue.

13. Spotify Chrome DOM Hydration Bridge (scripts/fetch_spotify_data.py, /okc-urlSpotify & FastMCP fetch_spotify_data)

An automated browser-driven extraction engine overcoming Widevine DRM encryption (MP4_128_CBCS) without paid third-party APIs (ADR 0011):

  • AppleScript Session Integration: Interacts directly with an active, user-authenticated Google Chrome session on macOS.
  • Automated Transcript View: Navigates to the Spotify episode, targets the native "Transcript" UI tab, and triggers live cue hydration.
  • Virtualized DOM Smooth-Scroll Harvesting: Executes a 35ms iterative smooth-scroll JavaScript loop over virtualized cue rows ([data-testid="transcript-cue-row"]), capturing 100% of timestamped spoken content while avoiding missing off-screen elements.
  • Zero-Loss Markdown Staging: Normalizes speaker cues, combines sequential sentences, and stages output to temp/fetched_data.json and temp/fetched_data.txt for atomic curation via commit_curated_note.

14. Instagram Saved Collections Batch Ingestion Engine (scripts/process_instagram_saved_batch.py, /okc-instagram)

A headless, high-throughput social video ingestion pipeline for curated Instagram collections (ADR 0011):

  • DOM Link Harvesting: Extracts saved reel URLs from the active browser session, bypassing anti-scraping blocks and login walls.
  • Audio Extraction & Offline Whisper Transcription: Uses yt-dlp to download reel audio streams and transcribes them offline using Whisper (base model via Buzz CLI or local Whisper) at $0.00 token cost.
  • Batch Atomic Curation & Master Plan Builder: Generates individual Obsidian notes conforming to vault standards in Leadership and Coach/raw/Instagram/, automatically computes creator statistics, and constructs a dedicated collection Master Plan (Master Plan β€” Santo Trabajo.md) with cross-links.

15. Modular Flashcards Subsystem & Declarative DeckRegistry (src/agent_tools/flashcards/, study_multi_choice.py)

A unified, AST-driven spaced repetition engine adhering to the "One Tool + N Inputs > N Scripts" standard (ADR 0011):

  • Unified AST & Regex Parsers (parsers.py): Robustly parses Basic TSV and 6-dimension Markdown multiple-choice cards (header, prompt, 4 options, isolated ? delimiter, answer key, and distractor analyses across options A–D).
  • Declarative DeckRegistry (registry.py): Centralizes canonical mappings for major technical courses and authors (Soledad Galli, Chip Huyen, ByteByteGo), infers categories dynamically, and resolves destination vault paths via resolve_study_path.
  • FastMCP Deck Inspection (inspect_flashcards_deck): Exposes live card counting, model validation, and subdeck inspection directly to the AI agent with zero terminal prompts.

πŸ— System Architecture

flowchart TD
    A["External Content (YouTube / Web / PDF)"] -->|Fetch Raw Data| B["Stage 1: Ingestion & Extraction"]
    B -->|3000-char Preview| C{"Technical Density Grader"}
    C -->|"Below Threshold (< MIN_TECHNICAL_SCORE)"| D["User Override Prompt (y/n)"]
    C -->|"Pass (>= MIN_TECHNICAL_SCORE)"| E["Antigravity Agent Context"]
    D -->|Approved| E
    
    E -->|Write Source Note| F["Zone 1: raw/"]
    E -->|Compile Concepts & Contradictions| G["Zone 2: wiki/"]
    
    F -->|Offline AST & Wikilink Extraction| H["Graphify Indexer (graphify_helper.py)"]
    G -->|Offline AST & Wikilink Extraction| H
    H -->|Update Structural Graph| I["graphify-out/graph.json (legacy)"]
    H -->|Regenerate Concept Cards| J["KNOWLEDGE.md Index Card"]
    
    F -.->|Live vault watch| M["graphify-daemon (resident process, MCP)"]
    M -.->|"query_graph / get_node -- always current"| N["Agents (Claude Code, Antigravity)"]
    M -.->|Periodic snapshot flush| Q["graphify-daemon/out/graph.json"]
    
    I -->|"GRAPHIFY_BACKEND=local/auto"| P["run_explore() / okc_doctor.py"]
    Q -->|"GRAPHIFY_BACKEND=remote/auto"| P
    
    K["Vault Linter / Health Check"] -.->|Scan Links & Orphans| G
    L["Sync Vault Pipeline"] -.->|Auto-Rebuild Master Plans| J
Loading

πŸš€ Getting Started & Agent Installation

1. System Requirements

This project runs locally and relies on Python 3.12+ and external command-line utilities.

  • Python Package Manager: uv (required for high-speed, isolated environment management).
  • Browser Scraping: Google Chrome or Chrome Canary (installed locally, required for dynamic CDP JS rendering).
  • Media Processing: ffmpeg (required by yt-dlp to extract audio streams).
  • Local Transcription Fallback: Buzz CLI (required for offline Whisper transcription fallback).

2. Installation Steps

Step 1: Install OS Prerequisites

macOS (via Homebrew):

# Install uv and ffmpeg
brew install uv ffmpeg

# Install Buzz (GUI + CLI)
brew install --cask buzz

# Install Graphify CLI
uv tool install "graphifyy[gemini]"

Linux (Ubuntu/Debian):

# Install uv
curl -LsSf https://astral.sh/uv/install.sh | sh

# Install ffmpeg
sudo apt update && sudo apt install -y ffmpeg

Step 2: Clone & Setup Workspace

# Clone the repository
git clone https://github.com/c-ibarra/obsidianKnowledgeCurator.git
cd obsidianKnowledgeCurator

# Install dependencies and setup virtual environment
uv sync

Step 3: Run Interactive Setup Wizard

Run the automated setup wizard to configure your local .env securely from .env.template, auto-detect your Obsidian Vault (~/Documents/Obsidian), and validate system dependencies:

# Interactive setup in terminal
uv run python scripts/setup_project.py

# Or within Antigravity chat:
/okc-setup

Alternatively, manually copy the template and edit your .env:

cp .env.template .env

Step 4: Initialize the Knowledge Graph & Skills

Build the initial structural graph across your vault and compile the KNOWLEDGE.md concept cards index:

# Verify vault health, rebuild Master Plans, and compile Graphify index
uv run python scripts/sync_vault.py --target-kb all

Step 5: Configure Anki Client & Anki MCP Server

To enable bi-directional spaced repetition flashcard generation and automatic synchronization (/okc-study):

  1. Install and Open Anki:

    • Download and install Anki.
    • Ensure Anki is running in the background during synchronization.
  2. Install the AnkiConnect Add-on:

    • In Anki, navigate to: Tools βž” Add-ons βž” Get Add-ons...
    • Enter the AnkiConnect code: 2055492159 and click OK.
    • Restart Anki.
  3. Configure AnkiConnect CORS / Allowed Origins:

    • Go to Tools βž” Add-ons, select AnkiConnect, and click Config.
    • Ensure webCorsOriginList includes local loopbacks and apiKey is empty (default) or matches your configuration:
      {
          "apiKey": null,
          "apiHost": "127.0.0.1",
          "apiPort": 8765,
          "webCorsOriginList": [
              "http://localhost",
              "http://127.0.0.1",
              "*"
          ]
      }
    • Restart Anki. Test connection by running curl http://127.0.0.1:8765 (should return "AnkiConnect").
  4. Configuring / Updating the Anki MCP Server in Antigravity / Claude Code:

    • The Anki MCP server is registered in your agent client configuration (mcp_config.json or Antigravity's settings).
    • If using npx or standard node MCP runner:
      {
        "mcpServers": {
          "anki": {
            "command": "npx",
            "args": ["-y", "@modelcontextprotocol/server-anki"],
            "env": {
              "ANKI_CONNECT_URL": "http://127.0.0.1:8765"
            }
          }
        }
      }
    • If using python-based anki-connect-mcp:
      {
        "mcpServers": {
          "anki": {
            "command": "uvx",
            "args": ["anki-connect-mcp"]
          }
        }
      }
    • To update the Anki MCP server, force cache refresh via npx -y @modelcontextprotocol/server-anki@latest or uvx --upgrade anki-connect-mcp.
    • Note: Even if the MCP server is idle or offline, scripts/study_deck.py automatically falls back to native direct HTTP calls against http://127.0.0.1:8765 to guarantee reliable synchronization without interruptions.

πŸ—‚ The "Zone" Knowledge Architecture

The vault is strictly divided into four zones to separate immutable sources from synthesized concepts:

Zone Purpose Agent Permissions
raw/ Immutable sources (Video transcripts, Web clippings). Append-Only. The agent saves summaries here but never modifies historical sources.
wiki/ Synthesized concepts and entities. Read-Write. Fully maintained by the LLM. The agent creates pages, injects wikilinks, and merges updates.
dev/ Architecture Decision Records (ADRs) and project files. Collaborative. The agent acts as a co-pilot but requires explicit human approval to modify.
dswok/ Protected external personal knowledge base. Read-Only / Indexed. The agent scans and indexes relationships into graph.json but never writes or modifies files.

πŸ”„ End-to-End Workflow

1. Multimedia, Video & Podcast Ingestion (fetch_youtube_data.py, fetch_podcast_data.py, fetch_twitter_data.py, fetch_spotify_data.py & FastMCP)

  • YouTube (/okc-urlYoutube): FastMCP tool fetch_youtube_data extracts audio streams via yt-dlp (--live-from-start), transcribing via youtube-transcript-api with fallback to local Buzz CLI Whisper. Curates in memory and commits via commit_curated_note.
  • Twitter/X (/okc-urlTwitter): FastMCP tool fetch_twitter_data extracts tweet text and attached video audio, transcribing via Buzz Whisper and committing atomically.
  • Podcasts & Audio (/okc-urlPodcast): FastMCP tool fetch_podcast_data ingests Siemens.FM, Apple Podcasts, RSS feeds, or direct .mp3/.m4a files via yt-dlp/curl and Buzz Whisper. Automatically delegates Spotify URLs to the Chrome DOM bridge.
  • Spotify Podcasts (/okc-urlSpotify): FastMCP tool fetch_spotify_data leverages AppleScript and Chrome DOM hydration scrolling to extract 100% of timestamped transcripts, bypassing Widevine DRM (MP4_128_CBCS) without audio decoding fees.
  • Pre-calculates "graphify_context" via GraphifyMapper to select raw target categories and link existing wiki concept notes.
  • Atomically writes curated notes to raw/ and compiled concepts to wiki/ via commit_curated_note.

2. YouTube Playlist Batch Ingestion (fetch_playlist_data.py, FastMCP & /okc-urlPlaylist)

  • Native FastMCP Batch Extractor (fetch_playlist_data): Extracts the full playlist architecture via yt-dlp --flat-playlist and pulls transcripts for all items into temp/fetched_playlist_data.json and temp/playlist_transcripts/<id>.txt in a single tool call.
  • In-Memory Batch Curation: Synthesizes all sequential raw source notes, compiles cross-video wiki concepts, and generates the Series Master Plan without context thrashing.
  • Atomic Multi-Note Commit (commit_curated_batch): Commits all raw notes, the Series Master Plan, and wiki concepts in a single zero-permission transaction, followed by instant --fast incremental vault synchronization.

3. High-Fidelity Web Article Ingestion (fetch_article_data.py, FastMCP & /okc-urlArticle)

  • Native FastMCP Tool (fetch_article_data): Headless, zero-terminal-popup execution via stdio FastMCP server (okc).
  • Full DOM Extraction: Targets semantic content containers (article, [itemprop='articleBody'], .post-content, .body.markup, etc.) with BeautifulSoup and markdownify, extracting 100% of article body without 5,000-character truncation limits.
  • Concurrent Image Preservation: Concurrently downloads infographics via ThreadPoolExecutor(max_workers=6), preserves <picture>/<noscript> images from Medium and Substack, filters out author avatars (<100px), sanitizes CDN URLs, and injects exact Obsidian wikilinks ![[assets/images/...]].
  • Integrated Quality Grading: Evaluates technical density and architectural relevance (0-100 score) before note generation.
  • Streamlined 2-Step Curation Flow:
    1. fetch_article_data(url="...", download_all_images=True) stages content and downloads media.
    2. commit_curated_note(raw_note_path="...", raw_content="...", wiki_notes={...}, fast_sync=True) atomically writes raw and wiki notes, updates SQLite index, and runs incremental synchronization.
  • Audio Auto-Detection: Automatically detects audio/podcast URLs and delegates to fetch_podcast_data.py.
  • Graphify Context Enrichment: Pre-calculates "graphify_context" via GraphifyMapper for raw target categories and wiki concept linking.

4. Non-Fiction Book Ingestion & Synthesis (fetch_book_data.py & okc-bookSummary)

  • High-Density Actionable Synthesis (HDAS Standard): Produces rich, 7-section modular chapter notes combining dense narrative prose with active learning tools (Central Thesis & 1-Sentence Insight, Inquiry Questions, Enriched Summary Development with Mental Models, Visual Metaphor / Key Quote / Common Pitfall Callouts, Smart Cross-Domain Commentary, Practical Application Guide with 15-minute challenges, and Executive 1-Sentence Takeaway).
  • Automated Structure Parsing: Ingests PDF, EPUB, DOCX, and TXT files, segmenting chapters and sanitizing text.
  • Default Obsidian Vault Output: Writes main book summaries and individual chapter notes (Chapter XX β€” <Title>.md) directly to VAULT_ROOT/dataScienceKnowledgeBase/<Category>/raw/books/.
  • Enforced Chapter Depth (1,600–2,650 words total): Requires dense explanatory narrative for Section 3 (Enriched Summary Development) to ensure thorough technical and conceptual depth.
  • Visual Content & Mermaid.js Diagrams: Reconstructs mindmaps, architecture flows, sequence diagrams, and embeds extracted figures (assets/images/).
  • Executive Master Note Hub: Compiles comprehensive master notes with full-book architecture mindmaps, mental models index, #flashcard spaced repetition cards, and specialized glossaries.
  • Automatic Temporary File Cleanup: Cleans up all working files in temp/ via --clean upon completion.

5. Office Document & Format Ingestion (fetch_doc_data.py & /okc-doc)

  • Multi-Format Extraction: Ingests .docx, .pptx, .xlsx, .epub, .pdf, .odt, and .csv using the unified AnyDoc engine (src/agent_tools/anydoc_engine.py).
  • Embedded Asset Extraction: Extracts embedded figures and charts directly to <VAULT_ROOT>/assets/images/ and links them with native Obsidian wikilinks.

6. Structural Graph Sync (graphify_helper.py, sync_vault.py & graphify-daemon)

  • Incrementally updates graphify-out/graph.json after every note edit (legacy pipeline).
  • Supports --fast / --incremental flag (sync_vault.py --fast) to synchronize SQLite indexes and dynamic Master Plans while performing surgical graph updates without crawling all 27,000+ nodes.
  • Automatically regenerates .agents/skills/obsidian-knowledge-curator/KNOWLEDGE.md with updated concept links.
  • Rebuilds Category Master Plans and audits wikilink health via vault_linter.py.
  • graphify-daemon (a separately-maintained resident process, ~/projects/graphify-daemon) serves the same concept graph over MCP from an always-current in-RAM snapshot β€” no rebuild step, republishes on every vault batch. Agents (Claude Code, Antigravity) query it directly via its MCP tools (query_graph, get_node, shortest_path, etc.) whenever it's running.
  • scripts/knowledge_commands.py's run_explore() and scripts/okc_doctor.py's dashboard node/edge counts read from the daemon by default too, controlled by GRAPHIFY_BACKEND in .env (auto/local/remote β€” see ADR 0005 and .agents/rules/graphify.md for the full architecture, including the scope/freshness trade-offs between the legacy pipeline and the daemon). vault_linter.py/update_master_plan.py remain local-only β€” the daemon doesn't index vault metadata (titles, links, contradictions), only the concept graph.

7. Full Diagnostics & Auto-Repair (scripts/okc_doctor.py, /okc-doctor)

  • Comprehensive 7-stage health check across the entire vault.
  • Runs SQLite index differential sync, multi-category Master Plan updates, dead wikilink & contradiction scans, invisible unicode (ZWSP) sanitation (--fix), visual assets inspection, and Graphify rebuild.

8. Note Normalization & Deduplication (scripts/normalize_notes.py, /okc-normalize)

  • Recursively audits target vault folders for duplicate files, empty stubs, and missing canonical headers.
  • Runs in safe dry-run mode by default, or with --fix to archive duplicates/stubs into _archive/ and inject canonical blockquotes.
  • Seamlessly integrates with --sync to trigger an atomic SQLite index rebuild and Master Plan update.

9. Study Decks, Flashcard Synthesis & Anki MCP Pipeline (scripts/study_deck.py, /okc-study)

A multi-modal spaced repetition and active recall engine:

  • create: Ingests notes, parses AST structures (markdown-it-py), extracts atomic KnowledgeUnits, applies SuperMemo 20-rules validation, writes <RootFolder>/study/<Deck Name>.md, and syncs directly to Anki via MCP.
    uv run python scripts/study_deck.py create --source "<folder_or_note>" --deck "<DeckName>" [--anki-deck "<AnkiName>"]
  • add: Dynamically binds new notes or source folders to an existing deck without duplicate card creation.
    uv run python scripts/study_deck.py add --source "<new_source>" --deck "<DeckName>"
  • update: Runs a Three-Way Merge across Markdown edits, SQLite persistence, and Anki. Automatically preserves human card modifications, suppresses user-deleted cards from reappearing, and pushes edits to Anki via updateNoteFields without resetting SRS review intervals.
    uv run python scripts/study_deck.py update --deck "<DeckName>"
  • sync-anki: Pushes pending or unsynced flashcards and frozen image media to Anki.
    uv run python scripts/study_deck.py sync-anki --deck "<DeckName>"
  • recover: Scans the Two-Phase Commit (PREPARED -> COMMITTED) journal to clean up or complete dangling transactions after system halts or unexpected crashes.
    uv run python scripts/study_deck.py recover [--clean-only]

10. High-Density Active Recall (HDAR) Book Flashcard Ingestion (scripts/book_flashcards_engine.py, /okc-bookFlashcards)

A specialized engine for processing entire technical books into active recall decks:

  • Full Book Ingestion: Extracts chapters, sections, and figures directly from EPUB or PDF non-fiction books:
    uv run python scripts/book_flashcards_engine.py --input "<path_to_epub_or_pdf>" --deck "<DeckName>"
  • Direct Anki Sync & Overwrite Protection:
    # Ingest and automatically push cards and media to Anki:
    uv run python scripts/book_flashcards_engine.py --input "<path>" --deck "<DeckName>" --sync-anki
    
    # Replace an existing deck cleanly:
    uv run python scripts/book_flashcards_engine.py --input "<path>" --deck "<DeckName>" --sync-anki --replace-deck
  • Resumed Execution & Incremental Processing:
    # Check current checkpoint status:
    uv run python scripts/book_flashcards_engine.py --deck "<DeckName>" --status
    
    # Re-compile multi-format outputs from existing checkpoint cards:
    uv run python scripts/book_flashcards_engine.py --deck "<DeckName>" --compile

11. Multiple-Choice Flashcard Engine & Validation Linter (scripts/study_multi_choice.py, scripts/validate_flashcards.py, /okc-study-multichoice)

A parameterized, idempotent active recall engine built on the DeckRegistry (ADR 0011):

  • validate: Enforces the 6-dimension rubric across multiple-choice decks (validating headers, isolated ? delimiters, exactly 4 options A–D, answer key parity, and comprehensive distractor analysis):
    uv run python scripts/study_multi_choice.py validate --file "<path_to_deck.md>"
  • sync-anki: Synchronizes multiple-choice questions to Anki with responsive styling:
    uv run python scripts/study_multi_choice.py sync-anki --file "<path_to_deck.md>"
  • status: Reports card counts, model types, and target subdeck mapping via DeckRegistry:
    uv run python scripts/study_multi_choice.py status --file "<path_to_deck.md>"

πŸ›  Technology Stack

Category Technology Purpose
Core AI Model Gemini 3.1 Pro Primary reasoning, text generation, and synthesis engine.
Agent Framework Google Antigravity SDK Tool calling, subagent orchestration, and skill workflows.
MCP Automation Server FastMCP (mcp SDK) Native stdio server exposing 22 zero-permission ingestion, study, and vault commit tools.
Knowledge Graph Graphify.net (graphifyy) Structural AST/wikilink graph indexing ($0.00 token cost).
Backend & Scripting Python 3.12+, uv High-speed, deterministic local execution environment.
Media & Scraping Chrome CDP/AppleScript, yt-dlp, Buzz CLI (Whisper), BeautifulSoup4 Dynamic CSR web scraping, Spotify/Instagram DOM bridges, audio extraction, and local speech-to-text.
Knowledge Base Obsidian Markdown-based local filesystem database.
Spaced Repetition & Study Anki, AnkiConnect, Anki MCP, DeckRegistry AST parsers, declarative deck routing, active recall flashcards, and SRS scheduling.
AST Parsing markdown-it-py Strict markdown tokenization and structure extraction.

πŸš€ Future Improvements & Lessons Learned

Future Improvements:

  1. Granular File-Level AST Caching: While --fast provides sub-second differential Master Plan and SQLite indexing, expand per-file hash caching across the full 27,000+ node AST graph parser to avoid any full graph rebuilding even on full syncs.
  2. Local Hybrid Search (AST Graph + Offline Embeddings): Incorporate a local vector database using a lightweight model (e.g., sentence-transformers via uv in Python) to enable semantic search alongside structural AST link-graphs at zero API token cost.
  3. Platform-Agnostic Whisper Fallback: Transition the macOS-only desktop Buzz CLI Whisper dependency to a standalone python library (e.g. faster-whisper or openai-whisper via uv) to enable headless/remote environment compatibility.
  4. Structured Note Schema Validation: Use a Python schema validator (e.g., Pydantic) to ensure generated note structures strictly conform to vault conventions before writing, eliminating any potential markdown structure formatting drifts.
  5. Autoshared Session Scraping via CDP: Allow the headless Chrome browser to load user Chrome profiles or cookies in order to fetch paywalled or subscriber-only technical publications (e.g. paid Medium or Substack newsletters).
  6. Rejected Source Audit Logs: Log low technical density scorecards to temp/rejected_sources.json for batch human audits instead of halting ingestion processes on immediate interactive prompts.

Lessons Learned:

  1. AST vs. LLM Indexing: Using local Python AST parsers (graphify_helper.py) to build graph relationships saves 95%+ in query latency and 100% in graph index cost compared to LLM-based graph extraction.
  2. Quality Gates Matter: Implementing the MIN_TECHNICAL_SCORE pre-fetch grader prevents low-quality web fluff from polluting the wiki/ concept graph.
  3. Decoupled Skill Architecture: Separating agent rules (SKILL.md) from static compiled knowledge (KNOWLEDGE.md) maintains lightweight system prompts while giving agents instant access to 500+ concept cards.
  4. Sandbox Boundaries Require Native FastMCP Handlers: Cortex agent environments intentionally restrict file modification tools (write_to_file) to brain artifact directories (.gemini/antigravity/brain/). Implementing atomic FastMCP tools (commit_curated_note) bypasses permission roadblocks and delivers reliable 1-shot vault writes without terminal interruptions or subshell escaping issues.
  5. Incremental vs. Full Rebuilds: Bypassing the 27,000-node AST graph crawl on single-note ingestions via --fast cuts synchronization latency by 87% (from ~20s to ~2.5s) while preserving complete SQLite and Master Plan integrity.
  6. Zero-Permission FastMCP Macro-Tools Eliminate Permission Fatigue: In high-governance agentic environments (Google Antigravity SDK, Claude Code), executing multi-command shell pipelines (python3 -c, bash heredocs, or multiple sequential scripts) triggers excessive security permission prompts and risks context fragmentation. Bundling batch ingestion (fetch_playlist_data) and atomic multi-file writing (commit_curated_batch) into dedicated FastMCP tools cuts agent workflow latency, eliminates 100% of terminal popups, and enforces transactional vault consistency.

About

Autonomous, modular knowledge lifecycle agent built on the Antigravity 2.0 SDK and Model Context Protocol (MCP) for deterministic Obsidian graph curation.

Topics

Resources

Contributing

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages