Skip to content

About

FastAPI backend for trustworthy RAG workflows, with document ingestion, PostgreSQL retrieval, source-grounded LLM answers, provider failure handling, Docker Compose, and answer-run audit readback.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

fastapi-rag-agent-backend

Production-shaped FastAPI backend for trustworthy RAG/agent workflows.

This project demonstrates the backend engineering behind controlled document ingestion, PostgreSQL retrieval, source-grounded answer generation, provider failure handling, and durable answer-run audit traces.

The goal is not to present RAG as magic document intelligence. The goal is to show an inspectable backend pipeline:

ingest documents -> chunk/store text -> retrieve evidence -> generate grounded answers -> persist/read back audit records

Current status

Implemented so far:

  • FastAPI application structure
  • GET /health health endpoint
  • GET /status/db database status endpoint
  • POST /documents/upload endpoint for single UTF-8 .txt document upload
  • Batch .txt folder importer via scripts/import_text_documents.py
  • Deterministic plain-text chunking
  • Content SHA-256 hashing
  • PostgreSQL persistence for documents and document chunks
  • Content-hash duplicate detection for ingested document content
  • Explicit document source_type provenance for upload and batch import paths
  • GET /documents/search endpoint using PostgreSQL full-text search over stored chunks
  • PostgreSQL full-text GIN index over chunk text
  • POST /answers endpoint for source-grounded answer generation
  • OpenAI-backed LLM client behind a small project-level LLMClient boundary
  • Source chunks returned with generated answers
  • Controlled LLM-provider failure handling
  • Durable answer_runs audit records for success, no_sources, and llm_error outcomes
  • answer_run_id returned from POST /answers
  • GET /answers/{answer_run_id} endpoint for persisted answer-run audit readback
  • SQLAlchemy ORM model structure
  • Alembic database migrations
  • Explicit runtime configuration validation
  • pytest coverage for health/status, document processing, persistence, retrieval, upload, answer generation, provider failure handling, and answer-run audit behaviour
  • Ruff lint configuration
  • Modern pyproject.toml dependency setup
flowchart TD
    upload["POST /documents/upload"] --> ingest["Validate UTF-8 text, Hash content, Chunk text"]
    batch["Batch .txt importer"] --> ingest

    ingest --> docs["documents table"]
    ingest --> chunks["document_chunks table, FTS GIN index"]

    search["GET /documents/search"] --> retrieval["PostgreSQL full-text retrieval"]
    retrieval --> chunks
    retrieval --> search_response["Matching chunks + metadata"]

    answer["POST /answers"] --> answer_service["Answer generation service"]
    answer_service --> retrieval
    answer_service --> prompt["Build grounded prompt"]
    prompt --> llm["LLM client boundary"]
    llm --> answer_response["Answer + source chunks + answer_run_id"]

    answer_service --> runs["answer_runs table"]
    readback["GET /answers/{answer_run_id}"] --> runs
Loading

Local development

Create and activate a virtual environment, then install the project with development dependencies:

python -m pip install -e ".[dev]"

Start the local PostgreSQL database:

docker compose up -d db

Configure local settings in .env:

DATABASE_URL=postgresql+psycopg://rag_agent:rag_agent_dev_password@localhost:5432/rag_agent
DATABASE_CONNECT_TIMEOUT_SECONDS=2
OPENAI_API_KEY=
OPENAI_MODEL=gpt-4o-mini
OPENAI_TIMEOUT_SECONDS=30

DATABASE_URL is required. The application and Alembic both load the database URL from the same settings source.

POST /answers requires OPENAI_API_KEY to be set. The key should be stored only in your local .env file and must not be committed to the repository. If the key is blank, non-LLM endpoints can still run, while answer generation returns a controlled 503 response.

Apply database migrations:

python -m alembic upgrade head

Run tests:

python -m pytest

Run linting:

python -m ruff check .

Run the API locally:

python -m uvicorn app.main:app --reload

Health check:

GET http://127.0.0.1:8000/health

Expected response:

{"status":"ok"}

Database status check:

GET http://127.0.0.1:8000/status/db

End-to-end local demo

This demo shows the core workflow:

ingest document -> search chunks -> generate grounded answer -> read back audit record

Start the local database and apply migrations:

docker compose up -d db
python -m alembic upgrade head

Import the sample policy document:

python scripts/import_text_documents.py --input-dir .\sample_docs

Start the API:

python -m uvicorn app.main:app --reload

Search the stored chunks:

GET http://127.0.0.1:8000/documents/search?query=refund&limit=5

Generate a source-grounded answer:

POST http://127.0.0.1:8000/answers

Example request body:

{
  "question": "Are digital products refundable?",
  "limit": 5
}

Example response shape:

{
  "answer_run_id": 1,
  "answer": "Digital products are only refundable if the download link has not been used.",
  "source_count": 1,
  "sources": [
    {
      "chunk_id": 1,
      "document_id": 1,
      "filename": "refund_policy.txt",
      "chunk_index": 0,
      "chunk_text": "Refund Policy...",
      "rank": 0.123
    }
  ]
}

Read back the persisted audit record:

GET http://127.0.0.1:8000/answers/1

The readback endpoint returns the stored answer-run record, including the original question, final answer text where applicable, retrieval limit, retrieved chunk count, source chunk IDs, outcome, error type where applicable, and timestamp.

Document ingestion

The current MVP supports two controlled ingestion paths.

1. Single document upload

The API accepts one UTF-8 plain text file at a time:

POST /documents/upload

Current boundary:

  • accepts .txt / text/plain
  • expects valid UTF-8 text
  • applies a 1 MB upload limit
  • stores the document and derived chunks in PostgreSQL
  • records source_type="upload"

The upload endpoint is intended for controlled operator/admin ingestion, not bulk browser-based document management.

2. Batch text document import

The batch importer imports .txt files from a target folder:

python scripts/import_text_documents.py --input-dir .\sample_docs

The importer:

  • scans the target folder for .txt files
  • reads each file as UTF-8 text
  • calculates the document content SHA-256 hash
  • compares hashes against already ingested documents
  • skips duplicate document content
  • stores new documents and chunks in PostgreSQL
  • records source_type="batch_import"

Duplicate detection is content-based. The importer is not a file synchronisation system: it does not track source file paths, replace old document versions, or watch folders continuously.

Retrieval and answer generation

Stored document chunks can be searched directly:

GET /documents/search?query=refund&limit=5

The search endpoint uses PostgreSQL full-text search over persisted document chunks and returns matching chunks with document/source metadata and rank information.

Source-grounded answers can be generated from retrieved chunks:

POST /answers

Example request body:

{
  "question": "What is the refund policy?",
  "limit": 5
}

The answer endpoint:

  • retrieves relevant stored chunks
  • builds a grounded prompt from those chunks
  • calls the configured LLM client
  • returns the generated answer
  • includes the source chunks used as grounding evidence
  • returns answer_run_id
  • persists an answer-run audit record

Successful and no-source answer runs persist the final answer text returned to the caller, along with the question, retrieval limit, retrieved chunk count, source chunk IDs, outcome, error type where applicable, and timestamp.

Failed LLM-provider runs are recorded with their outcome and error type, but without answer text because no answer was returned.

The persisted audit record can be read back by ID:

GET /answers/{answer_run_id}

If the LLM provider fails during answer generation, the API returns a controlled 502 response rather than exposing provider-specific exception details.

If the LLM client is not configured because the API key is blank or missing, the API returns a controlled 503 response for answer generation.

Database migrations

This project uses SQLAlchemy models and Alembic migrations for database schema management.

  • SQLAlchemy ORM models live in app/db/models.py.
  • The shared ORM metadata root is defined in app/db/base.py.
  • Alembic reads Base.metadata via alembic/env.py.
  • alembic/env.py imports the ORM models so autogenerate can compare the database schema against the mapped tables.
  • PostgreSQL schema changes are applied through versioned Alembic migration files in alembic/versions/.

After changing SQLAlchemy models, generate a migration:

python -m alembic revision --autogenerate -m "describe schema change"

Review the generated migration before applying it. If Alembic tries to drop unrelated project tables, stop and check model registration before applying the migration.

Apply pending migrations:

python -m alembic upgrade head

Check the current database revision:

python -m alembic current

Do not manually edit the live database schema outside Alembic during normal development.

Testing approach

The test suite is split around behaviour boundaries.

Real database integration tests are used where the database behaviour is the point of the test:

  • document persistence
  • duplicate content constraints
  • PostgreSQL full-text retrieval
  • answer-run persistence
  • answer-run readback
  • database status behaviour

Fakes are used where the external dependency is not the behaviour under test:

  • LLM provider calls
  • provider failure simulation
  • API orchestration around already-tested service boundaries

This keeps the tests focused on project behaviour rather than testing invented shapes inside the test suite.

Design rationale

The current MVP deliberately uses PostgreSQL full-text search before vector retrieval. Full-text search gives a deterministic, inspectable retrieval baseline for exact terms, policy phrases, IDs, and document vocabulary before adding embeddings, embedding costs, vector indexes, and semantic ranking behaviour.

The LLM is isolated behind a small LLMClient boundary. The answer-generation workflow depends on the project boundary, not directly on the OpenAI SDK.

Answer responses include source chunks, and answer-run audit records persist the final answer text where applicable plus source chunk IDs. This gives a reviewer a clear trace from API response to stored audit record to source evidence without duplicating source chunk text inside the audit table.

Planned Phase 1 scope

Phase 1 is focused on a practical RAG/LLM backend.

Implemented Phase 1 capabilities now include:

  • document ingestion
  • text chunking
  • PostgreSQL persistence
  • PostgreSQL full-text retrieval
  • LLM answer generation
  • source traceability
  • controlled provider error handling
  • answer-run audit persistence
  • answer-run audit readback
  • local end-to-end demo path

Remaining Phase 1 work:

  • API usage polish where needed
  • final review for stale wording or overclaims

Later enhancements:

  • pgvector semantic retrieval
  • hybrid full-text/vector retrieval
  • retry policy refinement
  • authentication
  • multi-tenancy
  • background queue workers
  • AWS deployment
  • formal PDP/MCP access-control layer
  • React dashboard

About

FastAPI backend for trustworthy RAG workflows, with document ingestion, PostgreSQL retrieval, source-grounded LLM answers, provider failure handling, Docker Compose, and answer-run audit readback.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages