AegisAI is a secure, enterprise-oriented knowledge platform in development. It is being built to let organizations ingest internal content, retrieve it safely, and eventually chat with it through a permission-aware RAG experience.
The backend foundation, document-ingestion boundary, background-processing runtime, and text-processing pipeline are complete: containerized FastAPI services, PostgreSQL, JWT authentication, database-backed RBAC, enterprise SSO, secure document management, Redis/Celery workers, and traceable extracted text/chunks. Embeddings, retrieval, and chat follow.
| Capability | Status | Outcome |
|---|---|---|
| API and local platform | Available | FastAPI API, PostgreSQL 16, Qdrant, Docker Compose, health checks, and Alembic migrations. |
| Local authentication | Available | User registration, bcrypt password hashing, short-lived access JWTs, rotatable refresh tokens, logout, and inactive-user protection. |
| Authorization | Available | Local roles and permissions, administrator bootstrap, and request-time RBAC enforcement. |
| Enterprise SSO | Available | Google OpenID Connect, GitHub OAuth, and Microsoft Entra ID adapters with PKCE, signed state, nonce validation, account linking, and local AegisAI sessions. |
| Document ingestion | Available | RBAC-protected upload, metadata management, local persistent original-file storage, SHA-256 integrity metadata, and soft deletion. |
| Background processing | Available | Redis/Celery workers verify durable uploaded sources outside HTTP requests, with PostgreSQL-backed job state, retries, cancellation, and failure handling. |
| Knowledge processing | Available | Workers safely extract supported files, normalize text, create deterministic chunks, and persist traceable output for later embedding. |
| Vector indexing | Available | Workers queue and process OpenAI embeddings into validated Qdrant collections with traceable PostgreSQL records, cleanup, and safe progress visibility. |
| Retrieval and RAG | Planned | Retrieval, citations, and RAG chat. |
Qdrant is already provisioned as local infrastructure. Phase 6 stores original document bytes in the persistent local document_data volume and metadata in PostgreSQL; Phase 9.6 automatically indexes document vectors after extraction when OPENAI_API_KEY is configured.
| Area | Technology |
|---|---|
| API | FastAPI and Uvicorn |
| Application and data layer | Python 3.12, SQLAlchemy 2.x, Alembic |
| Relational database | PostgreSQL 16 |
| Vector database | Qdrant |
| Identity and authorization | Passlib/bcrypt, python-jose JWT, local RBAC, OAuth 2.0/OpenID Connect adapters |
| Configuration and validation | Pydantic v2 and Pydantic Settings |
| Local platform | Docker and Docker Compose |
| Milestone | Status | Delivered or planned outcome |
|---|---|---|
| Phases 1–5 — Foundation, data, identity, and access control | Complete | Containerized backend, migrations, local authentication, RBAC, and enterprise SSO. |
| Phase 6 — Document ingestion | Complete | Secure local storage, upload validation, metadata lifecycle, RBAC enforcement, and document-management APIs. |
| Phase 7 — Background processing | Complete | Redis/Celery runtime, durable outbox delivery, worker integrity checks, job status, retry, and cancellation. |
| Phase 8 — Text extraction and chunking | Complete | Safe TXT/Markdown/PDF/DOCX extraction, normalized traceable chunks, worker lifecycle, reprocessing, and RBAC-protected inspection APIs. |
| Phase 9 — Embeddings and Qdrant indexing | Complete | OpenAI embedding boundary, Qdrant collection safety, durable indexing and cleanup jobs, traceable vector records, and safe status visibility. |
| Phases 10–12 — Retrieval and RAG | Planned | Metadata-filtered retrieval, streaming chat with citations, and permission-aware retrieval. |
| Phases 13–16 — Governance and product operations | Planned | Audit logging, administration UI, web frontend, and observability. |
| Phases 17–20 — Production scale | Planned | CI/CD, Kubernetes, multi-tenancy, API keys, rate limits, and retention controls. |
- RBAC design explains the current role and permission model.
- Document ingestion design defines the implemented Phase 6 storage, lifecycle, authorization, and API contract.
- Background processing design defines the implemented Phase 7 job, outbox, worker, and retry contract.
- Text extraction and chunking design defines the implemented Phase 8 format, lifecycle, traceability, safety, and manual-verification contract.
- Embeddings and Qdrant indexing design defines the Phase 9 vector, lifecycle, idempotency, and safety contract.
Available now
Browser, CLI, or future frontend
│
▼
FastAPI API :8000
│
┌─────────┼───────────────────────────────────────────────┐
│ │ │
▼ ▼ ▼
Local login Enterprise SSO Protected route
or refresh Google | GitHub | Entra dependency
│ │ │
└────┬────┘ ▼
▼ authenticate access JWT
AuthService / SsoAccountService │
│ ▼
▼ evaluate local RBAC policy
AegisAI access + refresh tokens │
│ ▼
└───────────────► PostgreSQL ◄──────────── allow or deny request
users
refresh_tokens
external_identities
roles / permissions
Documents ──► PostgreSQL outbox ──► Redis ──► Celery workers
│
▼
source integrity ──► extraction and chunks
│
▼
PostgreSQL extraction/chunk records
Planned next
embeddings ──► Qdrant ──► permission-aware retrieval + RAG chat
The backend follows a layered design so that HTTP, business rules, and persistence remain independently testable:
API routes and dependencies → services → repositories → PostgreSQL
│
schemas and security
Services own transaction boundaries. Repositories add, query, flush, and delete records but do not independently commit, so related changes either commit together or roll back together.
Authentication establishes a local AegisAI user. Authorization then decides whether that user may perform a specific action.
Password login or verified SSO identity
│
▼
local AegisAI user and session
access JWT + persisted refresh token
│
▼
Authorization: Bearer <access JWT>
│
▼
get_current_user validates token and loads user
│
▼
require_permission checks PostgreSQL-backed RBAC
│
┌────────┴────────┐
▼ ▼
HTTP 403 route handler
The access JWT contains identity and token metadata, not permissions. Each permission-aware request checks PostgreSQL, so a role or permission change applies immediately instead of waiting for an old JWT to expire.
RBAC is represented by the following relationships:
users ──< user_roles >── roles ──< role_permissions >── permissions
A user can have several roles, and a role can grant several permissions. The seeded administrator system role has every currently defined permission. Local roles are the only authorization source: SSO provider roles, groups, and access tokens are never copied into AegisAI authorization decisions.
SSO supports Google, GitHub, and Microsoft Entra ID. The browser flow uses short-lived signed state, PKCE, and OIDC nonce validation where applicable. Provider tokens are used only to verify identity; AegisAI issues and stores its own tokens.
The local-account policy is deliberately conservative:
- An existing unique
(provider, provider_subject)binding always resolves to its linked AegisAI user. - A new provider identity can link to an existing local user only when the provider supplies a verified email that exactly matches that user.
- Otherwise, a verified-email identity receives a just-in-time local AegisAI account and identity binding.
- An identity without a verified email is rejected rather than allowed to create or take over an account.
Just-in-time SSO users receive an unknown, cryptographically random password hash. This keeps the existing user model consistent without creating a password credential that anyone knows. An inactive user cannot create or refresh an AegisAI session.
- Docker Engine and Docker Compose
- A free local port for each of
8000,5432, and6333
Create the local configuration file once:
cp backend/.env.example backend/.envSet a long, unique JWT_SECRET_KEY in backend/.env before using the application outside local experimentation. Keep SSO disabled until at least one provider is configured.
Then run the one canonical startup command from the repository root:
docker compose up --build --force-recreateThat command builds the backend image, runs unit tests during image build, validates the Alembic migration chain, waits for PostgreSQL to become healthy, runs the tests again, applies alembic upgrade head to PostgreSQL, and starts Uvicorn. The backend does not start if its tests or migration fail.
When startup completes:
| Service | URL |
|---|---|
| API | http://localhost:8000 |
| Health check | http://localhost:8000/health |
| Interactive OpenAPI docs | http://localhost:8000/docs |
| PostgreSQL | localhost:5432 |
| Qdrant API | http://localhost:6333 |
# Follow backend startup and application logs
docker compose logs -f backend
# Check service state
docker compose ps
# Stop services while retaining PostgreSQL and Qdrant volumes
docker compose down
# Repeat the complete build, test, migrate, and start workflow
docker compose up --build --force-recreateDocker Compose is the supported local-development workflow. It is not yet a production deployment recipe; production hardening, observability, CI/CD, Kubernetes, and multi-tenancy are planned later.
Copy backend/.env.example to backend/.env. Do not commit backend/.env, OAuth client secrets, JWT secrets, refresh tokens, or access tokens.
| Setting | Purpose |
|---|---|
APP_NAME, APP_VERSION, APP_ENV |
Application identity and environment label. |
HOST, PORT |
Backend listener configuration. Compose exposes port 8000. |
DATABASE_URL |
PostgreSQL connection URL. Inside Compose, the hostname must remain postgres. |
QDRANT_URL, QDRANT_API_KEY, QDRANT_COLLECTION_NAME |
Qdrant connection and active derived-vector collection. A key is optional for local Docker. |
EMBEDDING_PROVIDER, EMBEDDING_MODEL, EMBEDDING_VECTOR_DIMENSION |
Active embedding shape. Changing the dimension requires a new collection and deliberate reindex. |
OPENAI_BASE_URL, OPENAI_API_KEY |
OpenAI embedding endpoint and secret. The key is needed only when Phase 9 worker indexing is enabled. |
DOCUMENT_STORAGE_PATH |
Local original-document storage path. Compose mounts the persistent document_data volume at this path. |
DOCUMENT_MAX_UPLOAD_BYTES |
Maximum streamed upload size. The default is 25 MiB and is enforced by the upload service. |
DOCUMENT_MAX_EXTRACTED_TEXT_CHARACTERS |
Maximum parser output retained from one document; default 5,000,000 characters. |
DOCUMENT_CHUNK_TARGET_CHARACTERS, DOCUMENT_CHUNK_OVERLAP_CHARACTERS |
Model-neutral Phase 8 chunking defaults: 1,200 target characters with 200 characters of context overlap. |
JWT_SECRET_KEY |
Long, unique secret used to sign AegisAI access and refresh JWTs. |
JWT_ALGORITHM |
JWT signing algorithm; the supplied configuration uses HS256. |
ACCESS_TOKEN_EXPIRE_MINUTES, REFRESH_TOKEN_EXPIRE_DAYS |
Local token lifetimes. |
SSO is disabled by default. Enable it only after configuring one provider application and registering its exact redirect URI.
| Setting | Purpose |
|---|---|
SSO_ENABLED |
Enables provider-based browser sign-in. |
SSO_CALLBACK_BASE_URL |
Public API base URL used to build provider redirect URIs. Use HTTPS in deployed environments. |
SSO_STATE_SECRET_KEY |
A distinct long random secret for signed, temporary SSO state. Do not reuse JWT_SECRET_KEY. |
SSO_TRANSACTION_EXPIRE_MINUTES |
Short expiry for state, PKCE, and nonce transaction data. |
GOOGLE_CLIENT_ID, GOOGLE_CLIENT_SECRET |
Google OpenID Connect web-application credentials. |
GITHUB_CLIENT_ID, GITHUB_CLIENT_SECRET |
GitHub OAuth application credentials. |
MICROSOFT_ENTRA_CLIENT_ID, MICROSOFT_ENTRA_CLIENT_SECRET |
Microsoft Entra ID application credentials. |
MICROSOFT_ENTRA_TENANT_ID |
A tenant ID to restrict Entra sign-in, or organizations only when multi-tenant organizational access is intended. |
Register one exact callback URL for each configured provider:
{SSO_CALLBACK_BASE_URL}/auth/sso/google/callback
{SSO_CALLBACK_BASE_URL}/auth/sso/github/callback
{SSO_CALLBACK_BASE_URL}/auth/sso/microsoft/callback
For production, use a publicly reachable HTTPS callback base URL, distinct random secrets, and a tenant-specific Entra ID whenever access should be limited to one organization.
OpenAPI documentation is available at http://localhost:8000/docs. It is the complete live endpoint contract; the summary below highlights the current API surface.
| Method | Path | Purpose |
|---|---|---|
GET |
/ |
Service metadata. |
GET |
/health |
Application health check. |
GET |
/database/health |
PostgreSQL connectivity check. |
POST |
/auth/register |
Create a local email/password user. |
POST |
/auth/login |
Exchange OAuth2 form credentials for AegisAI tokens. |
GET |
/auth/me |
Return the authenticated local user. |
POST |
/auth/refresh |
Rotate a valid refresh token and return a new pair. |
POST |
/auth/logout |
Soft-revoke a refresh token. |
GET |
/auth/sso/{provider} |
Start a browser SSO flow for google, github, or microsoft. |
GET |
/protected |
Minimal protected-route example. |
/rbac/* |
See the RBAC section below | Manage roles and assignments with administrator permissions. |
POST |
/documents |
Upload an allowed document with documents:write. |
GET |
/documents?offset=0&limit=25 |
List document metadata with documents:read. |
GET, PATCH, DELETE |
/documents/{document_id} |
Inspect with documents:read; rename or delete with documents:write. |
GET |
/documents/{document_id}/processing-jobs |
Inspect safe job history with documents:read. |
POST |
/documents/{document_id}/processing-jobs/{job_id}/retry |
Requeue one failed job with documents:write. |
GET |
/documents/{document_id}/indexing-status |
Inspect current vector progress and safe indexing state with documents:read. |
GET |
/documents/{document_id}/extraction |
Inspect safe extraction metadata with documents:read. |
GET |
/documents/{document_id}/extraction/chunks |
Inspect ordered, paginated chunks with documents:read. |
POST |
/documents/{document_id}/reprocess |
Queue replacement extraction with documents:write; returns 202 Accepted. |
- Register through
POST /auth/register, or create a local user through a verified SSO flow. - Use
POST /auth/loginwith OAuth2 form data (usernameis the email) for password login. - Send the returned access token to protected routes.
- Send the refresh token only to
/auth/refreshor/auth/logout; do not use it as a bearer token.
curl http://localhost:8000/auth/me \
-H 'Authorization: Bearer YOUR_ACCESS_TOKEN'POST /documents accepts one multipart file field. PDFs, DOCX, TXT, and
Markdown are supported, and the default streamed limit is 25 MiB. The server
derives the title, generates the storage key, and records the authenticated
uploader; clients never supply a storage path or uploader ID.
Document reads and writes use the existing global documents:read and
documents:write permissions. uploader_user_id is provenance for future
tenant and resource policies, not a current per-document access rule. See the
document ingestion design for the complete API,
lifecycle, and storage behavior.
After an upload, the durable outbox sends source validation and then text
extraction to the worker. A successful extraction makes the document READY.
Readers can inspect only extraction metadata and bounded chunk pages; original
storage keys, broker identifiers, and parser errors stay internal. See the
text extraction and chunking design
for the lifecycle and a manual verification walkthrough.
Set OPENAI_API_KEY in the uncommitted backend/.env to enable real vector
generation. After a document reaches READY, its worker queues an
embedding_indexing job. Inspect its safe progress without exposing vector,
Qdrant, or provider details:
curl http://localhost:8000/documents/DOCUMENT_ID/indexing-status \
-H 'Authorization: Bearer YOUR_ACCESS_TOKEN'The caller needs documents:read. A succeeded response means every current
chunk has a traceable vector in the configured collection. If indexing_status
is failed, inspect the safe job history and retry that specific failed job
with documents:write; do not retry while it is queued or running.
Changing EMBEDDING_MODEL or EMBEDDING_VECTOR_DIMENSION requires a new
QDRANT_COLLECTION_NAME and deliberate reprocessing. Never delete or alter an
existing collection just to make a changed configuration fit.
Start browser SSO by visiting, for example:
http://localhost:8000/auth/sso/google
After a successful provider sign-in, AegisAI returns its own access and refresh tokens, plus the local user and provider name. The callback response is marked Cache-Control: no-store, and the temporary SSO transaction cookie is cleared.
To test an SSO session in Swagger:
- Copy only the returned
access_token. - Open
/docsand select Authorize. - Choose AegisAI access token.
- Paste the raw JWT without the
Bearerprefix; Swagger adds that header prefix. - Call
GET /auth/meor another protected endpoint.
The separate OAuth2PasswordBearer option in Swagger is for local password login. Both documentation options reach the same JWT validation and RBAC checks.
The current permission catalogue contains:
documents:read documents:write
users:read users:manage
roles:read roles:manage roles:assign
The /rbac management API requires both an active local user and the indicated database-backed permission.
| Endpoint group | Required permission | Purpose |
|---|---|---|
GET /rbac/permissions, GET /rbac/roles, role-permission reads |
roles:read |
View the permission catalogue, roles, and grants. |
| Role creation/deletion and role-permission changes | roles:manage |
Manage non-system roles and their permissions. |
| User-role reads | users:read |
View a user's role assignments. |
| User-role assignment/removal | roles:assign |
Grant or revoke roles. |
Bootstrap the first administrator only after that user exists locally (through registration or SSO):
docker compose exec backend python -m scripts.bootstrap_administrator admin@example.comThe command is idempotent. New SSO users have no role by default, so protected RBAC management calls correctly return HTTP 403 until an administrator grants an appropriate local role.
Run backend Python commands from backend/; that directory makes the app package importable.
cd backend
venv/bin/python -m unittest discover -s tests -vThe unit suite uses isolated SQLite databases and mocks where appropriate. It covers API handlers, services, repositories, JWT handling, refresh-token rotation, RBAC enforcement, SSO provider adapters, account linking, session issuance, document cleanup, background-job state, extraction, chunking, reprocessing, embedding validation and idempotency, Qdrant collection safety, Swagger security schemes, migrations, and application startup.
The Dockerfile runs this suite during image build and produces the complete Alembic upgrade SQL. Compose runs the suite again before applying migrations and launching the API.
Compose normally applies migrations automatically at startup. For development work, run Alembic from backend/:
cd backend
venv/bin/alembic history
venv/bin/alembic revision --autogenerate -m "describe the change"
venv/bin/alembic upgrade head
venv/bin/alembic currentReview every generated migration before applying it, particularly constraint and index changes. The current migration chain creates users, refresh tokens, RBAC tables and seeded permissions, the administrator system role, external-identity bindings, documents, processing/outbox records, extraction/chunk records, embedding pointers, and durable vector-cleanup requests.
When running Alembic from the host, use a database URL reachable from the host—normally localhost, not Compose's internal postgres hostname. ALEMBIC_DATABASE_URL can override the configured database URL for that command.
.
├── backend/
│ ├── alembic/ # Migration environment and revisions
│ ├── app/
│ │ ├── api/ # HTTP routes and FastAPI dependencies
│ │ ├── core/ # Settings, logging, and domain exceptions
│ │ ├── db/ # Engine, sessions, and declarative base
│ │ ├── integrations/ # External provider adapters, including SSO
│ │ ├── models/ # SQLAlchemy models
│ │ ├── repositories/ # Database queries and persistence operations
│ │ ├── schemas/ # Request and response contracts
│ │ ├── security/ # Password hashing, JWTs, RBAC guards, SSO state
│ │ └── services/ # Transactional application use cases
│ ├── scripts/ # Startup and administrator bootstrap commands
│ ├── tests/ # Unit and API-boundary tests
│ ├── Dockerfile
│ └── .env.example
└── docker-compose.yaml # Local API, PostgreSQL, and Qdrant stack
- Keep
backend/.env, OAuth secrets, JWTs, and refresh tokens out of source control, issue trackers, screenshots, and shared terminal output. - Rotate a token if it is exposed.
/auth/logoutrevokes its refresh token; the associated access token remains valid only until its configured short expiry. - Keep
JWT_SECRET_KEYandSSO_STATE_SECRET_KEYdistinct to limit the impact of a compromised secret. - Register exact HTTPS OAuth callback URLs in deployed environments. OAuth providers reject mismatched redirect URIs.
- Provider access tokens are not returned as AegisAI session tokens and are not used for AegisAI authorization.
- The current project has no published vulnerability-reporting policy. Do not disclose security-sensitive material in a public issue.
The next implementation milestones are:
- Phase 10: retrieval and metadata filtering.
- Phase 11: RAG chat, streaming, and citations.
- Phase 12: permission-aware retrieval.
- Phase 13: audit logging.
- Phase 14: administration dashboard.
- Phase 15: Next.js frontend.
- Phase 16: observability.
- Phase 17: CI/CD.
- Phase 18: Kubernetes.
- Phase 19: multi-tenancy.
- Phase 20: enterprise API keys, rate limits, and retention policies.