BankBot-AI is a local demo banking assistant built with Streamlit, a rule-first query-policy layer, optional local Ollama generation, a JSON demo store, and a small Flask API. It is an interview project, not a production banking system.
- Hybrid chatbot routing: domain gating, deterministic account intents, and an Ollama fallback for banking-policy questions.
- Deterministic balance, transactions, spending, profile, transfer guidance, greeting, goodbye, and help responses using the signed-in demo account data.
- Streamed Ollama responses with timeout, unavailable-service, malformed-stream, empty-response, and unavailable-model handling.
- Plotly/Pandas Streamlit dashboard and persisted chat history in the demo JSON database.
- Flask endpoints for health, login, balance, transactions, transfer, and logout.
- bcrypt PIN verification, token sessions with inactivity expiry, login rate limiting, input validation, and JSON API error responses.
- Offline routing/factuality evaluation, a 700-record policy-derived benchmark, and a component end-to-end benchmark.
Streamlit UI (bankbot.py)
-> query policy (app/services/query_policy.py)
-> deterministic response from current demo account data
-> or one streamed Ollama request, then response-policy validation
-> JSON demo store (bank_db.json)
Flask API (app/api.py)
-> shared security helpers (security.py)
-> JSON demo store (isolated temporary stores in tests/benchmarks)
The classifier is rule-first, not a separate ML classifier. It normalizes common phrasings and typos, blocks strong off-topic signals, reserves banking policy questions for the LLM, and selects a single deterministic action only when safe.
Python 3.10+ is recommended. Create an environment, then install the one canonical dependency file:
python -m venv .venv
.\.venv\Scripts\Activate.ps1
pip install -r requirements.txtConfiguration is read from environment variables (and .env when present):
DATABASE_FILE=bank_db.json
OLLAMA_URL=http://localhost:11434
OLLAMA_MODEL=llama3.2
OLLAMA_TIMEOUT=120
SESSION_TIMEOUT_MINUTES=15
MAX_LOGIN_ATTEMPTS=5
LOCKOUT_MINUTES=15
Use a non-default SECRET_KEY outside local development. Do not commit .env
or a database containing real credentials or customer data.
For local LLM features, install Ollama separately and make the configured model available, for example:
ollama pull llama3.2
ollama serveStart the Streamlit application:
streamlit run bankbot.pyStart the API in another terminal:
python -m flask --app app.api run --host 127.0.0.1 --port 5000API endpoints:
| Method | Endpoint | Purpose |
|---|---|---|
GET |
/api/health |
Health response |
POST |
/api/login |
Issue a Bearer token from a 10-digit account and 4-digit PIN |
GET |
/api/balance |
Authenticated account balance |
GET |
/api/transactions?limit=N |
Authenticated transaction list; N must be positive |
POST |
/api/transfers |
Authenticated JSON transfer with recipient and numeric amount |
POST |
/api/logout |
Invalidate the supplied Bearer token |
Transfers are constrained to ₹1–₹50,000 and persist an atomic JSON-file update. They are demo operations only; recipients are not real bank accounts.
Validate the static benchmark schema and uniqueness:
python tests/validate_benchmark_700.pyRun the reproducible offline benchmark. It measures domain gating, intent routing, deterministic/mock-data answer checks, safety checks, and routing latency. It does not call Ollama:
python evaluate_metrics.py tests/benchmark_700.json --output benchmark_700_results.jsonRun the small, hand-authored account-response regression set. It covers date and limit phrasing, malformed ranges, ambiguous requests, and a third-party account request. It is a development regression set, not an independent held-out result:
python evaluate_metrics.py tests/account_accuracy_cases.json --output account_accuracy_results.jsonRun the component benchmark (deterministic chatbot, isolated transfer contract, in-process Flask API, security, and reliability checks):
python e2e_benchmark.py --output e2e_benchmark_results.jsonRun the small live reference evaluation only when Ollama is running:
python e2e_benchmark.py --eval-ollama --output e2e_live_results.jsonRun the selected 30- or 100-query live-routing subsets when a larger local-model sample is appropriate:
python evaluate_metrics.py tests/benchmark_ollama_smoke.json --eval-ollama --output ollama_smoke_results.json
python evaluate_metrics.py tests/benchmark_ollama_100.json --eval-ollama --output ollama_100_results.jsonRun all automated tests:
python -m unittest discover -s tests -vThe reports keep these dimensions separate:
- Routing accuracy/F1: expected domain and rule intent versus the current query-policy result.
- Deterministic answer correctness: only answers verified against the mock account data. LLM-bound records are excluded from this denominator.
- LLM transport and policy validation: successful streaming and a policy- accepted response are not semantic correctness.
- LLM semantic correctness: scored only where an explicit authoritative reference criterion exists; otherwise it remains unscored/human-review.
- Latency: routing latency is not model generation time. Live reports expose
time-to-first-token, completed generation, and LLM end-to-end timing separately.
The component benchmark separately reports isolated transfer logic and
in-process API request timing. Its
latencysection reports p50, p95, mean, and sample count for routing, deterministic balance/transaction responses, bcrypt login, warm authenticated API reads, and login-plus-balance requests.
The 700-query dataset is versioned in tests/benchmark_700.json; its manifest
states clearly that it is newly authored but policy-derived, not an independent
held-out benchmark. Do not cite its score as general real-world accuracy.
On the local single-process fixture described in the report, the offline 700-record run produced 99.00% intent-routing accuracy and 98.99% weighted F1. Its data-grounded answer denominator was 465 records: balance was 40/40 and transactions 31/31. These figures meet the stated balance/transaction targets for this policy-derived mock-data workload; they are not independent production accuracy estimates.
The repeated latency workload recorded routing p50/p95 of 0.95/1.36 ms (n=30), deterministic balance response p95 of 0.02 ms (n=10), deterministic transaction response p95 of 0.10 ms (n=10), and warm authenticated balance API p95 of 1.47 ms (n=10). Bcrypt login p95 was 472.15 ms and login-plus-balance API p95 was 463.63 ms (n=10); bcrypt's work factor was not reduced. No source or timing boundary for the historical 45 ms/120 ms figures exists in this repository. Treat the warm routing/API results as comparable local targets, not network-load or LLM latency claims.
- JSON storage, in-memory sessions, and in-memory rate limiting are suitable only for a single-process demo; use a transactional database and shared session/rate limit store for deployment.
- Flask test-client timings exclude network, TLS, reverse proxy, and concurrent load effects.
- The application is not a real payments system and has no recipient verification, idempotency key, audit service, or external banking integration.
- The LLM can still produce unsupported banking statements; response policy checks are guardrails, not a factuality guarantee.
- The benchmark does not prove security. It provides regression checks for the specific controls implemented here.