Summary
#249's test plan runs:
pytest tests/test_mcp_server.py tests/test_kubeflow_pipeline_v4.py tests/test_canonical_rag_ingest.py tests/test_hugo_ingest.py
Three of those four files don't exist, on main or anywhere in the PR branch's history. The module that production uses to build the kubeflow_docs v4 collection is therefore untested:
Module (docs-agent-mcp/pipelines/utils/) |
Lines |
Public functions |
Referenced by any test |
canonical_rag_ingest.py |
818 |
|
|
milvus_store.py |
412 |
|
|
hugo_ingest.py |
97 |
|
|
| total |
1,327 |
28 |
0 |
"Referenced" here means the function name appears anywhere under tests/ (git grep -w on dedbc84), so this isn't just about missing imports. No test touches any of them, directly or indirectly.
The fourth file, test_mcp_server.py, exists but errors on all 63 tests (module 'docs_agent_mcp_server' has no attribute 'client') after the server split. So the retrieval side of v4 is currently unverified as well.
Related findings from the same check
- Evaluation evidence isn't in the repo.
docs/RAG_V4_ARCHITECTURE.md bases its promotion decision on files that are all absent from main: tests/retrieval_golden.json, tests/retrieval_baseline_portforward_analysis.md, tests/retrieval_candidate_8Gi_repeat3.json, artifacts/v4_router_comparison.json, artifacts/local_hybrid_v4_embedded.jsonl, plus the linked docs/local-milvus-runbook.md and MILVUS_INFRA_UPGRADE.md. As things stand, the reported hit@5 numbers (e.g. auto-router 80.4% / 51.0%) can't be reproduced from the repository.
- Two cleaners have no caller.
utils/hugo_ingest.py::clean_hugo_markdown isn't used; canonical_rag_ingest imports only parse_frontmatter and process_html_table. utils/utils.py::clean_content has no caller either, yet 19 of the 26 tests in test_pipeline_utils.py cover it. Tests against these look like coverage but don't exercise production.
- Stale anchor. The doc lists
docs-agent-mcp/pipelines/canonical_rag_ingest.py; the file is at docs-agent-mcp/pipelines/utils/canonical_rag_ingest.py.
- Test dependencies.
python-frontmatter, html-table-rescuer, markdownify and beautifulsoup4 are in pipelines/requirements.txt and Dockerfile.pipeline but not in requirements-test.txt. Any new test for this path would silently skip in PR Safety unless they're added.
Question first
@SanthoshToorpu, do the three test files exist on your machine? If so, committing them would be much faster and more faithful than rewriting them, since you know the intended behaviour best. The same goes for the golden set and eval artifacts.
If not, I'd like to write them
Starting small and working from documented behaviour (the v4 doc, docstrings and inline comments), not from whatever the code happens to do:
test_hugo_ingest.py: only the production surface, parse_frontmatter (YAML/TOML/BOM/missing/malformed frontmatter; the title/description/weight keys parse_canonical_document reads) and process_html_table (rowspan/colspan expansion for release tables, no-op without a table).
test_canonical_rag_ingest.py: section paths, chunking bounds, release_date extraction, doc_type.
test_milvus_store.py: the v4 schema/index shape against the table in the doc, plus the clean-rebuild safety gates.
Each would be a separate small PR, adding the test dependencies to requirements-test.txt so they run in CI rather than skip.
Also related: #258 fixes the collection abort that currently stops the whole suite from running on main.
Summary
#249's test plan runs:
Three of those four files don't exist, on
mainor anywhere in the PR branch's history. The module that production uses to build thekubeflow_docsv4 collection is therefore untested:docs-agent-mcp/pipelines/utils/)canonical_rag_ingest.pymilvus_store.pyhugo_ingest.py"Referenced" here means the function name appears anywhere under
tests/(git grep -wondedbc84), so this isn't just about missing imports. No test touches any of them, directly or indirectly.The fourth file,
test_mcp_server.py, exists but errors on all 63 tests (module 'docs_agent_mcp_server' has no attribute 'client') after the server split. So the retrieval side of v4 is currently unverified as well.Related findings from the same check
docs/RAG_V4_ARCHITECTURE.mdbases its promotion decision on files that are all absent frommain:tests/retrieval_golden.json,tests/retrieval_baseline_portforward_analysis.md,tests/retrieval_candidate_8Gi_repeat3.json,artifacts/v4_router_comparison.json,artifacts/local_hybrid_v4_embedded.jsonl, plus the linkeddocs/local-milvus-runbook.mdandMILVUS_INFRA_UPGRADE.md. As things stand, the reported hit@5 numbers (e.g. auto-router 80.4% / 51.0%) can't be reproduced from the repository.utils/hugo_ingest.py::clean_hugo_markdownisn't used;canonical_rag_ingestimports onlyparse_frontmatterandprocess_html_table.utils/utils.py::clean_contenthas no caller either, yet 19 of the 26 tests intest_pipeline_utils.pycover it. Tests against these look like coverage but don't exercise production.docs-agent-mcp/pipelines/canonical_rag_ingest.py; the file is atdocs-agent-mcp/pipelines/utils/canonical_rag_ingest.py.python-frontmatter,html-table-rescuer,markdownifyandbeautifulsoup4are inpipelines/requirements.txtandDockerfile.pipelinebut not inrequirements-test.txt. Any new test for this path would silently skip inPR Safetyunless they're added.Question first
@SanthoshToorpu, do the three test files exist on your machine? If so, committing them would be much faster and more faithful than rewriting them, since you know the intended behaviour best. The same goes for the golden set and eval artifacts.
If not, I'd like to write them
Starting small and working from documented behaviour (the v4 doc, docstrings and inline comments), not from whatever the code happens to do:
test_hugo_ingest.py: only the production surface,parse_frontmatter(YAML/TOML/BOM/missing/malformed frontmatter; thetitle/description/weightkeysparse_canonical_documentreads) andprocess_html_table(rowspan/colspan expansion for release tables, no-op without a table).test_canonical_rag_ingest.py: section paths, chunking bounds,release_dateextraction,doc_type.test_milvus_store.py: the v4 schema/index shape against the table in the doc, plus the clean-rebuild safety gates.Each would be a separate small PR, adding the test dependencies to
requirements-test.txtso they run in CI rather than skip.Also related: #258 fixes the collection abort that currently stops the whole suite from running on
main.