Skip to content

bug(pipelines): three of four default code-crawl directories 404 and the run still succeeds #245

Description

@Neilblaze

Problem

code-pipeline.py builds the code_rag collection behind search_kubeflow_code. Three faults make it skip most of what it is pointed at, and none of them shows up as a failure.

1. Three of the four default directories no longer exist

directory_paths: str = "apps/pipeline/upstream,apps/katib,common/istio,apps/jupyter",

kubeflow/manifests moved apps/ to applications/ in June 2025:

apps/pipeline/upstream  -> 404      apps/katib   -> 404
apps/jupyter            -> 404      common/istio -> 10 entries

api_request prints API error: HTTP 404 and returns None; get_files_recursive turns that into an empty list; the loop prints Found 0 files in apps/katib/ and carries on. The component exits 0 with only common/istio indexed: 84 of the 570 files the corrected paths reach.

The same path swallows 5xx (never retried) and post-retry rate-limit exhaustion, so any transient failure during a long crawl leaves a partial index while the run stays green.

2. Every file over 1 MiB is stored empty

file_resp = api_request(item["url"])
if file_resp and "content" in file_resp:
    content = base64.b64decode(file_resp["content"]).decode("utf-8")

The Contents API only base64-encodes blobs under 1 MiB. Above that it answers HTTP 200 with "encoding": "none" and "content": "". The key is present, so the guard passes, b64decode("") gives b"", and the file is written to the dataset with no content. chunk_and_embed_code then drops it as Skipping tiny file.

$ gh api repos/kubeflow/manifests/contents/applications/kserve/kserve/upstream/kserve.yaml
size=7051337  encoding=none  content_len=0

Five YAML files in the repo are over the cap, 22.1 MB in total, all install manifests:
applications/kserve/.../kserve_kubeflow.yaml (7.06 MB), kserve.yaml (7.05 MB), experimental/ray/.../resources.yaml (3.98 MB), kserve-crds.yaml (2.31 MB), applications/spark/.../resources.yaml (1.72 MB). None sit under the current defaults, so this bites the moment anyone points the crawl at KServe, Spark or Ray.

3. Embedding responses are not length-checked

vectors = response.json()
for idx, vector in enumerate(vectors):
    batch[idx]["embedding"] = vector

utils.embed_texts raises when the service returns a different number of vectors than inputs. The inline copy dropped that check, so a short response leaves records with no embedding key, which surfaces one component later as a bare KeyError in store_code_milvus.

Suggested fix

Read blobs from the download_url the directory listing already returns, which has no size cap and removes the per-file API call; fail the component when a path cannot be read; refresh the defaults; restore the length check.


P.S.: I'll shortly open a PR addressing this. Thanks!

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

No labels
No labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions