Skip to content
Merged
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion requirements-cloud.txt
Original file line number Diff line number Diff line change
Expand Up @@ -2,4 +2,4 @@

# The optional cloud storage stack supports Python 3.10 through 3.12.
# Pin wt-data-platform-sdk so LanceDB dependencies remain reproducible.
wt-data-platform-sdk @ git+https://github.com/AI45Lab/wt-data-platform-sdk.git@v0.4.0
wt-data-platform-sdk @ git+https://github.com/AI45Lab/wt-data-platform-sdk.git@v0.4.1

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔒 Security & Privacy | 🟡 Minor | ⚡ Quick win

🧩 Analysis chain

🏁 Script executed:

#!/bin/bash
set -euo pipefail

git ls-remote --tags \
  https://github.com/AI45Lab/wt-data-platform-sdk.git \
  'refs/tags/v0.4.1*'

Repository: AI45Lab/SAfactory

Length of output: 274


🌐 Web query:

pip Git URL Git commit hash tag support @ tag commit

💡 Result:

Pip supports installing from Git repositories by specifying a Git reference (such as a branch, tag, or commit hash) using the @ symbol appended to the URL [1][2][3]. The standard syntax is: ProjectName @ git+https://git.example.com/MyProject.git@ Supported formats for include: - Branch names (e.g., @main, @develop) [1][4][5] - Tag names (e.g., @v1.0) [1][6] - Commit hashes (e.g., @da39a3ee5e6b4b0d3255bfef95601890afd80709) [1][2][3] - Other Git references (e.g., @refs/pull/123/head) [1][2][3] Key considerations: - Specification Logic: Pip follows a specific resolution order when encountering a reference [7]. It first attempts to resolve the provided string as a branch or tag name [7][8]. If no matching branch or tag is found, it proceeds to treat the string as a commit hash [7][8]. - Full vs. Partial Hashes: While pip accepts both, it is highly recommended to use the full 40-character commit hash [1][2][3]. Using a full hash is more efficient as it often allows pip to avoid extra network calls or redundant operations [1][2][3]. - Immutability: For repeatable builds and secure, deterministic installs, you should always prefer pinning to specific tags or full commit hashes rather than moving branch names [9][5]. - Requirements Files: You can use this same syntax directly within a requirements.txt file [9][5]. When used there, the Direct URL syntax (PackageName @ URL) is preferred for clarity and reproducibility [9][5]. Example usage: pip install "MyProject @ git+https://git.example.com/MyProject.git@v1.0" pip install "MyProject @ git+https://git.example.com/MyProject.git@da39a3ee5e6b4b0d3255bfef95601890afd80709"

Citations:


Pin the Git dependency to a fixed commit hash.

wt-data-platform-sdk currently installs from the mutable v0.4.1 Git tag. Change the requirement to the fixed commit and keep the version identifier in a comment for traceability.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@requirements-cloud.txt` at line 5, Update the wt-data-platform-sdk
requirement in requirements-cloud.txt to reference the requested fixed commit
hash instead of the mutable v0.4.1 tag, and retain v0.4.1 as an inline comment
for traceability.

Source: MCP tools


🗄️ Data Integrity & Integration | 🟠 Major | 🏗️ Heavy lift

🧩 Analysis chain

🏁 Script executed:

#!/bin/bash
set -euo pipefail

# Expect: every production deployment sets an explicit production profile
# or supplies explicit production table configuration.
rg -n -C 4 --hidden --glob '!.git' \
  'WT_SDK_PROFILE|TableConfig\s*\(|GatewayConfig\s*\(|WT_SDK_LANDING_TABLE|WT_SDK_SERVING_TABLE' .

Repository: AI45Lab/SAfactory

Length of output: 4774


🏁 Script executed:

#!/bin/bash
set -euo pipefail

printf '--- requirements-cloud.txt ---\n'
cat -n requirements-cloud.txt

printf '\n--- relevant cloud_strategy_impl.py sections ---\n'
sed -n '1,190p' core/data_manager/strategy/cloud_strategy_impl.py
printf '\n--- cloud_strategy_impl.py lines 330-380 ---\n'
sed -n '330,380p' core/data_manager/strategy/cloud_strategy_impl.py

printf '\n--- all WT_SDK references and files ---\n'
rg -n --hidden --glob '!.git' 'WT_SDK|wind_tunnel_(landing|serving)|landing_test|serving_test' .

printf '\n--- workflow/env files candidates ---\n'
git ls-files | rg '(^\.github/workflows/|\.env|env|docker-compose|compose|Dockerfile|deploy|deployment|readme|requirements)' || true

printf '\n--- diff stat/name status ---\n'
git diff --stat
git diff -- requirements-cloud.txt | sed -n '1,200p'

Repository: AI45Lab/SAfactory

Length of output: 246


🌐 Web query:

AI45Lab wt-data-platform-sdk v0.4.1 WT_SDK_PROFILE production test default

💡 Result:

The wt-data-platform-sdk is a component used within the AI45Lab ecosystem, specifically associated with the Safactory project, to handle the persistence of trajectory and environment data to an S3-backed LanceDB data platform [1]. When configuring the environment for this SDK, the WT_SDK_PROFILE environment variable is used to specify the operational environment [1]. The supported values for WT_SDK_PROFILE are: production: Selects the production landing and serving tables [1]. test: Selects the test tables [1]. default: (While often implied as a fallback in such configurations, the primary explicit documentation for Safactory highlights production and test as the two specific profile options for data platform interaction) [1]. To use the SDK, you define your data platform connection settings in a local.env file, including the WT_SDK_PROFILE and other required variables like WT_SDK_DB_URI and AWS credentials [1]. You must also set the gateway storage_type to cloud and launch the application with the --storage-type cloud flag to enable this functionality [1]. For detailed guidance on table configuration and data retrieval, you should refer to the official AI45Lab/wt-data-platform-sdk repository [1].

Citations:


🏁 Script executed:

#!/bin/bash
set -euo pipefail

printf '%s\n' '--- requirements-cloud.txt ---'
cat -n requirements-cloud.txt

printf '%s\n' ''
printf '%s\n' '--- relevant cloud_strategy_impl.py sections ---'
sed -n '1,190p' core/data_manager/strategy/cloud_strategy_impl.py
printf '%s\n' ''
printf '%s\n' '--- cloud_strategy_impl.py lines 330-380 ---'
sed -n '330,380p' core/data_manager/strategy/cloud_strategy_impl.py

printf '%s\n' ''
printf '%s\n' '--- all WT_SDK references and table/table-like references ---'
rg -n --hidden --glob '!.git' 'WT_SDK|wind_tunnel_(landing|serving)|wind_tunnel_landing|wind_tunnel_serving|landing_test|serving_test|TableConfig|GatewayConfig' . || true

printf '%s\n' ''
printf '%s\n' "--- workflow/env files candidates ---"
git ls-files | rg '(^\.github/workflows/|(^|/)(\.env|env\.|\.env\.|docker-compose|compose|Dockerfile|deploy|deployment|readme|requirements|\.gitignore)$|requirements.*\.txt$)' || true

printf '%s\n' ''
printf '%s\n' '--- diff stat/name status ---'
git diff --stat || true
git diff -- requirements-cloud.txt | sed -n '1,200p' || true

Repository: AI45Lab/SAfactory

Length of output: 12942


🌐 Web query:

GitHub AI45Lab wt-data-platform-sdk v0.4.0 config.py WT_SDK_PROFILE production default

💡 Result:

The wt-data-platform-sdk is a component used by the AI45Lab organization (GitHub organization AI45Lab) to manage persistent data storage—specifically for trajectory and environment data—within its agentic infrastructure projects, such as Safactory [1]. Regarding the configuration of WT_SDK_PROFILE, it is used to determine the environment context for the SDK's operations, such as selecting between production and test databases/tables [1]. In typical usage of the wt-data-platform-sdk, you configure the environment by setting the WT_SDK_PROFILE variable along with other connection parameters [1]: 1. WT_SDK_PROFILE: Set this to 'production' to target production landing or serving tables, or 'test' for test tables [1]. 2. Other associated environment variables typically include WT_SDK_DB_URI (e.g., s3://...), WT_SDK_ENV_CONFIG_DB_URI, and various AWS-related credentials (AWS_ACCESS_KEY_ID, AWS_SECRET_ACCESS_KEY) if utilizing S3-backed storage [1]. By default, the SDK may fall back to local storage (e.g., SQLite) if cloud-specific variables are not configured or if the application's storage type is set accordingly [1]. For specific implementation details, such as the contents of config.py or the exact behavior in v0.4.0, it is recommended to refer to the official repository documentation at https://github.com/AI45Lab/wt-data-platform-sdk [1].

Citations:


🌐 Web query:

GitHub AI45Lab wt-data-platform-sdk v0.4.1 config.py WT_SDK_PROFILE production default

💡 Result:

The wt-data-platform-sdk is associated with AI45Lab, specifically in the context of their SAfactory framework, which uses it to persist trajectory and environment data to an S3-backed LanceDB data platform [1][2]. In this SDK, the WT_SDK_PROFILE environment variable is used to manage configuration profiles, such as production or test [2]. Setting WT_SDK_PROFILE to production selects the production landing/serving tables, while other values (such as test) select their corresponding tables [2]. While exact file contents for version 0.4.1 are not indexed, typical configuration for this SDK involves setting several environment variables to establish the connection, such as: WT_SDK_PROFILE=production WT_SDK_DB_URI=s3://YOUR_DATA_DATABASE WT_SDK_ENV_CONFIG_DB_URI=s3://YOUR_ENV_CONFIG_DATABASE WT_SDK_S3_ENDPOINT=https://YOUR_S3_ENDPOINT WT_SDK_S3_ALLOW_HTTP=true AWS_ACCESS_KEY_ID=YOUR_ACCESS_KEY AWS_SECRET_ACCESS_KEY=YOUR_SECRET_KEY For further documentation, instructions on querying data, or specific details regarding version 0.4.1, you should refer to the official AI45Lab/wt-data-platform-sdk repository on GitHub [2].

Citations:


🏁 Script executed:

#!/bin/bash
set -euo pipefail

tmp="$(mktemp -d)"
cd "$tmp"

for ref in v0.4.0 v0.4.1; do
  url="https://raw.githubusercontent.com/AI45Lab/wt-data-platform-sdk/refs/tags/${ref}/wt_sdk/config.py"
  printf '%s\n' "--- ${ref} config.py relevant section ---"
  if command -v curl >/dev/null 2>&1; then
    curl -fsSL "$url" | sed -n '1,180p'
  elif command -v wget >/dev/null 2>&1; then
    wget -qO- "$url" | sed -n '1,180p'
  else
    echo "No curl or wget available"
  fi

  printf '%s\n' '--- defaults containing WT_SDK_PROFILE or profile assignment ---'
  curl -fsSL "$url" | rg -n -C 3 'WT_SDK_PROFILE|profile|profile=' || true
  printf '%s\n' ''
done

Repository: AI45Lab/SAfactory

Length of output: 12253


🏁 Script executed:

#!/bin/bash
set -euo pipefail

python3 - <<'PY'
from pathlib import Path
import ast

repo = Path.cwd()
for path in sorted(repo.glob("core/data_manager/strategy/*.py")):
    text = path.read_text()
    for node in ast.walk(ast.parse(text, filename=str(path))):
        if isinstance(node, ast.Call) and hasattr(node, 'func'):
            name = getattr(node.func, 'id', None) or getattr(getattr(node.func, 'attr', None), 'id', None)
            if name == 'GatewayConfig':
                print(path)
                print(f'  line {node.lineno}: explicit tables keyword={any(k.arg == "tables" for k in node.keywords)}')
                print('  keywords:', [k.arg or '**' for k in node.keywords])
PY

Repository: AI45Lab/SAfactory

Length of output: 345


Set the production SDK profile explicitly.

GatewayConfig() creates the table config without profile, so cloud storage now uses the SDK test profile from v0.4.1 when WT_SDK_PROFILE is unset. This selects landing_test and serving_test, while production storage should use wind_tunnel_landing and wind_tunnel_serving. Set WT_SDK_PROFILE=production or pass an explicit production table profile into GatewayConfig. Add a regression check for the production landing and serving tables.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@requirements-cloud.txt` at line 5, Configure the production SDK profile
explicitly for GatewayConfig by setting WT_SDK_PROFILE=production or supplying a
production table profile, ensuring cloud storage resolves wind_tunnel_landing
and wind_tunnel_serving instead of test tables. Add a regression check that
verifies both production landing and serving table names.

Source: MCP tools