adopt new version to include dldb #23 issue fix - #51
Conversation
📝 WalkthroughWalkthroughThe cloud requirements update the pinned ChangesSDK dependency update
Estimated code review effort: 1 (Trivial) | ~2 minutes Possibly related PRs
Suggested reviewers: 🚥 Pre-merge checks | ✅ 5✅ Passed checks (5 passed)
✨ Finishing Touches🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
There was a problem hiding this comment.
Actionable comments posted: 2
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@requirements-cloud.txt`:
- Line 5: Update the wt-data-platform-sdk requirement in requirements-cloud.txt
to reference the requested fixed commit hash instead of the mutable v0.4.1 tag,
and retain v0.4.1 as an inline comment for traceability.
- Line 5: Configure the production SDK profile explicitly for GatewayConfig by
setting WT_SDK_PROFILE=production or supplying a production table profile,
ensuring cloud storage resolves wind_tunnel_landing and wind_tunnel_serving
instead of test tables. Add a regression check that verifies both production
landing and serving table names.
🪄 Autofix
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: defaults
Review profile: CHILL
Plan: Pro Plus
Run ID: 8694563a-cf22-44a5-884a-1638f2b62ec4
📒 Files selected for processing (1)
requirements-cloud.txt
| # The optional cloud storage stack supports Python 3.10 through 3.12. | ||
| # Pin wt-data-platform-sdk so LanceDB dependencies remain reproducible. | ||
| wt-data-platform-sdk @ git+https://github.com/AI45Lab/wt-data-platform-sdk.git@v0.4.0 | ||
| wt-data-platform-sdk @ git+https://github.com/AI45Lab/wt-data-platform-sdk.git@v0.4.1 |
There was a problem hiding this comment.
🔒 Security & Privacy | 🟡 Minor | ⚡ Quick win
🧩 Analysis chain
🏁 Script executed:
#!/bin/bash
set -euo pipefail
git ls-remote --tags \
https://github.com/AI45Lab/wt-data-platform-sdk.git \
'refs/tags/v0.4.1*'Repository: AI45Lab/SAfactory
Length of output: 274
🌐 Web query:
pip Git URL Git commit hash tag support @ tag commit
💡 Result:
Pip supports installing from Git repositories by specifying a Git reference (such as a branch, tag, or commit hash) using the @ symbol appended to the URL [1][2][3]. The standard syntax is: ProjectName @ git+https://git.example.com/MyProject.git@ Supported formats for include: - Branch names (e.g., @main, @develop) [1][4][5] - Tag names (e.g., @v1.0) [1][6] - Commit hashes (e.g., @da39a3ee5e6b4b0d3255bfef95601890afd80709) [1][2][3] - Other Git references (e.g., @refs/pull/123/head) [1][2][3] Key considerations: - Specification Logic: Pip follows a specific resolution order when encountering a reference [7]. It first attempts to resolve the provided string as a branch or tag name [7][8]. If no matching branch or tag is found, it proceeds to treat the string as a commit hash [7][8]. - Full vs. Partial Hashes: While pip accepts both, it is highly recommended to use the full 40-character commit hash [1][2][3]. Using a full hash is more efficient as it often allows pip to avoid extra network calls or redundant operations [1][2][3]. - Immutability: For repeatable builds and secure, deterministic installs, you should always prefer pinning to specific tags or full commit hashes rather than moving branch names [9][5]. - Requirements Files: You can use this same syntax directly within a requirements.txt file [9][5]. When used there, the Direct URL syntax (PackageName @ URL) is preferred for clarity and reproducibility [9][5]. Example usage: pip install "MyProject @ git+https://git.example.com/MyProject.git@v1.0" pip install "MyProject @ git+https://git.example.com/MyProject.git@da39a3ee5e6b4b0d3255bfef95601890afd80709"
Citations:
- 1: https://pip.pypa.io/en/stable/topics/vcs-support/?highlight=git%2B
- 2: https://pip.pypa.io/en/latest/topics/vcs-support/
- 3: https://pip.pypa.io/en/stable/topics/vcs-support/
- 4: https://stackoverflow.com/questions/20101834/pip-install-from-git-repo-branch
- 5: https://codemia.io/knowledge-hub/path/how_to_state_in_requirementstxt_a_direct_github_source
- 6: https://stackoverflow.com/questions/13685920/install-specific-git-commit-with-pip
- 7: Install a specific commit, without looking for tags or branches. pypa/pip#4970
- 8: When installing from git and refering to a commit hash, pip tries to find branch or tag of a name equal to the hash string. pypa/pip#4507
- 9: https://pip.pypa.io/en/stable/user_guide/?highlight=
Pin the Git dependency to a fixed commit hash.
wt-data-platform-sdk currently installs from the mutable v0.4.1 Git tag. Change the requirement to the fixed commit and keep the version identifier in a comment for traceability.
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@requirements-cloud.txt` at line 5, Update the wt-data-platform-sdk
requirement in requirements-cloud.txt to reference the requested fixed commit
hash instead of the mutable v0.4.1 tag, and retain v0.4.1 as an inline comment
for traceability.
Source: MCP tools
🗄️ Data Integrity & Integration | 🟠 Major | 🏗️ Heavy lift
🧩 Analysis chain
🏁 Script executed:
#!/bin/bash
set -euo pipefail
# Expect: every production deployment sets an explicit production profile
# or supplies explicit production table configuration.
rg -n -C 4 --hidden --glob '!.git' \
'WT_SDK_PROFILE|TableConfig\s*\(|GatewayConfig\s*\(|WT_SDK_LANDING_TABLE|WT_SDK_SERVING_TABLE' .Repository: AI45Lab/SAfactory
Length of output: 4774
🏁 Script executed:
#!/bin/bash
set -euo pipefail
printf '--- requirements-cloud.txt ---\n'
cat -n requirements-cloud.txt
printf '\n--- relevant cloud_strategy_impl.py sections ---\n'
sed -n '1,190p' core/data_manager/strategy/cloud_strategy_impl.py
printf '\n--- cloud_strategy_impl.py lines 330-380 ---\n'
sed -n '330,380p' core/data_manager/strategy/cloud_strategy_impl.py
printf '\n--- all WT_SDK references and files ---\n'
rg -n --hidden --glob '!.git' 'WT_SDK|wind_tunnel_(landing|serving)|landing_test|serving_test' .
printf '\n--- workflow/env files candidates ---\n'
git ls-files | rg '(^\.github/workflows/|\.env|env|docker-compose|compose|Dockerfile|deploy|deployment|readme|requirements)' || true
printf '\n--- diff stat/name status ---\n'
git diff --stat
git diff -- requirements-cloud.txt | sed -n '1,200p'Repository: AI45Lab/SAfactory
Length of output: 246
🌐 Web query:
AI45Lab wt-data-platform-sdk v0.4.1 WT_SDK_PROFILE production test default
💡 Result:
The wt-data-platform-sdk is a component used within the AI45Lab ecosystem, specifically associated with the Safactory project, to handle the persistence of trajectory and environment data to an S3-backed LanceDB data platform [1]. When configuring the environment for this SDK, the WT_SDK_PROFILE environment variable is used to specify the operational environment [1]. The supported values for WT_SDK_PROFILE are: production: Selects the production landing and serving tables [1]. test: Selects the test tables [1]. default: (While often implied as a fallback in such configurations, the primary explicit documentation for Safactory highlights production and test as the two specific profile options for data platform interaction) [1]. To use the SDK, you define your data platform connection settings in a local.env file, including the WT_SDK_PROFILE and other required variables like WT_SDK_DB_URI and AWS credentials [1]. You must also set the gateway storage_type to cloud and launch the application with the --storage-type cloud flag to enable this functionality [1]. For detailed guidance on table configuration and data retrieval, you should refer to the official AI45Lab/wt-data-platform-sdk repository [1].
Citations:
🏁 Script executed:
#!/bin/bash
set -euo pipefail
printf '%s\n' '--- requirements-cloud.txt ---'
cat -n requirements-cloud.txt
printf '%s\n' ''
printf '%s\n' '--- relevant cloud_strategy_impl.py sections ---'
sed -n '1,190p' core/data_manager/strategy/cloud_strategy_impl.py
printf '%s\n' ''
printf '%s\n' '--- cloud_strategy_impl.py lines 330-380 ---'
sed -n '330,380p' core/data_manager/strategy/cloud_strategy_impl.py
printf '%s\n' ''
printf '%s\n' '--- all WT_SDK references and table/table-like references ---'
rg -n --hidden --glob '!.git' 'WT_SDK|wind_tunnel_(landing|serving)|wind_tunnel_landing|wind_tunnel_serving|landing_test|serving_test|TableConfig|GatewayConfig' . || true
printf '%s\n' ''
printf '%s\n' "--- workflow/env files candidates ---"
git ls-files | rg '(^\.github/workflows/|(^|/)(\.env|env\.|\.env\.|docker-compose|compose|Dockerfile|deploy|deployment|readme|requirements|\.gitignore)$|requirements.*\.txt$)' || true
printf '%s\n' ''
printf '%s\n' '--- diff stat/name status ---'
git diff --stat || true
git diff -- requirements-cloud.txt | sed -n '1,200p' || trueRepository: AI45Lab/SAfactory
Length of output: 12942
🌐 Web query:
GitHub AI45Lab wt-data-platform-sdk v0.4.0 config.py WT_SDK_PROFILE production default
💡 Result:
The wt-data-platform-sdk is a component used by the AI45Lab organization (GitHub organization AI45Lab) to manage persistent data storage—specifically for trajectory and environment data—within its agentic infrastructure projects, such as Safactory [1]. Regarding the configuration of WT_SDK_PROFILE, it is used to determine the environment context for the SDK's operations, such as selecting between production and test databases/tables [1]. In typical usage of the wt-data-platform-sdk, you configure the environment by setting the WT_SDK_PROFILE variable along with other connection parameters [1]: 1. WT_SDK_PROFILE: Set this to 'production' to target production landing or serving tables, or 'test' for test tables [1]. 2. Other associated environment variables typically include WT_SDK_DB_URI (e.g., s3://...), WT_SDK_ENV_CONFIG_DB_URI, and various AWS-related credentials (AWS_ACCESS_KEY_ID, AWS_SECRET_ACCESS_KEY) if utilizing S3-backed storage [1]. By default, the SDK may fall back to local storage (e.g., SQLite) if cloud-specific variables are not configured or if the application's storage type is set accordingly [1]. For specific implementation details, such as the contents of config.py or the exact behavior in v0.4.0, it is recommended to refer to the official repository documentation at https://github.com/AI45Lab/wt-data-platform-sdk [1].
Citations:
🌐 Web query:
GitHub AI45Lab wt-data-platform-sdk v0.4.1 config.py WT_SDK_PROFILE production default
💡 Result:
The wt-data-platform-sdk is associated with AI45Lab, specifically in the context of their SAfactory framework, which uses it to persist trajectory and environment data to an S3-backed LanceDB data platform [1][2]. In this SDK, the WT_SDK_PROFILE environment variable is used to manage configuration profiles, such as production or test [2]. Setting WT_SDK_PROFILE to production selects the production landing/serving tables, while other values (such as test) select their corresponding tables [2]. While exact file contents for version 0.4.1 are not indexed, typical configuration for this SDK involves setting several environment variables to establish the connection, such as: WT_SDK_PROFILE=production WT_SDK_DB_URI=s3://YOUR_DATA_DATABASE WT_SDK_ENV_CONFIG_DB_URI=s3://YOUR_ENV_CONFIG_DATABASE WT_SDK_S3_ENDPOINT=https://YOUR_S3_ENDPOINT WT_SDK_S3_ALLOW_HTTP=true AWS_ACCESS_KEY_ID=YOUR_ACCESS_KEY AWS_SECRET_ACCESS_KEY=YOUR_SECRET_KEY For further documentation, instructions on querying data, or specific details regarding version 0.4.1, you should refer to the official AI45Lab/wt-data-platform-sdk repository on GitHub [2].
Citations:
🏁 Script executed:
#!/bin/bash
set -euo pipefail
tmp="$(mktemp -d)"
cd "$tmp"
for ref in v0.4.0 v0.4.1; do
url="https://raw.githubusercontent.com/AI45Lab/wt-data-platform-sdk/refs/tags/${ref}/wt_sdk/config.py"
printf '%s\n' "--- ${ref} config.py relevant section ---"
if command -v curl >/dev/null 2>&1; then
curl -fsSL "$url" | sed -n '1,180p'
elif command -v wget >/dev/null 2>&1; then
wget -qO- "$url" | sed -n '1,180p'
else
echo "No curl or wget available"
fi
printf '%s\n' '--- defaults containing WT_SDK_PROFILE or profile assignment ---'
curl -fsSL "$url" | rg -n -C 3 'WT_SDK_PROFILE|profile|profile=' || true
printf '%s\n' ''
doneRepository: AI45Lab/SAfactory
Length of output: 12253
🏁 Script executed:
#!/bin/bash
set -euo pipefail
python3 - <<'PY'
from pathlib import Path
import ast
repo = Path.cwd()
for path in sorted(repo.glob("core/data_manager/strategy/*.py")):
text = path.read_text()
for node in ast.walk(ast.parse(text, filename=str(path))):
if isinstance(node, ast.Call) and hasattr(node, 'func'):
name = getattr(node.func, 'id', None) or getattr(getattr(node.func, 'attr', None), 'id', None)
if name == 'GatewayConfig':
print(path)
print(f' line {node.lineno}: explicit tables keyword={any(k.arg == "tables" for k in node.keywords)}')
print(' keywords:', [k.arg or '**' for k in node.keywords])
PYRepository: AI45Lab/SAfactory
Length of output: 345
Set the production SDK profile explicitly.
GatewayConfig() creates the table config without profile, so cloud storage now uses the SDK test profile from v0.4.1 when WT_SDK_PROFILE is unset. This selects landing_test and serving_test, while production storage should use wind_tunnel_landing and wind_tunnel_serving. Set WT_SDK_PROFILE=production or pass an explicit production table profile into GatewayConfig. Add a regression check for the production landing and serving tables.
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@requirements-cloud.txt` at line 5, Configure the production SDK profile
explicitly for GatewayConfig by setting WT_SDK_PROFILE=production or supplying a
production table profile, ensuring cloud storage resolves wind_tunnel_landing
and wind_tunnel_serving instead of test tables. Add a regression check that
verifies both production landing and serving table names.
Source: MCP tools
Summary by CodeRabbit