Skip to content

fix: preserve and retry simulation checkpoints after save failures - #169

Open
林SO (Linxiushen) wants to merge 1 commit into
microsoft:mainfrom
Linxiushen:fix/checkpoint-persistence
Open

林SO (Linxiushen) wants to merge 1 commit into
microsoft:mainfrom
Linxiushen:fix/checkpoint-persistence

Conversation

@Linxiushen

Copy link
Copy Markdown

When the simulation cache is on a different filesystem from the system temporary directory, Simulation.checkpoint() cannot replace the destination. The failure also clears has_unsaved_cache_changes, so another checkpoint silently skips the pending trace. Failed writes leave temporary files behind.

This change creates the temporary file beside the destination, clears the dirty flag only after successful replacement, and removes temporary files after failures. Existing cache contents remain intact if JSON serialization fails, and the next checkpoint can retry after the destination is repaired.

Validation:

  • Actual Windows C: temporary directory / D: cache reproduction: before, WinError 17, no cache file, and dirty flag cleared; after, the complete trace is saved successfully.
  • Regression tests use real filesystem operations for relative and absolute paths, missing parent directories, replacement errors, serialization errors, Unicode round trips, and retries. No LLM calls are made.
  • python -m pytest tests/unit/test_control_checkpoint.py -q --tb=short -o addopts='': 3 failed / 2 passed before the fix; 5 passed after the fix (Python 3.11, Windows).
  • git diff --check: passed.
  • The full LLM-dependent core suite was not run.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant