Repository navigation
Add allocate_from_rows: from a table of per-period counts to an allocation - #19
Merged
Merged
Conversation
…ation A scheduled job keeps no state between runs; it reads the history from a table and needs the next allocation. allocate_from_rows, allocate_cells_from_rows, batches_from_rows and replay do that, folding the periods in sorted order and arms and cells in name order so the result depends only on the table. Rows may be dicts, sqlite3 rows, or pandas or polars frames. examples/sql/period_counts.sql builds the table from an exposure log and a conversion log: one count per user, in the period of the first exposure, a fixed window, users with an open window left out. It is tested on SQLite against hand-built edge cases, and examples/from_warehouse.py runs it end to end. README sections in English and Korean. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What
A scheduled job keeps no state between runs: it reads the experiment's history from a table and needs the next allocation. This adds that path.
allocate_from_rows(rows, ...)returns anAllocation;allocate_cells_from_rowsreturns{cell: Allocation}from the experimental contextual model.batches_from_rowsandreplayare the two steps inside.sqlite3.Rowobjects, or a pandas or polars frame, without importing either library.seedmakes a run reproducible.examples/sql/period_counts.sqlbuilds the table from an exposure log and a conversion log, andexamples/from_warehouse.pyruns it end to end on SQLite.SQL choices
Each user is counted once, in the period of their first exposure, and a user seen in two variations is dropped. The conversion window is fixed, and a user whose window is still open is left out. That filter depends on when the user was exposed, never on what they did, so it does not bias the rates. Only the SQLite version is tested; the three dialect-specific spots are marked in the file.
Tests
tests/test_history.py: the helpers equal the manual fold, row order and table type do not matter, bad tables are rejected with a reason, and the SQL is run against hand-built edge cases (double exposure, open window, conversion before exposure, conversion after the window, repeated conversions, another experiment).Not merged yet
The README on
mainwould describe functions that the released 2.4.0 does not have, so this should merge together with the release that ships them.🤖 Generated with Claude Code