Skip to content

Commit a8186da

Browse files
committed
fix(stats): 明细清理不得截断小时汇总(边界小时缩水)
purge_expired 按「now - 90天」精确切点删明细,而 rollup_hourly 对整行是 REPLACE 语义。边界那一小时被删掉一半后,下一轮 rollup 会把汇总行覆盖成 「剩下的那一半」,被删部分永久丢失(明细已不在,无从找回)。 实测(同一小时内 100+10 tokens):purge 后 usage_hourly 从 110/2 变成 10/1。 - purge_expired 切点向上对齐到小时边界,只删「整个小时都已过期」的明细, 保证「仍有明细的小时保有全部明细」,REPLACE 语义下重算才精确 - 代价:明细最多多留 1 小时 - rollup_hourly docstring 写明这个前提,避免以后有人改回精确切点 - 回归测试:同一小时一条过期一条未过期 → 汇总保持 110/2
1 parent f0b217b commit a8186da

2 files changed

Lines changed: 34 additions & 2 deletions

File tree

‎src/stats/collector.py‎

Lines changed: 14 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -111,8 +111,17 @@ def record(
111111
)
112112

113113
def purge_expired(self, retention_days: int = 90, now: int | None = None) -> int:
114-
"""明细保留 90 天;小时汇总永久(PROPOSAL §8)。"""
115-
cutoff = int(now if now is not None else time.time()) - retention_days * 86400
114+
"""明细保留 90 天;小时汇总永久(PROPOSAL §8)。
115+
116+
切点向上对齐到小时边界:只删**整个小时都已过期**的明细。
117+
这不是保守取值而是正确性要求——`rollup_hourly` 对整行用 REPLACE
118+
语义,只有保证「仍有明细的小时保有全部明细」,重算才精确;
119+
若把边界小时只删一半,下一轮 rollup 会把汇总行覆盖为剩下的那一半,
120+
永久丢掉被删部分(已过期明细无法再从明细找回)。
121+
代价:明细最多多留 1 小时。
122+
"""
123+
raw_cutoff = int(now if now is not None else time.time()) - retention_days * 86400
124+
cutoff = (raw_cutoff // 3600 + 1) * 3600
116125
with self._db.transaction() as conn:
117126
cursor = conn.execute("DELETE FROM usage_events WHERE ts < ?", (cutoff,))
118127
return cursor.rowcount
@@ -122,6 +131,9 @@ def rollup_hourly(self, since: int | None = None) -> int:
122131
123132
默认汇总全部明细:小时汇总要永久保留,不能因为「只看最近 N 小时」
124133
而丢掉历史数据。需要增量时由调用方显式传 since。
134+
135+
整行 REPLACE 语义要求「有明细的小时保有全部明细」——
136+
`purge_expired` 按整小时删除来保这个前提(见其 docstring)。
125137
"""
126138
with self._db.transaction() as conn:
127139
cursor = conn.execute(

‎tests/test_m15_operations.py‎

Lines changed: 20 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -1049,6 +1049,26 @@ def test_stats_overview_empty_and_filters(stats):
10491049
assert query.by_provider() == []
10501050

10511051

1052+
def test_purge_keeps_partially_populated_hour_intact(stats):
1053+
"""边界小时不得被删一半:否则下一轮 rollup 把汇总行覆盖成缩小值。
1054+
1055+
rollup 对整行是 REPLACE 语义,purge 只删已完全过期的小时(切点向上
1056+
对齐到小时边界),保证「仍有明细的小时保有全部明细」。
1057+
"""
1058+
collector, _query = stats
1059+
now = int(time.time())
1060+
hour = (now // 3600) * 3600
1061+
# 同一小时内两条明细:早的已过期(200 天前同一小时),晚的还在
1062+
collector.record(username="u", provider="trae", model="m", ok=True,
1063+
input_tokens=100, now=hour - 200 * 86400 + 1)
1064+
collector.record(username="u", provider="trae", model="m", ok=True,
1065+
input_tokens=10, now=hour + 1)
1066+
RetentionTask(collector, retention_days=90).run_once()
1067+
row = collector._db.connect().execute(
1068+
"SELECT SUM(input_tokens) AS t, SUM(requests) AS r FROM usage_hourly").fetchone()
1069+
assert (row["t"], row["r"]) == (110, 2) # 两条都在,未被截断
1070+
1071+
10521072
def test_stats_retention_and_rollup(stats):
10531073
collector, query = stats
10541074
collector.record(username="u", provider="trae", model="m", ok=True, input_tokens=5,

0 commit comments

Comments
 (0)