[dev] 对话记录独立库与冷热查询 - #12
Conversation
一、独立数据库(分表)
- CHAT_LOG_SQL_DSN=local 时不再复用 one-api.db,默认使用独立的
chatlog.db(WAL + busy_timeout);支持 local:<path> 与
CHAT_LOG_SQLITE_PATH 自定义文件
- 启动迁移时清理主库(仅限 SQLite 主库,且确认与对话库非同一文件)
中遗留的 chat_logs/chat_turns/chat_sessions 旧表,其余表不受影响;
现有对话数据按约定不做迁移,新库从空开始
- CloseDB 补充关闭 CHATLOG_DB
二、冷热查询
- 新增进程内热缓存(model/chat_log_hot_cache.go):
* 近期会话元数据窗口(默认 1000 条,写透 + 每 60s 从库刷新,
覆盖多节点与冷启动场景)
* 每会话轮次缓存(含响应体),全局字节预算(默认 64MB)按会话
LRU 淘汰;轮次写入后不可变,缓存命中可安全直出
* 未命中即冷启动回源数据库,最新页回填缓存
* CHAT_LOG_HOT_CACHE_ENABLED / _SESSION_WINDOW / _TURN_BYTES /
_REFRESH_SECONDS 可调
- 查询性能:会话列表改为 keyset 游标分页(消除全表 COUNT 与深
OFFSET),新增 model_name 索引、last_active_at 时间范围过滤;
轮次按 id 分页(消除一次性加载全部 LONGTEXT)
- (session_id, turn_index) 复合索引
三、API 与前端
- GET /api/chat_logs/sessions:limit/cursor/时间范围参数,
返回 {items, has_more, next_cursor};无过滤首页优先走热缓存
- GET /api/chat_logs/sessions/:id:limit/before_id 轮次分页,
返回 {session, turns, has_more, next_turn_id}
- 前端会话列表改为"加载更多"式游标加载,详情抽屉支持"加载更早的
对话";i18n 补齐 7 种语言
- controller: limit 参数在热/冷路径分流前统一归一化(缺省 20、上限 100),修复缺省时热路径只返回 1 条、DB 路径返回 20 条的不一致 - 热缓存: listRecent 的 has_more 改用切片前的窗口长度计算,修复 恒为 false 导致"加载更多"不出现 - 热缓存: getTurns 计数校验改为 n < totalTurns 即回源数据库, 修复多节点场景下缓存落后时返回过期的"最新页"(改用计数而非 TurnIndex 校验:并行轮次下 TurnIndex 不可靠) - 热缓存: evictTurns 守卫放宽到可淘汰最后一个会话,字节预算成为 硬上限(原 >1 守卫下单条目可无限超出预算) - 热缓存: recordTurn/admitTurns 存入前拷贝 ChatTurn,缓存拥有 自己的数据,避免共享调用方结构体 - chooseDB: 日志带 envName 与真实条件,消除 "SQL_DSN not set" 在显式 local 及 LOG/CHAT_LOG 场景下的误导 - 回归测试:limit 缺省热路径、has_more 非满窗口翻页、落后缓存 回源、归一化表驱动用例
There was a problem hiding this comment.
🟡 Changes recommended
The current hot-cache turn serving condition prevents cache hits for long sessions (TurnCount > admitted page size), undermining the intended “latest page hot-path” behavior.
Once you've addressed the issues Copilot identified, you can request another Copilot review.
Review details
Suppressed comments (2)
Previously missed (2) — in code that hasn't changed since the last review.
model/chat_log.go:116
- Chat session cursor encoding uses base64.URLEncoding, which can introduce '=' padding in a query parameter and is inconsistent with other code in this repo that uses base64.RawURLEncoding for URL-safe tokens. Since this is a public cursor value, prefer RawURLEncoding for both encode/decode to avoid padding and reduce the chance of clients mangling the cursor.
web/src/features/chat-logs/components/session-detail-sheet.tsx:93 - The inline comment about React Query infinite pages is inaccurate:
fetchNextPage()appends older pages toquery.data.pagesrather than prepending them. The current.reverse().flatMap(...)logic is fine, but the comment should match React Query behavior to avoid future confusion.
- Files reviewed: 21/21 changed files
- Comments generated: 1
- Review effort level: Lite
| n := len(entry) | ||
| if n < totalTurns { | ||
| return nil, false, false // cache is behind the database: cold start | ||
| } |
复审指出:getTurns 的严格计数校验(n < totalTurns 即回源)使长会话 冷加载后永远 miss——AdmitChatTurns 只回填最新一页,n 恒小于总数, 缓存对长会话完全失效。 改为 covered 记账: - 每个轮次缓存条目记录 covered(条目同步到的 DB 总轮数) - 冷加载 admit 以 DB 新鲜 TurnCount 锚定 covered;总数不变期间 条目即持有最新轮次——即使只缓存了长会话的最新一页 - 本节点写透追加时 covered++(本节点刚落库一轮,总数精确 +1; 若条目本就落后于其他节点,等值校验仍失败,方向安全) - 命中条件:covered == totalTurns 且缓存范围完整覆盖请求页 (超大轮次不入缓存导致页不完整时仍回源,保证不漏数据) - has_more = covered > limit 回归测试:长会话冷 admit 后二次查看命中、其他节点追加后回源、 本节点写透追加保持锚定并作为最新轮次直出。
There was a problem hiding this comment.
🟡 Changes recommended
The hot-cache admission path can mark a page as fully covered even when oversized turns were skipped, which can cause serving incomplete “newest page” results from cache.
Once you've addressed the issues Copilot identified, you can request another Copilot review.
Review details
Suppressed comments (1)
Previously missed (1) — in code that hasn't changed since the last review.
web/src/features/chat-logs/components/session-detail-sheet.tsx:93
- Comment says older pages are “prepended”, but useInfiniteQuery appends additional pages as they’re fetched (and this code then reverses pages for rendering). Updating the comment avoids misleading future maintenance.
- Files reviewed: 21/21 changed files
- Comments generated: 1
- Review effort level: Lite
| bytes := 0 | ||
| kept := make([]*ChatTurn, 0, len(turns)) | ||
| for _, t := range turns { | ||
| bodyBytes := len(t.NewMessages) + len(t.ResponseBody) | ||
| if bodyBytes > maxCachedChatTurnBodyBytes { | ||
| continue | ||
| } | ||
| // the cache owns its copies: admitted turns are also handed to the | ||
| // HTTP handler for serialization | ||
| owned := *t | ||
| kept = append(kept, &owned) | ||
| bytes += bodyBytes | ||
| } | ||
| c.mu.Lock() | ||
| defer c.mu.Unlock() | ||
| c.replaceTurnsLocked(sessionId, kept, totalTurns, bytes) | ||
| c.evictTurnsLocked() | ||
| } |
复审指出:admitTurns 跳过页内超大轮次后仍以 covered=totalTurns 锚定条目,当被跳过的是最新一轮时,getTurns 会在 n>=limit 时直出 缺少最新轮次的"最新页"(n<limit 守卫覆盖不到该情形)。 修复:admit 页内任一轮次超过单条缓存上限时整页放弃缓存,该会话 保持由数据库服务;写透路径本就安全(跳过时 covered 不递增, 等值校验必然失败回源)。同步修正 entry 与 getTurns 的注释。 回归测试:最新轮超大的会话 admit 后任何 limit 均回源;已锚定 条目追加超大轮次后 covered 落后于 DB 总数而回源。 (验证备注:relay/channel 的 HTTP2 GOAWAY 重试测试在 develop 基线 上同样间歇失败,为存量时序 flaky,与本 PR 无关。)
一、独立数据库(分表)
chatlog.db(WAL + busy_timeout);支持 local: 与
CHAT_LOG_SQLITE_PATH 自定义文件
中遗留的 chat_logs/chat_turns/chat_sessions 旧表,其余表不受影响;
现有对话数据按约定不做迁移,新库从空开始
二、冷热查询
覆盖多节点与冷启动场景)
LRU 淘汰;轮次写入后不可变,缓存命中可安全直出
_REFRESH_SECONDS 可调
OFFSET),新增 model_name 索引、last_active_at 时间范围过滤;
轮次按 id 分页(消除一次性加载全部 LONGTEXT)
三、API 与前端
返回 {items, has_more, next_cursor};无过滤首页优先走热缓存
返回 {session, turns, has_more, next_turn_id}
对话";i18n 补齐 7 种语言
Important
📝 变更描述 / Description
(简述:做了什么?为什么这样改能生效?请基于你对代码逻辑的理解来写,避免粘贴未经整理的内容)
🚀 变更类型 / Type of change
🔗 关联任务 / Related Issue
✅ 提交前检查项 / Checklist
Bug fix,我已提交或关联对应 Issue,且不会将设计取舍、预期不一致或理解偏差直接归类为 bug。📸 运行证明 / Proof of Work
(请在此粘贴截图、关键日志或测试报告,以证明变更生效)