Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
10 changes: 7 additions & 3 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -127,7 +127,7 @@ That is what **authorizing the infrastructure once, at launch** means: on first

Capabilities that sediment yet more of the coordination you used to repeat every day into defaults:

- 📖 **A dictionary that learns.** Until now the dictionary only knew what you typed into it by hand. Now, when you correct a word OpenLess just wrote, it asks — once, on a small card — whether to remember it, and one click puts it in. Paired with **cursor context** (opt-in, macOS), which lets the polish model read what you are writing around your cursor, OpenLess stops being a transcriber that guesses at homophones and starts being an input method that knows your words. Every suggestion is reviewed by you; nothing is learned silently.
- 📖 **A dictionary that learns.** Until now the dictionary only knew what you typed into it by hand. Now, when you correct a word OpenLess just wrote, it asks — once, on a small card — whether to remember it, and one click puts it in. Paired with **cursor context** (opt-in, macOS / Windows), which lets the polish model read what you are writing around your cursor, OpenLess stops being a transcriber that guesses at homophones and starts being an input method that knows your words. Every suggestion is reviewed by you; nothing is learned silently.
- 🎨 **Style Pack Marketplace.** OpenLess no longer ships a single fixed "polish" voice. Build your own **style packs** with custom system prompts, switch between them with a hotkey, and **install community packs in one click** — or publish your own to share. When a style is tuned to your exact task (cold emails, commit messages, 小红书 posts, formal reports, your team's tone), the output is not merely cleaner — it is *noticeably better*, because the model is finally writing the way you intend.
- ⚡ **Streaming insertion.** Text now flows to your cursor **character by character** as it is polished, rather than making you wait for the complete result. Perceived latency drops sharply, so dictation feels nearly as fast as thinking — and it automatically falls back to a one-shot paste when an application cannot accept streamed keystrokes.

Expand Down Expand Up @@ -376,13 +376,17 @@ The dictionary handles your proper nouns, product names, names of people, and ne
- **Learn from corrections (experimental).** Open Settings → Experiments & extensions → Learn from corrections to enable it and configure observation duration (10–60 seconds, default 60), suggestion duration (5–60 seconds, default 10), and maximum automatic phrase length (2–32 characters, default 12). It defaults to off and is not enabled by cursor context or cloud sync. On macOS, Windows and Android, supported editors can be observed after insertion. Suggestions require confirmation before expiry to enter the dictionary. Changing parameters stops the current observation and clears pending suggestions; new values apply to the next dictation. Android requires accessibility. When observation is unavailable, use **Remember a word** in history details; Android IME result editing also offers an unchecked dictionary option. Every path requires explicit confirmation and does not create global replacement rules.
- **Entries that earn their keep get priority.** The hotword budget sent to ASR providers is finite (a few hundred characters). Entries are ranked by hit count, with a few reserved seats for words you just added by hand, so the terms you actually use keep their place instead of being pushed out by whatever you added most recently.

### Cursor context (opt-in, macOS)
### Cursor context (opt-in, macOS / Windows)

Settings → Privacy → Data storage → **Cursor context**. Off by default.

When on, each dictation reads a few hundred characters around your cursor **in the app you are writing in** and sends them with the polish request, so the model knows what you are writing about. Chinese homophones (接口/借口, 大鱼/大禹) are indistinguishable to an acoustic model but obvious from context. Local vocabulary learning has a separate switch and does not require cursor context. Observed text is not sent to a model; words explicitly added to the dictionary participate in future ASR and polish requests as described above.

Cursor context excludes password fields, macOS Secure Input, known password managers and terminals. Turning it off stops reading cursor context for polishing; other authorized accessibility features, including local vocabulary learning and insertion, work independently. Turning vocabulary learning off stops observation and clears pending suggestions.
On Windows, cursor context is read via UI Automation from the focused text control: it prefers `TextPattern2`'s caret range, with a `TextPattern` selection fallback for controls that only support collapsed selections. When the caret cannot be reliably located, it is skipped rather than guessed — wrong context is worse than none, and dictation continues normally either way.

Chromium-based editors (VS Code, Electron apps) only keep their accessibility tree's caret position in sync once they detect an active accessibility client, so with default settings the position UI Automation reads back can be stale. In VS Code, set `editor.accessibilitySupport` to `on` (Settings → search "accessibility support"; a window reload may be needed) to get accurate cursor context; the same applies to other Electron-based editors with similar settings. Plain Chrome/Edge text fields do not need this.

Cursor context excludes password fields, macOS Secure Input (or the Windows UIA password-control flag), known password managers and terminals. Turning it off stops reading cursor context for polishing; other authorized accessibility features, including local vocabulary learning and insertion, work independently. Turning vocabulary learning off stops observation and clears pending suggestions.

The main window is organized as Home / History / Dictionary / Settings. The Dictionary tab opens a separate editor window when you click "New". The Home tab shows total dictation time, total characters, average characters per minute, estimated time saved, and dictionary participation statistics.

Expand Down
10 changes: 7 additions & 3 deletions README.zh.md
Original file line number Diff line number Diff line change
Expand Up @@ -132,7 +132,7 @@ OpenLess 做的不是“更快的听写”,而是**消灭“想法 → 干净文

下面这些能力,把过去每天都要重复的协调,进一步沉降成了默认规则:

- 📖 **会自己长的词典。** 在此之前,词典里只有你亲手敲进去的东西。现在,当你改掉 OpenLess 刚写出来的某个词,它会在屏幕角落弹一张小卡片问一声要不要记住,点一下就进去了。配合**光标上下文**(需手动开启,仅 macOS)——让润色模型读得到你光标周围正在写的内容——OpenLess 不再是一个靠猜同音词的转写工具,而开始成为**一个认得你的词的输入法**。每一条建议都由你过目,没有任何东西是悄悄学走的。
- 📖 **会自己长的词典。** 在此之前,词典里只有你亲手敲进去的东西。现在,当你改掉 OpenLess 刚写出来的某个词,它会在屏幕角落弹一张小卡片问一声要不要记住,点一下就进去了。配合**光标上下文**(需手动开启,macOS / Windows)——让润色模型读得到你光标周围正在写的内容——OpenLess 不再是一个靠猜同音词的转写工具,而开始成为**一个认得你的词的输入法**。每一条建议都由你过目,没有任何东西是悄悄学走的。
- 🎨 **风格包市场(Style Pack Marketplace)。** OpenLess 不再只内置一种固定的“润色”语气。你可以用自定义系统提示词构建自己的**风格包**,用快捷键在它们之间切换,并**一键安装社区分享的风格包**——也可以发布自己的与他人分享。当风格与你的具体任务高度契合(冷启动邮件、commit message、小红书文案、正式报告、团队语气)时,产出的文本不只是更干净,而是*明显更好*,因为模型终于在按你真正想要的方式写作。
- ⚡ **流式插入。** 文本现在会随润色**逐字符**写入光标,而不必等待完整结果生成。感知延迟大幅下降,听写几乎和思考一样快——当某个应用无法接受流式按键时,它会自动回退为一次性粘贴。

Expand Down Expand Up @@ -383,13 +383,17 @@ OpenLess 的润色模型只重塑文本。它不回答问题、不执行任务
- **手改学词(实验)。** 在「设置 → 实验与扩展 → 手改学词」进入独立配置页,设置观察时长(10–60 秒,默认 60)、建议保留时长(5–60 秒,默认 10)及自动建议最大词长(2–32 个字符,默认 12)。功能默认关闭,不跟随光标上下文或云同步授权。macOS、Windows 和 Android 可在支持的编辑器中观察插入后的修改,候选需在有效期内逐条确认才加入词典。修改参数会结束当前观察并清空待确认建议,下次听写生效。Android 需要无障碍服务。无法观察时,可在历史详情点击「记住词汇」手动输入正确词;Android 输入法编辑结果也提供默认不勾选的加入词典选项。所有入口均需明确确认,不自动创建全局替换规则。
- **真正在用的词优先。** 发给 ASR 的热词预算是有限的(几百字符)。条目按命中次数排序,并给刚手动添加的词留几个保底席位——这样你天天在用的那些词不会被「最近刚加的」挤出去。

### 光标上下文(需手动开启,仅 macOS)
### 光标上下文(需手动开启,macOS / Windows)

设置 → 隐私 → 数据存储 → **光标上下文**。默认关闭。

开启后,每次听写会读取**你正在写的那个应用里**光标附近的几百个字,随润色请求一起发出,让模型知道你在写什么。中文同音词(接口/借口、大鱼/大禹)声学模型分不出来,但上下文能分。手改学词是独立的本地功能,不需要开启此设置。观察文本本身不会发送给模型;确认加入词典的词会按词典规则参与后续 ASR/润色。

光标上下文排除密码输入框、macOS Secure Input、已知密码管理器和终端。关闭此开关后不会为润色读取光标上下文;其他已授权的辅助功能(如手改学词或插入)独立工作。手改学词关闭后会停止观察并清空待确认建议。
Windows 使用 UI Automation 读取当前焦点文本控件的光标附近内容:优先使用 `TextPattern2` 的 caret range;控件不支持 `TextPattern2` 时,兼容使用 `TextPattern` 的 collapsed selection。无法可靠定位 caret 时直接跳过,不会猜测光标位置,也不会影响正常听写。

基于 Chromium 的编辑器(VS Code、Electron 应用)只有在检测到有无障碍客户端在用时,才会实时同步无障碍树里的光标位置,默认设置下 UI Automation 读到的位置可能是过时的。VS Code 里需要把 `editor.accessibilitySupport` 设为 `on`(设置里搜"accessibility support",改完可能需要重新加载窗口)才能拿到准确的光标上下文;其他基于 Electron 的编辑器一般也有类似设置。普通的 Chrome/Edge 文本框不需要这一步。

光标上下文排除密码输入框、macOS Secure Input(Windows 上为 UIA 密码控件标记)、已知密码管理器和终端。关闭此开关后不会为润色读取光标上下文;其他已授权的辅助功能(如手改学词或插入)独立工作。手改学词关闭后会停止观察并清空待确认建议。

主窗口组织为 首页 / 历史 / 词典 / 设置。点击“新建”时,词典页会打开一个独立的编辑窗口。首页展示总听写时长、总字数、平均每分钟字数、估算节省的时间,以及词典参与统计。

Expand Down
56 changes: 56 additions & 0 deletions openless-all/app/crates/openless-core/src/host_document/window.rs
Original file line number Diff line number Diff line change
Expand Up @@ -44,3 +44,59 @@ pub fn utf16_offset_to_char_offset(text: &str, utf16_offset: usize) -> usize {
}
text.chars().count()
}

#[cfg(test)]
mod tests {
use super::*;

/// Windows UIA 光标上下文方案(`windows_cursor_context开发方案.md` §6)的验收用例:
/// budget=600 时 before<=480、after<=120,且 before+after<=600。
#[test]
fn budget_splits_roughly_eighty_twenty() {
let long = "字".repeat(2000);
let span = plan_window(long.chars().count(), 1000, 600);
assert!(span.cursor_in_span <= 480, "before 不应超过预算的 80%");
assert!(span.len - span.cursor_in_span <= 120, "after 不应超过预算的 20%");
assert!(span.len <= 600);
}

/// 左侧文本不足时,剩余预算应让给右侧(方案 §6 "如果左侧不足 480 字,可以把剩余预算让给右侧")。
#[test]
fn insufficient_left_text_gives_remaining_budget_to_the_right() {
let text = "字".repeat(1000);
// cursor 只有 10 个字在左边,右边还有 990 个字可读。
let span = plan_window(text.chars().count(), 10, 600);
assert_eq!(span.cursor_in_span, 10, "左边全部给出,不应凭空截断");
assert_eq!(span.len, 600, "右侧应吃满剩余预算以补满 600");
}

/// 右侧文本不足时,剩余预算应让给左侧。
#[test]
fn insufficient_right_text_gives_remaining_budget_to_the_left() {
let text = "字".repeat(1000);
// cursor 在第 990 个字,右边只剩 10 个字。
let span = plan_window(text.chars().count(), 990, 600);
assert_eq!(span.len - span.cursor_in_span, 10, "右边全部给出");
assert_eq!(span.len, 600, "左侧应补满剩余预算到 600");
}

/// emoji / surrogate pair 不应被按 UTF-16 长度误判为 2 个 char 而切坏字符。
#[test]
fn window_around_cursor_does_not_split_emoji_or_surrogate_pairs() {
let text = "前文🙂😀后文";
// 光标紧跟在两个 emoji 之后(按 char 计数,不按 UTF-16 code unit 计数)。
let cursor_char = "前文🙂😀".chars().count();
let window = window_around_cursor(text, cursor_char, 600);
assert_eq!(window.before(), "前文🙂😀");
assert_eq!(window.after(), "后文");
// 两个 emoji 字符仍然完整,没有产生孤立的 surrogate / 替换字符。
assert!(window.text.chars().all(|c| c != '\u{FFFD}'));
}

#[test]
fn zero_budget_yields_an_empty_window_at_the_cursor() {
let span = plan_window(100, 50, 0);
assert_eq!(span.len, 0);
assert_eq!(span.cursor_in_span, 0);
}
}
35 changes: 28 additions & 7 deletions openless-all/app/src-tauri/src/core_adapters.rs
Original file line number Diff line number Diff line change
Expand Up @@ -3691,13 +3691,34 @@ impl openless_core::HostContextAdapter for TauriHostContextAdapter {
Box::pin(async move {
let front_app = crate::coordinator::capture_frontmost_app();
let cursor_context = if include_cursor {
crate::host_document::read_around_cursor(crate::host_document::DEFAULT_BUDGET_CHARS)
.await
.map(|window| {
let before = window.text.chars().take(window.cursor).collect::<String>();
let after = window.text.chars().skip(window.cursor).collect::<String>();
openless_core::prompts::cursor_context_input(&before, &after)
})
let started = std::time::Instant::now();
let window = crate::host_document::read_around_cursor(
crate::host_document::DEFAULT_BUDGET_CHARS,
)
.await;
// Metadata only — never the document body (design doc §18/§19). This is the
// same shape debug_read_cursor_context already logs; extending it to the real
// dictation path means "did this capture actually happen" is visible from the
// normal log file, with no devtools round-trip needed.
match &window {
Some(window) => log::info!(
"[cursor-context] status=ok chars_before={} chars_after={} elapsed_ms={} app={:?}",
window.before().chars().count(),
window.after().chars().count(),
started.elapsed().as_millis(),
front_app,
),
None => log::info!(
"[cursor-context] status=none elapsed_ms={} app={:?}",
started.elapsed().as_millis(),
front_app,
),
}
window.map(|window| {
let before = window.text.chars().take(window.cursor).collect::<String>();
let after = window.text.chars().skip(window.cursor).collect::<String>();
openless_core::prompts::cursor_context_input(&before, &after)
})
} else {
None
};
Expand Down
Loading
Loading