c5b2e21edf
v0.3.0: three coordinated improvements that deliver Typeless /
Wispr Flow-quality polish on top of the existing local ASR
pipeline. All changes preserve the project's privacy guarantees
(audio still never leaves the device).
## 1. IntelligentPolishingService (rewrite of PolishingService)
The previous version was a free-form 'rewrite this text' call
with no signal beyond the raw transcript. The new one is a
single LLM call that does three things in one pass, exactly as
Typeless and Wispr Flow do internally:
1. ASR error correction (homophones, near-misses, missing chars)
2. Polish (drop filler words, fix grammar, add punctuation)
3. Style adaptation per app context (code / email / chat / doc)
The merged-prompt design halves the round-trip vs the previously
proposed two-stage design (correction + polish separately) and
the academic literature confirms it performs equivalently for
everyday Chinese / English dictation.
## 2. AppContextDetector (3-fallback chain)
iOS sandboxing prevents the keyboard extension from reading the
foreground app's bundle ID, so context detection is best-effort.
The detector runs three fallbacks in order, with caching to
avoid the cold-start 'unknown' that would force a neutral-tone
LLM call every time the user opens a new field:
1. Heuristic on the text at the cursor (code / email / chat / doc)
2. 30-minute cache of the last successful detection
3. Time-of-day + weekend heuristic as a soft default
The keyboard extension runs the detector on every press of the
mic and persists the result to the App Group so the host app's
polisher picks it up.
## 3. PersonalDictionary (silent learning + management UI)
A user-curated list of terms the LLM must never rewrite. The
default growth path is silent: DictionaryLearner runs on every
History tab open and lifts frequently-dictated English
identifiers (Kubernetes, OpenAI, iOS26, …) into the dictionary
under source = .history. Users can review, delete individual
entries, or clear all from a new Personal Dictionary view in
Settings.
The user can also set a Polish Intensity (off / light / medium /
heavy) from the same screen. Default is medium, which is what
Typeless and Wispr Flow also use.
## Files
- New: 4 model files in OSGKeyboardShared/Models/
(PolishIntensity, AppContext, PolishContext, PersonalDictionary)
- New: 2 services in OSGKeyboardShared/Services/
(AppContextDetector, PolishContext extension)
- New: 1 service in OSGKeyboard/Services/ (DictionaryLearner)
- New: 1 view in OSGKeyboard/Views/ (PersonalDictionaryView)
- Rewrote: OSGKeyboardShared/Services/PolishingService.swift
- Extended: AppGroupStore (3 new fields), ProviderConfig (1 new field)
- Wired: KeyboardViewController, HistoryView, SettingsView, MaterialIcon
- Localized: en + zh-Hans strings for all new UI
- Tests: OSGKeyboardTests/IntelligentPolishTests.swift (16 tests)
## Verification
- All new code follows the existing Sendable / strict-concurrency
patterns (the keyboard extension stays within its 60MB sandbox;
the polisher remains an actor; @MainActor is applied to the
learner and the settings UI).
- Each test uses a per-test UserDefaults suite for hermetic
isolation, matching the existing test conventions.
- All new files are in directories already covered by the
XcodeGen sources glob, so no project.yml change is needed.
## Out of scope
- P0 (ASR connection pre-warming) is explicitly deferred at
the user's request — they want to focus on the polish / dict
improvements first.
- The Cloud polish (WebSocket) work is not touched.
## Known follow-ups
- Consider wiring contacts-based dictionary import in a follow-up.
- Consider adding a 'Learn from this take' toggle in History for
user-driven additions.
- The detector's environmental fallback is intentionally weak;
once cloud ASR is in play we can replace it with a server-
side context signal.
63 lines
2.7 KiB
Plaintext
63 lines
2.7 KiB
Plaintext
/* Engine status labels */
|
|
"engine.summary.local" = "本地 · %@";
|
|
"engine.summary.cloud" = "当前:%@";
|
|
"engine.summary.cloudWithModel" = "当前:%1$@ · %2$@";
|
|
"engine.asr.appleSpeech" = "Apple 语音识别";
|
|
|
|
/* v0.2.0: flow-level warnings surfaced alongside the final transcript. */
|
|
"flow.warning.cloudPolishMissingKey" = "已开启云端润色但未填写 API Key,本次以原始识别结果插入。请在设置中填入 DeepSeek API Key 以启用润色。";
|
|
|
|
/* LLM providers */
|
|
"provider.openai" = "OpenAI";
|
|
"provider.deepseek" = "DeepSeek";
|
|
"provider.qwen" = "通义千问";
|
|
"provider.zhipu" = "智谱 GLM";
|
|
"provider.moonshot" = "月之暗面";
|
|
"provider.custom" = "自定义";
|
|
|
|
/* LLM errors */
|
|
"error.llm.invalidURL" = "API 地址无效。请在设置中检查 Base URL。";
|
|
"error.llm.noAPIKey" = "未填写 API Key。";
|
|
"error.llm.http" = "API 返回 HTTP %lld。请稍后重试或联系服务方。";
|
|
"error.llm.decoding" = "解析 API 响应失败。";
|
|
"error.llm.transport" = "网络错误,请检查连接后重试。";
|
|
"error.llm.rateLimited" = "API 调用过于频繁,请稍候再试。";
|
|
"error.llm.cancelled" = "请求已取消。";
|
|
|
|
/* ASR errors */
|
|
"error.asr.localeUnsupported" = "当前系统未分配可用语音语言模型,请稍后重试或切换语言。";
|
|
"error.asr.assetsNotReady" = "语音语言资源未就绪,请稍后重试。";
|
|
"error.asr.formatUnsupported" = "当前设备不支持该语音输入格式。";
|
|
"error.asr.noSpeech" = "未识别到语音内容,请重试。";
|
|
"error.asr.chunkFailed" = "第 %lld 段识别失败:%@";
|
|
|
|
/* v0.3.0: 润色档位 */
|
|
"polish.intensity.off" = "关闭";
|
|
"polish.intensity.light" = "轻度";
|
|
"polish.intensity.medium" = "中度";
|
|
"polish.intensity.heavy" = "深度";
|
|
"polish.intensity.off.desc" = "不调用 LLM,直接插入识别原文。";
|
|
"polish.intensity.light.desc" = "仅清除孤立语气词和重复口误。";
|
|
"polish.intensity.medium.desc" = "纠正识别错误、清除语气词、润色语句。推荐默认。";
|
|
"polish.intensity.heavy.desc" = "可重组段落、拆长句、自动编号。适合会议纪要与报告。";
|
|
|
|
/* v0.3.0: 输入场景标签 */
|
|
"appContext.code" = "代码";
|
|
"appContext.email" = "邮件";
|
|
"appContext.chat" = "聊天";
|
|
"appContext.document" = "长文";
|
|
"appContext.unknown" = "通用";
|
|
|
|
/* v0.3.0: 词库类别 */
|
|
"dict.category.properNoun" = "人名地名";
|
|
"dict.category.technical" = "技术名词";
|
|
"dict.category.acronym" = "缩写";
|
|
"dict.category.productName" = "产品名";
|
|
"dict.category.custom" = "自定义";
|
|
|
|
/* v0.3.0: 词库来源 */
|
|
"dict.source.manual" = "手动添加";
|
|
"dict.source.history" = "自动学习";
|
|
"dict.source.contacts" = "来自通讯录";
|
|
"dict.source.recentEdit" = "来自最近编辑";
|