feat(keyboard): add English QuickType bar and system lexicon

Show verbatim/correction/completion slots, mmap a 40k-word list, and
use UITextChecker plus supplementary lexicon for conservative autocorrect.
This commit is contained in:
Rocky
2026-08-14 21:49:22 +08:00
parent 2c3a3f80f3
commit 4749a9cbf2
34 changed files with 45539 additions and 3395 deletions
+11 -1
View File
@@ -8,6 +8,10 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
## [Unreleased]
### Added
- **English QuickType bar**: while typing a word, three equal slots show the verbatim text (quoted when unknown), the unique Space correction, and a completion. The bar stays empty before typing and between committed words. / **英文 QuickType 栏**:输入单词时,三个等宽格显示原文(生词带引号)、空格会采用的唯一纠错和补全;尚未输入及单词提交后保持空白。
- **System English lexicon**: typing uses `UITextChecker` completions/guesses and `requestSupplementaryLexicon` contact names / text replacements. / **系统英文词库**:打字使用 `UITextChecker` 补全/猜测,以及 `requestSupplementaryLexicon` 的通讯录名与文本替换。
- **QWERTY proximity correction**: fat-finger substitutions on neighboring keys (for example `gppd``good`) outrank distant edit-distance neighbors. / **邻键纠错**:相邻键的胖手指替换(如 `gppd``good`)优先于远键编辑距离。
- **Larger English word list**: about 40k unigrams and truncated bigrams, derived from Peter Norvigs public-domain n-gram counts, shipped as an mmap binary so the keyboard extension does not parse them into Swift dictionaries. / **更大英文词表**:约 4 万 unigram 与截断 bigram,来自 Peter Norvig 公有领域 n-gram 计数,以 mmap 二进制随扩展加载,避免解析进 Swift 字典。
- **Overlapping key presses**: the typing grid tracks multiple fingers, so the next key can go down before the previous lifts. Pending letters commit in press order (not release order); Shift can be held with one finger while another types. / **叠指连打**:打字网格跟踪多指,上一键未松开也可按下下一键。未提交的字母按按下顺序出字(而非抬手顺序);一只手指按住 Shift 时另一只可打字。
- **Period shortcut**: in English, a second Space shortly after a Space that follows a word becomes `. ` and arms sentence Shift, matching the system "." Shortcut. / **句号快捷**:英文下,在单词后的空格上短时间内再按一次空格会变成 `. ` 并点亮句首 Shift,对齐系统「句号快捷」。
- **Return key labels**: Go / Search / Send / Done / Next / Join and the other `UIReturnKeyType` values show their system captions on the green action key instead of collapsing to Send or a return arrow. / **回车键文案**:前往 / 搜索 / 发送 / 完成 / 下一项 / 加入等 `UIReturnKeyType` 在绿色动作键上显示系统对应文案,不再一律变成「发送」或换行箭头。
@@ -17,9 +21,15 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
- **Navigate skill**: copying text and tapping Navigate asks the LLM for one address (or origin → destination) and the host app opens driving directions — Amap if installed, then Baidu Maps, then Apple Maps. No companion Shortcut. No address stays in the current app with a keyboard tip. / **导航技能**:复制文字并点「导航」会抽取一条地址(或起点→终点),由 App 直接打开驾车导航——已装高德则用高德,否则百度,再否则 Apple 地图。不需要配套捷径。没有地址则留在当前 App,键盘上给出提示。
- **Save to Notes skill**: the Skills tab opens a ready-made companion Shortcut named `OSGSaveToNotes` on the system Add page (one new Apple Note with an explicit title and body). After you tap Add, copying text and tapping Notes asks the LLM for a short title from the time and content; the body stays the original clipboard. / **存入备忘录技能**:技能页会打开已做好的配套捷径 `OSGSaveToNotes` 的系统添加页(新建一条带标题和正文的苹果备忘录)。点添加后,复制文字并点「备忘录」会按时间和内容生成短标题;正文保持剪贴板原文。
- **Skills clipboard guide**: when Clipboard History is off, the Skills tab shows a card that jumps to in-app Clipboard settings and to iOS Settings for paste authorization. / **技能页剪贴板指引**:未开启剪贴板历史时,技能页展示可点击卡片,分别跳转 App 内剪贴板设置和系统设置以完成粘贴授权。
- **Custom skills**: Skills tab `+` adds a user skill (name, about, SF Symbol, prompt, required iCloud Shortcut link with name lookup, independent Shortcut name, thinking off by default). Built-in thinking stays off and disabled. No cap on how many custom skills you can save; the keyboard still holds at most 8. / **自定义技能**:技能页右上角 `+` 可添加用户技能(名称、介绍、SF Symbol、提示词、必填 iCloud 捷径链接并自动读取名称、可与技能名分开的捷径名、思考默认关)。内置技能思考固定关闭且不可开。自定义数量不设上限,键盘仍最多启用 8 个。
- **Custom skills**: Skills tab `+` adds a user skill (name, about, SF Symbol, prompt, optional iCloud Shortcut link with name lookup, independent Shortcut name, thinking off by default). Without a link, the model result is reviewed and inserted directly; with a link, it keeps the Shortcut export flow. Built-in thinking stays off and disabled. No cap on how many custom skills you can save; the keyboard still holds at most 8. / **自定义技能**:技能页右上角 `+` 可添加用户技能(名称、介绍、SF Symbol、提示词、可选 iCloud 捷径链接并自动读取名称、可与技能名分开的捷径名、思考默认关)。不填链接时,模型结果经确认后直接插入;填写链接时继续走捷径导出流程。内置技能思考固定关闭且不可开。自定义数量不设上限,键盘仍最多启用 8 个。
### Changed
- **Keyboard input tabs**: the four-tab capsule is centered independently of the side controls, with equal 42 pt hit widths and a translucent-black light-mode track. AI and Voice use enlarged `sparkle` and `waveform.mid` symbols, and the leading logo is 16 pt tall. / **键盘输入标签**:四标签胶囊不受两侧控件影响并在键盘上独立居中,点击宽度统一为 42 pt,浅色模式轨道使用半透明黑色;AI 与语音使用放大的 `sparkle``waveform.mid` 图标,左侧 Logo 高度为 16 pt。
- **English lexicon mmap**: the 40k-word English table loads only in English and stays file-mapped; Chinese typing no longer pulls it in on keyboard appear. / **英文词表 mmap**:4 万词英文表仅在英文加载且走文件映射;中文打字不再在唤起键盘时一并灌入。
- **English autocorrect conservatism**: Title Case / short / ALL CAPS tokens are not replaced, except same-length transpositions (`Teh``The`). Machine-applied corrections no longer boost the replacement; rejecting them learns the original. / **英文自动更正更克制**Title Case / 短词 / 全大写默认不改,仅保留同长换位(`Teh``The`)。机器改写不再给新词加分;拒绝纠错会学会原文。
### Fixed
- **Keyboard switch crash**: `requestSupplementaryLexicon` completion hops to the main actor before writing session state, so switching to OSG Keyboard no longer traps on `com.apple.TextInput.lexicon-request`. / **切换键盘崩溃**`requestSupplementaryLexicon` 回调先回到主线程再写会话状态,切换到 OSG Keyboard 不再在 `com.apple.TextInput.lexicon-request` 上触发隔离断言。
- **Save to Notes Shortcut**: bind the combined first-line title and clipboard body through iPhone Create Notes real `WFCreateNoteInput` field. The previous `contents` binding was ignored and left an enter-content sheet or a title-only note. / **存入备忘录捷径**:通过 iPhone「创建备忘录」真正的 `WFCreateNoteInput` 字段绑定首行标题与剪贴板正文;旧版 `contents` 绑定会被忽略,导致弹出内容填写框或只生成标题。
- **Extract tasks Shortcut**: receive Shortcut Input as Text, then split lines and add each title to Reminders — the previous recipe could finish successfully without creating items. / **提取待办捷径**:先把快捷指令输入收成文本,再按行写入提醒;旧配方会成功跑完但不创建条目。
- **Skill reorder feedback**: long-press lifts a skill card and the grid slides live under the finger, matching Home Screen rearrange. / **技能拖动排序**:长按拎起技能卡片,网格随手指实时让位,接近主屏幕图标重排。
+10 -5
View File
@@ -30,11 +30,16 @@ commit, SHA-256 and output SHA-256.
- Eight opt-in fuzzy-pinyin rule groups.
- Offline English typing data under
`OSGKeyboardShared/Resources/Typing/English/`:
- `english_lexicon.tsv`curated word list with synthetic relative
frequency ranks for autocomplete / autocorrect
- `english_bigrams.tsv` — light next-word candidates
- Not derived from GPL/LGPL dictionaries; ranks are ordering weights,
not a single third-party corpus dump
- `english_lexicon.bin`mmap-friendly binary of the top 40k alphabetic
unigrams (log-scaled ranks) plus truncated bigrams; this is what the
keyboard extension loads
- `english_lexicon.tsv` / `english_bigrams.tsv` — build inputs derived from
Peter Norvigs public-domain `count_1w.txt` / `count_2w.txt`
(https://norvig.com/ngrams/; not GPL/LGPL dictionaries). Not copied into
the app bundle.
- Rebuild with `python3 Scripts/typing/build_english_lexicon.py`
(add `--from-tsv` to compile the binary from existing TSV without network)
## Reference only
+27 -1
View File
@@ -31,6 +31,7 @@ struct OSGKeyboardApp: App {
#if DEBUG
if ProcessInfo.processInfo.arguments.contains("--whats-new-host") {
// Approach A: Notes-like host only; real keyboard extension overlays it.
// Also used by `--keyboard-appear-stress=` (pass both flags).
Self.makeWhatsNewHostView()
} else if ProcessInfo.processInfo.arguments.contains("--edit-demo") {
EditDemoView()
@@ -79,8 +80,13 @@ struct OSGKeyboardApp: App {
let scenario = whatsNewScenario(from: args) ?? .edit
let language = whatsNewLanguage(from: args)
let seed = whatsNewSeedText(for: scenario, language: language)
let appearStressCount = keyboardAppearStressCount(from: args)
WhatsNewDemoScenario.clear()
// Stress must not arm What's New playback that drives keys on the
// extension while we are tearing it down.
if appearStressCount == 0 {
WhatsNewDemoScenario.arm(scenario, seedText: seed, language: language)
}
if let defaults = AppGroup.defaultsIfAvailable {
defaults.set(true, forKey: AppGroupConfiguration.Keys.hasCompletedOnboarding)
// Force extension ExtL10n / SharedL10n into the demo language.
@@ -98,13 +104,33 @@ struct OSGKeyboardApp: App {
}
defaults.synchronize()
}
if appearStressCount > 0, let defaults = AppGroup.defaultsIfAvailable {
// Hit the crash path: typing surface + English supplementary lexicon.
defaults.set("english", forKey: "typing.input.defaultInputMode")
defaults.set(true, forKey: "typing.input.rememberLastSurface")
defaults.set("typing", forKey: "typing.input.lastSurface")
defaults.set("english", forKey: "typing.input.lastTypingLanguage")
defaults.synchronize()
}
return NotesHostDemoView(
scenario: scenario,
seedText: seed,
language: language
language: language,
appearStressCount: appearStressCount
)
}
private static func keyboardAppearStressCount(from args: [String]) -> Int {
if let paired = args.first(where: { $0.hasPrefix("--keyboard-appear-stress=") }) {
return Int(paired.dropFirst("--keyboard-appear-stress=".count)) ?? 0
}
if let idx = args.firstIndex(of: "--keyboard-appear-stress"),
args.index(after: idx) < args.endIndex {
return Int(args[args.index(after: idx)]) ?? 0
}
return 0
}
private static func whatsNewScenario(from args: [String]) -> WhatsNewDemoScenario? {
if let paired = args.first(where: { $0.hasPrefix("--whats-new-scenario=") }) {
let raw = String(paired.dropFirst("--whats-new-scenario=".count))
@@ -0,0 +1,7 @@
Peter Norvigs n-gram count files (https://norvig.com/ngrams/)
Norvig states: “I hereby release all these files into the public domain.”
OSGKeyboard does not redistribute the raw count files. `Scripts/typing/build_english_lexicon.py`
derives compact log-scaled unigram ranks and a truncated bigram list, then
compiles `english_lexicon.bin` for the keyboard extension to mmap.
@@ -30,6 +30,7 @@ all required copyright and permission notices must remain with distributions.
English typing lexicon
----------------------
english_lexicon.tsv and english_bigrams.tsv are OSG-curated word lists with
synthetic relative frequency ranks for offline autocomplete / autocorrect /
next-word ranking. They are not derived from GPL/LGPL dictionaries.
english_lexicon.bin (from english_lexicon.tsv / english_bigrams.tsv) is an
OSG-curated mmap ranking table with synthetic relative frequency ranks for
offline autocomplete / autocorrect / next-word ranking. It is not derived
from GPL/LGPL dictionaries.
@@ -105,7 +105,7 @@ enum OpenSourceLicenseCatalog {
id: "english-typing-lexicon",
name: "OSG English typing lexicon",
licenseName: "Project-owned notice",
purpose: "Offline English autocomplete, autocorrect, and next-word ranking lists curated by OSGKeyboard (english_lexicon.tsv / english_bigrams.tsv). Not derived from GPL/LGPL dictionaries; relative ranks are ordering weights only.",
purpose: "Offline English autocomplete, autocorrect, and next-word ranking (english_lexicon.bin, compiled from TSV). Log-scaled ranks derived from Peter Norvigs public-domain n-gram counts; not GPL/LGPL dictionaries.",
url: URL(string: "https://github.com/hkgood/OSGKeyboard/blob/main/NOTICE-TYPING.md"),
licenseText: englishLexiconNoticeText,
platforms: [.iOS]
@@ -191,16 +191,17 @@ enum OpenSourceLicenseCatalog {
static let englishLexiconNoticeText = """
OSG English typing lexicon (project-owned notice)
english_lexicon.tsv and english_bigrams.tsv are curated by OSGKeyboard for
offline English autocomplete, autocorrect, and next-word ranking inside the
iOS keyboard extension.
english_lexicon.bin (compiled from english_lexicon.tsv / english_bigrams.tsv)
is the mmap ranking table for offline English autocomplete, autocorrect,
and next-word prediction.
These lists are not derived from GPL or LGPL dictionaries. Relative
frequency values are synthetic ordering weights for ranking only, not
verbatim counts from a single third-party corpus.
Unigram ranks and truncated bigrams are derived from Peter Norvigs
public-domain n-gram count files (https://norvig.com/ngrams/). OSGKeyboard
does not ship the raw corpus. Relative frequency values are log-scaled
ordering weights, not verbatim Google counts.
See NOTICE-TYPING.md in the OSGKeyboard repository for the full typing
keyboard attribution map (Chinese Rime stack vs OSG-owned English data).
keyboard attribution map (Chinese Rime stack vs English data).
"""
static let bsd3Text = """
+22 -6
View File
@@ -334,16 +334,25 @@ struct AIAgentSkillsView: View {
}
private func saveDraft(_ draft: SkillEditorDraft) throws {
guard let url = AIShortcutShareLink.parse(draft.shortcutLink) else {
let rawShortcutLink = draft.shortcutLink.trimmingCharacters(
in: .whitespacesAndNewlines
)
let shortcutURL: URL?
if rawShortcutLink.isEmpty {
shortcutURL = nil
} else {
guard let parsedURL = AIShortcutShareLink.parse(rawShortcutLink) else {
throw AIUserSkillValidationError.invalidShortcutLink
}
shortcutURL = parsedURL
}
let skill = AIUserSkill(
id: draft.id,
name: draft.name,
summary: draft.summary,
systemImage: draft.systemImage,
prompt: draft.prompt,
shortcutICloudURL: url,
shortcutICloudURL: shortcutURL,
shortcutName: draft.shortcutName,
thinkingEnabled: draft.thinkingEnabled
)
@@ -565,7 +574,7 @@ private struct SkillEditorDraft: Identifiable, Equatable {
summary: skill.summary,
systemImage: skill.systemImage,
prompt: skill.prompt,
shortcutLink: skill.shortcutICloudURL.absoluteString,
shortcutLink: skill.shortcutICloudURL?.absoluteString ?? "",
shortcutName: skill.shortcutName,
thinkingEnabled: skill.thinkingEnabled
)
@@ -957,11 +966,18 @@ private struct SkillEditorSheet: View {
}
private var canSave: Bool {
!name.trimmingCharacters(in: .whitespacesAndNewlines).isEmpty
let trimmedShortcutLink = shortcutLink.trimmingCharacters(
in: .whitespacesAndNewlines
)
let validShortcutConfiguration = trimmedShortcutLink.isEmpty
|| (
AIShortcutShareLink.parse(trimmedShortcutLink) != nil
&& !shortcutName.trimmingCharacters(in: .whitespacesAndNewlines).isEmpty
)
return !name.trimmingCharacters(in: .whitespacesAndNewlines).isEmpty
&& !prompt.trimmingCharacters(in: .whitespacesAndNewlines).isEmpty
&& prompt.count <= AIUserSkillLimits.maximumPromptCharacters
&& !shortcutName.trimmingCharacters(in: .whitespacesAndNewlines).isEmpty
&& AIShortcutShareLink.parse(shortcutLink) != nil
&& validShortcutConfiguration
}
private var currentDraft: SkillEditorDraft {
+142 -1
View File
@@ -14,6 +14,7 @@ struct NotesHostDemoView: View {
let scenario: WhatsNewDemoScenario
let seedText: String
let language: WhatsNewDemoScenario.Language
var appearStressCount: Int = 0
private var title: String {
switch (scenario, language) {
@@ -43,7 +44,7 @@ struct NotesHostDemoView: View {
.fill(Color(uiColor: .secondarySystemGroupedBackground))
)
} else {
NotesHostTextView(text: seedText)
NotesHostTextView(text: seedText, appearStressCount: appearStressCount)
.padding(16)
.frame(maxWidth: .infinity, maxHeight: .infinity, alignment: .topLeading)
.background(
@@ -62,6 +63,7 @@ struct NotesHostDemoView: View {
language == .en ? Locale(identifier: "en") : Locale(identifier: "zh-Hans")
)
.task {
guard appearStressCount == 0 else { return }
// Refresh TTL while armed; stop once the extension consumes / plays.
WhatsNewDemoScenario.arm(scenario, seedText: seedText, language: language)
for _ in 0..<25 {
@@ -78,6 +80,7 @@ struct NotesHostDemoView: View {
/// the real custom keyboard extension.
private struct NotesHostTextView: UIViewRepresentable {
let text: String
var appearStressCount: Int = 0
func makeUIView(context: Context) -> UITextView {
let view = UITextView()
@@ -91,9 +94,16 @@ private struct NotesHostTextView: UIViewRepresentable {
view.textContainer.lineFragmentPadding = 0
view.returnKeyType = .default
view.delegate = context.coordinator
view.accessibilityIdentifier = "notes.host.textView"
context.coordinator.appearStressCount = appearStressCount
context.coordinator.textView = view
DispatchQueue.main.asyncAfter(deadline: .now() + 0.35) {
if context.coordinator.appearStressCount > 0 {
context.coordinator.startAppearStressIfNeeded()
} else {
view.becomeFirstResponder()
}
}
return view
}
@@ -101,6 +111,8 @@ private struct NotesHostTextView: UIViewRepresentable {
if uiView.text != text, !context.coordinator.userEdited {
uiView.text = text
}
// Stress owns first-responder; don't fight resignFirstResponder.
guard appearStressCount == 0 else { return }
if !uiView.isFirstResponder {
DispatchQueue.main.async {
_ = uiView.becomeFirstResponder()
@@ -112,10 +124,139 @@ private struct NotesHostTextView: UIViewRepresentable {
final class Coordinator: NSObject, UITextViewDelegate {
var userEdited = false
var appearStressCount = 0
weak var textView: UITextView?
private var started = false
private var waitingForShow = false
private var waitingForHide = false
private var showWaiter: CheckedContinuation<Bool, Never>?
private var hideWaiter: CheckedContinuation<Bool, Never>?
/// Invalidates leftover timeout tasks from a finished wait.
private var waitGeneration = 0
func textViewDidChange(_ textView: UITextView) {
userEdited = true
}
func startAppearStressIfNeeded() {
guard appearStressCount > 0, !started else { return }
started = true
NotificationCenter.default.addObserver(
self,
selector: #selector(keyboardDidShow),
name: UIResponder.keyboardDidShowNotification,
object: nil
)
NotificationCenter.default.addObserver(
self,
selector: #selector(keyboardDidHide),
name: UIResponder.keyboardDidHideNotification,
object: nil
)
Task { @MainActor [weak self] in
await self?.runAppearStress()
}
}
deinit {
NotificationCenter.default.removeObserver(self)
}
@objc private func keyboardDidShow(_ notification: Notification) {
finishWait(show: true, success: true)
}
@objc private func keyboardDidHide(_ notification: Notification) {
finishWait(show: false, success: true)
}
private func finishWait(show: Bool, success: Bool) {
if show {
guard waitingForShow, let pending = showWaiter else { return }
waitingForShow = false
showWaiter = nil
waitGeneration += 1
pending.resume(returning: success)
} else {
guard waitingForHide, let pending = hideWaiter else { return }
waitingForHide = false
hideWaiter = nil
waitGeneration += 1
pending.resume(returning: success)
}
}
@MainActor
private func runAppearStress() async {
let total = appearStressCount
OSGDiag.log("keyboard.stress begin count=\(total)", category: "boot")
guard let textView else {
OSGDiag.log("keyboard.stress FAIL textView gone", category: "boot")
return
}
guard await becomeAndWaitForShow(textView, timeoutNanoseconds: 8_000_000_000) else {
OSGDiag.log("keyboard.stress FAIL first-show timeout", category: "boot")
return
}
OSGDiag.log("keyboard.stress first-show ok", category: "boot")
var passed = 0
for cycle in 1...total {
guard await resignAndWaitForHide(textView, timeoutNanoseconds: 5_000_000_000) else {
OSGDiag.log("keyboard.stress FAIL cycle=\(cycle) hide timeout", category: "boot")
break
}
try? await Task.sleep(nanoseconds: 350_000_000)
guard await becomeAndWaitForShow(textView, timeoutNanoseconds: 8_000_000_000) else {
OSGDiag.log("keyboard.stress FAIL cycle=\(cycle) show timeout", category: "boot")
break
}
passed += 1
OSGDiag.log("keyboard.stress cycle=\(passed)/\(total) ok", category: "boot")
try? await Task.sleep(nanoseconds: 200_000_000)
}
OSGDiag.log("keyboard.stress done passed=\(passed)/\(total)", category: "boot")
try? await Task.sleep(nanoseconds: 250_000_000)
exit(passed == total ? 0 : 1)
}
@MainActor
private func becomeAndWaitForShow(_ textView: UITextView, timeoutNanoseconds: UInt64) async -> Bool {
await waitForKeyboard(show: true, timeoutNanoseconds: timeoutNanoseconds) {
textView.becomeFirstResponder()
}
}
@MainActor
private func resignAndWaitForHide(_ textView: UITextView, timeoutNanoseconds: UInt64) async -> Bool {
await waitForKeyboard(show: false, timeoutNanoseconds: timeoutNanoseconds) {
textView.resignFirstResponder()
}
}
@MainActor
private func waitForKeyboard(
show: Bool,
timeoutNanoseconds: UInt64,
trigger: () -> Void
) async -> Bool {
await withCheckedContinuation { continuation in
waitGeneration += 1
let generation = waitGeneration
if show {
waitingForShow = true
showWaiter = continuation
} else {
waitingForHide = true
hideWaiter = continuation
}
trigger()
Task { @MainActor [weak self] in
try? await Task.sleep(nanoseconds: timeoutNanoseconds)
guard let self, generation == waitGeneration else { return }
self.finishWait(show: show, success: false)
}
}
}
}
}
+3 -3
View File
@@ -477,10 +477,10 @@
"skills.editor.prompt" = "Processing prompt";
"skills.editor.thinking" = "Thinking";
"skills.editor.thinkingHint" = "Off by default. Turn on only when you want slower, deeper reasoning for this skill.";
"skills.editor.shortcut" = "Shortcut";
"skills.editor.shortcut" = "Shortcut (Optional)";
"skills.editor.linkPlaceholder" = "https://www.icloud.com/shortcuts/…";
"skills.editor.shortcutNamePlaceholder" = "Shortcut name (can differ from the skill name)";
"skills.editor.shortcutHint" = "Paste an iCloud share link. The published Shortcut name is filled in automatically and you can change it. Dont rename it in the Shortcuts app after adding.";
"skills.editor.shortcutNamePlaceholder" = "Shortcut name (required with a link)";
"skills.editor.shortcutHint" = "Leave the link empty to process text and insert the result after review. With a link, the Shortcut name is filled in automatically. Dont rename it after adding.";
"skills.editor.lookingUp" = "Looking up Shortcut name…";
"skills.editor.resolvedName" = "Will run: %@";
"skills.editor.lookupFailed" = "Couldnt read the Shortcut name. Check the link, or type the name yourself.";
@@ -476,10 +476,10 @@
"skills.editor.prompt" = "文本处理提示词";
"skills.editor.thinking" = "思考";
"skills.editor.thinkingHint" = "默认关闭。仅在需要该技能更慢、更深的推理时开启。";
"skills.editor.shortcut" = "捷径";
"skills.editor.shortcut" = "捷径(可选)";
"skills.editor.linkPlaceholder" = "https://www.icloud.com/shortcuts/…";
"skills.editor.shortcutNamePlaceholder" = "捷径名称(可与技能名称不同";
"skills.editor.shortcutHint" = "粘贴 iCloud 分享链接。发布名称会自动填入,也可以自行修改。添加到「快捷指令」后请勿改名。";
"skills.editor.shortcutNamePlaceholder" = "捷径名称(填写链接时必填";
"skills.editor.shortcutHint" = "不填链接时仅处理文字,结果确认后插入。填写链接时会自动读取捷径名称;添加到「快捷指令」后请勿改名。";
"skills.editor.lookingUp" = "正在读取捷径名称…";
"skills.editor.resolvedName" = "将运行:%@";
"skills.editor.lookupFailed" = "无法读取捷径名称。请检查链接,或手动填写名称。";
@@ -42,6 +42,7 @@ public final class KeyboardViewController: UIInputViewController {
private var typingSession: TypingSessionController {
if let typingSessionStorage { return typingSessionStorage }
let created = TypingSessionController()
created.systemLexicon = UIKitEnglishSystemLexicon()
typingSessionStorage = created
return created
}
@@ -233,6 +234,7 @@ public final class KeyboardViewController: UIInputViewController {
if state.surface == .typing {
OSGDiag.log("KVC.viewWillAppear enterTypingMode", category: "boot")
typingSession.enterTypingMode()
refreshEnglishSupplementaryLexicon()
}
clipboardCapture.keyboardDidAppear()
OSGDiag.log(
@@ -633,6 +635,7 @@ public final class KeyboardViewController: UIInputViewController {
state.surface = surface
if surface == .typing {
typingSession.enterTypingMode()
refreshEnglishSupplementaryLexicon()
} else {
typingSession.leaveTypingMode()
}
@@ -754,6 +757,33 @@ public final class KeyboardViewController: UIInputViewController {
}
}
/// Contacts and user text replacements, without a Contacts permission.
///
/// `requestSupplementaryLexicon` replies on `com.apple.TextInput.lexicon-request`,
/// not the main actor. Touching `TypingSessionController` there traps in Swift 6
/// (`_dispatch_assert_queue_fail`) and the extension is killed on appear.
private func refreshEnglishSupplementaryLexicon() {
requestSupplementaryLexicon { @Sendable lexicon in
Task { @MainActor [weak self] in
self?.applySupplementaryLexicon(lexicon)
}
}
}
private func applySupplementaryLexicon(_ lexicon: UILexicon) {
typingSessionStorage?.supplementaryWords = lexicon.entries.compactMap { entry -> String? in
let text = entry.documentText.trimmingCharacters(in: .whitespacesAndNewlines)
guard !text.isEmpty else { return nil }
let parts = text.split { $0.isWhitespace || $0 == "," }
guard parts.count == 1 else { return nil }
let token = String(parts[0])
guard token.allSatisfy({ $0.isLetter || $0 == "'" || $0 == "" || $0 == "-" }) else {
return nil
}
return token
}
}
private static func typingAutocapitalizationMode(
for type: UITextAutocapitalizationType
) -> TypingAutocapitalizationMode {
+64 -14
View File
@@ -142,6 +142,14 @@ struct TypingRootView: View {
}
private var idleTopBar: some View {
ZStack {
KeyboardTopControls(
state: state,
typing: typing,
palette: palette,
onInsert: onInsert
)
HStack(spacing: Spacing.xs) {
KeyboardBrandLogo(action: state.openSettings)
// Globe key now lives at the bottom-left of the keyboard (matching
@@ -158,13 +166,7 @@ struct TypingRootView: View {
}
Spacer(minLength: 0)
KeyboardTopControls(
state: state,
typing: typing,
palette: palette,
onInsert: onInsert
)
}
}
.padding(.horizontal, KeyboardTopBarMetrics.nestedHorizontalInset)
}
@@ -206,6 +208,61 @@ struct TypingRootView: View {
editingToolbar
.padding(.leading, KeyboardTopBarMetrics.nestedHorizontalInset)
}
if typing.language == .english {
englishQuickTypeBar
} else {
chineseCandidateStrip
}
if typing.canExpandCandidatePanel {
expandChevronButton
} else {
Color.clear.frame(width: KeyboardTopBarMetrics.nestedHorizontalInset)
}
}
}
/// Three equal QuickType slots. Space applies only `role == .correction`.
private var englishQuickTypeBar: some View {
HStack(spacing: 0) {
ForEach(
Array(
typing.composition.candidates
.prefix(EnglishSuggestionEngine.slotCount)
.enumerated()
),
id: \.element.id
) { index, candidate in
if index > 0 {
Rectangle()
.fill(palette.dividerStrong)
.frame(width: 1, height: 18)
}
englishQuickTypeSlot(candidate, index: index)
}
}
.frame(maxWidth: .infinity)
.padding(.leading, KeyboardTopBarMetrics.nestedHorizontalInset)
}
private func englishQuickTypeSlot(_ candidate: TypingCandidate, index: Int) -> some View {
let label = candidate.isQuoted ? "\"\(candidate.text)\"" : candidate.text
let weight: Font.Weight = candidate.role == .correction ? .semibold : .regular
return Text(label)
.font(.system(size: 17, weight: weight))
.foregroundStyle(palette.textPrimary)
.lineLimit(1)
.minimumScaleFactor(0.75)
.frame(maxWidth: .infinity, minHeight: 40)
.contentShape(Rectangle())
.onTapGesture {
apply(typing.selectCandidate(at: index))
}
.accessibilityAddTraits(.isButton)
.accessibilityLabel(candidate.text)
}
private var chineseCandidateStrip: some View {
ScrollView(.horizontal, showsIndicators: false) {
HStack(spacing: Spacing.xs) {
if typing.composition.candidates.isEmpty {
@@ -247,13 +304,6 @@ struct TypingRootView: View {
.padding(.trailing, Spacing.xs)
}
.scrollBounceBehavior(.basedOnSize, axes: .horizontal)
if typing.canExpandCandidatePanel {
expandChevronButton
} else {
Color.clear.frame(width: KeyboardTopBarMetrics.nestedHorizontalInset)
}
}
}
/// Opaque chip like the translation control so never shares pixels with text.
+5 -3
View File
@@ -132,15 +132,17 @@ struct AIKeyboardView: View {
)
.padding(.horizontal, KeyboardTopBarMetrics.nestedHorizontalInset)
} else {
HStack(spacing: Spacing.xs) {
KeyboardBrandLogo(action: state.openSettings)
Spacer(minLength: 0)
ZStack {
KeyboardTopControls(
state: state,
typing: typing,
palette: palette,
onInsert: onInsert
)
HStack {
KeyboardBrandLogo(action: state.openSettings)
Spacer(minLength: 0)
}
}
.padding(.horizontal, KeyboardTopBarMetrics.nestedHorizontalInset)
}
+7 -5
View File
@@ -238,17 +238,19 @@ public struct KeyboardRootView: View {
onDismiss: state.dismissClipboardSuggestion
)
} else {
HStack(spacing: Spacing.xs) {
KeyboardBrandLogo(action: state.openSettings)
// Globe key now lives at the bottom-left of the keyboard (matching
// iOS system layout); see micActionRow's bottom HStack.
Spacer(minLength: 0)
ZStack {
KeyboardTopControls(
state: state,
typing: typing,
palette: palette,
onInsert: onInsert
)
HStack {
KeyboardBrandLogo(action: state.openSettings)
// Globe key now lives at the bottom-left of the keyboard (matching
// iOS system layout); see micActionRow's bottom HStack.
Spacer(minLength: 0)
}
}
}
}
+42 -22
View File
@@ -23,8 +23,10 @@ enum KeyboardTopBarMetrics {
static let horizontalInset: CGFloat = 12
/// TypingRootView already contributes 8 pt around the entire key surface.
static let nestedHorizontalInset: CGFloat = horizontalInset - KeyboardChromeLayout.horizontalInset
static let logoHeight: CGFloat = 22
static let logoHeight: CGFloat = 16
static let logoWidth: CGFloat = logoHeight * 952 / 291
/// Equal hit width for AI / Voice / Chinese / English input tabs.
static let inputTabWidth: CGFloat = 42
/// Shared footprint for top-trailing chips (clipboard, cancel/X, translation).
static let trailingChipSize: CGFloat = 34
static let trailingChipIconSize: CGFloat = 15
@@ -113,8 +115,24 @@ struct KeyboardTopControls: View {
let onInsert: (String) -> Void
var body: some View {
HStack(spacing: 6) {
// /
ZStack {
inputTabSwitcher
if state.canShowClipboardEntry {
HStack {
Spacer(minLength: 0)
KeyboardClipboardMenuButton(
palette: palette,
action: state.openClipboardPanel
)
.equatable()
}
}
}
.frame(maxWidth: .infinity)
}
private var inputTabSwitcher: some View {
// Logo /
HStack(spacing: 2) {
ForEach(KeyboardInputTab.allCases, id: \.self) { tab in
tabButton(tab)
@@ -125,27 +143,17 @@ struct KeyboardTopControls: View {
.overlay(
Capsule().stroke(palette.divider, lineWidth: 0.5)
)
if state.canShowClipboardEntry {
KeyboardClipboardMenuButton(
palette: palette,
action: state.openClipboardPanel
)
.equatable()
}
}
}
private func tabButton(_ tab: KeyboardInputTab) -> some View {
let selected = isSelected(tab)
let width: CGFloat = tab == .english || tab == .ai ? 34 : 42
return Button {
withAnimation(Motion.soft) {
select(tab)
}
} label: {
tabLabel(tab, selected: selected, width: width)
tabLabel(tab, selected: selected)
}
.buttonStyle(TopControlPressStyle(pressedFill: pressedFill))
.disabled(tab != .voice && !state.canEnterTypingSurface)
@@ -157,13 +165,11 @@ struct KeyboardTopControls: View {
@ViewBuilder
private func tabLabel(
_ tab: KeyboardInputTab,
selected: Bool,
width: CGFloat
selected: Bool
) -> some View {
let label = Text(tab.title)
.font(.system(size: 12, weight: selected ? .semibold : .medium))
let label = tabContent(tab, selected: selected)
.foregroundStyle(selected ? palette.textPrimary : palette.textSecondary)
.frame(width: width, height: 30)
.frame(width: KeyboardTopBarMetrics.inputTabWidth, height: 30)
if selected {
let namespace = sharedSelectionNamespace ?? fallbackSelectionNamespace
@@ -179,6 +185,20 @@ struct KeyboardTopControls: View {
}
}
@ViewBuilder
private func tabContent(_ tab: KeyboardInputTab, selected: Bool) -> some View {
if tab == .ai {
Image(systemName: "sparkle")
.font(.system(size: 15, weight: selected ? .semibold : .medium))
} else if tab == .voice {
Image(systemName: "waveform.mid")
.font(.system(size: 15, weight: selected ? .semibold : .medium))
} else {
Text(tab.title)
.font(.system(size: 12, weight: selected ? .semibold : .medium))
}
}
private func tabOpacity(_ tab: KeyboardInputTab) -> Double {
guard tab != .voice, !state.canEnterTypingSurface else { return 1 }
if case .recording = state.phase {
@@ -191,10 +211,10 @@ struct KeyboardTopControls: View {
colorScheme == .dark ? Color(white: 0.22) : Color(white: 0.84)
}
/// NativeKeyboardKeyColors.fill
///
/// NativeKeyboardKeyColors.fill
/// 宿
private var tabTrackFill: Color {
colorScheme == .dark ? Color(white: 0.12) : Color(white: 0.87)
colorScheme == .dark ? Color(white: 0.12) : Color.black.opacity(0.12)
}
private func isSelected(_ tab: KeyboardInputTab) -> Bool {
+225 -20
View File
@@ -10,18 +10,42 @@ final class EnglishTypingTests: XCTestCase {
func testLexiconLoadsAndCompletesPrefix() {
let lexicon = EnglishLexicon()
lexicon.prepare()
XCTAssertGreaterThan(lexicon.wordCount, 1_000)
XCTAssertTrue(lexicon.isLoaded)
XCTAssertGreaterThan(lexicon.wordCount, 10_000)
let hits = lexicon.completions(prefix: "hel", limit: 5)
XCTAssertTrue(hits.contains("hello") || hits.contains("help") || hits.contains("held"))
lexicon.unload()
XCTAssertFalse(lexicon.isLoaded)
XCTAssertEqual(lexicon.wordCount, 0)
}
@MainActor
func testChineseTypingDoesNotLoadEnglishLexicon() {
EnglishLexicon.shared.unload()
let typing = TypingSessionController()
_ = typing.setLanguage(.chinese)
typing.enterTypingMode()
XCTAssertFalse(EnglishLexicon.shared.isLoaded)
_ = typing.setLanguage(.english)
XCTAssertTrue(EnglishLexicon.shared.isLoaded)
_ = typing.setLanguage(.chinese)
XCTAssertFalse(EnglishLexicon.shared.isLoaded)
typing.leaveTypingMode()
}
func testCorrectionFindsNearbyWord() {
let lexicon = EnglishLexicon()
lexicon.prepare()
// "teh" is a classic typo for "the".
let correction = lexicon.bestCorrection(for: "teh")
XCTAssertEqual(correction, "the")
XCTAssertNil(lexicon.bestCorrection(for: "the"))
let engine = EnglishSuggestionEngine()
engine.prepare()
// "teh" leaks into web unigrams; the engine must still treat it as a typo.
let decision = engine.correctionDecision(
for: "teh",
personalTerms: [],
learnedBoosts: [:]
)
XCTAssertEqual(decision?.replacement, "the")
XCTAssertNil(
engine.correctionDecision(for: "the", personalTerms: [], learnedBoosts: [:])
)
}
func testSuggestionEngineSkipsPersonalDictionaryTypos() {
@@ -45,7 +69,22 @@ final class EnglishTypingTests: XCTestCase {
learnedBoosts: [:]
)
)
XCTAssertEqual(composition.candidates.first?.text, "OSGKeyboard")
XCTAssertEqual(composition.candidates.first?.role, .verbatim)
XCTAssertEqual(composition.candidates.first?.text, "osg")
XCTAssertTrue(composition.candidates.contains { $0.text == "OSGKeyboard" })
}
func testSuggestionEngineReturnsNoCandidatesWithoutCurrentWord() {
let engine = EnglishSuggestionEngine()
engine.prepare()
let composition = engine.compositionWhileTyping(
EnglishSuggestionContext(
previousWord: "hello",
personalTerms: ["OSGKeyboard"]
)
)
XCTAssertEqual(composition, .empty)
}
func testAutocapitalizationAtFieldStartAndAfterSentence() {
@@ -208,7 +247,7 @@ final class EnglishTypingTests: XCTestCase {
XCTAssertEqual(typing.composition.preedit, "boa")
XCTAssertTrue(
typing.composition.candidates.contains {
$0.text.compare("boat", options: .caseInsensitive) == .orderedSame
$0.text.compare("board", options: .caseInsensitive) == .orderedSame
}
)
@@ -232,14 +271,14 @@ final class EnglishTypingTests: XCTestCase {
preceding += output.text
typing.syncAutocapitalization(accountingForInsert: output.text)
}
guard let boatIndex = typing.composition.candidates.firstIndex(where: {
$0.text.compare("boat", options: .caseInsensitive) == .orderedSame
guard let boardIndex = typing.composition.candidates.firstIndex(where: {
$0.text.compare("board", options: .caseInsensitive) == .orderedSame
}) else {
return XCTFail("expected boat completion")
return XCTFail("expected board completion")
}
preceding = "board"
let output = typing.selectCandidate(at: boatIndex)
let output = typing.selectCandidate(at: boardIndex)
XCTAssertEqual(output, .none)
XCTAssertEqual(typing.composition.preedit.lowercased(), "board")
@@ -283,23 +322,18 @@ final class EnglishTypingTests: XCTestCase {
@MainActor
func testAutocorrectUndoRestoresOriginal() {
let typing = TypingSessionController()
typing.suggestionsEnabled = true
_ = typing.setLanguage(.english)
typing.enterTypingMode()
let typing = makeIsolatedEnglishSession(suite: "english.undo.test")
for ch in ["t", "e", "h"] {
_ = typing.handleKey(ch)
}
let spaced = typing.handleSpace()
// Either corrected to "the " or left as-is if lexicon missing in test bundle.
if spaced.deleteCount > 0 {
XCTAssertEqual(spaced.deleteCount, 3)
XCTAssertTrue(spaced.text.hasPrefix("the"))
let undone = typing.handleKey("")
XCTAssertEqual(undone.text, "teh")
XCTAssertEqual(undone.deleteCount, spaced.text.count)
}
}
@MainActor
func testPeriodShortcutReplacesDoubleSpace() {
@@ -370,6 +404,177 @@ final class EnglishTypingTests: XCTestCase {
XCTAssertFalse(PeriodShortcut.shouldArm(afterSpaceFollowing: "hello "))
}
@MainActor
func testEnglishQuickTypePutsVerbatimFirstAndMarksCorrection() {
let typing = makeIsolatedEnglishSession(suite: "english.quicktype.bar.test")
for character in ["t", "e", "h"] {
_ = typing.handleKey(character)
}
XCTAssertEqual(typing.composition.candidates.first?.role, .verbatim)
XCTAssertEqual(
typing.composition.candidates.first?.text.lowercased(),
"teh"
)
XCTAssertTrue(
typing.composition.candidates.contains {
$0.role == .correction && $0.text.lowercased() == "the"
}
)
XCTAssertLessThanOrEqual(typing.composition.candidates.count, EnglishSuggestionEngine.slotCount)
}
@MainActor
func testEnglishSpaceAppliesCorrectionSlotOnly() {
let typing = makeIsolatedEnglishSession(suite: "english.quicktype.space.test")
for character in ["t", "e", "h"] {
_ = typing.handleKey(character)
}
let spaced = typing.handleSpace()
XCTAssertEqual(spaced.deleteCount, 3)
XCTAssertTrue(spaced.text.lowercased().hasPrefix("the"))
XCTAssertTrue(typing.composition.candidates.isEmpty)
}
@MainActor
func testEnglishSpaceKeepsVerbatimWhenNoCorrection() {
let typing = makeIsolatedEnglishSession(suite: "english.quicktype.verbatim.test")
for character in ["h", "e", "l"] {
_ = typing.handleKey(character)
}
let spaced = typing.handleSpace()
XCTAssertEqual(spaced, .insert(" "))
XCTAssertTrue(typing.composition.candidates.isEmpty)
}
func testTitleCaseNamesAreNotAutocorrected() {
let engine = EnglishSuggestionEngine()
engine.prepare()
XCTAssertNil(engine.correctionDecision(for: "Rocky", personalTerms: [], learnedBoosts: [:]))
XCTAssertNil(engine.correctionDecision(for: "Wang", personalTerms: [], learnedBoosts: [:]))
XCTAssertNil(engine.correctionDecision(for: "Chen", personalTerms: [], learnedBoosts: [:]))
XCTAssertNil(engine.correctionDecision(for: "Li", personalTerms: [], learnedBoosts: [:]))
}
func testTitleCaseTranspositionStillCorrects() {
let engine = EnglishSuggestionEngine()
engine.prepare()
XCTAssertEqual(
engine.correctionDecision(for: "Teh", personalTerms: [], learnedBoosts: [:])?.replacement,
"The"
)
}
func testProximityCorrectsAdjacentKeyTypos() {
let engine = EnglishSuggestionEngine()
engine.prepare()
XCTAssertEqual(
engine.correctionDecision(for: "gppd", personalTerms: [], learnedBoosts: [:])?.replacement,
"good"
)
}
func testRealWordFormIsNotCorrectedToFrom() {
let engine = EnglishSuggestionEngine()
engine.prepare()
XCTAssertNil(engine.correctionDecision(for: "form", personalTerms: [], learnedBoosts: [:]))
}
func testSupplementaryLexiconBlocksAutocorrect() {
let engine = EnglishSuggestionEngine()
engine.prepare()
XCTAssertNil(
engine.correctionDecision(
for: "teh",
personalTerms: [],
learnedBoosts: [:],
systemWords: ["teh"]
)
)
}
func testLearnedDefenseBlocksAutocorrect() {
let engine = EnglishSuggestionEngine()
engine.prepare()
XCTAssertNil(
engine.correctionDecision(
for: "teh",
personalTerms: [],
learnedBoosts: ["teh": 5]
)
)
}
@MainActor
func testAutocorrectDoesNotBoostReplacement() {
let suite = "english.learning.polarity.test"
let defaults = UserDefaults(suiteName: suite)!
defaults.removePersistentDomain(forName: suite)
let store = EnglishLearningStore(defaults: defaults)
let typing = TypingSessionController(learningStore: store)
typing.suggestionsEnabled = true
_ = typing.setLanguage(.english)
typing.enterTypingMode()
for character in ["t", "e", "h"] {
_ = typing.handleKey(character)
}
_ = typing.handleSpace()
XCTAssertEqual(store.boost(for: "the"), 0)
XCTAssertEqual(store.boost(for: "teh"), 0)
}
@MainActor
func testRejectingAutocorrectLearnsOriginal() {
let suite = "english.learning.defense.test"
let defaults = UserDefaults(suiteName: suite)!
defaults.removePersistentDomain(forName: suite)
let store = EnglishLearningStore(defaults: defaults)
let typing = TypingSessionController(learningStore: store)
typing.suggestionsEnabled = true
_ = typing.setLanguage(.english)
typing.enterTypingMode()
for character in ["t", "e", "h"] {
_ = typing.handleKey(character)
}
let spaced = typing.handleSpace()
if spaced.deleteCount > 0 {
_ = typing.handleKey("")
XCTAssertGreaterThanOrEqual(store.boost(for: "teh"), 5)
}
}
func testQWERTYNeighborsIncludeDiagonals() {
let aroundG = EnglishQWERTYProximity.neighbors(of: "g", includingSelf: true)
XCTAssertTrue(aroundG.contains("t"))
XCTAssertTrue(aroundG.contains("f"))
XCTAssertTrue(aroundG.contains("h"))
XCTAssertTrue(aroundG.contains("b"))
XCTAssertFalse(aroundG.contains("q"))
}
func testBigramsPredictNextWords() {
let lexicon = EnglishLexicon()
lexicon.prepare()
let next = lexicon.nextWords(after: "thank", limit: 4)
XCTAssertTrue(next.contains("you"))
}
@MainActor
private func makeIsolatedEnglishSession(suite: String) -> TypingSessionController {
let defaults = UserDefaults(suiteName: suite)!
defaults.removePersistentDomain(forName: suite)
let store = EnglishLearningStore(defaults: defaults)
let typing = TypingSessionController(learningStore: store)
typing.suggestionsEnabled = true
_ = typing.setLanguage(.english)
typing.enterTypingMode()
return typing
}
@MainActor
private func apply(_ typing: TypingSessionController, _ output: TypingOutput) {
typing.syncAutocapitalization(
+14 -8
View File
@@ -14,7 +14,8 @@ public struct AIUserSkill: Codable, Equatable, Identifiable, Sendable {
public var summary: String
public var systemImage: String
public var prompt: String
public var shortcutICloudURL: URL
/// Optional iCloud share URL. Nil means the skill only transforms text.
public var shortcutICloudURL: URL?
/// Name used by `shortcuts://run-shortcut?name=`. Independent of `name`.
public var shortcutName: String
/// Per-skill reasoning. Built-in skills are always off; custom defaults off.
@@ -28,8 +29,8 @@ public struct AIUserSkill: Codable, Equatable, Identifiable, Sendable {
summary: String = "",
systemImage: String = AIUserSkillLimits.defaultSystemImage,
prompt: String,
shortcutICloudURL: URL,
shortcutName: String,
shortcutICloudURL: URL? = nil,
shortcutName: String = "",
thinkingEnabled: Bool = false,
createdAt: Date = Date(),
updatedAt: Date? = nil
@@ -49,15 +50,16 @@ public struct AIUserSkill: Codable, Equatable, Identifiable, Sendable {
public var isUserCreated: Bool { id.hasPrefix("user.") }
public func asClipboardSkill() -> AIClipboardSkill {
AIClipboardSkill(
let exportsToShortcut = shortcutICloudURL != nil
return AIClipboardSkill(
id: id,
systemImage: systemImage,
titleKey: "",
cardTitleKey: "",
descriptionKey: "",
kind: .export,
kind: exportsToShortcut ? .export : .transform,
isDefault: false,
shortcutName: shortcutName,
shortcutName: exportsToShortcut ? shortcutName : nil,
shortcutICloudURL: shortcutICloudURL,
customName: name,
customSummary: summary,
@@ -149,10 +151,14 @@ public struct AIUserSkillCatalog: Codable, Equatable, Sendable {
maximum: AIUserSkillLimits.maximumPromptCharacters
)
}
guard !shortcutName.isEmpty else { throw AIUserSkillValidationError.emptyShortcutName }
guard AIShortcutShareLink.isValid(skill.shortcutICloudURL) else {
if let shortcutURL = skill.shortcutICloudURL {
guard !shortcutName.isEmpty else {
throw AIUserSkillValidationError.emptyShortcutName
}
guard AIShortcutShareLink.isValid(shortcutURL) else {
throw AIUserSkillValidationError.invalidShortcutLink
}
}
guard !icon.isEmpty else { throw AIUserSkillValidationError.emptyIcon }
var saved = skill
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
@@ -130,13 +130,26 @@ public final class AIAgentSkillLayoutStore: ObservableObject {
}
public func saveUserSkill(_ skill: AIUserSkill) throws {
let previousURL = userCatalog.skill(id: skill.id)?.shortcutICloudURL
let previousSkill = userCatalog.skill(id: skill.id)
let previousURL = previousSkill?.shortcutICloudURL
let previousLayout = layout.sanitized(catalog: mergedCatalog)
var catalog = userCatalog
try catalog.upsert(skill)
commitUserCatalog(catalog)
if previousURL != nil, previousURL != skill.shortcutICloudURL {
dropShortcutConfirmation(for: skill.id)
guard previousSkill != nil, previousURL != skill.shortcutICloudURL else {
return
}
let keepsKeyboardSlot = skill.shortcutICloudURL == nil
commitLayout(
AIAgentSkillLayout(
enabledIDs: keepsKeyboardSlot
? previousLayout.enabledIDs
: previousLayout.enabledIDs.filter { $0 != skill.id },
confirmedShortcutIDs: previousLayout.confirmedShortcutIDs.filter {
$0 != skill.id
}
)
)
}
public func deleteUserSkill(id: String) {
@@ -152,16 +165,6 @@ public final class AIAgentSkillLayoutStore: ObservableObject {
)
}
private func dropShortcutConfirmation(for id: String) {
let current = layout.sanitized(catalog: mergedCatalog)
commitLayout(
AIAgentSkillLayout(
enabledIDs: current.enabledIDs.filter { $0 != id },
confirmedShortcutIDs: current.confirmedShortcutIDs.filter { $0 != id }
)
)
}
private func commitLayout(_ layout: AIAgentSkillLayout) {
persistLayout(layout)
self.layout = loadLayout()
+418 -139
View File
@@ -2,180 +2,359 @@
// OSGKeyboard · Shared
//
// Offline English word list + bigrams for the typing extension.
// Loaded once, kept compact for the keyboard RSS budget.
// The 40k-word table is a mmap'd binary (`english_lexicon.bin`); dirty heap
// stays near zero until a lookup materializes a handful of result strings.
// TSV files in the repo are the build input, not the runtime format.
import Foundation
public struct EnglishScoredCorrection: Equatable, Sendable {
public var word: String
public var spatialCost: Int
public var frequency: Int
public var isTransposition: Bool
public var isShortening: Bool
public init(
word: String,
spatialCost: Int,
frequency: Int,
isTransposition: Bool,
isShortening: Bool
) {
self.word = word
self.spatialCost = spatialCost
self.frequency = frequency
self.isTransposition = isTransposition
self.isShortening = isShortening
}
}
/// Ranked English lexicon used by autocomplete / autocorrect / next-word.
public final class EnglishLexicon: @unchecked Sendable {
public static let shared = EnglishLexicon()
/// Lowercased word relative frequency (higher is more common).
private var frequencies: [String: Int] = [:]
/// Sorted lowercased words for prefix binary search.
private var sortedWords: [String] = []
/// previous(lower) next-word candidates (lower).
private var bigrams: [String: [String]] = [:]
private var mapped: Data?
private var header: FileHeader?
private var loaded = false
private let lock = NSLock()
public init() {}
/// True after a successful mmap. Tests use this to prove Chinese typing
/// does not pull the English table into the extension.
public var isLoaded: Bool {
lock.lock()
defer { lock.unlock() }
return loaded
}
public func prepare() {
lock.lock()
defer { lock.unlock() }
guard !loaded else { return }
loadLexicon()
loadBigrams()
loaded = true
loadMappedLexicon()
}
/// Release in-memory tables when leaving the typing surface (jetsam recovery).
/// Release the mapped file when leaving English / the typing surface.
public func unload() {
lock.lock()
defer { lock.unlock() }
frequencies.removeAll(keepingCapacity: false)
sortedWords.removeAll(keepingCapacity: false)
bigrams.removeAll(keepingCapacity: false)
mapped = nil
header = nil
loaded = false
}
public var wordCount: Int {
prepareIfNeeded()
return sortedWords.count
lock.lock()
defer { lock.unlock() }
return header?.unigramCount ?? 0
}
public func frequency(of word: String) -> Int {
prepareIfNeeded()
return frequencies[word.lowercased()] ?? 0
withMap { buf, header in
guard let index = lookupIndex(asciiLowered(word), header: header, buf: buf) else {
return 0
}
return frequency(at: index, header: header, buf: buf)
} ?? 0
}
public func contains(_ word: String) -> Bool {
prepareIfNeeded()
return frequencies[word.lowercased()] != nil
withMap { buf, header in
lookupIndex(asciiLowered(word), header: header, buf: buf) != nil
} ?? false
}
/// Highest-frequency unigrams, for next-word fallback when no bigram hits.
public func topWords(limit: Int = 6) -> [String] {
guard limit > 0 else { return [] }
return withMap { buf, header in
let count = min(limit, header.unigramCount)
var words: [String] = []
words.reserveCapacity(count)
for rank in 0..<count {
let index = Int(
readU16(buf, header.freqRankOffset + rank * 2)
)
guard index < header.unigramCount else { continue }
if let word = string(at: index, header: header, buf: buf) {
words.append(word)
}
}
return words
} ?? []
}
/// Prefix completions, highest frequency first.
public func completions(prefix: String, limit: Int = 8) -> [String] {
prepareIfNeeded()
let needle = prefix.lowercased()
let needle = asciiLowered(prefix)
guard !needle.isEmpty, limit > 0 else { return [] }
var results: [(String, Int)] = []
var index = lowerBound(needle)
while index < sortedWords.count {
let word = sortedWords[index]
guard word.hasPrefix(needle) else { break }
if word != needle {
results.append((word, frequencies[word] ?? 0))
return withMap { buf, header in
var scored: [(Int, Int)] = []
var index = lowerBound(needle, header: header, buf: buf)
while index < header.unigramCount {
guard let bytes = wordBytes(at: index, header: header, buf: buf) else { break }
guard hasPrefix(bytes, needle) else { break }
if !bytesEqual(bytes, needle) {
scored.append((index, frequency(at: index, header: header, buf: buf)))
}
index += 1
// Soft cap scan to keep keystroke path cheap.
if results.count >= limit * 8 { break }
if scored.count >= limit * 8 { break }
}
results.sort { lhs, rhs in
scored.sort { lhs, rhs in
if lhs.1 != rhs.1 { return lhs.1 > rhs.1 }
return lhs.0 < rhs.0
}
return Array(results.prefix(limit).map(\.0))
return scored.prefix(limit).compactMap { pair in
string(at: pair.0, header: header, buf: buf)
}
} ?? []
}
/// Best edit-distance 2 correction, or nil when the typed word is fine.
/// Uses DamerauLevenshtein so adjacent swaps (teh the) count as 1.
/// Scans only same-initial-letter candidates (not the full frequency table).
/// Nearby words scored by QWERTY proximity + frequency. Does not decide
/// whether autocorrect should fire the suggestion engine does.
public func scoredCorrections(for typed: String, limit: Int = 6) -> [EnglishScoredCorrection] {
let needle = asciiLowered(typed)
guard needle.count >= 3, let firstByte = needle.first, limit > 0 else { return [] }
let first = Character(UnicodeScalar(firstByte))
var initials = Set(EnglishQWERTYProximity.neighbors(of: first, includingSelf: true))
initials.insert(first)
return withMap { buf, header in
var best: [ScoredIndex] = []
best.reserveCapacity(limit)
for initial in initials {
guard let letter = initial.asciiLetterIndex else { continue }
let rangeOffset = header.initialOffset + letter * 4
let start = Int(readU16(buf, rangeOffset))
let count = Int(readU16(buf, rangeOffset + 2))
guard start >= 0, count >= 0, start + count <= header.unigramCount else { continue }
for index in start..<(start + count) {
guard let bytes = wordBytes(at: index, header: header, buf: buf) else { continue }
let delta = abs(bytes.count - needle.count)
guard delta <= 2, !bytesEqual(bytes, needle) else { continue }
guard let alignment = EnglishQWERTYProximity.align(
typedASCII: needle,
candidateASCII: bytes
) else { continue }
guard alignment.cost > 0 else { continue }
insertBest(
ScoredIndex(
index: index,
spatialCost: alignment.cost,
frequency: frequency(at: index, header: header, buf: buf),
isTransposition: alignment.isTransposition,
isShortening: alignment.isShortening
),
into: &best,
limit: limit
)
}
}
return best.compactMap { scored in
guard let word = string(at: scored.index, header: header, buf: buf) else {
return nil
}
return EnglishScoredCorrection(
word: word,
spatialCost: scored.spatialCost,
frequency: scored.frequency,
isTransposition: scored.isTransposition,
isShortening: scored.isShortening
)
}
} ?? []
}
/// Best proximity correction, or nil when the typed word is already known.
public func bestCorrection(for typed: String) -> String? {
prepareIfNeeded()
let needle = typed.lowercased()
guard needle.count >= 2, let first = needle.first else { return nil }
if frequencies[needle] != nil { return nil }
var best: (word: String, distance: Int, freq: Int)?
var index = lowerBound(String(first))
while index < sortedWords.count {
let word = sortedWords[index]
guard word.first == first else { break }
defer { index += 1 }
guard abs(word.count - needle.count) <= 2 else { continue }
let freq = frequencies[word] ?? 0
let distance = damerauLevenshtein(needle, word, max: 2)
guard distance > 0, distance <= 2 else { continue }
if let current = best {
if distance < current.distance
|| (distance == current.distance && freq > current.freq) {
best = (word, distance, freq)
}
} else {
best = (word, distance, freq)
}
}
guard let best else { return nil }
// Distance-2 corrections need a common word so rare near-misses don't win.
if best.distance == 2, best.freq < 200 { return nil }
return best.word
if contains(typed) { return nil }
return scoredCorrections(for: typed, limit: 1).first?.word
}
public func nextWords(after previous: String, limit: Int = 6) -> [String] {
prepareIfNeeded()
let key = previous.lowercased()
guard let list = bigrams[key] else { return [] }
return Array(list.prefix(limit))
guard limit > 0 else { return [] }
let needle = asciiLowered(previous)
return withMap { buf, header in
guard let prevIndex = lookupIndex(needle, header: header, buf: buf) else {
return []
}
guard let group = lookupBigramGroup(prevIndex: prevIndex, header: header, buf: buf) else {
return []
}
let count = min(limit, group.nextCount)
var words: [String] = []
words.reserveCapacity(count)
for offset in 0..<count {
let index = Int(readU16(buf, header.bigramNextOffset + (group.firstNext + offset) * 2))
if let word = string(at: index, header: header, buf: buf) {
words.append(word)
}
}
return words
} ?? []
}
// MARK: - Private
// MARK: - Mapped file
private func prepareIfNeeded() {
if !loaded { prepare() }
private struct FileHeader {
var unigramCount: Int
var bigramGroupCount: Int
var stringPoolOffset: Int
var stringPoolSize: Int
var unigramOffset: Int
var freqRankOffset: Int
var initialOffset: Int
var bigramIndexOffset: Int
var bigramNextOffset: Int
var fileSize: Int
static let magic = "OSGENG01"
static let version = 1
static let headerSize = 64
static let initialCount = 26
static func parse(_ data: Data) -> FileHeader? {
guard data.count >= headerSize else { return nil }
return data.withUnsafeBytes { buf -> FileHeader? in
let magicBytes = UnsafeRawBufferPointer(rebasing: buf[0..<8])
let magic = String(bytes: magicBytes, encoding: .ascii)
guard magic == Self.magic else { return nil }
guard Int(readU32(buf, 8)) == version else { return nil }
let unigramCount = Int(readU32(buf, 12))
let bigramGroupCount = Int(readU32(buf, 16))
let stringPoolOffset = Int(readU32(buf, 20))
let stringPoolSize = Int(readU32(buf, 24))
let unigramOffset = Int(readU32(buf, 28))
let freqRankOffset = Int(readU32(buf, 32))
let initialOffset = Int(readU32(buf, 36))
let bigramIndexOffset = Int(readU32(buf, 40))
let bigramNextOffset = Int(readU32(buf, 44))
let fileSize = data.count
guard unigramCount >= 0, unigramCount <= 200_000 else { return nil }
guard bigramGroupCount >= 0, bigramGroupCount <= 100_000 else { return nil }
guard region(unigramOffset, unigramCount * 8, in: fileSize),
region(freqRankOffset, unigramCount * 2, in: fileSize),
region(initialOffset, initialCount * 4, in: fileSize),
region(bigramIndexOffset, bigramGroupCount * 8, in: fileSize),
region(stringPoolOffset, stringPoolSize, in: fileSize)
else {
return nil
}
private func loadLexicon() {
return FileHeader(
unigramCount: unigramCount,
bigramGroupCount: bigramGroupCount,
stringPoolOffset: stringPoolOffset,
stringPoolSize: stringPoolSize,
unigramOffset: unigramOffset,
freqRankOffset: freqRankOffset,
initialOffset: initialOffset,
bigramIndexOffset: bigramIndexOffset,
bigramNextOffset: bigramNextOffset,
fileSize: fileSize
)
}
}
private static func region(_ offset: Int, _ size: Int, in fileSize: Int) -> Bool {
offset >= 0 && size >= 0 && offset <= fileSize && size <= fileSize - offset
}
}
private struct ScoredIndex {
var index: Int
var spatialCost: Int
var frequency: Int
var isTransposition: Bool
var isShortening: Bool
}
private struct BigramGroup {
var nextCount: Int
var firstNext: Int
}
private func loadMappedLexicon() {
guard let url = Bundle(for: EnglishLexicon.self)
.url(forResource: "english_lexicon", withExtension: "tsv", subdirectory: nil)
?? Bundle(for: EnglishLexicon.self)
.url(forResource: "english_lexicon", withExtension: "tsv")
?? Bundle.main.url(forResource: "english_lexicon", withExtension: "tsv")
.url(forResource: "english_lexicon", withExtension: "bin")
?? Bundle.main.url(forResource: "english_lexicon", withExtension: "bin")
else {
return
}
guard let data = try? String(contentsOf: url, encoding: .utf8) else { return }
var map: [String: Int] = [:]
for line in data.split(whereSeparator: \.isNewline) {
let parts = line.split(separator: "\t", maxSplits: 1)
guard parts.count == 2,
let freq = Int(parts[1]) else { continue }
let word = String(parts[0]).lowercased()
guard !word.isEmpty else { continue }
map[word] = freq
}
frequencies = map
sortedWords = map.keys.sorted()
}
private func loadBigrams() {
guard let url = Bundle(for: EnglishLexicon.self)
.url(forResource: "english_bigrams", withExtension: "tsv")
?? Bundle.main.url(forResource: "english_bigrams", withExtension: "tsv")
// `.mappedIfSafe` keeps the 40k table on file-backed pages. Jetsam
// charges dirty heap, not these clean mapped pages.
guard let data = try? Data(contentsOf: url, options: [.mappedIfSafe]),
let parsed = FileHeader.parse(data)
else {
return
}
guard let data = try? String(contentsOf: url, encoding: .utf8) else { return }
var map: [String: [String]] = [:]
for line in data.split(whereSeparator: \.isNewline) {
let parts = line.split(separator: "\t", maxSplits: 1)
guard parts.count == 2 else { continue }
let prev = String(parts[0]).lowercased()
let nexts = parts[1].split(whereSeparator: \.isWhitespace).map { String($0).lowercased() }
guard !prev.isEmpty, !nexts.isEmpty else { continue }
map[prev] = nexts
}
bigrams = map
mapped = data
header = parsed
loaded = true
}
private func lowerBound(_ prefix: String) -> Int {
private func withMap<T>(_ body: (UnsafeRawBufferPointer, FileHeader) -> T) -> T? {
lock.lock()
defer { lock.unlock() }
guard loaded, let data = mapped, let header else { return nil }
return data.withUnsafeBytes { buf in
body(buf, header)
}
}
private func lookupIndex(
_ needle: [UInt8],
header: FileHeader,
buf: UnsafeRawBufferPointer
) -> Int? {
let index = lowerBound(needle, header: header, buf: buf)
guard index < header.unigramCount,
let bytes = wordBytes(at: index, header: header, buf: buf),
bytesEqual(bytes, needle)
else {
return nil
}
return index
}
private func lowerBound(
_ needle: [UInt8],
header: FileHeader,
buf: UnsafeRawBufferPointer
) -> Int {
var low = 0
var high = sortedWords.count
var high = header.unigramCount
while low < high {
let mid = (low + high) / 2
if sortedWords[mid] < prefix {
guard let bytes = wordBytes(at: mid, header: header, buf: buf) else {
high = mid
continue
}
if compare(bytes, needle) < 0 {
low = mid + 1
} else {
high = mid
@@ -184,40 +363,140 @@ public final class EnglishLexicon: @unchecked Sendable {
return low
}
/// DamerauLevenshtein with early exit when distance would exceed `max`.
private func damerauLevenshtein(_ a: String, _ b: String, max: Int) -> Int {
let aChars = Array(a)
let bChars = Array(b)
let aCount = aChars.count
let bCount = bChars.count
if abs(aCount - bCount) > max { return max + 1 }
var prevPrev = [Int](repeating: 0, count: bCount + 1)
var prev = Array(0...bCount)
for i in 1...aCount {
var current = [Int](repeating: 0, count: bCount + 1)
current[0] = i
var rowMin = current[0]
for j in 1...bCount {
let cost = aChars[i - 1] == bChars[j - 1] ? 0 : 1
var value = min(
prev[j] + 1,
current[j - 1] + 1,
prev[j - 1] + cost
private func lookupBigramGroup(
prevIndex: Int,
header: FileHeader,
buf: UnsafeRawBufferPointer
) -> BigramGroup? {
var low = 0
var high = header.bigramGroupCount
while low < high {
let mid = (low + high) / 2
let midPrev = Int(readU16(buf, header.bigramIndexOffset + mid * 8))
if midPrev < prevIndex {
low = mid + 1
} else {
high = mid
}
}
guard low < header.bigramGroupCount else { return nil }
let offset = header.bigramIndexOffset + low * 8
guard Int(readU16(buf, offset)) == prevIndex else { return nil }
return BigramGroup(
nextCount: Int(readU16(buf, offset + 2)),
firstNext: Int(readU32(buf, offset + 4))
)
// Adjacent transposition
if i > 1, j > 1,
aChars[i - 1] == bChars[j - 2],
aChars[i - 2] == bChars[j - 1] {
value = min(value, prevPrev[j - 2] + 1)
}
current[j] = value
rowMin = min(rowMin, value)
private func frequency(at index: Int, header: FileHeader, buf: UnsafeRawBufferPointer) -> Int {
Int(readU16(buf, header.unigramOffset + index * 8 + 6))
}
if rowMin > max { return max + 1 }
prevPrev = prev
prev = current
private func wordBytes(
at index: Int,
header: FileHeader,
buf: UnsafeRawBufferPointer
) -> UnsafeBufferPointer<UInt8>? {
guard index >= 0, index < header.unigramCount else { return nil }
let record = header.unigramOffset + index * 8
let poolOff = Int(readU32(buf, record))
let length = Int(buf[record + 4])
let start = header.stringPoolOffset + poolOff
guard length >= 0,
start >= header.stringPoolOffset,
start + length <= header.stringPoolOffset + header.stringPoolSize,
start + length <= header.fileSize,
let base = buf.baseAddress
else {
return nil
}
return prev[bCount]
return UnsafeBufferPointer(
start: base.advanced(by: start).assumingMemoryBound(to: UInt8.self),
count: length
)
}
private func string(
at index: Int,
header: FileHeader,
buf: UnsafeRawBufferPointer
) -> String? {
guard let bytes = wordBytes(at: index, header: header, buf: buf) else { return nil }
return String(bytes: bytes, encoding: .ascii)
}
private func insertBest(_ scored: ScoredIndex, into best: inout [ScoredIndex], limit: Int) {
if let existing = best.firstIndex(where: { $0.index == scored.index }) {
if isOrderedBefore(scored, best[existing]) {
best[existing] = scored
best.sort(by: isOrderedBefore)
}
return
}
if best.count < limit {
best.append(scored)
best.sort(by: isOrderedBefore)
return
}
if let last = best.last, isOrderedBefore(scored, last) {
best[best.count - 1] = scored
best.sort(by: isOrderedBefore)
}
}
private func isOrderedBefore(_ lhs: ScoredIndex, _ rhs: ScoredIndex) -> Bool {
if lhs.spatialCost != rhs.spatialCost { return lhs.spatialCost < rhs.spatialCost }
if lhs.frequency != rhs.frequency { return lhs.frequency > rhs.frequency }
return lhs.index < rhs.index
}
}
private func readU16(_ buf: UnsafeRawBufferPointer, _ offset: Int) -> UInt16 {
UInt16(littleEndian: buf.loadUnaligned(fromByteOffset: offset, as: UInt16.self))
}
private func readU32(_ buf: UnsafeRawBufferPointer, _ offset: Int) -> UInt32 {
UInt32(littleEndian: buf.loadUnaligned(fromByteOffset: offset, as: UInt32.self))
}
private func asciiLowered(_ string: String) -> [UInt8] {
string.utf8.map { byte in
(byte >= 65 && byte <= 90) ? byte + 32 : byte
}
}
private func compare(_ word: UnsafeBufferPointer<UInt8>, _ needle: [UInt8]) -> Int {
let count = min(word.count, needle.count)
for index in 0..<count {
let left = word[index]
let right = needle[index]
if left < right { return -1 }
if left > right { return 1 }
}
if word.count < needle.count { return -1 }
if word.count > needle.count { return 1 }
return 0
}
private func hasPrefix(_ word: UnsafeBufferPointer<UInt8>, _ prefix: [UInt8]) -> Bool {
guard word.count >= prefix.count else { return false }
for index in prefix.indices where word[index] != prefix[index] {
return false
}
return true
}
private func bytesEqual(_ word: UnsafeBufferPointer<UInt8>, _ needle: [UInt8]) -> Bool {
guard word.count == needle.count else { return false }
for index in needle.indices where word[index] != needle[index] {
return false
}
return true
}
private extension Character {
var asciiLetterIndex: Int? {
guard let value = utf8.first, value >= 97, value <= 122 else { return nil }
return Int(value - 97)
}
}
@@ -0,0 +1,210 @@
// EnglishQWERTYProximity.swift
// OSGKeyboard · Shared
//
// Spatial cost for English autocorrect. Adjacent (including diagonal) keys
// are cheap; far substitutions are expensive. Inspired by AOSP LatinIME's
// proximity weighting formula only, no Android code.
import Foundation
public struct EnglishAlignment: Equatable, Sendable {
/// Weighted edit cost. `0` means identical.
public var cost: Int
public var isTransposition: Bool
public var isShortening: Bool
}
public enum EnglishQWERTYProximity: Sendable {
/// Two adjacent substitutions, or one farther miss, still eligible.
public static let maxAutocorrectCost = 34
public static let adjacentCost = 10
public static let nearCost = 22
public static let farCost = 34
public static let insDelCost = 18
public static let transpositionCost = 10
/// US QWERTY, staggered rows matching the on-screen letter grid.
private static let coordinates: [Character: (x: Double, y: Double)] = {
let rows: [[Character]] = [
Array("qwertyuiop"),
Array("asdfghjkl"),
Array("zxcvbnm")
]
let offsets: [Double] = [0, 0.5, 1.5]
var map: [Character: (x: Double, y: Double)] = [:]
for (rowIndex, row) in rows.enumerated() {
let origin = offsets[rowIndex]
for (column, letter) in row.enumerated() {
map[letter] = (origin + Double(column), Double(rowIndex))
}
}
return map
}()
public static func neighbors(of letter: Character, includingSelf: Bool) -> [Character] {
let needle = Character(letter.lowercased())
guard let origin = coordinates[needle] else {
return includingSelf ? [needle] : []
}
var hits: [Character] = []
for (candidate, point) in coordinates {
let distance = chebyshev(origin, point)
if distance == 0 {
if includingSelf { hits.append(candidate) }
} else if distance <= 1.01 {
hits.append(candidate)
}
}
return hits
}
public static func keyDistance(_ a: Character, _ b: Character) -> Int {
let left = Character(a.lowercased())
let right = Character(b.lowercased())
if left == right { return 0 }
guard let origin = coordinates[left], let other = coordinates[right] else {
return farCost
}
let distance = chebyshev(origin, other)
if distance <= 1.01 { return adjacentCost }
if distance <= 2.01 { return nearCost }
return farCost
}
public static func align(typed: String, candidate: String) -> EnglishAlignment? {
let source = asciiLowered(typed)
let targetBytes = asciiLowered(candidate)
return targetBytes.withUnsafeBufferPointer { pointer in
align(typedASCII: source, candidateASCII: pointer)
}
}
/// Same cost model as `align(typed:candidate:)`, but the candidate stays in
/// a mapped file no Swift `String` per scanned word.
public static func align(
typedASCII: [UInt8],
candidateASCII: UnsafeBufferPointer<UInt8>
) -> EnglishAlignment? {
let source = typedASCII
let target = candidateASCII
let delta = abs(source.count - target.count)
guard delta <= 2 else { return nil }
if delta == 0, bytesEqual(source, target) {
return EnglishAlignment(cost: 0, isTransposition: false, isShortening: false)
}
if source.count == target.count, isAdjacentTransposition(source, target) {
return EnglishAlignment(
cost: transpositionCost,
isTransposition: true,
isShortening: false
)
}
if source.count == target.count {
var cost = 0
for index in source.indices {
cost += keyDistance(source[index], target[index])
if cost > maxAutocorrectCost { return nil }
}
return EnglishAlignment(
cost: cost,
isTransposition: false,
isShortening: false
)
}
let cost = bandedEditCost(source, target)
guard cost <= maxAutocorrectCost else { return nil }
return EnglishAlignment(
cost: cost,
isTransposition: false,
isShortening: target.count < source.count
)
}
private static func keyDistance(_ a: UInt8, _ b: UInt8) -> Int {
if a == b { return 0 }
guard a >= 97, a <= 122, b >= 97, b <= 122 else { return farCost }
return keyDistance(Character(UnicodeScalar(a)), Character(UnicodeScalar(b)))
}
private static func isAdjacentTransposition(
_ source: [UInt8],
_ target: UnsafeBufferPointer<UInt8>
) -> Bool {
guard source.count == target.count, source.count >= 2 else { return false }
var mismatch = -1
for index in source.indices where source[index] != target[index] {
if mismatch == -1 {
mismatch = index
} else if index == mismatch + 1,
source[mismatch] == target[index],
source[index] == target[mismatch] {
for rest in (index + 1)..<source.count where source[rest] != target[rest] {
return false
}
return true
} else {
return false
}
}
return false
}
/// Banded Levenshtein with proximity substitutions and a Damerau swap.
private static func bandedEditCost(
_ source: [UInt8],
_ target: UnsafeBufferPointer<UInt8>
) -> Int {
let aCount = source.count
let bCount = target.count
var previous = Array(0...bCount).map { $0 * insDelCost }
var older = previous
for i in 1...aCount {
var current = [Int](repeating: 0, count: bCount + 1)
current[0] = i * insDelCost
var rowMin = current[0]
for j in 1...bCount {
let substitution = previous[j - 1] + keyDistance(source[i - 1], target[j - 1])
var value = min(
previous[j] + insDelCost,
current[j - 1] + insDelCost,
substitution
)
if i > 1, j > 1,
source[i - 1] == target[j - 2],
source[i - 2] == target[j - 1] {
value = min(value, older[j - 2] + transpositionCost)
}
current[j] = value
rowMin = min(rowMin, value)
}
if rowMin > maxAutocorrectCost { return maxAutocorrectCost + 1 }
older = previous
previous = current
}
return previous[bCount]
}
private static func asciiLowered(_ string: String) -> [UInt8] {
string.utf8.map { byte in
(byte >= 65 && byte <= 90) ? byte + 32 : byte
}
}
private static func bytesEqual(_ source: [UInt8], _ target: UnsafeBufferPointer<UInt8>) -> Bool {
guard source.count == target.count else { return false }
for index in source.indices where source[index] != target[index] {
return false
}
return true
}
private static func chebyshev(
_ a: (x: Double, y: Double),
_ b: (x: Double, y: Double)
) -> Double {
max(abs(a.x - b.x), abs(a.y - b.y))
}
}
@@ -1,8 +1,8 @@
// EnglishSuggestionEngine.swift
// OSGKeyboard · Shared
//
// Builds TypingComposition for English: completions while composing,
// high-confidence corrections on commit, next-word predictions after.
// Builds a 3-slot English QuickType board: verbatim / correction / completion
// (or next-word after commit). Space applies only the correction slot.
import Foundation
@@ -12,19 +12,29 @@ public struct EnglishSuggestionContext: Sendable {
public var personalTerms: [String]
public var learnedBoosts: [String: Int]
public var includeOriginalAfterCorrection: String?
/// Contacts / text replacements from `UILexicon`.
public var systemWords: [String]
public var systemCompletions: [String]
public var systemGuesses: [String]
public init(
currentWord: String = "",
previousWord: String = "",
personalTerms: [String] = [],
learnedBoosts: [String: Int] = [:],
includeOriginalAfterCorrection: String? = nil
includeOriginalAfterCorrection: String? = nil,
systemWords: [String] = [],
systemCompletions: [String] = [],
systemGuesses: [String] = []
) {
self.currentWord = currentWord
self.previousWord = previousWord
self.personalTerms = personalTerms
self.learnedBoosts = learnedBoosts
self.includeOriginalAfterCorrection = includeOriginalAfterCorrection
self.systemWords = systemWords
self.systemCompletions = systemCompletions
self.systemGuesses = systemGuesses
}
}
@@ -47,6 +57,10 @@ public struct EnglishCorrectionDecision: Equatable, Sendable {
/// Pure ranking / candidate builder no UITextDocumentProxy access.
public struct EnglishSuggestionEngine: Sendable {
public static let slotCount = 3
/// In-vocabulary words only yield to a much more common transposition / neighbor.
public static let inVocabularyFrequencyGap = 250
private let lexicon: EnglishLexicon
public init(lexicon: EnglishLexicon = .shared) {
@@ -57,112 +71,238 @@ public struct EnglishSuggestionEngine: Sendable {
lexicon.prepare()
}
/// Suggestions while the user is mid-word.
/// Suggestions only while the user is actively typing an English word.
public func compositionWhileTyping(_ context: EnglishSuggestionContext) -> TypingComposition {
let prefix = context.currentWord
guard !prefix.isEmpty else {
return nextWordComposition(context)
}
var ranked: [(text: String, score: Int, id: String)] = []
var seen = Set<String>()
func append(_ raw: String, baseScore: Int, tag: String, preserveCase: Bool = false) {
let display = preserveCase ? raw : matchCase(of: prefix, to: raw)
let key = display.lowercased()
guard seen.insert(key).inserted else { return }
let boost = context.learnedBoosts[key] ?? 0
let personalBoost = context.personalTerms.contains { $0.lowercased() == key } ? 5_000 : 0
ranked.append((display, baseScore + boost + personalBoost, "\(tag)|\(key)"))
}
for term in context.personalTerms where term.lowercased().hasPrefix(prefix.lowercased())
&& term.lowercased() != prefix.lowercased() {
append(term, baseScore: 8_000 + term.count, tag: "personal", preserveCase: true)
}
for word in lexicon.completions(prefix: prefix, limit: 12) {
append(word, baseScore: lexicon.frequency(of: word), tag: "complete")
}
ranked.sort { lhs, rhs in
if lhs.score != rhs.score { return lhs.score > rhs.score }
return lhs.text.count < rhs.text.count
}
let candidates = ranked.prefix(8).map {
TypingCandidate(id: $0.id, text: $0.text, engineIndex: 0)
}
return TypingComposition(preedit: prefix, candidates: Array(candidates))
guard !prefix.isEmpty else { return .empty }
return makeBoard(context).composition
}
/// Decide whether to autocorrect on space / punctuation.
public func correctionDecision(
for typed: String,
personalTerms: [String],
learnedBoosts: [String: Int]
learnedBoosts: [String: Int],
previousWord: String = "",
systemWords: [String] = [],
systemGuesses: [String] = []
) -> EnglishCorrectionDecision? {
let trimmed = typed
guard trimmed.count >= 2 else { return nil }
let lower = trimmed.lowercased()
if personalTerms.contains(where: { $0.lowercased() == lower }) { return nil }
if (learnedBoosts[lower] ?? 0) >= 5 { return nil }
if shouldSkipAutocorrect(trimmed) { return nil }
if lexicon.contains(lower) { return nil }
guard let correction = lexicon.bestCorrection(for: lower) else { return nil }
// Personal dictionary wins over lexicon corrections.
if personalTerms.contains(where: { $0.lowercased() == correction }) {
return EnglishCorrectionDecision(original: trimmed, replacement: matchCase(of: trimmed, to: correction))
}
let typedBoost = learnedBoosts[lower] ?? 0
let correctionFreq = lexicon.frequency(of: correction) + (learnedBoosts[correction] ?? 0)
// High-confidence gate: correction must clearly beat defending the typo.
guard correctionFreq >= 80, correctionFreq > typedBoost + 40 else { return nil }
return EnglishCorrectionDecision(
original: trimmed,
replacement: matchCase(of: trimmed, to: correction)
let context = EnglishSuggestionContext(
currentWord: typed,
previousWord: previousWord,
personalTerms: personalTerms,
learnedBoosts: learnedBoosts,
systemWords: systemWords,
systemGuesses: systemGuesses
)
return makeBoard(context).decision
}
public func nextWordComposition(_ context: EnglishSuggestionContext) -> TypingComposition {
var ranked: [(text: String, score: Int, id: String)] = []
var ranked: [(text: String, score: Int, role: TypingCandidateRole, quoted: Bool)] = []
var seen = Set<String>()
func append(_ raw: String, baseScore: Int, tag: String) {
func append(_ raw: String, baseScore: Int, role: TypingCandidateRole, quoted: Bool = false) {
let key = raw.lowercased()
guard seen.insert(key).inserted else { return }
let boost = context.learnedBoosts[key] ?? 0
let personalBoost = context.personalTerms.contains { $0.lowercased() == key } ? 2_000 : 0
ranked.append((raw, baseScore + boost + personalBoost, "\(tag)|\(key)"))
let personalBoost = isPersonal(key, in: context) ? 2_000 : 0
ranked.append((raw, baseScore + boost + personalBoost, role, quoted))
}
if let original = context.includeOriginalAfterCorrection {
append(original, baseScore: 20_000, tag: "original")
append(original, baseScore: 20_000, role: .verbatim, quoted: true)
}
if !context.previousWord.isEmpty {
for (index, word) in lexicon.nextWords(after: context.previousWord, limit: 8).enumerated() {
append(word, baseScore: 1_000 - index * 10, tag: "next")
append(word, baseScore: 1_200 - index * 10, role: .nextWord)
}
}
for term in context.personalTerms.prefix(4) {
append(term, baseScore: 500, tag: "personal")
append(term, baseScore: 500, role: .nextWord)
}
if ranked.filter({ $0.role == .nextWord }).isEmpty {
for (index, word) in lexicon.topWords(limit: 6).enumerated() {
append(word, baseScore: 200 - index, role: .nextWord)
}
}
ranked.sort { $0.score > $1.score }
let candidates = ranked.prefix(8).map {
TypingCandidate(id: $0.id, text: $0.text, engineIndex: 0)
let candidates = ranked.prefix(Self.slotCount).map {
TypingCandidate(
id: "\($0.role.rawValue)|\($0.text.lowercased())",
text: $0.text,
role: $0.role,
isQuoted: $0.quoted
)
}
return TypingComposition(preedit: "", candidates: Array(candidates))
}
// MARK: - Helpers
public func isKnownWord(_ word: String, personalTerms: [String], systemWords: [String]) -> Bool {
let lower = word.lowercased()
if lexicon.contains(lower) { return true }
if personalTerms.contains(where: { $0.lowercased() == lower }) { return true }
if systemWords.contains(where: { $0.lowercased() == lower }) { return true }
return false
}
private func shouldSkipAutocorrect(_ typed: String) -> Bool {
if typed.count <= 1 { return true }
// MARK: - Board
private struct Board {
var composition: TypingComposition
var decision: EnglishCorrectionDecision?
}
private func makeBoard(_ context: EnglishSuggestionContext) -> Board {
let typed = context.currentWord
let decision = makeCorrectionDecision(context)
var slots: [TypingCandidate] = []
var seen = Set<String>()
func add(_ text: String, role: TypingCandidateRole, quoted: Bool = false) {
let key = text.lowercased()
guard seen.insert(key).inserted else { return }
slots.append(
TypingCandidate(
id: "\(role.rawValue)|\(key)",
text: text,
role: role,
isQuoted: quoted
)
)
}
let known = isKnownWord(
typed,
personalTerms: context.personalTerms,
systemWords: context.systemWords
)
add(typed, role: .verbatim, quoted: !known)
if let decision {
add(decision.replacement, role: .correction)
}
for term in context.personalTerms where term.lowercased().hasPrefix(typed.lowercased())
&& term.lowercased() != typed.lowercased() {
add(term, role: .completion)
if slots.count >= Self.slotCount { break }
}
for word in context.systemCompletions {
let display = matchCase(of: typed, to: word)
add(display, role: .completion)
if slots.count >= Self.slotCount { break }
}
for word in lexicon.completions(prefix: typed, limit: 8) {
add(matchCase(of: typed, to: word), role: .completion)
if slots.count >= Self.slotCount { break }
}
let composition = TypingComposition(
preedit: typed,
candidates: Array(slots.prefix(Self.slotCount))
)
return Board(composition: composition, decision: decision)
}
private func makeCorrectionDecision(_ context: EnglishSuggestionContext) -> EnglishCorrectionDecision? {
let typed = context.currentWord
guard typed.count >= 3 else { return nil }
let lower = typed.lowercased()
if isProtectedToken(typed) { return nil }
if isPersonal(lower, in: context) { return nil }
if context.systemWords.contains(where: { $0.lowercased() == lower }) { return nil }
if (context.learnedBoosts[lower] ?? 0) >= 5 { return nil }
let inLexicon = lexicon.contains(lower)
let typedFreq = lexicon.frequency(of: lower) + (context.learnedBoosts[lower] ?? 0)
var pool = lexicon.scoredCorrections(for: lower, limit: 8)
for guess in context.systemGuesses {
let word = guess.lowercased()
guard word != lower else { continue }
if pool.contains(where: { $0.word == word }) { continue }
guard let alignment = EnglishQWERTYProximity.align(typed: lower, candidate: word) else { continue }
pool.append(
EnglishScoredCorrection(
word: word,
spatialCost: alignment.cost,
frequency: max(lexicon.frequency(of: word), 1),
isTransposition: alignment.isTransposition,
isShortening: alignment.isShortening
)
)
}
var best: (EnglishScoredCorrection, Int)?
for candidate in pool {
guard allowsAutocorrect(
typed: typed,
replacement: candidate.word,
inLexicon: inLexicon,
typedFreq: typedFreq,
candidate: candidate
) else { continue }
var score = candidate.frequency * 2 - candidate.spatialCost
if isPersonal(candidate.word, in: context) { score += 5_000 }
score += context.learnedBoosts[candidate.word] ?? 0
if lexicon.nextWords(after: context.previousWord).contains(candidate.word) {
score += 80
}
if let current = best {
if score > current.1 { best = (candidate, score) }
} else {
best = (candidate, score)
}
}
guard let best else { return nil }
let keepScore = inLexicon ? typedFreq * 2 : 0
guard best.1 > keepScore + 40 else { return nil }
return EnglishCorrectionDecision(
original: typed,
replacement: matchCase(of: typed, to: best.0.word)
)
}
private func allowsAutocorrect(
typed: String,
replacement: String,
inLexicon: Bool,
typedFreq: Int,
candidate: EnglishScoredCorrection
) -> Bool {
if isTitleCase(typed) {
// Teh The is a same-length transposition. Rocky Rock is not.
guard candidate.isTransposition, !candidate.isShortening else { return false }
}
if inLexicon {
let gap = candidate.frequency - typedFreq
// Web-corpus dumps leak typos (`teh`, `adn`) at the floor of the
// list. Real words like `form` sit much higher and must not yield
// to `from`.
let looksLikeLeakedTypo = typedFreq <= 680
if candidate.isTransposition {
return looksLikeLeakedTypo && gap >= 40
}
if typed.count == replacement.count,
candidate.spatialCost <= EnglishQWERTYProximity.adjacentCost {
return looksLikeLeakedTypo && gap >= Self.inVocabularyFrequencyGap
}
return false
}
return candidate.frequency > 0
}
private func isProtectedToken(_ typed: String) -> Bool {
if typed.count <= 2 { return true }
if typed.allSatisfy(\.isUppercase) { return true }
if typed.contains(where: \.isNumber) { return true }
if typed.contains("@") || typed.contains(".") || typed.contains("/") { return true }
@@ -170,6 +310,16 @@ public struct EnglishSuggestionEngine: Sendable {
return false
}
private func isTitleCase(_ typed: String) -> Bool {
guard let first = typed.first, first.isUppercase else { return false }
let rest = typed.dropFirst()
return !rest.isEmpty && rest.allSatisfy(\.isLowercase)
}
private func isPersonal(_ key: String, in context: EnglishSuggestionContext) -> Bool {
context.personalTerms.contains { $0.lowercased() == key }
}
private func matchCase(of sample: String, to word: String) -> String {
if sample.allSatisfy(\.isUppercase) {
return word.uppercased()
@@ -0,0 +1,67 @@
// EnglishSystemLexicon.swift
// OSGKeyboard · Shared
//
// Apple's sanctioned English sources for a custom keyboard: UITextChecker
// completions / guesses, plus UILexicon names from
// `requestSupplementaryLexicon`. The engine stays pure; the keyboard
// extension fills these fields on each refresh.
import Foundation
#if canImport(UIKit)
import UIKit
#endif
@MainActor
public protocol EnglishSystemLexiconProviding: AnyObject {
func completions(prefix: String, limit: Int) -> [String]
func guesses(for word: String, limit: Int) -> [String]
}
@MainActor
public final class EmptyEnglishSystemLexicon: EnglishSystemLexiconProviding {
public init() {}
public func completions(prefix: String, limit: Int) -> [String] {
[]
}
public func guesses(for word: String, limit: Int) -> [String] {
[]
}
}
#if canImport(UIKit)
/// System spellchecker. Always called from `TypingSessionController` (@MainActor).
@MainActor
public final class UIKitEnglishSystemLexicon: EnglishSystemLexiconProviding {
public var language: String
public init(language: String = "en_US") {
self.language = language
}
public func completions(prefix: String, limit: Int) -> [String] {
guard !prefix.isEmpty, limit > 0 else { return [] }
let checker = UITextChecker()
let range = NSRange(location: 0, length: (prefix as NSString).length)
let hits = checker.completions(forPartialWordRange: range, in: prefix, language: language) ?? []
return Array(hits.prefix(limit))
}
public func guesses(for word: String, limit: Int) -> [String] {
guard word.count >= 3, limit > 0 else { return [] }
let checker = UITextChecker()
let range = NSRange(location: 0, length: (word as NSString).length)
let hits = checker.guesses(forWordRange: range, in: word, language: language) ?? []
return Array(hits.prefix(limit))
}
public static func learnWord(_ word: String) {
let trimmed = word.trimmingCharacters(in: .whitespacesAndNewlines)
guard !trimmed.isEmpty else { return }
if !UITextChecker.hasLearnedWord(trimmed) {
UITextChecker.learnWord(trimmed)
}
}
}
#endif
@@ -21,6 +21,18 @@ public enum TypingInputLanguage: String, CaseIterable, Identifiable, Sendable {
}
}
/// Role of an English QuickType slot. Chinese candidates stay `.completion`.
public enum TypingCandidateRole: String, Equatable, Sendable {
/// The word currently being typed. Space does not replace it.
case verbatim
/// The unique slot Space will apply when autocorrect is armed.
case correction
/// Prefix completion; tap to accept, Space ignores it.
case completion
/// Next-word prediction after a committed word; tap to insert.
case nextWord
}
/// One candidate row item after composing.
public struct TypingCandidate: Identifiable, Equatable, Sendable {
public let id: String
@@ -28,17 +40,24 @@ public struct TypingCandidate: Identifiable, Equatable, Sendable {
public let annotation: String?
/// Absolute engine index for Chinese selection (may differ from display order).
public let engineIndex: Int
public let role: TypingCandidateRole
/// Unknown verbatim shown in quotes, matching the system / KeyboardKit contract.
public let isQuoted: Bool
public init(
id: String = UUID().uuidString,
text: String,
annotation: String? = nil,
engineIndex: Int = 0
engineIndex: Int = 0,
role: TypingCandidateRole = .completion,
isQuoted: Bool = false
) {
self.id = id
self.text = text
self.annotation = annotation
self.engineIndex = engineIndex
self.role = role
self.isQuoted = isQuoted
}
}
@@ -27,6 +27,10 @@ public final class TypingSessionController: ObservableObject {
/// When true, English suggestions / autocorrect stay off (secure fields).
@Published public var suggestionsEnabled: Bool = true
/// `UITextChecker` completions / guesses. Empty in unit tests.
public var systemLexicon: EnglishSystemLexiconProviding = EmptyEnglishSystemLexicon()
/// Names and text replacements from `requestSupplementaryLexicon`.
public var supplementaryWords: [String] = []
/// Chevron appears only for Chinese composition with at least two candidates.
public var canExpandCandidatePanel: Bool {
@@ -128,9 +132,15 @@ public final class TypingSessionController: ObservableObject {
)
TypingInputConfiguration.shared.reload()
refreshPersonalTerms()
// English lexicon is small; load when entering typing (not at KVC init).
// mmap the English table only while English is active. Chinese typing
// already has Rime; loading both on appear is what jetsams the extension.
if language == .english {
englishEngine.prepare()
OSGDiag.log("typing.enter after englishPrepare \(OSGDiag.memoryTag())", category: "boot")
} else {
EnglishLexicon.shared.unload()
OSGDiag.log("typing.enter skip englishPrepare lang=\(language.rawValue) \(OSGDiag.memoryTag())", category: "boot")
}
syncAutocapitalization()
if FlowSessionBridge.isHostHeavy() {
OSGDiag.log("typing.enter defer rime hostHeavy=1 — retry scheduled", category: "boot")
@@ -216,6 +226,7 @@ public final class TypingSessionController: ObservableObject {
synchronizeEnglishDocumentContext(caretMoved: true)
} else {
clearEnglishWordState(keepPrevious: false)
EnglishLexicon.shared.unload()
composition = engine.composition
}
return output
@@ -427,6 +438,9 @@ public final class TypingSessionController: ObservableObject {
pendingAutocorrection = nil
englishCurrentWord = pending.original
learningStore.recordDefense(of: pending.original)
#if canImport(UIKit)
UIKitEnglishSystemLexicon.learnWord(pending.original)
#endif
refreshEnglishSuggestions()
return .replace(deleteCount: deleteCount, with: pending.original)
}
@@ -461,15 +475,18 @@ public final class TypingSessionController: ObservableObject {
var decision = englishEngine.correctionDecision(
for: word,
personalTerms: personalTermsCache,
learnedBoosts: learningStore.snapshot()
learnedBoosts: learningStore.snapshot(),
previousWord: englishPreviousWord,
systemWords: supplementaryWords,
systemGuesses: systemLexicon.guesses(for: word, limit: 6)
) {
decision.appliedSuffix = suffix
pendingAutocorrection = decision
englishPreviousWord = decision.replacement
englishCurrentWord = ""
learningStore.recordAcceptance(of: decision.replacement)
// Suggestions stay hidden until the user starts the next word.
composition = .empty
// Machine-applied correction does not count as the user accepting
// the replacement otherwise names train the wrong word.
refreshEnglishSuggestions(afterCommittedWord: decision.replacement)
return .replace(
deleteCount: word.count,
with: decision.replacement + suffix
@@ -479,7 +496,17 @@ public final class TypingSessionController: ObservableObject {
englishPreviousWord = word
englishCurrentWord = ""
pendingAutocorrection = nil
learningStore.recordAcceptance(of: word, amount: 1)
// Learn OOV / names the user actually committed; skip common words.
if !englishEngine.isKnownWord(
word,
personalTerms: personalTermsCache,
systemWords: supplementaryWords
) {
learningStore.recordDefense(of: word, amount: 2)
#if canImport(UIKit)
UIKitEnglishSystemLexicon.learnWord(word)
#endif
}
refreshEnglishSuggestions(afterCommittedWord: word)
return suffix.isEmpty ? .none : .insert(suffix)
}
@@ -487,7 +514,8 @@ public final class TypingSessionController: ObservableObject {
private func selectEnglishCandidate(at index: Int) -> TypingOutput {
guard composition.candidates.indices.contains(index) else { return .none }
guard englishCandidateAnchorMatchesDocument() else { return .none }
let chosen = composition.candidates[index].text
let candidate = composition.candidates[index]
let chosen = candidate.text
// Restoring original after autocorrect (no current word).
if englishCurrentWord.isEmpty,
@@ -498,16 +526,27 @@ public final class TypingSessionController: ObservableObject {
englishPreviousWord = pending.original
englishCurrentWord = ""
learningStore.recordDefense(of: pending.original)
#if canImport(UIKit)
UIKitEnglishSystemLexicon.learnWord(pending.original)
#endif
refreshEnglishSuggestions(afterCommittedWord: pending.original)
return .replace(deleteCount: deleteCount, with: pending.original + " ")
}
if candidate.role == .verbatim {
learningStore.recordDefense(of: chosen)
#if canImport(UIKit)
UIKitEnglishSystemLexicon.learnWord(chosen)
#endif
} else {
learningStore.recordAcceptance(of: chosen)
}
if !englishCurrentWord.isEmpty {
let deleteCount = englishCurrentWord.count
englishPreviousWord = chosen
englishCurrentWord = ""
pendingAutocorrection = nil
learningStore.recordAcceptance(of: chosen)
refreshEnglishSuggestions(afterCommittedWord: chosen)
return .replace(deleteCount: deleteCount, with: chosen + " ")
}
@@ -516,7 +555,6 @@ public final class TypingSessionController: ObservableObject {
englishPreviousWord = chosen
englishCurrentWord = ""
pendingAutocorrection = nil
learningStore.recordAcceptance(of: chosen)
refreshEnglishSuggestions(afterCommittedWord: chosen)
return .insert(chosen + " ")
}
@@ -537,19 +575,23 @@ public final class TypingSessionController: ObservableObject {
composition = .empty
return
}
// Idle / between words: no candidate bar. Completions start after
// the first letter of the current word.
// With no active English word, keep the candidate bar empty. This also
// prevents next-word predictions from appearing between committed words.
guard !englishCurrentWord.isEmpty else {
composition = .empty
return
}
let previous = word ?? englishPreviousWord
let typed = englishCurrentWord
let context = EnglishSuggestionContext(
currentWord: englishCurrentWord,
currentWord: typed,
previousWord: previous,
personalTerms: personalTermsCache,
learnedBoosts: learningStore.snapshot(),
includeOriginalAfterCorrection: nil
includeOriginalAfterCorrection: pendingAutocorrection?.original,
systemWords: supplementaryWords,
systemCompletions: typed.isEmpty ? [] : systemLexicon.completions(prefix: typed, limit: 6),
systemGuesses: typed.count >= 3 ? systemLexicon.guesses(for: typed, limit: 6) : []
)
composition = englishEngine.compositionWhileTyping(context)
}
+121
View File
@@ -62,6 +62,86 @@ final class AIUserSkillTests: XCTestCase {
XCTAssertEqual(catalog.entries.first?.shortcutName, "My Tasks")
}
func testTextOnlySkillAllowsEmptyShortcutConfiguration() throws {
var catalog = AIUserSkillCatalog()
let skill = AIUserSkill(
name: "Rewrite",
prompt: "Rewrite the clipboard"
)
try catalog.upsert(skill)
let saved = try XCTUnwrap(catalog.entries.first)
let clipboardSkill = saved.asClipboardSkill()
XCTAssertNil(saved.shortcutICloudURL)
XCTAssertEqual(saved.shortcutName, "")
XCTAssertEqual(clipboardSkill.kind, .transform)
XCTAssertFalse(clipboardSkill.requiresShortcut)
XCTAssertNil(clipboardSkill.shortcutName)
}
func testShortcutNameWithoutLinkDoesNotExport() throws {
var catalog = AIUserSkillCatalog()
let skill = AIUserSkill(
name: "Rewrite",
prompt: "Rewrite the clipboard",
shortcutName: "Ignored without a link"
)
try catalog.upsert(skill)
let saved = try XCTUnwrap(catalog.entries.first)
XCTAssertEqual(saved.shortcutName, "Ignored without a link")
XCTAssertEqual(saved.asClipboardSkill().kind, .transform)
XCTAssertNil(saved.asClipboardSkill().shortcutName)
}
func testShortcutSkillRequiresNameAndValidShareLink() {
var catalog = AIUserSkillCatalog()
XCTAssertThrowsError(
try catalog.upsert(
AIUserSkill(
name: "Export",
prompt: "Export it",
shortcutICloudURL: sampleURL
)
)
) { error in
XCTAssertEqual(error as? AIUserSkillValidationError, .emptyShortcutName)
}
XCTAssertThrowsError(
try catalog.upsert(
AIUserSkill(
name: "Export",
prompt: "Export it",
shortcutICloudURL: URL(string: "https://example.com/not-a-shortcut"),
shortcutName: "Run Me"
)
)
) { error in
XCTAssertEqual(error as? AIUserSkillValidationError, .invalidShortcutLink)
}
}
func testExistingShortcutSkillEncodingDecodesWithOptionalURL() throws {
let original = AIUserSkill(
name: "Export",
prompt: "Export it",
shortcutICloudURL: sampleURL,
shortcutName: "Run Me"
)
let decoded = try JSONDecoder().decode(
AIUserSkill.self,
from: JSONEncoder().encode(original)
)
XCTAssertEqual(decoded.shortcutICloudURL, sampleURL)
XCTAssertEqual(decoded.asClipboardSkill().kind, .export)
XCTAssertTrue(decoded.asClipboardSkill().requiresShortcut)
}
func testThinkingDefaultsOffAndBuiltinCannotEnable() {
let user = AIUserSkill(
name: "Custom",
@@ -186,4 +266,45 @@ final class AIUserSkillStoreTests: XCTestCase {
XCTAssertFalse(store.layout.hasConfirmedShortcut(skill.id))
XCTAssertFalse(store.layout.isEnabled(skill.id))
}
func testRemovingShortcutLinkKeepsEnabledTextSkill() throws {
let store = AIAgentSkillLayoutStore(defaults: makeDefaults())
var skill = AIUserSkill(
name: "Custom",
prompt: "Do it",
shortcutICloudURL: URL(
string: "https://www.icloud.com/shortcuts/65bf33ba4206484ba78d582eaf1e9c44"
),
shortcutName: "Run Me"
)
try store.saveUserSkill(skill)
XCTAssertEqual(store.confirmShortcutAndEnable(skill.id), .enabled)
skill.shortcutICloudURL = nil
try store.saveUserSkill(skill)
XCTAssertTrue(store.layout.isEnabled(skill.id))
XCTAssertFalse(store.layout.hasConfirmedShortcut(skill.id))
XCTAssertEqual(store.userSkill(id: skill.id)?.asClipboardSkill().kind, .transform)
}
func testAddingShortcutLinkDisablesTextSkillUntilConfirmed() throws {
let store = AIAgentSkillLayoutStore(defaults: makeDefaults())
var skill = AIUserSkill(
name: "Custom",
prompt: "Do it"
)
try store.saveUserSkill(skill)
XCTAssertEqual(store.enable(skill.id), .enabled)
skill.shortcutICloudURL = URL(
string: "https://www.icloud.com/shortcuts/65bf33ba4206484ba78d582eaf1e9c44"
)
skill.shortcutName = "Run Me"
try store.saveUserSkill(skill)
XCTAssertFalse(store.layout.isEnabled(skill.id))
XCTAssertFalse(store.layout.hasConfirmedShortcut(skill.id))
XCTAssertEqual(store.enable(skill.id), .needsShortcut)
}
}
@@ -0,0 +1,77 @@
// EnglishTypingOnDeviceTests.swift
// OSGKeyboardTests
//
// Hosted in the main app so these can run on a physical iPhone.
// ExtTests stay tool-hosted (simulator-only).
import XCTest
@testable import OSGKeyboardShared
final class EnglishTypingOnDeviceTests: XCTestCase {
func testLexiconLoadsFortyThousandWords() {
let lexicon = EnglishLexicon()
lexicon.prepare()
XCTAssertTrue(lexicon.isLoaded)
XCTAssertGreaterThanOrEqual(lexicon.wordCount, 30_000)
XCTAssertTrue(lexicon.contains("the"))
XCTAssertTrue(lexicon.contains("hello"))
XCTAssertTrue(lexicon.contains("definitely"))
lexicon.unload()
XCTAssertFalse(lexicon.isLoaded)
}
func testTehAutocorrectsToThe() {
let engine = EnglishSuggestionEngine()
engine.prepare()
XCTAssertEqual(
engine.correctionDecision(for: "teh", personalTerms: [], learnedBoosts: [:])?.replacement,
"the"
)
}
func testTitleCaseNamesStay() {
let engine = EnglishSuggestionEngine()
engine.prepare()
XCTAssertNil(engine.correctionDecision(for: "Rocky", personalTerms: [], learnedBoosts: [:]))
XCTAssertNil(engine.correctionDecision(for: "Wang", personalTerms: [], learnedBoosts: [:]))
}
func testProximityGppdBecomesGood() {
let engine = EnglishSuggestionEngine()
engine.prepare()
XCTAssertEqual(
engine.correctionDecision(for: "gppd", personalTerms: [], learnedBoosts: [:])?.replacement,
"good"
)
}
func testFormIsNotCorrectedToFrom() {
let engine = EnglishSuggestionEngine()
engine.prepare()
XCTAssertNil(engine.correctionDecision(for: "form", personalTerms: [], learnedBoosts: [:]))
}
func testThankYouBigram() {
let lexicon = EnglishLexicon()
lexicon.prepare()
XCTAssertTrue(lexicon.nextWords(after: "thank").contains("you"))
}
@MainActor
func testUITextCheckerCompletionsAvailable() {
let system = UIKitEnglishSystemLexicon()
let hits = system.completions(prefix: "hel", limit: 6)
XCTAssertFalse(hits.isEmpty, "device UITextChecker should complete hel")
}
@MainActor
func testQuickTypeBoardMarksCorrection() {
let engine = EnglishSuggestionEngine()
engine.prepare()
let composition = engine.compositionWhileTyping(
EnglishSuggestionContext(currentWord: "teh")
)
XCTAssertEqual(composition.candidates.first?.role, .verbatim)
XCTAssertTrue(composition.candidates.contains { $0.role == .correction && $0.text == "the" })
}
}
@@ -0,0 +1,43 @@
// EnglishKeyboardDeviceUITests.swift
// OSGKeyboardUITests
//
// Physical-device pass: Notes-like host + the real keyboard extension.
// Skips if OSGKeyboard is not enabled as the current keyboard.
import XCTest
final class EnglishKeyboardDeviceUITests: XCTestCase {
func testOSGKeyboardAppearsOnNotesHost() throws {
let app = XCUIApplication()
app.launchArguments = [
"--whats-new-host",
"--whats-new-lang=en",
"--whats-new-scenario=edit"
]
app.launch()
let textView = app.textViews["notes.host.textView"]
XCTAssertTrue(
textView.waitForExistence(timeout: 12),
"Notes host text view should appear"
)
if !textView.exists {
return
}
textView.tap()
let keyboard = XCUIApplication(bundleIdentifier: "com.osgkeyboard.ios.keyboard")
let appeared = keyboard.wait(for: .runningForeground, timeout: 8)
|| keyboard.windows.firstMatch.waitForExistence(timeout: 8)
if !appeared {
throw XCTSkip(
"OSGKeyboard extension is not the active keyboard on this device. Enable it in Settings ▸ Keyboard, then re-run."
)
}
XCTAssertTrue(
keyboard.windows.firstMatch.exists,
"OSGKeyboard extension window should be on screen"
)
}
}
+131
View File
@@ -0,0 +1,131 @@
#!/usr/bin/env bash
# Physical-device keyboard appear/hide stress.
# Shows and dismisses the real OSGKeyboard extension N times, then checks
# crash reports.
#
# Usage:
# ./Scripts/keyboard-appear-stress.sh [COUNT=50] [UDID]
set -euo pipefail
ROOT="$(cd "$(dirname "$0")/.." && pwd)"
cd "$ROOT"
COUNT="${1:-50}"
UDID="${2:-00008130-001C249C0E52001C}"
BUNDLE="com.osgkeyboard.ios"
OUT_DIR="${ROOT}/.tmp/keyboard-appear-stress-$(date +%Y%m%d-%H%M%S)"
mkdir -p "$OUT_DIR"
CONSOLE="$OUT_DIR/console.log"
SUMMARY="$OUT_DIR/summary.txt"
DERIVED="${ROOT}/.derivedData-device-stress"
APP="$DERIVED/Build/Products/Debug-iphoneos/OSGKeyboard.app"
echo "==> Out: $OUT_DIR"
echo "==> Count: $COUNT Device: $UDID"
echo "==> Building Debug-iphoneos"
xcodebuild build \
-project "$ROOT/OSGKeyboard.xcodeproj" \
-scheme OSGKeyboard \
-destination "platform=iOS,id=$UDID" \
-configuration Debug \
-derivedDataPath "$DERIVED" \
-allowProvisioningUpdates \
-onlyUsePackageVersionsFromResolvedFile \
CODE_SIGNING_ALLOWED=YES \
>/dev/null
echo "==> Installing"
xcrun devicectl device install app --device "$UDID" "$APP" --timeout 180 >/dev/null
crash_list() {
xcrun devicectl device info files \
--device "$UDID" \
--domain-type systemCrashLogs \
--timeout 30 2>/dev/null \
| awk '/OSGKeyboardExt-/{print $1}'
}
BEFORE="$OUT_DIR/crashes-before.txt"
AFTER="$OUT_DIR/crashes-after.txt"
crash_list | sort >"$BEFORE"
echo "==> Launching appear-stress count=$COUNT"
# 50 hide/show cycles plus first show; ~0.5s each + timeouts.
TIMEOUT_SECS=$((COUNT * 3 + 40))
set +e
python3 - "$TIMEOUT_SECS" "$UDID" "$BUNDLE" "$COUNT" "$CONSOLE" <<'PY'
import subprocess, sys, time, os, signal
timeout, udid, bundle, count, console = sys.argv[1:6]
cmd = [
"xcrun", "devicectl", "device", "process", "launch",
"--device", udid,
"--console",
"--terminate-existing",
bundle,
"--whats-new-host",
"--whats-new-lang=en",
"--whats-new-scenario=edit",
f"--keyboard-appear-stress={count}",
]
with open(console, "w") as out:
proc = subprocess.Popen(cmd, stdout=out, stderr=subprocess.STDOUT)
try:
proc.wait(timeout=int(timeout))
except subprocess.TimeoutExpired:
proc.send_signal(signal.SIGTERM)
try:
proc.wait(timeout=8)
except subprocess.TimeoutExpired:
proc.kill()
sys.exit(124)
sys.exit(proc.returncode or 0)
PY
LAUNCH_STATUS=$?
set -e
crash_list | sort >"$AFTER"
NEW_CRASHES="$OUT_DIR/crashes-new.txt"
comm -13 "$BEFORE" "$AFTER" >"$NEW_CRASHES"
PASSED="$(python3 - "$CONSOLE" <<'PY'
import re, sys
text = open(sys.argv[1], errors="replace").read()
hits = re.findall(r"keyboard\.stress done passed=(\d+)/(\d+)", text)
print(hits[-1][0] if hits else "")
PY
)"
TOTAL="$(python3 - "$CONSOLE" <<'PY'
import re, sys
text = open(sys.argv[1], errors="replace").read()
hits = re.findall(r"keyboard\.stress done passed=(\d+)/(\d+)", text)
print(hits[-1][1] if hits else "")
PY
)"
FAIL_LINE="$(grep 'keyboard.stress FAIL' "$CONSOLE" | tail -1 || true)"
{
echo "Keyboard appear/hide stress"
echo "device=$UDID count=$COUNT"
echo "launch_exit=$LAUNCH_STATUS"
echo "passed=${PASSED:-0}/${TOTAL:-$COUNT}"
echo "fail_line=${FAIL_LINE:-none}"
echo "new_OSGKeyboardExt_crashes:"
if [[ -s "$NEW_CRASHES" ]]; then
cat "$NEW_CRASHES"
else
echo " (none)"
fi
echo "console=$CONSOLE"
} | tee "$SUMMARY"
if [[ -s "$NEW_CRASHES" ]]; then
echo "FAIL: new OSGKeyboardExt crash reports" >&2
exit 1
fi
if [[ "${PASSED:-0}" != "$COUNT" ]]; then
echo "FAIL: expected $COUNT cycles, got ${PASSED:-0}" >&2
tail -40 "$CONSOLE" >&2
exit 1
fi
echo "PASS: $COUNT/$COUNT appear-hide cycles, no new extension crashes"
+427
View File
@@ -0,0 +1,427 @@
#!/usr/bin/env python3
"""Build compact English unigram + bigram TSVs from Peter Norvig's public-domain
n-gram counts (https://norvig.com/ngrams/).
Norvig: I hereby release all these files into the public domain.
We store log-scaled ranks (not raw counts) so the keyboard extension stays small
and we are not redistributing the full Google Web Trillion Word Corpus dump.
"""
from __future__ import annotations
import argparse
import math
import re
import struct
import sys
import urllib.request
from pathlib import Path
ROOT = Path(__file__).resolve().parents[2]
OUT_DIR = ROOT / "OSGKeyboardShared" / "Resources" / "Typing" / "English"
UNIGRAM_URL = "https://norvig.com/ngrams/count_1w.txt"
BIGRAM_URL = "https://norvig.com/ngrams/count_2w.txt"
WORD_RE = re.compile(r"^[a-z]+(?:'[a-z]+)?$")
MAX_UNIGRAMS = 40_000
MAX_BIGRAMS = 8_000
MAX_WORD_LEN = 20
# Stop reading the 2-gram file once we have enough accepted rows; the file is
# already sorted by descending count.
BIGRAM_SCAN_LIMIT = 80_000
# Seed collocations so next-word still works if the 2-gram download fails.
FALLBACK_BIGRAMS: list[tuple[str, list[str]]] = [
("the", "of and to in is for that with on a".split()),
("of", "the a this my our their these those course them".split()),
("to", "the be a do make see get go have my".split()),
("and", "the a then also other more so we you".split()),
("a", "lot few new good great little bit way time person".split()),
("in", "the a my this order fact front time case addition".split()),
("is", "a the not to that this it one more also".split()),
("for", "the a example me you us this that now sure".split()),
("that", "the is was I you it we they this are".split()),
("i", "am have will would can do think know want was".split()),
("it", "is was would will can be has had to not".split()),
("on", "the a my this time top of it you that".split()),
("you", "are can will would have do know want to should".split()),
("with", "the a my you it this that him her them".split()),
("as", "a the well much soon far long if of to".split()),
("this", "is was the a time one way thing point case".split()),
("we", "are have will can would do need want should were".split()),
("have", "a the been to been a been the time been".split()),
("be", "a the able to in on there here with as".split()),
("are", "a the not you we they going to in on".split()),
("not", "a the be to sure only yet even really the".split()),
("but", "the I a it is also then we you not".split()),
("from", "the a my this that it you now here there".split()),
("at", "the a my this time least home work school night".split()),
("by", "the a this that now then far me you email".split()),
("or", "the a not so to it you we they this".split()),
("an", "hour example email idea issue update account apple app".split()),
("if", "you the I we it that this not so a".split()),
("will", "be you I we the not have to a get".split()),
("can", "be you I we not the a help see get".split()),
("would", "be you I we like have not the a to".split()),
("do", "you not the I we it that this a".split()),
("there", "is are was were a the no not been have".split()),
("their", "own new first last time way work house car".split()),
("what", "is the a you I we do time about if".split()),
("when", "the I you we it is a this that not".split()),
("which", "is the a you we they of in to that".split()),
("who", "is are was were the a you I we".split()),
("how", "to much many long about is the a you".split()),
("about", "the a this that it you to time me".split()),
("into", "the a this that my it you a new".split()),
("just", "a the like to be now want wanted got".split()),
("like", "a the to this that it you I we".split()),
("so", "I the a you we that this much many".split()),
("than", "the a I you we this that it to".split()),
("then", "the I you we a it to is was".split()),
("them", "to a the in on with for and I".split()),
("these", "are is the a days things people ones two".split()),
("those", "are is the a who were days people ones".split()),
("my", "own new first last time way email phone name".split()),
("your", "own new email phone name time way account".split()),
("our", "own new first last time team way house".split()),
("going", "to be the a in on for with".split()),
("want", "to a the you I we it".split()),
("need", "to a the you I we it".split()),
("let", "me you us the a".split()),
("please", "let me you the a".split()),
("thank", "you so much".split()),
("thanks", "for so much".split()),
("looking", "forward to for at".split()),
("let", "me you us know".split()),
]
def fetch_lines(url: str, max_lines: int | None = None) -> list[str]:
req = urllib.request.Request(
url,
headers={"User-Agent": "OSGKeyboard-lexicon-builder/1.0"},
)
with urllib.request.urlopen(req, timeout=120) as response:
lines: list[str] = []
for raw in response:
line = raw.decode("utf-8", errors="ignore").strip()
if not line:
continue
lines.append(line)
if max_lines is not None and len(lines) >= max_lines:
break
return lines
def parse_count_line(line: str) -> tuple[str, int] | None:
parts = line.split()
if len(parts) < 2:
return None
token = parts[0].lower()
try:
count = int(parts[-1])
except ValueError:
return None
return token, count
def log_rank(count: int) -> int:
return max(1, int(round(math.log10(count) * 100)))
def build_unigrams(lines: list[str]) -> dict[str, int]:
ranked: list[tuple[str, int]] = []
seen: set[str] = set()
for line in lines:
parsed = parse_count_line(line)
if parsed is None:
continue
word, count = parsed
if word in seen:
continue
if not WORD_RE.match(word) or len(word) > MAX_WORD_LEN:
continue
seen.add(word)
ranked.append((word, count))
if len(ranked) >= MAX_UNIGRAMS:
break
return {word: log_rank(count) for word, count in ranked}
def build_bigrams(
lines: list[str],
unigrams: dict[str, int],
) -> dict[str, list[str]]:
grouped: dict[str, list[tuple[str, int]]] = {}
accepted = 0
for line in lines:
parsed = parse_count_line(line)
if parsed is None:
continue
token, count = parsed
parts = token.split("_")
if len(parts) != 2:
# Norvig 2-grams are "word1 word2 count"
bits = line.lower().split()
if len(bits) < 3:
continue
left, right, count_s = bits[0], bits[1], bits[-1]
try:
count = int(count_s)
except ValueError:
continue
else:
left, right = parts
if left not in unigrams or right not in unigrams:
continue
if left == right:
continue
bucket = grouped.setdefault(left, [])
if any(word == right for word, _ in bucket):
continue
bucket.append((right, count))
accepted += 1
if accepted >= MAX_BIGRAMS * 3:
break
result: dict[str, list[str]] = {}
used = 0
for left, pairs in grouped.items():
pairs.sort(key=lambda item: item[1], reverse=True)
nxt = [word for word, _ in pairs[:8]]
if not nxt:
continue
result[left] = nxt
used += len(nxt)
if used >= MAX_BIGRAMS:
break
return result
def merge_fallback(bigrams: dict[str, list[str]]) -> dict[str, list[str]]:
merged = dict(bigrams)
for left, rights in FALLBACK_BIGRAMS:
existing = merged.get(left, [])
seen = set(existing)
for word in rights:
if word not in seen:
existing.append(word)
seen.add(word)
merged[left] = existing[:10]
return merged
def write_unigrams(path: Path, unigrams: dict[str, int]) -> None:
rows = sorted(unigrams.items(), key=lambda item: (-item[1], item[0]))
path.write_text("".join(f"{word}\t{freq}\n" for word, freq in rows), encoding="utf-8")
def write_bigrams(path: Path, bigrams: dict[str, list[str]]) -> None:
rows = sorted(bigrams.items(), key=lambda item: item[0])
path.write_text(
"".join(f"{left}\t{' '.join(rights)}\n" for left, rights in rows),
encoding="utf-8",
)
# mmap binary (`english_lexicon.bin`), little-endian. Layout:
# 64-byte header, then unigram records, freq-rank indices, 26 initial
# ranges, bigram groups, packed next-word indices, ASCII string pool.
# The keyboard maps this file; it must not parse TSV into Swift dictionaries.
BIN_MAGIC = b"OSGENG01"
BIN_VERSION = 1
BIN_HEADER_SIZE = 64
BIN_INITIAL_COUNT = 26
def _align4(offset: int) -> int:
return (offset + 3) & ~3
def write_binary(
path: Path,
unigrams: dict[str, int],
bigrams: dict[str, list[str]],
) -> None:
words = sorted(unigrams.keys())
index_of = {word: index for index, word in enumerate(words)}
pool = bytearray()
records: list[tuple[int, int, int]] = []
for word in words:
encoded = word.encode("ascii")
if len(encoded) > 255:
continue
freq = min(int(unigrams[word]), 65_535)
records.append((len(pool), len(encoded), freq))
pool.extend(encoded)
initials = [(0, 0)] * BIN_INITIAL_COUNT
cursor = 0
while cursor < len(words):
first = words[cursor][0]
if "a" <= first <= "z":
start = cursor
while cursor < len(words) and words[cursor][0] == first:
cursor += 1
initials[ord(first) - ord("a")] = (start, cursor - start)
else:
cursor += 1
freq_order = sorted(
range(len(words)),
key=lambda index: (-unigrams[words[index]], words[index]),
)
groups: list[tuple[int, int, int]] = []
next_indices: list[int] = []
for left in sorted(bigrams.keys()):
prev_index = index_of.get(left)
if prev_index is None:
continue
rights = [index_of[word] for word in bigrams[left] if word in index_of]
if not rights:
continue
groups.append((prev_index, len(rights), len(next_indices)))
next_indices.extend(rights)
unigram_offset = _align4(BIN_HEADER_SIZE)
freq_offset = _align4(unigram_offset + len(records) * 8)
initial_offset = _align4(freq_offset + len(freq_order) * 2)
bigram_index_offset = _align4(initial_offset + BIN_INITIAL_COUNT * 4)
bigram_next_offset = _align4(bigram_index_offset + len(groups) * 8)
pool_offset = _align4(bigram_next_offset + len(next_indices) * 2)
total = pool_offset + len(pool)
blob = bytearray(total)
struct.pack_into(
"<8s14I",
blob,
0,
BIN_MAGIC,
BIN_VERSION,
len(records),
len(groups),
pool_offset,
len(pool),
unigram_offset,
freq_offset,
initial_offset,
bigram_index_offset,
bigram_next_offset,
0,
0,
0,
0,
)
for index, (pool_off, length, freq) in enumerate(records):
struct.pack_into(
"<IBBH",
blob,
unigram_offset + index * 8,
pool_off,
length,
0,
freq,
)
for index, word_index in enumerate(freq_order):
struct.pack_into("<H", blob, freq_offset + index * 2, word_index)
for letter, (start, count) in enumerate(initials):
struct.pack_into("<HH", blob, initial_offset + letter * 4, start, count)
for index, (prev_index, count, first_next) in enumerate(groups):
struct.pack_into(
"<HHI",
blob,
bigram_index_offset + index * 8,
prev_index,
count,
first_next,
)
for index, word_index in enumerate(next_indices):
struct.pack_into("<H", blob, bigram_next_offset + index * 2, word_index)
blob[pool_offset : pool_offset + len(pool)] = pool
path.write_bytes(blob)
def read_unigrams_tsv(path: Path) -> dict[str, int]:
result: dict[str, int] = {}
for line in path.read_text(encoding="utf-8").splitlines():
if not line.strip():
continue
word, freq_s = line.split("\t", 1)
result[word.lower()] = int(freq_s)
return result
def read_bigrams_tsv(path: Path) -> dict[str, list[str]]:
result: dict[str, list[str]] = {}
for line in path.read_text(encoding="utf-8").splitlines():
if not line.strip():
continue
left, rights = line.split("\t", 1)
result[left.lower()] = [word.lower() for word in rights.split() if word]
return result
def emit_outputs(unigrams: dict[str, int], bigrams: dict[str, list[str]]) -> None:
OUT_DIR.mkdir(parents=True, exist_ok=True)
write_unigrams(OUT_DIR / "english_lexicon.tsv", unigrams)
write_bigrams(OUT_DIR / "english_bigrams.tsv", bigrams)
write_binary(OUT_DIR / "english_lexicon.bin", unigrams, bigrams)
print(f"Wrote {OUT_DIR / 'english_lexicon.tsv'}", file=sys.stderr)
print(f"Wrote {OUT_DIR / 'english_bigrams.tsv'}", file=sys.stderr)
print(
f"Wrote {OUT_DIR / 'english_lexicon.bin'} "
f"({(OUT_DIR / 'english_lexicon.bin').stat().st_size} bytes)",
file=sys.stderr,
)
def main() -> int:
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument(
"--from-tsv",
action="store_true",
help="Compile english_lexicon.bin from existing TSV files (no network).",
)
args = parser.parse_args()
OUT_DIR.mkdir(parents=True, exist_ok=True)
if args.from_tsv:
unigram_path = OUT_DIR / "english_lexicon.tsv"
bigram_path = OUT_DIR / "english_bigrams.tsv"
if not unigram_path.is_file() or not bigram_path.is_file():
print("Missing english_lexicon.tsv / english_bigrams.tsv", file=sys.stderr)
return 1
unigrams = read_unigrams_tsv(unigram_path)
bigrams = read_bigrams_tsv(bigram_path)
write_binary(OUT_DIR / "english_lexicon.bin", unigrams, bigrams)
print(
f"Wrote {OUT_DIR / 'english_lexicon.bin'} "
f"({(OUT_DIR / 'english_lexicon.bin').stat().st_size} bytes) "
f"from {len(unigrams)} unigrams / {len(bigrams)} bigram keys",
file=sys.stderr,
)
return 0
print(f"Fetching unigrams from {UNIGRAM_URL}", file=sys.stderr)
unigram_lines = fetch_lines(UNIGRAM_URL, max_lines=200_000)
unigrams = build_unigrams(unigram_lines)
print(f"Kept {len(unigrams)} unigrams", file=sys.stderr)
bigrams: dict[str, list[str]] = {}
try:
print(f"Fetching bigrams from {BIGRAM_URL}", file=sys.stderr)
bigram_lines = fetch_lines(BIGRAM_URL, max_lines=BIGRAM_SCAN_LIMIT)
bigrams = build_bigrams(bigram_lines, unigrams)
print(f"Kept {sum(len(v) for v in bigrams.values())} bigram edges", file=sys.stderr)
except Exception as exc: # noqa: BLE001 — fallback is intentional
print(f"Bigram download failed ({exc}); using fallback collocations", file=sys.stderr)
bigrams = merge_fallback(bigrams)
emit_outputs(unigrams, bigrams)
return 0
if __name__ == "__main__":
raise SystemExit(main())
+1
View File
@@ -120,6 +120,7 @@
"platform": "ios",
"tests": [
"OSGKeyboardExtTests/EnglishTypingTests",
"OSGKeyboardTests/EnglishTypingOnDeviceTests",
"OSGKeyboardExtTests/CandidatePanelExpandTests",
"OSGKeyboardExtTests/ClipboardSuggestionLifecycleTests",
"OSGKeyboardExtTests/KeyboardStateTests",
+4 -3
View File
@@ -139,6 +139,7 @@ targets:
- path: OSGKeyboard/Resources/Typing/Licenses/LICENSE-JIEBA-MIT.txt
- path: OSGKeyboard/Resources/Typing/Licenses/LICENSE-PHRASE-PINYIN-DATA-MIT.txt
- path: OSGKeyboard/Resources/Typing/Licenses/LICENSE-PINYIN-DATA-MIT.txt
- path: OSGKeyboard/Resources/Typing/Licenses/LICENSE-NORVIG-PD.txt
- path: OSGKeyboard/Resources/Typing/Licenses/THIRD_PARTY_NOTICES.md
- path: OSGKeyboard/Resources/Typing/Licenses/third-party-notices.zip
- path: OSGKeyboard/Resources/Typing/Licenses/LIBRIME-COMBINED-NOTICES.txt
@@ -348,9 +349,7 @@ targets:
buildPhase: resources
# Prefer sources+buildPhase over target-level `resources:` — the latter
# was not emitting Copy Bundle Resources entries for this framework.
- path: OSGKeyboardShared/Resources/Typing/English/english_lexicon.tsv
buildPhase: resources
- path: OSGKeyboardShared/Resources/Typing/English/english_bigrams.tsv
- path: OSGKeyboardShared/Resources/Typing/English/english_lexicon.bin
buildPhase: resources
- path: OSGKeyboardShared/Resources/PolishStyles
buildPhase: resources
@@ -582,6 +581,8 @@ targets:
buildPhase: resources
- path: OSGKeyboard/Resources/Typing/Licenses/LICENSE-PINYIN-DATA-MIT.txt
buildPhase: resources
- path: OSGKeyboard/Resources/Typing/Licenses/LICENSE-NORVIG-PD.txt
buildPhase: resources
- path: OSGKeyboard/Resources/Typing/Licenses/LICENSE.txt
buildPhase: resources
- path: OSGKeyboard/Resources/Typing/Licenses/NOTICE.txt