0401f67a89e3bbdb44cc8cb9aeb2d82212405dad
7 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
c07cf4db9f |
refactor: drop Qwen3 CoreML ASR, add local-engine cloud polish toggle
Rolls back the v0.2.0 Qwen3 CoreML on-device ASR stack and replaces the
'local engine' UX with iOS 26 SpeechAnalyzer + DictationTranscriber only.
The 'Cloud polish after ASR' toggle (ProviderConfig.localModeCloudPolishEnabled)
lets users opt into a post-ASR DeepSeek round-trip from the local engine.
Defaults to off so the local engine stays genuinely local. New PolishError.missingAPIError
surfaces an inline 'fill in your key' warning when the toggle is on but the
Keychain is empty. DeepSeek preset default model bumped to deepseek-v4-flash.
Deleted:
- OSGKeyboard/ThirdParty/Qwen3Speech/ (74 files, ~16k LoC)
- OSGKeyboard/Services/ModelManager.swift (492)
- OSGKeyboard/Services/OnDeviceModelWarmup.swift (197)
- OSGKeyboard/Services/Qwen3ASRService.swift (257)
- OSGKeyboard/Services/ModelDownloadSourcePicker.swift (126)
- OSGKeyboard/Views/OnDeviceModelsView.swift (184)
- OSGKeyboard/Views/DownloadConfirmSheet.swift (96)
- OSGKeyboardShared/Models/OnDeviceModel.swift (140)
- OSGKeyboardShared/Services/OnDeviceModelStatus.swift (104)
- Qwen3ASRServiceProvider registration in OSGKeyboardApp
- Qwen3Speech package declaration in project.yml
- 5 .qwen3ASR enum / branch reference sites in HomeView, OnboardingView,
LocalEngineSettingsRows, FlowSessionManager, ASRService, EngineServiceLabel
- Two pre-existing Swift 6 strict-concurrency errors in
LiveDictationController + FlowSessionManager (the weak [weak self] in
detached-task MainActor.run blocks) that were blocking clean builds
Added:
- LocalModelsGroup: 'Built-in iOS SpeechAnalyzer' badge + 'Cloud polish
after ASR' Switch toggle
- PolishingService: honour localModeCloudPolishEnabled; new .missingAPIKey
error case with localised warning
- AppGroupStore.localModeCloudPolishEnabled (mirrored into App Group
so the keyboard extension honours the toggle during live dictation)
- SettingsView: show provider/api sections when local-mode cloud polish
is on so the user can paste a DeepSeek key
- FlowSessionManager: route through PolishingService for local + polish-on
flow; translate missingAPIKey into a polished warning
- KeyboardViewController: handle PolishingService.PolishError.missingAPIKey
in the keyboard-side live polish path
- CHANGELOG v0.2.1: documents the rollback + new toggle
- README.md / README.zh.md: engine matrix section, data flow note
Verified: xcodebuild -scheme OSGKeyboard -destination 'generic/platform=iOS Simulator'
build succeeds under SWIFT_STRICT_CONCURRENCY=complete.
|
||
|
|
df1c5ff32c |
feat: migrate on-device Qwen3 ASR to CoreML for background Flow dictation
Replace MLX GPU inference with CoreML bundles so transcription continues while the host app is backgrounded. Adds model download and warm-up, vendored Qwen3Speech, and updates onboarding, settings, and copy for the ~1.6 GB CoreML package (iOS 18+). |
||
|
|
7f059dbd45 |
feat: TypeWhisper Flow sessions, Phase 4 UX, and GitHub Pages privacy site
Migrate keyboard dictation to continuous Flow sessions with auto-start, tap-to-toggle recording, 60s countdown, five-step onboarding, and App Group IPC. Add docs/ GitHub Pages site with en/zh privacy policy for App Store compliance. |
||
|
|
6148d05093 |
chore: align iOS 26 capability docs and UI localization
Keep repository messaging consistent with the implemented iOS 26 SpeechAnalyzer path, and remove mixed hardcoded copy by routing remaining UI/error text through localized string keys. |
||
|
|
a227309059 |
fix: feed DictationTranscriber Int16 PCM, not Float32
The keyboard preview crashed on first record with a
`__abort_with_payload` deep inside Speech's
`DictationTranscriber`. The disassembly surfaced three
preconditions checked before a `brk #0x1`:
+620 "Audio sample data must be 16-bit signed integers"
+848 "Multi-channel audio is not supported"
+1072 "Client info not fully initialized"
We hit the first one. `DictationTranscriber` (iOS 26's new
`SpeechAnalyzer`-backed engine) is strict about its input
format: only Int16 PCM, not the Float32 PCM that the iOS 18
`SFSpeechRecognizer` path accepted. Our audio-tap and
`AudioBufferSnapshot.samples: [Float]` are Float32 all the
way down — that was the SFSpeech shape, and the previous
`AppleSpeechASR` adapted internally. With iOS 26 as the
deployment target, the only ASR backend is
`SpeechAnalyzerASR`, and the conversion needed to happen at
the `AnalyzerInput` boundary.
Fix:
- `transcribe` builds the `AVAudioFormat` as
`.pcmFormatInt16, 16 kHz, 1 ch, interleaved: true` (the
canonical layout for Int16 Speech input).
- `makeInputStream` runs the per-sample conversion
`Int16(round(clamp(s * 32767, -32768, 32767)))` into the
`AVAudioPCMBuffer`'s `int16ChannelData[0]`. The explicit
clip is required (a `s == 1.5` from a gain-overflow at the
audio-engine boundary would otherwise wrap to a negative
Int16 after the implicit truncation). `round()` (not
truncate) preserves DC balance — `0.5` quantises to
`+16384`, not `+16383`, matching what audio DAWs expect.
- The conversion helper is exposed as
`ASRServiceFactory.convertFloat32ToInt16` so unit tests
can lock the math without instantiating the full pipeline.
Why not change `AudioBufferSnapshot` to `[Int16]` instead
(see earlier first-principles discussion): the snapshot is a
transport format that both `AudioCaptureService` (in the
ext) and `PreviewASRController` (in the main app) produce.
Float32 is the natural shape coming out of `AVAudioEngine`,
and pushing the conversion to the ASR service keeps the
transport contract platform-agnostic — a future second
backend with different format needs can have its own
adaptation without dragging everyone else.
Tests:
- `testFloat32ToInt16EdgeCases` — 0, ±1, ±0.5, ±1.5
(gain-overflow case).
- `testFloat32ToInt16RoundTrip` — quantisation step is
1/32767 (so the asymmetric Int16 range is honoured: -32768
has no exact Float source).
- `testFloat32ToInt16Empty` — `sourceCount == 0` with nil
pointers is a no-op (function guards on count before
dereferencing).
- All 25 tests pass (22 existing + 3 new).
- BUILD SUCCEEDED.
🤖 Generated with Claude Code
|
||
|
|
81581f0e5f |
chore: drop iOS 18–25 / non-iPhone support, require iOS 26
Two related cleanups the user asked for in one shot:
1. iPhone-only is now enforced at every target — Mac Catalyst and
visionOS were never configured in `project.yml`, but the
`OSGKeyboardShared` framework and the two test bundles were
still defaulting to `TARGETED_DEVICE_FAMILY = "1,2"` (iPhone +
iPad). All four targets now explicitly set `"1"`. SDK is
`iphoneos` for everyone, no `xros` / `macosx`.
2. Deployment target bumped from iOS 18.0 to iOS 26.0 across the
board (`project.yml` + the ext's per-target setting). With
iOS 26 as the floor, the iOS 18–25 SFSpeechRecognizer path
became dead code and several `#available` checks became
always-true. Removed:
- `AppleSpeechASR` (the entire SFSpeechRecognizer-based ASR
backend) and the `#available(iOS 26.0, *)` factory branch.
`ASRServiceFactory.make()` now returns `SpeechAnalyzerASR()`
directly. SpeechAnalyzer is always fully on-device, which
also made the `requiresOnDevice` flag meaningless.
- `requiresOnDevice` from the `ASRService.transcribe` protocol
signature, from `ProviderConfig`, `AppGroupStore`,
`KeyboardState`, `AppGroupPersistor`, and the ext's
`KeyboardViewController` (`state.requiresOnDevice`,
`state.setRequiresOnDevice`, `persistRequiresOnDevice`).
- `#available(iOS 17.0, *)` branch in
`PreviewASRController.requestMicrophonePermission` and the
ext's `PermissionManager.requestMicPermission` — both now
just call the iOS 17+ `AVAudioApplication` API directly.
- The `else` (iOS 18–25) branch in `SettingsView.asrEngineRow`
— the on-device-only toggle is gone, the row is a static
"SpeechAnalyzer active" badge. Same for the `else` branch
in `EnginePickerSection.localSubtitle`.
- `makeMicAuthHandler` (the iOS < 17 mic permission callback
wrapper) from `PreviewASRController`.
No `#available` / `@available` checks remain in the codebase
except for the SpeechAnalyzer class itself (now unnecessary
too, but kept for clarity — `AVAudioApplication` and
`SpeechAnalyzer` are both iOS 17+ / iOS 26+ respectively,
and the deployment target of 26 makes the explicit
`@available` redundant; I left the SpeechAnalyzer class
un-`@available` and removed the `@available(iOS 26.0, *)`
decoration since it's no longer needed).
The Keyboard ext's existing ASRService usage
(`asr.transcribe(stream:locale:)`) is unchanged at the
call-site level — just the third argument is gone.
3. Updated `info.plist` UISupportedInterfaceOrientations is
already `[UIInterfaceOrientationPortrait]` only, which is
correct for an iPhone-only app; no change needed.
Build: BUILD SUCCEEDED.
Tests: 22/22 pass.
Verified: `TARGETED_DEVICE_FAMILY = 1` on all four targets,
`SDKROOT = iphoneos` on all four.
🤖 Generated with Claude Code
|
||
|
|
aeb28f4b36 |
fix: green accent + real iOS ASR in preview + bilingual audit
Three user-flagged fixes, scoped tightly to the files each affects.
The uncommitted Engine-mode wiring / locale picker / etc. from a
prior agent pass is intentionally not included in this commit.
1. Revert accent to brand green (#3AA05A).
Last commit flipped `AccentColor` + `Palette.light.accent` to
Apple system blue (#007AFF) on the assumption that "green CTA
on a near-white surface looks wrong." Review pushed back:
the brand *is* the green, and the system tint should match it
so that NavStack Done buttons, Toggles, and our custom
`primaryButton()` modifier all read as the same colour. Restored
`#3AA05A` in both `AccentColor.colorset/Contents.json` and
`Palette.light.accent` (plus the muted / glow variants).
2. Real iOS ASR in `KeyboardPreviewSheet`.
The previous fix only swapped the static placeholder for a
`TextField` and routed a hardcoded stub through it — review
asked, fairly, "are you actually calling `SFSpeechRecognizer`?"
Answer: no. This change makes the preview *run real ASR*:
- `ASRService` (+ the iOS 26 `SpeechAnalyzer` path) moves from
`OSGKeyboardExt/Services/` to `OSGKeyboardShared/Services/`,
so the host app can import the same `ASRServiceFactory.make()`
the keyboard extension uses.
- `OSGKeyboardShared` gains `Speech.framework` and
`AVFoundation.framework` as SDK dependencies in `project.yml`.
- New `OSGKeyboard/Views/PreviewASRController.swift` (the
extension's `AudioCaptureService` is app-extension-only, so
the preview owns its own `AVAudioEngine` + `AVAudioSession`
and downsamples to 16 kHz mono Float32 via `AVAudioConverter`).
- `KeyboardPreviewSheet.cyclePhase()` now calls
`asr.start(locale:)` / `asr.stop()` instead of toggling a
`StubPhase`. The disc's level meter is driven by RMS from the
actual audio tap; the transcript line under the chips shows
the live `SFSpeechRecognizer` partial; the top textbox
receives the `.final` transcript via `onChange(of: lastFinal)`.
3. Record disc in BOTH stubs now uses the green accent for idle.
Was `Color(white: 0.22)` dark-gray, which read as "inert
surface" rather than "tap me". The keyboard extension and the
in-app preview now share the same brand-green disc gradient so
the keyboard's primary CTA is the same colour in both modes.
4. Bilingual audit across every user-facing string.
Every Text() in the main-app and extension views is now
`中文 · English` or carries an English secondary line. Covered:
- OnboardingView (Back, Next, Done, Continue, step instructions,
PrivacyFootnote rows)
- HomeView (status header, hero label, accessibility)
- APISettingsCard (Base URL, API Key, Model, "Get an API key",
"Connection", test-connection states + error messages)
- KeyboardPreviewSheet (title, subtitle, TextField placeholder,
clear button accessibility)
- KeyboardPreviewStub (mode/locale chips, REC badge, space bar)
- KeyboardRootView (gear accessibility, space bar, requesting
state, denied messages, local-engine chip)
- RecordButton (accessibility label)
- SettingsView (Done, Reset dialog, language section subtitle,
engine section, local-engine subtitle, on-device label)
- AppGroupErrorView (title, body, three remediation steps)
- LLMProvider preset names and blurbs
Where the prior pattern was "Chinese headline + English footnote"
(e.g. OnboardingView's 启用 OSGKeyboard / Enable OSGKeyboard),
that pattern was preserved — bilingual coverage means every
screen reads as both, not that every line is rigidly `中 · EN`.
Build: BUILD SUCCEEDED on iPhone 17 Pro / iOS 26 simulator.
Tests: 21/21 pass (no test changes).
Visual: light-mode home shot at /tmp/osgk_light_green.png shows
the brand-green CTA restored across Next button, mic icon, and
page dot.
🤖 Generated with Claude Code
|