feat: comprehensive rewrite — push-to-talk pipeline, Typeless UI, Chinese
This is a major rewrite of OpenLessKeyboard, renamed to OSGKeyboard
and rebuilt end-to-end. 59 files changed (+3205/-1550).
Architecture
------------
- Rename project, targets, directories from OpenLess* to OSGKeyboard*
(OpenLess / OpenLessKeyboard / OpenLessShared / OpenLessTests).
- AudioCaptureService rewritten as @unchecked Sendable class with
OSAllocatedUnfairLock instead of an actor, so it survives Swift 6
strict-concurrency checks while still serialising engine + converter
state correctly.
- Single design system (Palette / Spacing / Radius / TypeStyle /
Motion) lifted into OSGKeyboardShared so the host app and the
keyboard extension stay in lock-step.
Push-to-talk — first-principles fix
-----------------------------------
- App Group + audio-input entitlements were stripped by Xcode's
Automatic Signing. They are now declared in project.yml so
'xcodegen generate' re-emits them every time. iOS Developer
Account is untouched; only the App Group capability was added.
- State machine uses a real stored `phase` (was a derived shim
that locked out every press after the first because
recordStream was never nilled after the pipeline finished).
- Microphone permission is requested inside pressBegan (async
Task) so the press flow optimistically enters .recording;
permission denial surfaces a short error and returns to idle.
- Replaced LongPressGesture(0.15s) with a DragGesture +
TapGesture pair separated by pressArmed, so a single tap no
longer fires both onPressBegan and onTap simultaneously.
- Real RMS / peak level meter from the AVAudioEngine tap (was a
pseudo-random walk); the visible waveform is now driven by
actual audio.
- SFSpeechRecognizer(locale:) with selectable ASR locales
(auto / zh-Hans / zh-Hant / en-US / ja-JP / ko-KR) for
first-class Chinese / English / Japanese / Korean dictation,
with on-device recognition when supported.
- AVAudioSession now deactivates on stop so other apps' audio
routing is restored.
Keyboard UI — Typeless-inspired layout
---------------------------------------
- Hero area is 280 pt with a 96 pt record disc, breathing outer
ring, and a 12-bar waveform driven by the real RMS.
- inputView.allowsSelfSizing + a heightAnchor constraint so iOS
no longer crops the keyboard under the Spotlight bar / home
indicator.
- Top bar: mode chip (Off / 转写 / 润色) + locale chip
(Auto / 简体 / 繁體 / EN / 日 / 한) + status badge + ⚙.
- Bottom bar: globe / delete / 空格 / return — all 40 pt and
balanced.
- RecordButton onPressEnded is now safe to fire from a quick
press; pressArmed prevents double-firing.
LLM / Polishing
---------------
- LLMClient: stopped leaking the server response body in errors
(server body is now logged at debug, never surfaced to UI);
added a dedicated .rateLimited case for 429.
- PolishingService timeout 8s → 12s to accommodate slower
domestic LLM providers.
- AppGroupStore.defaultSystemPrompt is now provider-aware
(Chinese for zhipu/moonshot/qwen/deepseek, English otherwise).
Onboarding & Settings
---------------------
- Re-themed OnboardingView / HomeView / SettingsView on the
new design system.
- ProviderPickerSection now shows 6 providers (OpenAI, DeepSeek,
Qwen DashScope, 智谱 GLM, 月之暗面 Moonshot, Custom) with
blurb + selected accent.
- PickerRow for Mode and ASR locale; System Prompt editor with
reset-to-default.
- API settings page "Get an API key" used SwiftUI Link, which
has a hit-test bug on iOS 18 that ate gestures from adjacent
TextFields (manifested as "typing jumps to a website"). It is
now an explicit Button + contentShape + .submitLabel(.done) on
the fields.
Polish & tests
--------------
- LLMClientTests: 4 unit tests passing (ProviderConfig
persistence + OpenAI request/response + HTTP error + missing
key); test App Group renamed to the correct identifier.
- ProviderConfig.apply now captures the previous provider id
*before* mutating, so switching providers actually resets the
system prompt to the new default.
Build
-----
- Swift 6 strict concurrency, iOS 18.0 deployment target.
- Tested on Xcode 26 + iPhone 17 Pro simulator. A real device on
iOS 27 beta aborts with __abort_with_payload (dispatch
library ABI mismatch); use an iOS 18 real device or the
iOS 26 simulator for now.
🤖 Generated with Claude Code
This commit is contained in:
@@ -0,0 +1,151 @@
|
||||
// ASRService.swift
|
||||
// OSGKeyboard · Keyboard Extension
|
||||
//
|
||||
// Speech-to-text abstraction over Apple's `SFSpeechRecognizer`.
|
||||
// Honours a user-selected locale (auto / zh-CN / en-US / ja-JP …) so
|
||||
// dictation is first-class for non-English languages.
|
||||
|
||||
import Foundation
|
||||
import AVFoundation
|
||||
import Speech
|
||||
import os.lock
|
||||
import OSGKeyboardShared
|
||||
|
||||
// MARK: - Sendable conformance
|
||||
|
||||
// `AVAudioPCMBuffer` and `SFSpeechRecognitionTask` are not Sendable. We
|
||||
// only ever access them serially — the PCM buffer is built and consumed
|
||||
// inside a single Task, and the recogniser task is cancelled but never
|
||||
// shared concurrently — so an unchecked conformance is sound here.
|
||||
extension AVAudioPCMBuffer: @unchecked @retroactive Sendable {}
|
||||
extension SFSpeechRecognitionTask: @unchecked @retroactive Sendable {}
|
||||
|
||||
// MARK: - Protocol
|
||||
|
||||
public protocol ASRService: Sendable {
|
||||
/// Start a transcription session. The returned stream emits `.partial`
|
||||
/// updates and exactly one `.final` (or `.error`) before finishing.
|
||||
func transcribe(
|
||||
stream: AsyncStream<AudioBufferSnapshot>,
|
||||
locale: Locale
|
||||
) -> AsyncStream<ASREvent>
|
||||
|
||||
/// Cancel any in-flight recognition and tear down its tasks.
|
||||
func cancel()
|
||||
}
|
||||
|
||||
public enum ASREvent: Sendable, Equatable {
|
||||
case partial(String)
|
||||
case final(String)
|
||||
case error(String)
|
||||
}
|
||||
|
||||
// MARK: - Factory
|
||||
|
||||
public enum ASRServiceFactory {
|
||||
public static func make() -> ASRService {
|
||||
AppleSpeechASR()
|
||||
}
|
||||
}
|
||||
|
||||
// MARK: - Apple Speech implementation
|
||||
|
||||
final class AppleSpeechASR: ASRService, @unchecked Sendable {
|
||||
|
||||
private let lock = OSAllocatedUnfairLock()
|
||||
private var recognizerTask: SFSpeechRecognitionTask?
|
||||
private var feedTask: Task<Void, Never>?
|
||||
|
||||
func transcribe(
|
||||
stream: AsyncStream<AudioBufferSnapshot>,
|
||||
locale: Locale
|
||||
) -> AsyncStream<ASREvent> {
|
||||
AsyncStream { continuation in
|
||||
let recognizer = SFSpeechRecognizer(locale: locale)
|
||||
?? SFSpeechRecognizer(locale: .current)
|
||||
guard let recognizer, recognizer.isAvailable else {
|
||||
continuation.yield(.error("Speech recognizer unavailable for \(locale.identifier)"))
|
||||
continuation.finish()
|
||||
return
|
||||
}
|
||||
recognizer.defaultTaskHint = .dictation
|
||||
|
||||
let request = SFSpeechAudioBufferRecognitionRequest()
|
||||
request.shouldReportPartialResults = true
|
||||
request.requiresOnDeviceRecognition = recognizer.supportsOnDeviceRecognition
|
||||
|
||||
let task = recognizer.recognitionTask(with: request) { result, error in
|
||||
if let error {
|
||||
let nsErr = error as NSError
|
||||
// Codes 203 / 1110 = "no speech detected" — a normal exit.
|
||||
if nsErr.code == 203 || nsErr.code == 1110 {
|
||||
continuation.yield(.final(""))
|
||||
} else {
|
||||
continuation.yield(.error(error.localizedDescription))
|
||||
}
|
||||
continuation.finish()
|
||||
return
|
||||
}
|
||||
guard let result else { return }
|
||||
if result.isFinal {
|
||||
continuation.yield(.final(result.bestTranscription.formattedString))
|
||||
continuation.finish()
|
||||
} else {
|
||||
continuation.yield(.partial(result.bestTranscription.formattedString))
|
||||
}
|
||||
}
|
||||
|
||||
self.lock.withLock { self.recognizerTask = task }
|
||||
|
||||
// Feed audio: for each snapshot, build a 16 kHz mono Float32
|
||||
// PCM buffer and immediately `request.append(pcm)`. The PCM
|
||||
// buffer never leaves this task, so it doesn't need to be
|
||||
// Sendable.
|
||||
let feedFormat = AVAudioFormat(
|
||||
commonFormat: .pcmFormatFloat32,
|
||||
sampleRate: 16_000,
|
||||
channels: 1,
|
||||
interleaved: false
|
||||
)!
|
||||
self.feedTask = Task { [request] in
|
||||
for await snap in stream {
|
||||
if Task.isCancelled { break }
|
||||
guard !snap.samples.isEmpty,
|
||||
let pcm = AVAudioPCMBuffer(
|
||||
pcmFormat: feedFormat,
|
||||
frameCapacity: AVAudioFrameCount(snap.samples.count)
|
||||
)
|
||||
else { continue }
|
||||
pcm.frameLength = AVAudioFrameCount(snap.samples.count)
|
||||
if let dst = pcm.floatChannelData?[0] {
|
||||
snap.samples.withUnsafeBufferPointer { src in
|
||||
if let base = src.baseAddress {
|
||||
memcpy(dst, base, snap.samples.count * MemoryLayout<Float>.size)
|
||||
}
|
||||
}
|
||||
}
|
||||
request.append(pcm)
|
||||
}
|
||||
if !Task.isCancelled {
|
||||
request.endAudio()
|
||||
}
|
||||
}
|
||||
|
||||
continuation.onTermination = { @Sendable [weak self] _ in
|
||||
self?.cancel()
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
func cancel() {
|
||||
let (recTask, feedT) = lock.withLock { () -> (SFSpeechRecognitionTask?, Task<Void, Never>?) in
|
||||
let r = self.recognizerTask
|
||||
let f = self.feedTask
|
||||
self.recognizerTask = nil
|
||||
self.feedTask = nil
|
||||
return (r, f)
|
||||
}
|
||||
recTask?.cancel()
|
||||
feedT?.cancel()
|
||||
}
|
||||
}
|
||||
Reference in New Issue
Block a user