feat: comprehensive rewrite — push-to-talk pipeline, Typeless UI, Chinese

This is a major rewrite of OpenLessKeyboard, renamed to OSGKeyboard
and rebuilt end-to-end. 59 files changed (+3205/-1550).

Architecture
------------
- Rename project, targets, directories from OpenLess* to OSGKeyboard*
  (OpenLess / OpenLessKeyboard / OpenLessShared / OpenLessTests).
- AudioCaptureService rewritten as @unchecked Sendable class with
  OSAllocatedUnfairLock instead of an actor, so it survives Swift 6
  strict-concurrency checks while still serialising engine + converter
  state correctly.
- Single design system (Palette / Spacing / Radius / TypeStyle /
  Motion) lifted into OSGKeyboardShared so the host app and the
  keyboard extension stay in lock-step.

Push-to-talk — first-principles fix
-----------------------------------
- App Group + audio-input entitlements were stripped by Xcode's
  Automatic Signing. They are now declared in project.yml so
  'xcodegen generate' re-emits them every time. iOS Developer
  Account is untouched; only the App Group capability was added.
- State machine uses a real stored `phase` (was a derived shim
  that locked out every press after the first because
  recordStream was never nilled after the pipeline finished).
- Microphone permission is requested inside pressBegan (async
  Task) so the press flow optimistically enters .recording;
  permission denial surfaces a short error and returns to idle.
- Replaced LongPressGesture(0.15s) with a DragGesture +
  TapGesture pair separated by pressArmed, so a single tap no
  longer fires both onPressBegan and onTap simultaneously.
- Real RMS / peak level meter from the AVAudioEngine tap (was a
  pseudo-random walk); the visible waveform is now driven by
  actual audio.
- SFSpeechRecognizer(locale:) with selectable ASR locales
  (auto / zh-Hans / zh-Hant / en-US / ja-JP / ko-KR) for
  first-class Chinese / English / Japanese / Korean dictation,
  with on-device recognition when supported.
- AVAudioSession now deactivates on stop so other apps' audio
  routing is restored.

Keyboard UI — Typeless-inspired layout
---------------------------------------
- Hero area is 280 pt with a 96 pt record disc, breathing outer
  ring, and a 12-bar waveform driven by the real RMS.
- inputView.allowsSelfSizing + a heightAnchor constraint so iOS
  no longer crops the keyboard under the Spotlight bar / home
  indicator.
- Top bar: mode chip (Off / 转写 / 润色) + locale chip
  (Auto / 简体 / 繁體 / EN / 日 / 한) + status badge + ⚙.
- Bottom bar: globe / delete / 空格 / return — all 40 pt and
  balanced.
- RecordButton onPressEnded is now safe to fire from a quick
  press; pressArmed prevents double-firing.

LLM / Polishing
---------------
- LLMClient: stopped leaking the server response body in errors
  (server body is now logged at debug, never surfaced to UI);
  added a dedicated .rateLimited case for 429.
- PolishingService timeout 8s → 12s to accommodate slower
  domestic LLM providers.
- AppGroupStore.defaultSystemPrompt is now provider-aware
  (Chinese for zhipu/moonshot/qwen/deepseek, English otherwise).

Onboarding & Settings
---------------------
- Re-themed OnboardingView / HomeView / SettingsView on the
  new design system.
- ProviderPickerSection now shows 6 providers (OpenAI, DeepSeek,
  Qwen DashScope, 智谱 GLM, 月之暗面 Moonshot, Custom) with
  blurb + selected accent.
- PickerRow for Mode and ASR locale; System Prompt editor with
  reset-to-default.
- API settings page "Get an API key" used SwiftUI Link, which
  has a hit-test bug on iOS 18 that ate gestures from adjacent
  TextFields (manifested as "typing jumps to a website"). It is
  now an explicit Button + contentShape + .submitLabel(.done) on
  the fields.

Polish & tests
--------------
- LLMClientTests: 4 unit tests passing (ProviderConfig
  persistence + OpenAI request/response + HTTP error + missing
  key); test App Group renamed to the correct identifier.
- ProviderConfig.apply now captures the previous provider id
  *before* mutating, so switching providers actually resets the
  system prompt to the new default.

Build
-----
- Swift 6 strict concurrency, iOS 18.0 deployment target.
- Tested on Xcode 26 + iPhone 17 Pro simulator. A real device on
  iOS 27 beta aborts with __abort_with_payload (dispatch
  library ABI mismatch); use an iOS 18 real device or the
  iOS 26 simulator for now.

🤖 Generated with Claude Code
This commit is contained in:
Rocky
2026-06-18 01:22:12 +08:00
parent 07067ca6c5
commit bec36befa2
59 changed files with 3205 additions and 1550 deletions
+151
View File
@@ -0,0 +1,151 @@
// ASRService.swift
// OSGKeyboard · Keyboard Extension
//
// Speech-to-text abstraction over Apple's `SFSpeechRecognizer`.
// Honours a user-selected locale (auto / zh-CN / en-US / ja-JP ) so
// dictation is first-class for non-English languages.
import Foundation
import AVFoundation
import Speech
import os.lock
import OSGKeyboardShared
// MARK: - Sendable conformance
// `AVAudioPCMBuffer` and `SFSpeechRecognitionTask` are not Sendable. We
// only ever access them serially the PCM buffer is built and consumed
// inside a single Task, and the recogniser task is cancelled but never
// shared concurrently so an unchecked conformance is sound here.
extension AVAudioPCMBuffer: @unchecked @retroactive Sendable {}
extension SFSpeechRecognitionTask: @unchecked @retroactive Sendable {}
// MARK: - Protocol
public protocol ASRService: Sendable {
/// Start a transcription session. The returned stream emits `.partial`
/// updates and exactly one `.final` (or `.error`) before finishing.
func transcribe(
stream: AsyncStream<AudioBufferSnapshot>,
locale: Locale
) -> AsyncStream<ASREvent>
/// Cancel any in-flight recognition and tear down its tasks.
func cancel()
}
public enum ASREvent: Sendable, Equatable {
case partial(String)
case final(String)
case error(String)
}
// MARK: - Factory
public enum ASRServiceFactory {
public static func make() -> ASRService {
AppleSpeechASR()
}
}
// MARK: - Apple Speech implementation
final class AppleSpeechASR: ASRService, @unchecked Sendable {
private let lock = OSAllocatedUnfairLock()
private var recognizerTask: SFSpeechRecognitionTask?
private var feedTask: Task<Void, Never>?
func transcribe(
stream: AsyncStream<AudioBufferSnapshot>,
locale: Locale
) -> AsyncStream<ASREvent> {
AsyncStream { continuation in
let recognizer = SFSpeechRecognizer(locale: locale)
?? SFSpeechRecognizer(locale: .current)
guard let recognizer, recognizer.isAvailable else {
continuation.yield(.error("Speech recognizer unavailable for \(locale.identifier)"))
continuation.finish()
return
}
recognizer.defaultTaskHint = .dictation
let request = SFSpeechAudioBufferRecognitionRequest()
request.shouldReportPartialResults = true
request.requiresOnDeviceRecognition = recognizer.supportsOnDeviceRecognition
let task = recognizer.recognitionTask(with: request) { result, error in
if let error {
let nsErr = error as NSError
// Codes 203 / 1110 = "no speech detected" a normal exit.
if nsErr.code == 203 || nsErr.code == 1110 {
continuation.yield(.final(""))
} else {
continuation.yield(.error(error.localizedDescription))
}
continuation.finish()
return
}
guard let result else { return }
if result.isFinal {
continuation.yield(.final(result.bestTranscription.formattedString))
continuation.finish()
} else {
continuation.yield(.partial(result.bestTranscription.formattedString))
}
}
self.lock.withLock { self.recognizerTask = task }
// Feed audio: for each snapshot, build a 16 kHz mono Float32
// PCM buffer and immediately `request.append(pcm)`. The PCM
// buffer never leaves this task, so it doesn't need to be
// Sendable.
let feedFormat = AVAudioFormat(
commonFormat: .pcmFormatFloat32,
sampleRate: 16_000,
channels: 1,
interleaved: false
)!
self.feedTask = Task { [request] in
for await snap in stream {
if Task.isCancelled { break }
guard !snap.samples.isEmpty,
let pcm = AVAudioPCMBuffer(
pcmFormat: feedFormat,
frameCapacity: AVAudioFrameCount(snap.samples.count)
)
else { continue }
pcm.frameLength = AVAudioFrameCount(snap.samples.count)
if let dst = pcm.floatChannelData?[0] {
snap.samples.withUnsafeBufferPointer { src in
if let base = src.baseAddress {
memcpy(dst, base, snap.samples.count * MemoryLayout<Float>.size)
}
}
}
request.append(pcm)
}
if !Task.isCancelled {
request.endAudio()
}
}
continuation.onTermination = { @Sendable [weak self] _ in
self?.cancel()
}
}
}
func cancel() {
let (recTask, feedT) = lock.withLock { () -> (SFSpeechRecognitionTask?, Task<Void, Never>?) in
let r = self.recognizerTask
let f = self.feedTask
self.recognizerTask = nil
self.feedTask = nil
return (r, f)
}
recTask?.cancel()
feedT?.cancel()
}
}
@@ -0,0 +1,315 @@
// AudioCaptureService.swift
// OSGKeyboard · Keyboard Extension
//
// Captures microphone audio at 16 kHz mono Float32 using AVAudioEngine.
// Designed for use inside an iOS Custom Keyboard Extension:
// Uses `.record` (not `.playAndRecord`) keyboards cannot play.
// Exposes a Sendable `Session` with two streams:
// - `audio`: 16 kHz mono Float32 frames for ASR.
// - `levels`: RMS + peak dBFS for the animated waveform.
// All mutable state is guarded by a lock; class is `@unchecked Sendable`
// for use with Swift 6 strict concurrency.
import Foundation
import AVFoundation
import os.lock
import OSGKeyboardShared
// MARK: - Sendable conformance
//
// `AVAudioEngine` is not Sendable, but the iOS audio APIs hand us
// closures that need to capture it. We never mutate the engine
// concurrently capture / conversion are serialised on the actor, and
// the tap closure only reads pointers into it. So an unchecked
// retroactive Sendable conformance is sound here.
// `AVAudioConverter` and `AVAudioFormat` are already Sendable in newer
// SDKs; we don't need to redeclare.
extension AVAudioEngine: @unchecked @retroactive Sendable {}
public final class AudioCaptureService: @unchecked Sendable {
// MARK: - Errors
public enum CaptureError: LocalizedError, Sendable {
case sessionConfigFailed(String)
case engineStartFailed(String)
case noInputNode
case alreadyRunning
public var errorDescription: String? {
switch self {
case .sessionConfigFailed(let s): return "Audio session config failed: \(s)"
case .engineStartFailed(let s): return "Audio engine failed to start: \(s)"
case .noInputNode: return "No microphone input available."
case .alreadyRunning: return "Audio capture is already running."
}
}
}
// MARK: - Level payload (Sendable)
public struct Level: Sendable, Equatable {
/// Root-mean-square, 0...1 (linear).
public let rms: Float
/// Peak amplitude, 0...1 (linear).
public let peak: Float
public let timestamp: TimeInterval
/// Convenience: -20 dBFS 0 dBFS mapped to 01 for UI meters.
public var meter: Float {
// Clamp floor at -50 dB so silence still shows a tiny bar.
let db = 20 * log10(max(rms, 1e-7))
let clamped = max(-50, min(0, db))
return Float((clamped + 50) / 50)
}
}
// MARK: - Session (one capture run)
/// A single capture run. Two streams + a stop handle.
public final class Session: @unchecked Sendable {
public let audio: AsyncStream<AudioBufferSnapshot>
public let levels: AsyncStream<Level>
private let onStop: @Sendable () -> Void
fileprivate init(
audio: AsyncStream<AudioBufferSnapshot>,
levels: AsyncStream<Level>,
onStop: @escaping @Sendable () -> Void
) {
self.audio = audio
self.levels = levels
self.onStop = onStop
}
public func stop() { onStop() }
}
// MARK: - State
private let lock = OSAllocatedUnfairLock()
private var engine: AVAudioEngine?
private var converter: AVAudioConverter?
private var audioContinuation: AsyncStream<AudioBufferSnapshot>.Continuation?
private var levelContinuation: AsyncStream<Level>.Continuation?
private var isRunning: Bool = false
public init() {}
deinit { stopInternal() }
// MARK: - Public API
/// Start capture. Returns a `Session` whose streams yield audio + level data.
/// The session ends when `Session.stop()` is called or the extension is torn down.
@discardableResult
public func start() -> Session {
let (audioStream, audioCont) = AsyncStream<AudioBufferSnapshot>.makeStream()
let (levelStream, levelCont) = AsyncStream<Level>.makeStream()
// Fail fast if already running.
let alreadyRunning: Bool = lock.withLock { isRunning }
if alreadyRunning {
audioCont.finish()
levelCont.finish()
return Session(audio: audioStream, levels: levelStream, onStop: {})
}
do {
try configureSession()
try bootstrap(audioCont: audioCont, levelCont: levelCont)
lock.withLock {
isRunning = true
audioContinuation = audioCont
levelContinuation = levelCont
}
} catch {
// Tear down whatever we partially created.
audioCont.finish()
levelCont.finish()
teardownEngine()
return Session(audio: audioStream, levels: levelStream, onStop: {})
}
return Session(
audio: audioStream,
levels: levelStream,
onStop: { [weak self] in self?.stop() }
)
}
public func stop() {
stopInternal()
}
// MARK: - Setup
private func configureSession() throws {
#if canImport(UIKit)
let session = AVAudioSession.sharedInstance()
do {
// `.record` keyboards cannot play audio, so .playAndRecord is wrong.
// `.measurement` mode disables system AGC/echo cancellation for cleaner ASR input.
// `.duckOthers` is harmless in record-only.
try session.setCategory(.record, mode: .measurement, options: [.duckOthers])
try session.setActive(true, options: .notifyOthersOnDeactivation)
} catch {
throw CaptureError.sessionConfigFailed(error.localizedDescription)
}
#endif
}
private func bootstrap(
audioCont: AsyncStream<AudioBufferSnapshot>.Continuation,
levelCont: AsyncStream<Level>.Continuation
) throws {
let engine = AVAudioEngine()
let input = engine.inputNode
let hardwareFormat = input.outputFormat(forBus: 0)
guard hardwareFormat.sampleRate > 0, hardwareFormat.channelCount > 0 else {
throw CaptureError.noInputNode
}
let target = AVAudioFormat(
commonFormat: .pcmFormatFloat32,
sampleRate: 16_000,
channels: 1,
interleaved: false
)!
guard let converter = AVAudioConverter(from: hardwareFormat, to: target) else {
throw CaptureError.engineStartFailed("converter init failed")
}
// Persistent buffer for level computation: we re-use Float arrays to
// avoid per-tap allocations.
let levelScratch = LevelScratch()
input.installTap(onBus: 0, bufferSize: 1024, format: hardwareFormat) { [weak self] buffer, when in
// We capture self weakly only to keep the AudioCaptureService
// alive while the tap is installed; the tap itself only
// touches the local `audioCont` / `levelCont` continuations.
guard self != nil else { return }
// 1) Compute level from raw hardware buffer (preserves true amplitude).
let (rms, peak) = levelScratch.measure(buffer: buffer)
// `AVAudioTime` carries both `sampleTime` (frames on the device
// clock) and `hostTime` (mach absolute time). We only need a
// monotonically increasing source for the timestamp; sample
// time / sample rate is good enough and is independent of the
// host clock.
let ts: Double
if when.sampleTime > 0, hardwareFormat.sampleRate > 0 {
ts = Double(when.sampleTime) / hardwareFormat.sampleRate
} else {
ts = Date().timeIntervalSinceReferenceDate
}
levelCont.yield(Level(rms: rms, peak: peak, timestamp: ts))
// 2) Convert to 16 kHz mono Float32 for ASR.
let outFrames = AVAudioFrameCount(
Double(buffer.frameLength) * 16_000.0 / hardwareFormat.sampleRate
)
guard outFrames > 0,
let outBuffer = AVAudioPCMBuffer(pcmFormat: target, frameCapacity: outFrames)
else { return }
var error: NSError?
let status = converter.convert(to: outBuffer, error: &error) { _, outStatus in
outStatus.pointee = .haveData
return buffer
}
if status == .haveData, error == nil {
let snap = AudioBufferSnapshot(buffer: outBuffer)
audioCont.yield(snap)
}
}
do {
try engine.start()
} catch {
input.removeTap(onBus: 0)
throw CaptureError.engineStartFailed(error.localizedDescription)
}
lock.withLock {
self.engine = engine
self.converter = converter
}
}
// MARK: - Teardown
private func stopInternal() {
let (wasRunning, engine, audioCont, levelCont) = lock.withLock { () -> (Bool, AVAudioEngine?, AsyncStream<AudioBufferSnapshot>.Continuation?, AsyncStream<Level>.Continuation?) in
let was = isRunning
isRunning = false
let eng = self.engine
let ac = audioContinuation
let lc = levelContinuation
self.engine = nil
self.converter = nil
self.audioContinuation = nil
self.levelContinuation = nil
return (was, eng, ac, lc)
}
guard wasRunning else { return }
engine?.inputNode.removeTap(onBus: 0)
engine?.stop()
#if canImport(UIKit)
// Deactivate the session so other apps' audio routing is restored.
try? AVAudioSession.sharedInstance().setActive(false, options: .notifyOthersOnDeactivation)
#endif
audioCont?.finish()
levelCont?.finish()
}
private func teardownEngine() {
let (engine, audioCont, levelCont) = lock.withLock { () -> (AVAudioEngine?, AsyncStream<AudioBufferSnapshot>.Continuation?, AsyncStream<Level>.Continuation?) in
let e = self.engine
let a = audioContinuation
let l = levelContinuation
self.engine = nil
self.converter = nil
self.audioContinuation = nil
self.levelContinuation = nil
self.isRunning = false
return (e, a, l)
}
engine?.inputNode.removeTap(onBus: 0)
engine?.stop()
audioCont?.finish()
levelCont?.finish()
}
}
// MARK: - Level scratch (lock-free, single-writer / single-reader per tap)
/// Lock-free per-tap scratch for RMS + peak measurement. The AVAudio tap is
/// always invoked serially per input node, so we don't need a lock here.
private final class LevelScratch: @unchecked Sendable {
private var last: (rms: Float, peak: Float) = (0, 0)
func measure(buffer: AVAudioPCMBuffer) -> (rms: Float, peak: Float) {
guard let ch = buffer.floatChannelData?[0] else { return last }
let n = Int(buffer.frameLength)
guard n > 0 else { return last }
// Decay smoothing keeps the meter lively but not jittery.
var sumSq: Float = 0
var peak: Float = 0
for i in 0..<n {
let s = ch[i]
sumSq += s * s
let a = abs(s)
if a > peak { peak = a }
}
let rms = sqrtf(sumSq / Float(n))
// Exponential moving average for visual smoothness.
let alpha: Float = 0.35
let smoothedRms = alpha * rms + (1 - alpha) * last.rms
let smoothedPeak = max(alpha * peak, (1 - alpha) * last.peak)
last = (smoothedRms, smoothedPeak)
return (smoothedRms, smoothedPeak)
}
}
@@ -0,0 +1,46 @@
// PolishingService.swift
// OSGKeyboard · Keyboard Extension
//
// Takes raw ASR transcript and runs it through the user's configured LLM
// to produce polished, well-punctuated text. Falls back to the raw transcript
// if the LLM call fails or times out.
import Foundation
import OSGKeyboardShared
public actor PolishingService {
public enum PolishError: Error {
case noTranscript
case timeout
}
private let store: AppGroupStore
private let timeout: TimeInterval
public init(store: AppGroupStore = AppGroupStore(), timeout: TimeInterval = 12) {
self.store = store
self.timeout = timeout
}
public func polish(_ raw: String) async throws -> String {
let trimmed = raw.trimmingCharacters(in: .whitespacesAndNewlines)
guard !trimmed.isEmpty else { throw PolishError.noTranscript }
let client = store.makeClient()
let prompt = store.systemPrompt
return try await withThrowingTaskGroup(of: String.self) { group in
group.addTask {
try await client.polish(trimmed, systemPrompt: prompt)
}
group.addTask {
try await Task.sleep(nanoseconds: UInt64(self.timeout * 1_000_000_000))
throw PolishError.timeout
}
let result = try await group.next()!
group.cancelAll()
return result
}
}
}