fix: green accent + real iOS ASR in preview + bilingual audit
Three user-flagged fixes, scoped tightly to the files each affects.
The uncommitted Engine-mode wiring / locale picker / etc. from a
prior agent pass is intentionally not included in this commit.
1. Revert accent to brand green (#3AA05A).
Last commit flipped `AccentColor` + `Palette.light.accent` to
Apple system blue (#007AFF) on the assumption that "green CTA
on a near-white surface looks wrong." Review pushed back:
the brand *is* the green, and the system tint should match it
so that NavStack Done buttons, Toggles, and our custom
`primaryButton()` modifier all read as the same colour. Restored
`#3AA05A` in both `AccentColor.colorset/Contents.json` and
`Palette.light.accent` (plus the muted / glow variants).
2. Real iOS ASR in `KeyboardPreviewSheet`.
The previous fix only swapped the static placeholder for a
`TextField` and routed a hardcoded stub through it — review
asked, fairly, "are you actually calling `SFSpeechRecognizer`?"
Answer: no. This change makes the preview *run real ASR*:
- `ASRService` (+ the iOS 26 `SpeechAnalyzer` path) moves from
`OSGKeyboardExt/Services/` to `OSGKeyboardShared/Services/`,
so the host app can import the same `ASRServiceFactory.make()`
the keyboard extension uses.
- `OSGKeyboardShared` gains `Speech.framework` and
`AVFoundation.framework` as SDK dependencies in `project.yml`.
- New `OSGKeyboard/Views/PreviewASRController.swift` (the
extension's `AudioCaptureService` is app-extension-only, so
the preview owns its own `AVAudioEngine` + `AVAudioSession`
and downsamples to 16 kHz mono Float32 via `AVAudioConverter`).
- `KeyboardPreviewSheet.cyclePhase()` now calls
`asr.start(locale:)` / `asr.stop()` instead of toggling a
`StubPhase`. The disc's level meter is driven by RMS from the
actual audio tap; the transcript line under the chips shows
the live `SFSpeechRecognizer` partial; the top textbox
receives the `.final` transcript via `onChange(of: lastFinal)`.
3. Record disc in BOTH stubs now uses the green accent for idle.
Was `Color(white: 0.22)` dark-gray, which read as "inert
surface" rather than "tap me". The keyboard extension and the
in-app preview now share the same brand-green disc gradient so
the keyboard's primary CTA is the same colour in both modes.
4. Bilingual audit across every user-facing string.
Every Text() in the main-app and extension views is now
`中文 · English` or carries an English secondary line. Covered:
- OnboardingView (Back, Next, Done, Continue, step instructions,
PrivacyFootnote rows)
- HomeView (status header, hero label, accessibility)
- APISettingsCard (Base URL, API Key, Model, "Get an API key",
"Connection", test-connection states + error messages)
- KeyboardPreviewSheet (title, subtitle, TextField placeholder,
clear button accessibility)
- KeyboardPreviewStub (mode/locale chips, REC badge, space bar)
- KeyboardRootView (gear accessibility, space bar, requesting
state, denied messages, local-engine chip)
- RecordButton (accessibility label)
- SettingsView (Done, Reset dialog, language section subtitle,
engine section, local-engine subtitle, on-device label)
- AppGroupErrorView (title, body, three remediation steps)
- LLMProvider preset names and blurbs
Where the prior pattern was "Chinese headline + English footnote"
(e.g. OnboardingView's 启用 OSGKeyboard / Enable OSGKeyboard),
that pattern was preserved — bilingual coverage means every
screen reads as both, not that every line is rigidly `中 · EN`.
Build: BUILD SUCCEEDED on iPhone 17 Pro / iOS 26 simulator.
Tests: 21/21 pass (no test changes).
Visual: light-mode home shot at /tmp/osgk_light_green.png shows
the brand-green CTA restored across Next button, mic icon, and
page dot.
🤖 Generated with Claude Code
This commit is contained in:
@@ -98,23 +98,20 @@ public enum Palette {
|
||||
/// Light palette — iOS system light mode defaults. Used by the main app
|
||||
/// when the user is in light mode; the keyboard extension stays dark.
|
||||
///
|
||||
/// Accent: Apple system blue (#007AFF), matching the AccentColor asset
|
||||
/// and the iOS HIG default. We deliberately do NOT use the dark-mode
|
||||
/// green here — a bright green CTA on a near-white background reads as
|
||||
/// "go to a garden centre" rather than "tap me to enable your
|
||||
/// keyboard", and a Typeless/Apple-style design system expects the
|
||||
/// accent to follow the system tint in light mode. The keyboard
|
||||
/// extension always renders dark and keeps its green accent for the
|
||||
/// "polish / on-device" affordances, where green-on-dark is the more
|
||||
/// legible pairing.
|
||||
/// Accent is the same brand green (#3AA05A) used in the dark palette —
|
||||
/// "one accent" is core to the design system, and the recording-state
|
||||
/// CTA needs to read as the same brand colour in both modes. The
|
||||
/// keyboard's record disc also picks up `palette.accent` (see
|
||||
/// `KeyboardPreviewStub.recordDisc`), so this single token drives
|
||||
/// every interactive surface.
|
||||
public static let light = ThemePalette(
|
||||
background: Color(red: 0.980, green: 0.980, blue: 0.988), // #FAFAFC
|
||||
surface: Color(red: 1.000, green: 1.000, blue: 1.000), // #FFFFFF
|
||||
surfaceElevated: Color(red: 0.941, green: 0.941, blue: 0.961), // #F0F0F5
|
||||
surfaceMuted: Color(red: 0.953, green: 0.953, blue: 0.965), // #F3F3F6
|
||||
accent: Color(red: 0.000, green: 0.478, blue: 1.000), // #007AFF
|
||||
accentMuted: Color(red: 0.000, green: 0.478, blue: 1.000).opacity(0.12),
|
||||
accentGlow: Color(red: 0.000, green: 0.478, blue: 1.000).opacity(0.28),
|
||||
accent: Color(red: 0.227, green: 0.627, blue: 0.353), // #3AA05A
|
||||
accentMuted: Color(red: 0.227, green: 0.627, blue: 0.353).opacity(0.14),
|
||||
accentGlow: Color(red: 0.227, green: 0.627, blue: 0.353).opacity(0.32),
|
||||
danger: Color(red: 1.000, green: 0.231, blue: 0.188), // #FF3B30
|
||||
success: Color(red: 0.157, green: 0.812, blue: 0.412), // #28CF69
|
||||
warning: Color(red: 1.000, green: 0.620, blue: 0.094), // #FF9E18
|
||||
|
||||
@@ -38,7 +38,7 @@ public struct LLMProvider: Identifiable, Codable, Hashable, Sendable {
|
||||
defaultBaseURL: "https://api.openai.com/v1",
|
||||
defaultModel: "gpt-4o-mini",
|
||||
apiKeyURL: URL(string: "https://platform.openai.com/api-keys"),
|
||||
blurb: "GPT-4o mini · 多语言"
|
||||
blurb: "GPT-4o mini · 多语言 · Multilingual"
|
||||
),
|
||||
.init(
|
||||
id: "deepseek",
|
||||
@@ -46,7 +46,7 @@ public struct LLMProvider: Identifiable, Codable, Hashable, Sendable {
|
||||
defaultBaseURL: "https://api.deepseek.com/v1",
|
||||
defaultModel: "deepseek-chat",
|
||||
apiKeyURL: URL(string: "https://platform.deepseek.com/api_keys"),
|
||||
blurb: "deepseek-chat · 中文友好"
|
||||
blurb: "deepseek-chat · 中文友好 · Chinese-friendly"
|
||||
),
|
||||
.init(
|
||||
id: "qwen",
|
||||
@@ -54,15 +54,15 @@ public struct LLMProvider: Identifiable, Codable, Hashable, Sendable {
|
||||
defaultBaseURL: "https://dashscope.aliyuncs.com/compatible-mode/v1",
|
||||
defaultModel: "qwen-plus",
|
||||
apiKeyURL: URL(string: "https://dashscope.console.aliyun.com/apiKey"),
|
||||
blurb: "通义千问 · OpenAI 兼容"
|
||||
blurb: "通义千问 · OpenAI 兼容 · OpenAI-compatible"
|
||||
),
|
||||
.init(
|
||||
id: "zhipu",
|
||||
name: "智谱 GLM",
|
||||
name: "智谱 GLM · Zhipu",
|
||||
defaultBaseURL: "https://open.bigmodel.cn/api/paas/v4",
|
||||
defaultModel: "glm-4-flash",
|
||||
apiKeyURL: URL(string: "https://bigmodel.cn/usercenter/apikeys"),
|
||||
blurb: "GLM-4-Flash · 中文优化"
|
||||
blurb: "GLM-4-Flash · 中文优化 · Chinese-optimized"
|
||||
),
|
||||
.init(
|
||||
id: "moonshot",
|
||||
@@ -70,14 +70,14 @@ public struct LLMProvider: Identifiable, Codable, Hashable, Sendable {
|
||||
defaultBaseURL: "https://api.moonshot.cn/v1",
|
||||
defaultModel: "moonshot-v1-8k",
|
||||
apiKeyURL: URL(string: "https://platform.moonshot.cn/console/api-keys"),
|
||||
blurb: "Kimi · 长上下文"
|
||||
blurb: "Kimi · 长上下文 · Long context"
|
||||
),
|
||||
.init(
|
||||
id: "custom",
|
||||
name: "Custom (OpenAI-compatible)",
|
||||
name: "Custom · 自定义",
|
||||
defaultBaseURL: "",
|
||||
defaultModel: "",
|
||||
blurb: "自建 / 任意 OpenAI 兼容端点"
|
||||
blurb: "自建 / 任意 OpenAI 兼容端点 · Any OpenAI-compatible endpoint"
|
||||
)
|
||||
]
|
||||
|
||||
|
||||
@@ -0,0 +1,317 @@
|
||||
// ASRService.swift
|
||||
// OSGKeyboard · Shared
|
||||
//
|
||||
// Speech-to-text abstraction.
|
||||
// • iOS 26+: uses `SpeechAnalyzer` + `DictationTranscriber` — always on-device.
|
||||
// • iOS 18–25: uses `SFSpeechRecognizer`, with optional requiresOnDevice flag.
|
||||
// Honours a user-selected locale (auto / zh-CN / en-US / ja-JP …) so
|
||||
// dictation is first-class for non-English languages.
|
||||
//
|
||||
// Lives in `OSGKeyboardShared` (not the keyboard extension target) so
|
||||
// that the host app's `KeyboardPreviewSheet` can run the same ASR
|
||||
// pipeline against real iOS audio — without it, the in-app preview
|
||||
// was a static mock that never actually called `SFSpeechRecognizer`,
|
||||
// and "did you actually wire up ASR?" was a fair review note.
|
||||
|
||||
import Foundation
|
||||
import AVFoundation
|
||||
import Speech
|
||||
import os
|
||||
|
||||
// MARK: - Sendable conformance
|
||||
|
||||
// `AVAudioPCMBuffer` and `SFSpeechRecognitionTask` are not Sendable. We
|
||||
// only ever access them serially — the PCM buffer is built and consumed
|
||||
// inside a single Task, and the recogniser task is cancelled but never
|
||||
// shared concurrently — so an unchecked conformance is sound here.
|
||||
extension AVAudioPCMBuffer: @unchecked @retroactive Sendable {}
|
||||
extension SFSpeechRecognitionTask: @unchecked @retroactive Sendable {}
|
||||
|
||||
// MARK: - Protocol
|
||||
|
||||
public protocol ASRService: Sendable {
|
||||
/// Start a transcription session. The returned stream emits `.partial`
|
||||
/// updates and exactly one `.final` (or `.error`) before finishing.
|
||||
/// - Parameters:
|
||||
/// - stream: Audio buffer stream from `AudioCaptureService`.
|
||||
/// - locale: Target recognition locale.
|
||||
/// - requiresOnDevice: When `true`, forces on-device recognition only
|
||||
/// (SFSpeechRecognizer path). Ignored on iOS 26+ where
|
||||
/// `SpeechAnalyzer` is always fully on-device.
|
||||
func transcribe(
|
||||
stream: AsyncStream<AudioBufferSnapshot>,
|
||||
locale: Locale,
|
||||
requiresOnDevice: Bool
|
||||
) -> AsyncStream<ASREvent>
|
||||
|
||||
/// Cancel any in-flight recognition and tear down its tasks.
|
||||
func cancel()
|
||||
}
|
||||
|
||||
public enum ASREvent: Sendable, Equatable {
|
||||
/// Emitted exactly once at the start of every `transcribe` call, so
|
||||
/// the UI can flag non-on-device locales (e.g. ja-JP on devices that
|
||||
/// only ship on-device ASR for en/zh). The ASR session continues
|
||||
/// either way — we fall back to cloud automatically.
|
||||
case capability(onDeviceSupported: Bool)
|
||||
case partial(String)
|
||||
case final(String)
|
||||
case error(String)
|
||||
}
|
||||
|
||||
// MARK: - Factory
|
||||
|
||||
public enum ASRServiceFactory {
|
||||
/// Returns the best available ASR backend for the current OS:
|
||||
/// `SpeechAnalyzerASR` on iOS 26+ (always on-device), `AppleSpeechASR`
|
||||
/// on older OS versions.
|
||||
public static func make() -> ASRService {
|
||||
if #available(iOS 26.0, *) {
|
||||
return SpeechAnalyzerASR()
|
||||
}
|
||||
return AppleSpeechASR()
|
||||
}
|
||||
}
|
||||
|
||||
// MARK: - Apple Speech implementation (iOS 18–25)
|
||||
|
||||
final class AppleSpeechASR: ASRService, @unchecked Sendable {
|
||||
|
||||
private let lock = OSAllocatedUnfairLock()
|
||||
private var recognizerTask: SFSpeechRecognitionTask?
|
||||
private var feedTask: Task<Void, Never>?
|
||||
|
||||
func transcribe(
|
||||
stream: AsyncStream<AudioBufferSnapshot>,
|
||||
locale: Locale,
|
||||
requiresOnDevice: Bool
|
||||
) -> AsyncStream<ASREvent> {
|
||||
AsyncStream { continuation in
|
||||
let recognizer = SFSpeechRecognizer(locale: locale)
|
||||
?? SFSpeechRecognizer(locale: .current)
|
||||
guard let recognizer, recognizer.isAvailable else {
|
||||
continuation.yield(.error("Speech recognizer unavailable for \(locale.identifier)"))
|
||||
continuation.finish()
|
||||
return
|
||||
}
|
||||
recognizer.defaultTaskHint = .dictation
|
||||
|
||||
let request = SFSpeechAudioBufferRecognitionRequest()
|
||||
request.shouldReportPartialResults = true
|
||||
// Honour the user's "force on-device" preference; fall back to
|
||||
// whatever the device natively supports if the flag is off.
|
||||
request.requiresOnDeviceRecognition = requiresOnDevice || recognizer.supportsOnDeviceRecognition
|
||||
let onDeviceSupported = recognizer.supportsOnDeviceRecognition
|
||||
if !onDeviceSupported {
|
||||
#if DEBUG
|
||||
print("⚠️ 设备不支持 \(locale.identifier) 端侧 ASR, 回退云端。")
|
||||
#endif
|
||||
}
|
||||
// Tell the UI about the capability *before* any partials so
|
||||
// the StatusBadge can light up the cloud-fallback indicator
|
||||
// as soon as the user presses the mic.
|
||||
continuation.yield(.capability(onDeviceSupported: onDeviceSupported))
|
||||
|
||||
let task = recognizer.recognitionTask(with: request) { result, error in
|
||||
if let error {
|
||||
let nsErr = error as NSError
|
||||
// Codes 203 / 1110 = "no speech detected" — a normal exit.
|
||||
if nsErr.code == 203 || nsErr.code == 1110 {
|
||||
continuation.yield(.final(""))
|
||||
} else {
|
||||
continuation.yield(.error(error.localizedDescription))
|
||||
}
|
||||
continuation.finish()
|
||||
return
|
||||
}
|
||||
guard let result else { return }
|
||||
if result.isFinal {
|
||||
continuation.yield(.final(result.bestTranscription.formattedString))
|
||||
continuation.finish()
|
||||
} else {
|
||||
continuation.yield(.partial(result.bestTranscription.formattedString))
|
||||
}
|
||||
}
|
||||
|
||||
self.lock.withLock { self.recognizerTask = task }
|
||||
|
||||
// Feed audio: for each snapshot, build a 16 kHz mono Float32
|
||||
// PCM buffer and immediately `request.append(pcm)`. The PCM
|
||||
// buffer never leaves this task, so it doesn't need to be
|
||||
// Sendable.
|
||||
let feedFormat = AVAudioFormat(
|
||||
commonFormat: .pcmFormatFloat32,
|
||||
sampleRate: 16_000,
|
||||
channels: 1,
|
||||
interleaved: false
|
||||
)!
|
||||
self.feedTask = Task { [request] in
|
||||
for await snap in stream {
|
||||
if Task.isCancelled { break }
|
||||
guard !snap.samples.isEmpty,
|
||||
let pcm = AVAudioPCMBuffer(
|
||||
pcmFormat: feedFormat,
|
||||
frameCapacity: AVAudioFrameCount(snap.samples.count)
|
||||
)
|
||||
else { continue }
|
||||
pcm.frameLength = AVAudioFrameCount(snap.samples.count)
|
||||
if let dst = pcm.floatChannelData?[0] {
|
||||
snap.samples.withUnsafeBufferPointer { src in
|
||||
if let base = src.baseAddress {
|
||||
memcpy(dst, base, snap.samples.count * MemoryLayout<Float>.size)
|
||||
}
|
||||
}
|
||||
}
|
||||
request.append(pcm)
|
||||
}
|
||||
if !Task.isCancelled {
|
||||
request.endAudio()
|
||||
}
|
||||
}
|
||||
|
||||
continuation.onTermination = { @Sendable [weak self] _ in
|
||||
self?.cancel()
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
func cancel() {
|
||||
let (recTask, feedT) = lock.withLock { () -> (SFSpeechRecognitionTask?, Task<Void, Never>?) in
|
||||
let r = self.recognizerTask
|
||||
let f = self.feedTask
|
||||
self.recognizerTask = nil
|
||||
self.feedTask = nil
|
||||
return (r, f)
|
||||
}
|
||||
recTask?.cancel()
|
||||
feedT?.cancel()
|
||||
}
|
||||
}
|
||||
|
||||
// MARK: - SpeechAnalyzer implementation (iOS 26+)
|
||||
|
||||
/// ASR backend that uses the iOS 26 `SpeechAnalyzer` + `DictationTranscriber`
|
||||
/// APIs. This engine is always fully on-device — `requiresOnDevice` has no
|
||||
/// effect and `.capability(onDeviceSupported: true)` is always emitted.
|
||||
@available(iOS 26.0, *)
|
||||
final class SpeechAnalyzerASR: ASRService, @unchecked Sendable {
|
||||
|
||||
private let lock = OSAllocatedUnfairLock()
|
||||
private var analyzer: SpeechAnalyzer?
|
||||
private var analyzerTask: Task<Void, Never>?
|
||||
|
||||
func transcribe(
|
||||
stream: AsyncStream<AudioBufferSnapshot>,
|
||||
locale: Locale,
|
||||
requiresOnDevice: Bool // ignored — SpeechAnalyzer is always on-device
|
||||
) -> AsyncStream<ASREvent> {
|
||||
AsyncStream { continuation in
|
||||
// SpeechAnalyzer is always fully on-device.
|
||||
continuation.yield(.capability(onDeviceSupported: true))
|
||||
|
||||
let transcriber = DictationTranscriber(locale: locale, preset: .progressiveShortDictation)
|
||||
let newAnalyzer = SpeechAnalyzer(modules: [transcriber])
|
||||
self.lock.withLock { self.analyzer = newAnalyzer }
|
||||
|
||||
let audioFormat = AVAudioFormat(
|
||||
commonFormat: .pcmFormatFloat32,
|
||||
sampleRate: 16_000,
|
||||
channels: 1,
|
||||
interleaved: false
|
||||
)!
|
||||
|
||||
let task = Task { [weak self] in
|
||||
guard let self else { return }
|
||||
do {
|
||||
try await newAnalyzer.prepareToAnalyze(in: audioFormat)
|
||||
|
||||
let inputStream = self.makeInputStream(from: stream, format: audioFormat)
|
||||
|
||||
// Feed audio in a child task so we can concurrently
|
||||
// iterate `transcriber.results` on the outer task.
|
||||
// After the audio stream ends, finalize so the results
|
||||
// sequence can drain and complete.
|
||||
let feedTask = Task {
|
||||
do {
|
||||
try await newAnalyzer.start(inputSequence: inputStream)
|
||||
try await newAnalyzer.finalizeAndFinishThroughEndOfInput()
|
||||
} catch {}
|
||||
}
|
||||
defer { feedTask.cancel() }
|
||||
|
||||
var lastText = ""
|
||||
do {
|
||||
for try await result in transcriber.results {
|
||||
if Task.isCancelled { break }
|
||||
// `result.text` is an AttributedString; extract plain text.
|
||||
let text = result.text.characters.map(String.init).joined()
|
||||
guard !text.isEmpty, text != lastText else { continue }
|
||||
lastText = text
|
||||
continuation.yield(.partial(text))
|
||||
}
|
||||
} catch {
|
||||
// Results sequence threw — likely cancellation.
|
||||
}
|
||||
|
||||
if !Task.isCancelled {
|
||||
continuation.yield(.final(lastText))
|
||||
}
|
||||
continuation.finish()
|
||||
} catch is CancellationError {
|
||||
continuation.finish()
|
||||
} catch {
|
||||
continuation.yield(.error(error.localizedDescription))
|
||||
continuation.finish()
|
||||
}
|
||||
}
|
||||
self.lock.withLock { self.analyzerTask = task }
|
||||
|
||||
continuation.onTermination = { @Sendable [weak self] _ in
|
||||
self?.cancel()
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
func cancel() {
|
||||
let (task, currentAnalyzer) = lock.withLock { () -> (Task<Void, Never>?, SpeechAnalyzer?) in
|
||||
let t = analyzerTask
|
||||
let a = analyzer
|
||||
analyzerTask = nil
|
||||
analyzer = nil
|
||||
return (t, a)
|
||||
}
|
||||
task?.cancel()
|
||||
if let a = currentAnalyzer {
|
||||
Task { await a.cancelAndFinishNow() }
|
||||
}
|
||||
}
|
||||
|
||||
/// Maps the `AudioBufferSnapshot` stream into the `AnalyzerInput` stream
|
||||
/// that `SpeechAnalyzer` consumes.
|
||||
private func makeInputStream(
|
||||
from stream: AsyncStream<AudioBufferSnapshot>,
|
||||
format: AVAudioFormat
|
||||
) -> AsyncStream<AnalyzerInput> {
|
||||
AsyncStream { continuation in
|
||||
Task {
|
||||
for await snap in stream {
|
||||
guard !snap.samples.isEmpty,
|
||||
let pcm = AVAudioPCMBuffer(
|
||||
pcmFormat: format,
|
||||
frameCapacity: AVAudioFrameCount(snap.samples.count)
|
||||
)
|
||||
else { continue }
|
||||
pcm.frameLength = AVAudioFrameCount(snap.samples.count)
|
||||
if let dst = pcm.floatChannelData?[0] {
|
||||
snap.samples.withUnsafeBufferPointer { src in
|
||||
guard let base = src.baseAddress else { return }
|
||||
memcpy(dst, base, snap.samples.count * MemoryLayout<Float>.size)
|
||||
}
|
||||
}
|
||||
continuation.yield(AnalyzerInput(buffer: pcm))
|
||||
}
|
||||
continuation.finish()
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
Reference in New Issue
Block a user