Sync project documentation with the current v0.2.1 implementation (Tap-to-toggle recording, 5-step onboarding, Flow session model, SpeechAnalyzer-only local engine, deepseek-v4-flash default). README.md / README.zh.md - Replace 'press-and-hold / 按住说话' with the v0.2.0+ tap-to-toggle interaction; add the 60-second per-take cap to feature bullets. - Update 3-step onboarding to the real 5-step flow. - Correct architecture diagram: AudioCaptureService in the extension is legacy/unused, PolishingService lives in OSGKeyboardShared, and the FlowSession* / LiveDictation* services are now in the tree. - Replace the legacy single-pipeline data flow with the actual Flow session data flow (keyboard -> App Group -> host app -> chunked ASR -> LLM polish -> App Group -> insertText). - Move the stale 'main -> 0.2 branch rename' banner into a project status section under the new 'v0.2.1' badge. - Expand 'Known limitations' with 60s/3min caps, 60MB sandbox note, URL scheme caveat, and the v0.2.0->v0.2.1 on-device LLM rollback. - Add a 'Development' section pointing to tests, CI, logging policy. - Add version badge, license link, privacy policy link. docs/index.html (GitHub Pages landing) - Mirror the same copy fixes in both EN and ZH i18n tables (verified: 50 keys each, no missing translations). - Grow feature grid from 6 to 8 cards (add 'Flow session' and 'Local + cloud polish'), grow steps from 3 to 5. - Bump footer to 'v0.2.1 · source available, non-commercial', add License link. docs/privacy.html + docs/privacy/index.html - Update 'Last updated' to July 3, 2026, tag v0.2.1. - Add explicit iOS 26+ requirement, Flow session explanation, children's privacy section, policy-changes section, license reference, and rocky.hk@gmail.com contact. - Make the 'no raw audio upload to any server' claim explicit. Verification - grep confirms no leftover 'press and hold' / '按住' / 'press to'. - HTML structure validated (well-formed on all three pages). - EN/ZH i18n keys are symmetric (50 each).
12 KiB
OSGKeyboard
Tap to talk, tap to stop — AI-polished text appears at your cursor in any app. A source-available, custom-keyboard-based voice input tool for iOS 26+, inspired by Typeless and OpenLess.
What is it?
OSGKeyboard is a free, source-available alternative to commercial voice-input tools. It runs as a Custom Keyboard Extension on iOS, so you can use it in any app — Messages, Notes, Mail, WeChat, ChatGPT, Claude, Cursor, you name it.
- Tap the mic to start recording
- Speak naturally (up to 60 seconds per take)
- Tap again to stop — the AI polishes your words into clean text and inserts at the cursor
Audio is transcribed on-device by Apple's SpeechAnalyzer + DictationTranscriber (iOS 26+). Only the polished transcript is sent to your chosen cloud LLM. No audio ever leaves your phone.
Under the hood, OSGKeyboard uses a Flow session model: a long-lived audio session runs in the host app, the keyboard extension writes tiny "start / stop" signals to the App Group, and the polished text is delivered back to the keyboard for insertion. You do not need to jump back to the host app between recordings.
Features
- 🎙 Tap-to-toggle recording with a Typeless-style circular mic button, 60-second per-take cap with live countdown
- 🧠 On-device ASR (
SpeechAnalyzer+DictationTranscriber, iOS 26+) - ✍️ AI polishing — adds structure, punctuation, fixes grammar, optionally produces lists
- 🧩 Local + cloud polish toggle — local engine is ASR-only by default; opt into a post-ASR cloud polish step (DeepSeek by default) when the iOS speech recognition isn't strong enough for your environment (noisy far-field audio, strong accents, etc.)
- 🔌 Bring-your-own API — works with any OpenAI-compatible endpoint (OpenAI, DeepSeek, Qwen DashScope, Moonshot, Zhipu, your own self-hosted server, …)
- 🔒 Privacy first — audio never leaves your device; only the final transcript is sent to the LLM you choose
- 🎨 Native SwiftUI — dark theme, frosted glass, pure Swift 6, ~3,600 lines of code
- 🪶 Zero dependencies — no SwiftPM packages, no CocoaPods, no Carthage
- 🔁 Flow session — keep recording across multiple takes without bouncing back to the host app
Quick start
Requirements
- macOS with Xcode 26 (matches
project.ymldeployment target iOS 26) - iPhone running iOS 26.0+
- XcodeGen:
brew install xcodegen - An OpenAI-compatible API key (e.g. from OpenAI, DeepSeek, or Qwen DashScope). Not needed if you stay on the "local ASR only" engine.
Build & run
git clone https://github.com/hkgood/OSGKeyboard.git
cd OSGKeyboard
./Scripts/generate-xcodeproj.sh # generates OSGKeyboard.xcodeproj via XcodeGen
open OSGKeyboard.xcodeproj # or build via CLI:
xcodebuild -project OSGKeyboard.xcodeproj -scheme OSGKeyboard \
-destination 'generic/platform=iOS Simulator' build
The
OSGKeyboard.xcodeprojis not committed — it is regenerated fromproject.ymlbyScripts/generate-xcodeproj.sh. Always re-run the script aftergit pullifproject.ymlhas changed.
Enable the keyboard in iOS
The host app walks you through a 5-step onboarding:
- Welcome — intro to OSGKeyboard
- Microphone — request mic access
- Speech recognition — request on-device speech recognition access
- Enable keyboard + Full Access — open iOS Settings to add OSGKeyboard and allow Full Access
- Engine + API — pick the local or cloud engine, then paste your API key (cloud / cloud-polish only)
After onboarding, in any text field, tap 🌐 to switch to OSGKeyboard, then tap the circular mic to start, speak, and tap again to stop.
"Allow Full Access" is required. Without it, iOS blocks the keyboard from using the microphone and from making network requests. We never log, store, or transmit your keystrokes — see
PrivacyInfo.xcprivacyand our Privacy Policy.
Architecture
OSGKeyboard/
├── OSGKeyboard/ # Main iOS app (host of the Flow session)
│ ├── Services/ # FlowSessionManager, AppPermissions, SpeechHistoryStore, …
│ ├── Views/ # SwiftUI: OnboardingView, HomeView, SettingsView, HistoryView, …
│ ├── OSGKeyboardApp.swift # @main entry, owns the FlowSessionManager
│ ├── PrivacyInfo.xcprivacy # Required privacy manifest
│ └── OSGKeyboard.entitlements # App Group + Keychain Group
├── OSGKeyboardExt/ # Custom Keyboard Extension
│ ├── KeyboardViewController.swift # Principal class (drives SwiftUI)
│ ├── Services/ # AppGroupPersistor, HostAppLauncher, AudioCaptureService (legacy, unused)
│ ├── Views/ # KeyboardRootView, RecordButton, WaveformView
│ └── PrivacyInfo.xcprivacy
├── OSGKeyboardShared/ # Framework shared by app + extension (APPLICATION_EXTENSION_API_ONLY=YES)
│ ├── Services/ # FlowSessionBridge, FlowSessionDarwin, LLMClient, PolishingService, ASRService, Keychain, AppGroupStore, …
│ ├── Models/ # LLMProvider, ProviderConfig, TranscriptionDelivery, AudioBufferSnapshot, …
│ ├── DesignSystem/ # Theme, ThemedRoot
│ └── Constants/ # AppGroup identifier
├── OSGKeyboardTests/ # XCTest unit tests (LLM, Keychain, ASR, Flow bridge, …)
├── OSGKeyboardExtTests/ # Keyboard-extension-side unit tests
├── Scripts/ # generate-xcodeproj.sh, patch-icon-composer.sh
├── docs/ # GitHub Pages site (privacy policy + landing)
├── project.yml # XcodeGen project definition (source of truth)
└── .github/workflows/ci.yml # Lint + build CI
Data flow — Flow session model
[Tap mic in keyboard]
└─► KeyboardViewController.pressBegan
└─► FlowSessionBridge.setRecordingState(.recording) [App Group UserDefaults]
└─► Darwin notification: "recordingState changed"
└─► FlowSessionManager (host app) sees the signal
└─► FlowContinuousCapture feeds 16 kHz PCM into ChunkedUtterancePipeline
└─► ASRService.transcribe (iOS 26 SpeechAnalyzer)
└─► ASREvent.partial / .final
└─► UtteranceTranscriptStitcher stitches the chunks
└─► PolishingService (LLMClient) [optional, configurable]
└─► FlowSessionBridge.storeTranscriptionResult
[Keyboard polls + Darwin notif]
└─► KeyboardViewController sees the result
└─► textDocumentProxy.insertText(polished)
Engine modes:
cloud(default) — on-device ASR viaSpeechAnalyzer, transcript is sent to your configured LLM for polish.local— on-device ASR viaSpeechAnalyzeronly; transcript is inserted as-is. No network round-trip.local+ "Cloud polish after ASR" toggle (Settings → Engine) — same on-device ASR, but the transcript is routed through your configured LLM before insertion. Useful when iOS speech recognition isn't accurate enough in your environment.
Cross-process plumbing (host app ↔ keyboard extension):
- App Group
group.com.osgkeyboard.shared—UserDefaultsfor the live Flow session state, recording state, audio levels, transcription delivery, and most preferences. - Shared Keychain group
com.osgkeyboard.shared— the LLM API key is written by the host app's Settings, read by both processes before every LLM call. - Darwin notifications (
CFNotificationCenter) — light-weight "something changed" pings; payloads still travel through the App Group.
Adding a new LLM provider
Open OSGKeyboardShared/Models/LLMProvider.swift and append a new LLMProvider to the presets array. The default OpenAICompatibleClient handles any endpoint that speaks the POST /chat/completions protocol.
LLMProvider(
id: "groq",
name: "Groq",
defaultBaseURL: "https://api.groq.com/openai/v1",
defaultModel: "llama-3.1-70b-versatile",
apiKeyURL: URL(string: "https://console.groq.com/keys")
)
That's it — no other code changes required.
To set it as the new default for first-time users, also bump the defaultProviderId constant used by ProviderConfig.
Known limitations
- iOS 26+ only. Earlier iOS versions are not supported. We dropped the pre-26 SFSpeechRecognizer / AVAudioSession branching so the entire ASR path can use the iOS 26
SpeechAnalyzerAPI exclusively. - ~60 MB memory cap for the keyboard extension (iOS sandbox). The Flow session is hosted in the main app, so audio buffers and ASR models live there, not in the extension.
- "Allow Full Access" required. Without it, the keyboard can't reach the microphone or make network requests for cloud polish.
- Password fields and some
WKWebViewtextareas are blocked by iOS itself — not something we can work around. - 60-second per-take cap. A long take is automatically stopped and dispatched for transcription; a new take can be started immediately.
- 3-minute per-utterance ASR cap. If you exceed it, the pipeline gracefully splits into multiple stitched chunks.
- No on-device LLM polish. The local engine is ASR-only; "AI polish" is always cloud-based and configurable. On-device model support was explored in v0.2.0 and rolled back in v0.2.1 to keep the dependency surface at zero SPM packages.
- URL scheme
osgkeyboard://can be opened by any app on the device. We don't trust it for anything beyond "wake the host app and (re)start the Flow session"; it never carries your API key or other secrets.
Development
- Build setup — see the Build Setup section at the top of this file. Run
./Scripts/generate-xcodeproj.shafter anyproject.ymlchange. - Tests —
xcodebuild test -project OSGKeyboard.xcodeproj -scheme OSGKeyboard -destination 'platform=iOS Simulator,name=iPhone 17'runs bothOSGKeyboardTestsandOSGKeyboardExtTeststargets. - CI —
.github/workflows/ci.ymlruns SwiftLint, a clean Debug build, and the test suite on every push to0.1/0.2and PRs. - Logging —
printis debug-only; release builds useNSLogfor the few cross-process status messages.
Project status
- Current release: v0.2.1 (2026-06-24)
- Default branch:
0.2(renamed frommainon 2026-06-24; the previousmainis preserved as0.1). - See
CHANGELOG.mdfor the full release history andTYPEWHISPER_FLOW_MIGRATION_TRACKER.mdfor the architecture-decision log behind the Flow session model.
License
OSGKeyboard Source Available License — personal learning and non-commercial local use only. No commercial use, redistribution, or public forks without permission. Commercial licensing: rocky.hk@gmail.com.
Acknowledgements
- Inspired by Typeless and the desktop open-source OpenLess
- Built with XcodeGen
- Powered by Apple's SpeechAnalyzer and SFSpeechRecognizer