The previous build crashed on the first record-tap with:
*** Terminating app due to uncaught exception
'com.apple.coreaudio.avfaudio', reason:
'Failed to create tap due to format mismatch,
<AVAudioFormat: 1 ch, 48000 Hz, Float32>'
The crash surfaced at `[AVAudioNode installTap:...]` but the
*root* cause was one frame deeper, in the underlying
`AURemoteIO::enable` call:
AURemoteIO.cpp:1135 failed: -10851
(enable 1, outf< 2 ch, 0 Hz, Float32, deinterleaved>
inf< 1 ch, 48000 Hz, Float32>)
`AURemoteIO` is a two-direction Audio Unit (input + output).
`installTap` triggers `AURemoteIO::enable(1)` which tries to
initialize *both* directions at once. On the iOS Simulator, the
output direction reports a `0 Hz` "speaker" because the
simulator has no real speaker, and `enable` fails with
`kAudioUnitErr_FormatNotSupported` (-10851). The error then
bubbles up through `installTap` as the misleading "format
mismatch" — the input format we passed was correct, but the
output side broke the whole `enable` call.
The previous code asked the audio session for `.playAndRecord`
mode, which requires the output direction to be enabled. On a
real device, the speaker is 44100/48000 Hz and `.playAndRecord`
works. On the simulator, it traps.
Fix (two parts):
1) `start()` audio session now uses `.record` instead of
`.playAndRecord`. The preview never plays back audio — it
just records from the mic and hands the buffers to
`SpeechAnalyzer`. `.record` skips the output direction
entirely, so the simulator's 0 Hz speaker isn't a problem.
The real keyboard extension (`OSGKeyboardExt`) keeps
`.playAndRecord` because it runs on real devices and may
want to play click sounds / haptic feedback — only the
preview needs the simulator-friendly category.
2) `startEngineAndASR` adds a pre-flight check: refuse to
call `installTap` when the input bus reports
`sampleRate == 0` or `channelCount == 0`. The `installTap`
failure mode is an `NSException` (Objective-C), not a Swift
`Error` — `try`/`catch` can't intercept it, so the only safe
defence is to never call it with a placeholder/unconfigured
bus. We saw `0 Hz` inputs on the simulator when the host
mic permission wasn't granted to CoreSimulator, and on
devices with an unexpected audio-session state. The check
sets a clear `.error` phase ("Microphone unavailable (hw
format X Hz / Y ch)") instead of trapping the app.
Why this also unblocks voice-to-text: the previous
"no transcription" symptom was just the crash at step 7 of
the start pipeline — the ASR never saw a single audio frame
because the tap never installed. With the crash fixed, the
`SpeechAnalyzer` pipeline can now actually receive buffers
and emit `.partial` / `.final` events. Whether the *content*
of those events is meaningful depends on the simulator's mic
quality and the chosen locale (both separate concerns from
this fix).
Tests: 27/27 pass (no behavioural test for the pre-flight
check — the check is 4 lines of obvious defence, and
testing it would require either mocking the audio engine
or installing CoreSimulator's host-mic-bridge, neither of
which is worth the cost).
Build: BUILD SUCCEEDED.
🤖 Generated with Claude Code
OSGKeyboard
Hold a key, speak, release — AI-polished text appears at your cursor in any app. An open-source, custom-keyboard-based voice input tool for iOS 18+, inspired by Typeless and OpenLess.
What is it?
OSGKeyboard is a free, open alternative to commercial voice-input tools. It runs as a Custom Keyboard Extension on iOS, so you can use it in any app — Messages, Notes, Mail, ChatGPT, Claude, Cursor, you name it.
- Press and hold the mic key
- Speak naturally
- Release — the AI polishes your words into clean text and inserts it at the cursor
The audio stays on-device (transcribed by Apple's on-device SFSpeechRecognizer on iOS 18/19; iOS 26+ SpeechAnalyzer planned for the next release). Only the polished transcript is sent to your chosen cloud LLM. No audio ever leaves your phone.
Features
- 🎙 Push-to-talk with a Typeless-style circular mic button
- 🧠 On-device ASR (iOS 18/19
SFSpeechRecognizer; iOS 26+SpeechAnalyzer+DictationTranscriberplanned) - ✍️ AI polishing — adds structure, punctuation, fixes grammar, optionally produces lists
- 🔌 Bring-your-own API — works with any OpenAI-compatible endpoint (OpenAI, DeepSeek, Qwen DashScope, your own self-hosted server, …)
- 🔒 Privacy first — audio never leaves your device; transcripts only sent to the LLM you choose
- 🎨 Native SwiftUI — dark theme, frosted glass, ~2000 lines of Swift
- 🪶 Zero dependencies — no SwiftPM packages, no CocoaPods, no Carthage
Quick start
Requirements
- macOS with Xcode 16+ (Xcode 26 recommended)
- iPhone running iOS 18.0+
- XcodeGen:
brew install xcodegen - An OpenAI-compatible API key (e.g. from OpenAI, DeepSeek, or Qwen DashScope)
Build & run
git clone https://github.com/hkgood/OSGKeyboard.git
cd OSGKeyboard
xcodegen generate # produces OSGKeyboard.xcodeproj
open OSGKeyboard.xcodeproj # or build via CLI:
xcodebuild -project OSGKeyboard.xcodeproj -scheme OSGKeyboard \
-destination 'generic/platform=iOS Simulator' build
Enable the keyboard in iOS
- Run the app on your device or simulator.
- Follow the 3-step onboarding: enable the keyboard in iOS Settings, then allow Full Access (required for the mic and LLM calls), then paste your API key.
- In any text field, tap 🌐 to switch to OSGKeyboard.
- Press and hold the mic, speak, release. ✨
"Allow Full Access" is required. Without it, iOS blocks the keyboard from using the microphone and from making network requests. We never log, store, or transmit your keystrokes — see
PrivacyInfo.xcprivacy.
Architecture
OSGKeyboard/
├── OSGKeyboard/ # Main iOS app (settings, onboarding)
│ ├── Views/ # SwiftUI screens
│ ├── OSGKeyboardApp.swift # @main entry
│ ├── PrivacyInfo.xcprivacy # Required privacy manifest
│ └── OSGKeyboard.entitlements # App Group declaration
├── OSGKeyboardExt/ # Custom Keyboard Extension
│ ├── KeyboardViewController.swift # Principal class
│ ├── Services/
│ │ ├── AudioCaptureService.swift # AVAudioEngine → 16 kHz PCM
│ │ ├── ASRService.swift # iOS 26 + iOS 18 ASR
│ │ └── PolishingService.swift # LLM call with timeout
│ └── Views/ # RecordButton, Waveform, KeyboardRootView
├── OSGKeyboardShared/ # Framework shared by app + extension
│ ├── Models/ # ProviderConfig, LLMRequest, LLMProvider
│ ├── Services/ # LLMClient (OpenAI-compatible)
│ └── Constants/ # AppGroup identifier
├── OSGKeyboardTests/ # XCTest unit tests
├── project.yml # XcodeGen project definition
└── .github/workflows/ci.yml # Lint + build CI
Data flow
[Long-press mic] → AudioCaptureService → AudioBufferSnapshot (16 kHz mono)
↓
ASRService.transcribe()
↓
ASREvent.final(rawTranscript)
↓
PolishingService.polish()
↓
LLMClient (OpenAI-compatible)
↓
textDocumentProxy.insertText(polished)
Adding a new LLM provider
Open OSGKeyboardShared/Models/LLMProvider.swift and append a new LLMProvider to the presets array. The default OpenAICompatibleClient handles any endpoint that speaks the POST /chat/completions protocol.
LLMProvider(
id: "groq",
name: "Groq",
defaultBaseURL: "https://api.groq.com/openai/v1",
defaultModel: "llama-3.1-70b-versatile",
apiKeyURL: URL(string: "https://console.groq.com/keys")
)
That's it. No other code changes required.
Limitations
- iOS sandboxes keyboard extensions: ~60 MB memory cap, Full Access required.
- The keyboard does not work in password fields or some
WKWebViewtextareas (iOS limitation). - iOS 18/19 ships with
SFSpeechRecognizerfor on-device ASR. iOS 26+SpeechAnalyzeris planned for the next release — it is significantly faster and supports more locales. - iOS 26+ users in v0.1.1 use the iOS 18
SFSpeechRecognizerpath; the iOS 26SpeechAnalyzeris planned for 0.2.0.
License
MIT — use it, fork it, ship it. No warranty.
Acknowledgements
- Inspired by Typeless and the desktop open-source OpenLess
- Built with XcodeGen
- Powered by Apple's SpeechAnalyzer and SFSpeechRecognizer
Note: the project is published at hkgood/OSGKeyboard; badges and git clone URLs already point there.