zhongshu 4d92999347 feat(translation-ui): tighten settings/onboarding UX around the translation feature
UI refinements on top of the translation pipeline (feature/translation@HEAD):

1. Onboarding engine page now hosts a translation row.
   APISetupPage renders the same TranslationPickerRow used in the
   language tab, so first-time users can pick a target language
   before they ever see the keyboard. Same persisted bindings; same
   'needs cloud' hint when the local engine is active.

2. Local engine hides the provider / API card unconditionally.
   Removed the 'local + cloud polish on → show API fields' branch
   from SettingsView. Provider/base URL/API key/model controls have
   no use in local mode (translation is cloud-only anyway), and
   exposing them invited users to fill in a DeepSeek key they
   can't use.

3. 'Cloud polish after ASR' toggle loses its long subtitle.
   The descriptive copy in LocalEngineSettingsRows.cloudPolishRow
   was a wall of text that explained things visible elsewhere in
   Settings. Title + switch is enough; the CloudPolishDisclosureBanner
   (rendered by EnginePickerSection when cloud is active) already
   covers the 'this sends text to your API' disclosure.

4. Translation row becomes a single dropdown with a 'Don't
   translate' default.
   TranslationPickerRow replaced with a one-row Menu picker:
   '不翻译 / English / 中文 (简体) / 中文 (繁體) / 日本語 / 한국어 /
   Français / Deutsch / Español / Русский / Português'.
   '不翻译' maps to translationEnabled=false; any locale maps to
   translationEnabled=true + translationTargetLocaleId=<id>.
   TranslationLanguageCatalog gains an 'off' sentinel so the picker's
   single binding stays a plain String.

5. Language tab reorder.
   SettingsView.languageAndModelsSection: ASR locale ('识别语言')
   now sits above the local-models block; translation row sits at
   the bottom. The reading order follows the pipeline direction
   (input → post-processing → post-post-processing).

Localization:
  - 'settings.translation.title' → '翻译' / 'Translation'
  - new 'settings.translation.off' / 'settings.translation.hint.needsCloud'
  - dropped unused subtitle / target-language keys

xcodebuild scheme=OSGKeyboard config=Debug destination=iPhone 17
Simulator: BUILD SUCCEEDED (0 warning, 0 error).
2026-06-25 13:23:54 +08:00
2026-06-17 23:08:58 +08:00

⚠️ 2026-06-24: Default branch renamed from main to 0.2. Old main is now 0.1; refactor/drop-qwen3-coreml is now 0.2. See CHANGELOG for details.

Build Setup

This project uses XcodeGenproject.yml is the source of truth, and OSGKeyboard.xcodeproj is regenerated, not committed (it's in .gitignore). New clones must generate the Xcode project before opening in Xcode, otherwise Xcode will show an empty project with no source files.

Prerequisites

  • macOS with Xcode 26 (matches project.yml deployment target iOS 26)
  • Homebrew

One-time setup

brew install xcodegen

Generate OSGKeyboard.xcodeproj

From the repository root:

./Scripts/generate-xcodeproj.sh

This script:

  1. Runs xcodegen generate to produce OSGKeyboard.xcodeproj from project.yml.
  2. Calls Scripts/patch-icon-composer.sh to patch the generated project for Xcode 26 Icon Composer bundles — OSGKeyboard/AppIcon.icon is treated as a single folder.iconcomposer.icon reference (see project.yml note on XcodeGen fileTypes.icon).

Then open the project:

open OSGKeyboard.xcodeproj

Troubleshooting: If Xcode shows an empty project, re-run ./Scripts/generate-xcodeproj.sh. If the build fails on icon assets, make sure OSGKeyboard/AppIcon.icon/ exists with Assets/ and icon.json inside (do not expand the folder manually).

Re-running

Run ./Scripts/generate-xcodeproj.sh any time project.yml changes (e.g. after git pull).


OSGKeyboard

Hold a key, speak, release — AI-polished text appears at your cursor in any app. A source-available, custom-keyboard-based voice input tool for iOS 26+, inspired by Typeless and OpenLess.

Platform Swift License CI

中文 README


What is it?

OSGKeyboard is a free, source-available alternative to commercial voice-input tools. It runs as a Custom Keyboard Extension on iOS, so you can use it in any app — Messages, Notes, Mail, ChatGPT, Claude, Cursor, you name it.

  1. Press and hold the mic key
  2. Speak naturally
  3. Release — the AI polishes your words into clean text and inserts it at the cursor

The audio stays on-device (transcribed by Apple's on-device SpeechAnalyzer + DictationTranscriber on iOS 26+). Only the polished transcript is sent to your chosen cloud LLM. No audio ever leaves your phone.


Features

  • 🎙 Push-to-talk with a Typeless-style circular mic button
  • 🧠 On-device ASR (SpeechAnalyzer + DictationTranscriber, iOS 26+)
  • ✍️ AI polishing — adds structure, punctuation, fixes grammar, optionally produces lists
  • 🧩 Local + cloud polish toggle — local engine is ASR-only by default; opt into a post-ASR cloud polish step (DeepSeek by default) when the iOS speech recognition isn't strong enough for your environment (noisy far-field audio, strong accents, etc.)
  • 🔌 Bring-your-own API — works with any OpenAI-compatible endpoint (OpenAI, DeepSeek, Qwen DashScope, your own self-hosted server, …)
  • 🔒 Privacy first — audio never leaves your device; transcripts only sent to the LLM you choose
  • 🎨 Native SwiftUI — dark theme, frosted glass, ~2000 lines of Swift
  • 🪶 Zero dependencies — no SwiftPM packages, no CocoaPods, no Carthage

Quick start

Requirements

Build & run

git clone https://github.com/hkgood/OSGKeyboard.git
cd OSGKeyboard
xcodegen generate          # produces OSGKeyboard.xcodeproj
open OSGKeyboard.xcodeproj # or build via CLI:
xcodebuild -project OSGKeyboard.xcodeproj -scheme OSGKeyboard \
  -destination 'generic/platform=iOS Simulator' build

Enable the keyboard in iOS

  1. Run the app on your device or simulator.
  2. Follow the 3-step onboarding: enable the keyboard in iOS Settings, then allow Full Access (required for the mic and LLM calls), then paste your API key.
  3. In any text field, tap 🌐 to switch to OSGKeyboard.
  4. Press and hold the mic, speak, release.

"Allow Full Access" is required. Without it, iOS blocks the keyboard from using the microphone and from making network requests. We never log, store, or transmit your keystrokes — see PrivacyInfo.xcprivacy.


Architecture

OSGKeyboard/
├── OSGKeyboard/                 # Main iOS app (settings, onboarding)
│   ├── Views/                   # SwiftUI screens
│   ├── OSGKeyboardApp.swift     # @main entry
│   ├── PrivacyInfo.xcprivacy    # Required privacy manifest
│   └── OSGKeyboard.entitlements # App Group declaration
├── OSGKeyboardExt/              # Custom Keyboard Extension
│   ├── KeyboardViewController.swift   # Principal class
│   ├── Services/
│   │   ├── AudioCaptureService.swift  # AVAudioEngine → 16 kHz PCM
│   │   ├── ASRService.swift           # iOS 26 SpeechAnalyzer ASR
│   │   └── PolishingService.swift     # LLM call with timeout
│   └── Views/                   # RecordButton, Waveform, KeyboardRootView
├── OSGKeyboardShared/           # Framework shared by app + extension
│   ├── Models/                  # ProviderConfig, LLMRequest, LLMProvider
│   ├── Services/                # LLMClient (OpenAI-compatible)
│   └── Constants/               # AppGroup identifier
├── OSGKeyboardTests/            # XCTest unit tests
├── project.yml                  # XcodeGen project definition
└── .github/workflows/ci.yml     # Lint + build CI

Data flow

[Long-press mic] → AudioCaptureService → AudioBufferSnapshot (16 kHz mono)
                                          ↓
                                  ASRService.transcribe()
                                          ↓
                                ASREvent.final(rawTranscript)
                                          ↓
                              PolishingService.polish()
                                          ↓
                            LLMClient (OpenAI-compatible)
                                          ↓
                          textDocumentProxy.insertText(polished)

Engine modes:

  • cloud (default): ASR runs on-device via SpeechAnalyzer, the transcript is sent to your configured LLM for polish.
  • local: ASR runs on-device via SpeechAnalyzer. The transcript is inserted as-is — no cloud round-trip.
  • local + "Cloud polish after ASR" toggle (Settings → On-device models): same as cloud in spirit, but routed via the local-engine pipeline so the keyboard extension can still use the on-device ASR; the transcript is sent to the configured LLM before insertion. Requires a DeepSeek (or other OpenAI-compatible) API key in the Keychain.

Adding a new LLM provider

Open OSGKeyboardShared/Models/LLMProvider.swift and append a new LLMProvider to the presets array. The default OpenAICompatibleClient handles any endpoint that speaks the POST /chat/completions protocol.

LLMProvider(
    id: "groq",
    name: "Groq",
    defaultBaseURL: "https://api.groq.com/openai/v1",
    defaultModel: "llama-3.1-70b-versatile",
    apiKeyURL: URL(string: "https://console.groq.com/keys")
)

That's it. No other code changes required.


Limitations

  • iOS sandboxes keyboard extensions: ~60 MB memory cap, Full Access required.
  • The keyboard does not work in password fields or some WKWebView textareas (iOS limitation).
  • iOS 26+ only. Earlier iOS versions are not supported.

License

OSGKeyboard Source Available License — personal learning and non-commercial local use only. No commercial use, redistribution, or public forks without permission. Commercial licensing: rocky.hk@gmail.com.


Acknowledgements


Note: the project is published at hkgood/OSGKeyboard; badges and git clone URLs already point there.

S
Description
No description provided
Readme 57 MiB
Languages
Swift 93.7%
Python 4.6%
Shell 0.8%
HTML 0.5%
Objective-C++ 0.3%