Guidelines for AI coding agents working in this repository.
AudioType is a native macOS menu bar app for voice-to-text. Users hold the fn key to record, release to transcribe, and the result is typed into the focused app. It supports three transcription backends: Groq Whisper (cloud), OpenAI Whisper (cloud), and Apple Speech (on-device). If no cloud API key is configured, the app falls back to Apple's on-device SFSpeechRecognizer automatically. It runs as an LSUIElement (no dock icon), built with Swift Package Manager (not Xcode projects).
# Debug build (used during development)
swift build
# Release build
swift build -c release
# Build + create .app bundle (debug)
make app
# Full dev cycle: kill running app, reset Accessibility, rebuild, install to /Applications, launch
make dev
# Clean all build artifacts
make cleanThere is no Xcode project. The app is built entirely via Package.swift + Makefile. Do not create .xcodeproj or .xcworkspace files.
# Run SwiftLint (must pass before merge — CI blocks on failure)
swiftlint lint AudioType
# Format code with swift-format
swift-format -i -r AudioType
# Install lint tools
make setupSwiftLint config is in .swiftlint.yml. Key rules:
force_castis an error — always useas?with guard/if-letopening_brace— opening braces must be on the same line, preceded by a spacetrailing_whitespace,line_length,function_body_length,file_length,type_body_lengthare disabled- See
.swiftlint.ymlfor the full opt-in rule list
# Run all tests
swift test
# Run a single test (by filter)
swift test --filter TestClassName
swift test --filter TestClassName/testMethodNameNote: the test target may be empty. CI runs swift test with continue-on-error: true.
CI runs on every push/PR to main (.github/workflows/ci.yml):
- Lint —
swiftlint lint AudioType(must pass; blocks Build and Test) - Build —
swift build(debug) +swift build -c release+make app+ codesign verify - Test —
swift test(soft failure allowed)
Releases (.github/workflows/release.yml) trigger on v* tags and produce AudioType.dmg + AudioType.zip.
The app uses a protocol-based engine abstraction with a shared base class to support multiple speech-to-text backends:
TranscriptionEngine (protocol)
├── WhisperAPIEngine (base class) — shared WAV encoding, multipart HTTP, response parsing
│ ├── GroqEngine — Cloud, Groq Whisper API, requires API key
│ └── OpenAIEngine — Cloud, OpenAI Whisper/GPT-4o API, requires API key
└── AppleSpeechEngine — On-device, Apple SFSpeechRecognizer, no API key needed
EngineResolver selects the active engine at runtime based on user preference (TranscriptionEngineType):
| Mode | Behavior |
|---|---|
| Auto (default) | Groq if key exists → OpenAI if key exists → Apple Speech |
| Groq Whisper | Always use Groq (fails if no key) |
| OpenAI Whisper | Always use OpenAI (fails if no key) |
| Apple Speech | Always use on-device recognition |
All engines implement a single method: transcribe(samples: [Float]) async throws -> String — accepting 16 kHz mono Float32 PCM samples from AudioRecorder.
fn key held → HotKeyManager → TranscriptionManager.startRecording()
↓
AudioRecorder (AVAudioEngine, 16kHz mono PCM)
↓
┌─ while held (Live Typing on): SpeechPauseDetector watches mic
│ levels; at a natural pause the buffer is drained as a chunk ─┐
↓ │
fn key released → TranscriptionManager.stopRecordingAndTranscribe() │
(remaining audio becomes the final chunk) │
↓ ←─────────────────────────────────┘
per chunk: EngineResolver.resolve() → TranscriptionEngine
↓ ↓ ↓
GroqEngine OpenAIEngine AppleSpeechEngine
(WhisperAPIEngine (WhisperAPIEngine (SFSpeechAudioBuffer-
→ Groq API) → OpenAI API) RecognitionRequest)
↓ ↓ ↓
transcribed text
(chunk API calls may overlap; insertion is
serialized so text lands in spoken order)
↓
TextPostProcessor (corrections)
↓
TextInserter (CGEvent keyboard simulation)
↓
text typed into focused app
Live Typing (chunked transcription while the key is held) is off by default and toggleable in Settings → General. With it off, the session is a single final chunk, the pre-existing behavior.
| Permission | Required for | Plist key |
|---|---|---|
| Microphone | Audio recording | NSMicrophoneUsageDescription |
| Accessibility | Keyboard simulation (TextInserter) | Granted via System Settings |
| Speech Recognition | Apple Speech engine (on-device) | NSSpeechRecognitionUsageDescription |
Speech recognition permission is requested on-demand the first time the Apple Speech engine is used. Cloud engines (Groq, OpenAI) do not require this permission.
AudioType/
App/ # App entry point, menu bar controller, transcription orchestration
AudioTypeApp.swift # @main, AppDelegate, onboarding flow
MenuBarController.swift # NSStatusItem, state-driven icon tinting, overlay windows
TranscriptionManager.swift # State machine (idle→recording→processing→idle/error)
Core/ # Business logic & transcription engines
AudioRecorder.swift # AVAudioEngine capture, PCM→16kHz resampling, RMS level
TranscriptionEngine.swift # TranscriptionEngine protocol, TranscriptionEngineType, EngineResolver
WAVEncoder.swift # WhisperAPIEngine base class, WAVEncoder, WhisperAPIConfig, Data helpers
GroqEngine.swift # GroqEngine subclass, GroqModel enum, TranscriptionLanguage
OpenAIEngine.swift # OpenAIEngine subclass, OpenAIModel enum
AppleSpeechEngine.swift # Apple SFSpeechRecognizer on-device transcription
HotKeyManager.swift # CGEventTap for fn key hold detection
TextInserter.swift # CGEvent keyboard simulation to type into focused app
TextPostProcessor.swift # Post-transcription corrections (tech terms, punctuation)
UI/ # SwiftUI views
RecordingOverlay.swift # Floating waveform (recording) / thinking dots (processing)
OnboardingView.swift # First-launch permission setup (API key optional)
SettingsView.swift # Engine picker, API keys, models, language, permissions
Theme.swift # Brand color system (coral palette, adaptive dark/light)
Utilities/
Permissions.swift # Microphone, Accessibility, Speech Recognition permission helpers
KeychainHelper.swift # macOS Keychain-based secret storage
Resources/
Assets.xcassets/ # Asset catalog (currently empty)
Resources/
Info.plist # Bundle config (LSUIElement, mic + speech recognition usage descriptions)
AppIcon.icns # App icon (coral gradient)
- Sort alphabetically:
import AppKit,import Foundation,import SwiftUI - Use specific submodule imports where appropriate:
import os.log(notimport os) - Only import what the file actually uses
- 2-space indentation (no tabs)
- Opening braces on the same line as the declaration
- No trailing whitespace (rule disabled in linter, but keep it clean)
- Use
// MARK: -sections to organize classes (// MARK: - Private,// MARK: - Transcription)
- Protocols for abstractions with multiple implementations:
TranscriptionEngine - Classes for stateful objects with reference semantics:
TranscriptionManager,AudioRecorder,WhisperAPIEngine,GroqEngine,OpenAIEngine,AppleSpeechEngine - Enums for namespaced constants and error types:
AudioTypeTheme,WhisperAPIError,AppleSpeechError,TranscriptionEngineType,KeychainHelper - Structs for SwiftUI views and config:
RecordingOverlay,SettingsView,WhisperAPIConfig - camelCase for properties/methods, PascalCase for types
- Identifier names: min 1 char, max 50 chars;
x,y,i,j,kare allowed
- Define domain-specific error enums conforming to
Error, LocalizedError - Provide human-readable
errorDescriptionfor every case - Use
do/catchortry?— nevertry! - Never force-cast (
as!) — useas?with guard/if-let - Errors shown to user go through
TranscriptionState.error(String)
- Protocol abstraction:
TranscriptionEnginewithWhisperAPIEnginebase class andAppleSpeechEngine - Inheritance for shared logic:
WhisperAPIEnginebase class handles WAV encoding, multipart HTTP, response parsing;GroqEngineandOpenAIEngineare thin subclasses supplying config - Resolver pattern:
EngineResolver.resolve()picks the engine at runtime based on config - Singleton:
TranscriptionManager.shared,TextPostProcessor.shared,AudioLevelMonitor.shared @MainActoronTranscriptionManager— all state mutations on main thread- NotificationCenter for decoupled state communication (
transcriptionStateChanged); audio levels write directly toAudioLevelMonitor.sharedon the main queue @Published+ ObservableObject for SwiftUI reactivity- Closures for callbacks:
HotKeyManager(callback:),audioRecorder.onLevelUpdate os.logLogger with subsystem"com.audiotype"— use per-class categories
- Create a new subclass of
WhisperAPIEngineinAudioType/Core/ - Override
config(withWhisperAPIConfig) andcurrentModel— that's it for the engine - Add a model enum if the provider has multiple models
- Add static convenience methods (
isConfigured,setApiKey,clearApiKey) - Add a case to
TranscriptionEngineTypeand updateEngineResolver.resolve() - Update
EngineResolver.anyEngineAvailableif the engine has standalone availability - Add UI in
SettingsView.swift(API key field, model picker)
- Create a new class conforming to
TranscriptionEngineinAudioType/Core/ - Implement
displayName,isAvailable, andtranscribe(samples:) - Add a case to
TranscriptionEngineTypeand updateEngineResolver.resolve() - Update
EngineResolver.anyEngineAvailableif the engine has standalone availability - Add any needed permissions to
Permissions.swiftandInfo.plist
All colors live in AudioType/UI/Theme.swift (AudioTypeTheme enum). Never use hardcoded color literals in views. The palette:
- Coral
#FF6B6B— brand color, waveform bars, accents, checkmarks - Amber
#FFB84D— processing state (thinking dots, menu bar icon) - Recording red
#FF4D4D— menu bar icon while recording - Adaptive variants for dark mode (coralLight, amberLight)
- Idle: SF Symbol
waveform.circle.fill,isTemplate = true(follows OS appearance) - Recording: same symbol, tinted
nsRecordingRed,isTemplate = false - Processing:
ellipsis.circle.fill, tintednsAmber,isTemplate = false - Error:
exclamationmark.triangle.fill, tinted.systemRed
- API keys stored in macOS Keychain via
KeychainHelper(Security framework) - Never commit
.env, credentials, or API keys - Audio is recorded in-memory only — never written to disk
The .app bundle is assembled by make app (not Xcode):
- Binary:
.build/debug/AudioType→AudioType.app/Contents/MacOS/AudioType - Plist:
Resources/Info.plist→AudioType.app/Contents/Info.plist - Icon:
Resources/AppIcon.icns→AudioType.app/Contents/Resources/AppIcon.icns - Ad-hoc codesigned:
codesign --force --deep --sign -
When updating the version, change both CFBundleVersion and CFBundleShortVersionString in Resources/Info.plist and the display string in SettingsView.swift.