A fast, local speech-to-text transcription app for Windows. Press a shortcut, speak, get text.
- Toggle recording — Press a global shortcut to start recording, press again to stop
- Local transcription — Powered by whisper.cpp, everything runs on your machine. No cloud, no API keys
- AI polish — Optionally clean up transcriptions with a local Ollama LLM (spelling, grammar, filler words)
- Configurable AI prompt — Customize the system prompt to control how the AI polishes your transcriptions
- Multiple Whisper models — Choose from Tiny to Large V3 Turbo depending on your speed/accuracy needs
- Clipboard & paste — Copy to clipboard, paste at cursor, or both
- Floating launcher — A tiny always-on-top pill that shows recording/transcription status
- System tray — Runs quietly in the tray with quick access to settings
- Transcription history — Browse and manage all past transcriptions
- Built-in API server — The runtime exposes a local HTTP API (
localhost:15000) that other apps can use for transcription and model management
- Press Ctrl+Shift+Space (configurable) to start recording
- Press the same shortcut again to stop recording
- Loon sends the audio to a local runtime server, which runs whisper.cpp
- The transcribed text is copied to your clipboard (or pasted at your cursor)
- Optionally, Ollama polishes the text for better readability
| Layer | Technology |
|---|---|
| Desktop framework | Tauri v2 (Rust + WebView) |
| Frontend | React 19 + TypeScript + Vite + Tailwind CSS |
| Transcription | whisper.cpp (bundled CLI) |
| AI polish | Ollama (local LLM, optional) |
| Audio recording | cpal |
| Database | SQLite via rusqlite |
| Clipboard | arboard + enigo (simulate paste) |
- Windows 10/11
- Ollama (only if you want AI polish — not required for basic transcription)
Download the latest release from the Releases page.
Prerequisites:
# Clone the repo
git clone https://github.com/Ravish-Vishwakarma/loon.git
cd loon
# Install frontend dependencies
npm install
# Run in dev mode
npm run tauri dev
# Build for production
npm run tauri buildLoon runs a local HTTP server on http://localhost:15000 that other apps can use to transcribe audio or manage models.
| Endpoint | Method | Description |
|---|---|---|
/v1/health |
GET | Health check |
/v1/transcribe |
POST | Transcribe audio (multipart form: file + model_id) |
/v1/models/available |
GET | List all available models |
/v1/models/downloaded |
GET | List downloaded models |
/v1/models/download |
POST | Download a model ({ "model_id": "..." }) |
/v1/models/{model_id} |
DELETE | Delete a model |
Example — transcribe a WAV file with curl:
curl -X POST http://localhost:15000/v1/transcribe \
-F "file=@recording.wav" \
-F "model_id=whisper-base"Open Settings by right-clicking the system tray icon → Setting.
| Setting | Description |
|---|---|
| Keyboard Shortcut | Global shortcut to start/stop recording (default: Ctrl+Shift+Space) |
| Output Mode | Copy to clipboard, paste at cursor, or both |
| Active Model | Whisper model used for transcription — download and select one |
| AI Model | Ollama model for polishing (requires Ollama running) |
| Auto Polish | Automatically polish transcriptions after recording |
| Polish Prompt | Custom prompt template for the AI polish step |
| Model | Size | Speed | Accuracy |
|---|---|---|---|
| whisper-tiny | 75 MB | Fastest | Basic |
| whisper-base | 142 MB | Fast | Good |
| whisper-small | 466 MB | Medium | Better |
| whisper-medium | 1.5 GB | Slow | Great |
| whisper-large-v3 | 3.1 GB | Slowest | Best |
| whisper-large-v3-turbo | 1.5 GB | Medium | Best (optimized) |
Models are downloaded on-demand from Hugging Face and stored in the app data directory.
loon/
├── src/ # Frontend (React + TypeScript)
│ ├── pages/
│ │ ├── launcher.tsx # Floating pill launcher UI
│ │ ├── homepage.tsx # Transcription history
│ │ └── settings.tsx # Settings page
│ └── components/ # Shared UI components
├── src-tauri/ # Tauri backend (Rust)
│ └── src/
│ ├── lib.rs # App setup, tray, window management
│ ├── shortcut.rs # Global shortcut + transcription flow
│ ├── recorder.rs # Audio recording + WAV processing
│ ├── config.rs # Config load/save
│ ├── db.rs # SQLite database
│ ├── ollama.rs # Ollama AI polish client
│ ├── clipboard.rs # Clipboard + simulated paste
│ └── proxy.rs # IPC proxy for runtime server
├── runtime/ # Local transcription server (Rust/Axum)
│ └── src/
│ ├── api/ # HTTP endpoints (transcribe, models, health)
│ └── backends/ # whisper.cpp CLI backend
└── whisper/ # Bundled whisper.cpp binaries
MIT
