Private, system-wide dictation for Windows.
Fast local CPU/GPU transcription, optional cloud services, and safe text insertion in one desktop workflow.
PrimeDictate records from a global hotkey, transcribes with the selected local or cloud engine, optionally cleans or rewrites the result, and safely inserts it into the window that was active when recording started. Speech-to-text and text processing are independent stages: using local STT never requires a cloud text model, and selecting cloud text processing does not upload audio.
- System-wide dictation with configurable Toggle and Push-to-Talk (Hold) modes.
- Local Whisper on CPU, NVIDIA CUDA, or Vulkan-capable AMD/Intel/NVIDIA GPUs.
- Persistent loopback-only Vulkan inference with background model warmup for fast consecutive dictation.
- Separate managed models for faster-whisper (CPU/CUDA) and whisper.cpp GGML (Vulkan).
- Optional Groq, OpenAI, and Gemini Audio cloud transcription.
- Optional per-session Windows output muting during live dictation, with prior mute states restored when recording stops.
- Optional rule-based cleanup, local Ollama/LM Studio, or cloud LLM processing.
- Turkish and English interfaces backed by complete, matching locale catalogs.
- Adaptive voice activity detection, bounded recordings, and background finalization.
- Chunked, cancellable media transcription with overlap de-duplication and TXT/SRT/VTT/JSON export.
- Focus-safe paste, clipboard restoration, searchable local history, and a draggable multi-monitor overlay with explicit processing/ready feedback.
- Credentials in Windows Credential Manager, redacted rotating logs, and privacy-safe diagnostic ZIPs.
- State-aware system tray controls, including on-demand CUDA/Vulkan VRAM release for gaming, and single-instance application lifecycle.
| Speech-to-text and model management | Hotkey and audio settings |
|---|---|
![]() |
![]() |
Floating dictation control:
GPU model memory controls:
| Model loaded — release VRAM for gaming | VRAM released — reload the dictation model |
|---|---|
![]() |
![]() |
Microphone or media file
|
v
1. Speech to Text (required)
CPU / CUDA / Vulkan / cloud STT
|
v
2. Text Processing (optional)
Rules / local LLM / cloud LLM
|
v
Safe paste / history / file export
Cloud fallback is disabled by default and requires explicit consent. When enabled, audio is sent to the configured cloud STT provider only after the chosen local engine fails.
| Engine | Best fit | Runtime and models | Audio leaves device |
|---|---|---|---|
| Local CPU | Compatibility, short dictation, modern fast CPUs | faster-whisper; managed CPU model cache | No |
| NVIDIA CUDA | Sustained high-throughput local transcription | faster-whisper; managed CUDA model cache | No |
| Vulkan | AMD/Intel GPU acceleration and supported NVIDIA systems | whisper.cpp; separate GGML models | No |
| Groq/OpenAI/Gemini STT | Low local resource use or cloud preference | Provider-managed models | Yes |
CPU and GPU can be changed later in Speech to Text; the first-run choice is not permanent. PrimeDictate validates the chosen backend, reports the detected device, and prevents incompatible model/backend combinations.
PrimeDictate warms the selected local engine in the background after startup. The bundled Vulkan path keeps the selected GGML model in a loopback-only whisper-server process, avoiding repeated process and model-loading cost during consecutive dictation. If the persistent server cannot start or answer, PrimeDictate automatically falls back to its verified one-shot CLI path. CPU may still win on some systems, and CUDA is typically preferred on a supported NVIDIA GPU; measure on the target hardware because model size, driver, CPU, GPU, and recording length all matter.
For CUDA or Vulkan, the system tray provides a state-aware Game Mode — Release VRAM action. CPU provides the equivalent Release RAM action for memory-constrained systems. PrimeDictate keeps the selected local model resident by default for minimum dictation latency; after an explicit release, the same action changes to Load Dictation Model. Memory controls are disabled during recording, transcription, model loading, and startup warmup, and are hidden only for cloud engines.
The floating control shows Transcribing while a result is pending. As soon as the result is available, dictation returns to idle immediately and the Play control becomes available; no artificial success cooldown is imposed. Diagnostic logs include Dictation stop-to-result latency=... for end-to-end measurement.
Supported local sizes depend on the runtime. CPU/CUDA use faster-whisper model packages; Vulkan uses compatible whisper.cpp GGML packages and therefore stores a separate copy. large-v3 is not offered where the bundled Vulkan catalog has no compatible artifact; large-v3-turbo is the high-end Vulkan option.
- Open Speech to Text and choose CPU, CUDA, Vulkan, or cloud STT.
- For a local engine, choose and download a compatible model.
- Select the spoken language or automatic detection.
- Optionally configure cleanup or LLM processing under Text Processing - API.
- In Settings, select the microphone, assign a safe global hotkey, and choose Toggle or Hold mode.
- Save, focus any text field, and use the configured shortcut.
The default shortcut is Ctrl+Alt+D. Ordinary unmodified keys are rejected to prevent accidental global capture; function keys may be assigned alone. Hotkeys are registered live and invalid saved values fall back safely.
Installed and portable editions use the current Windows profile rather than writing personal data beside the executable:
%APPDATA%\PrimeDictate\
|-- config.json
|-- history.json
|-- logs\PrimeDictate.log
`-- models\
|-- faster-whisper\ CPU/CUDA packages
`-- whisper.cpp\ Vulkan GGML packages
API credentials are stored in Windows Credential Manager. Diagnostic bundles redact tokens, provider secrets, and Windows user-profile paths; they do not include transcript history, recordings, or API credentials.
Clipboard insertion remembers one foreground target window and its owning process before recording, then restores focus only when safe. If that target cannot be restored or no external target was captured, the result remains on the clipboard instead of being pasted into an unintended window. A sent paste command cannot prove that every target application accepted it. When the Qt clipboard is available, restoration preserves its advertised MIME formats (including rich text, images, and file lists); restoration is skipped if another application or the user changes the clipboard after injection.
Installed and portable builds both run with standard user rights by default. Enable Run PrimeDictate as administrator in Settings only when dictating into elevated applications. PrimeDictate then registers a highest-privilege Task Scheduler entry: Windows asks for UAC once while the setting is saved, and later manual launches run that pre-authorized task without another prompt. Combining it with Start with Windows adds an at-logon trigger and keeps the launch hidden. Disabling administrator mode returns subsequent launches to standard rights and removes the task. The installer may still request elevation to write under Program Files, but it does not force the installed application to remain elevated.
Supported containers include .mp3, .wav, .mp4, .m4a, .mkv, .flac, and .ogg. Long media is decoded incrementally, processed in overlapping bounded chunks, and de-duplicated at chunk boundaries. Jobs can be cancelled; CPU/CUDA stop at safe segment boundaries, a persistent Vulkan or cloud HTTP request may first need to return, and the one-shot Vulkan fallback terminates its active CLI process.
Exports include plain text plus timestamp-aware SRT, VTT, and JSON formats.
For packaged releases:
- Windows 10 or 11, 64-bit.
- A microphone for live dictation.
- Current compatible GPU drivers when using CUDA or Vulkan.
- Enough storage for each selected backend's model files.
For source development: Python 3.12, Git, and Inno Setup 6 when producing the installer.
git clone https://github.com/MaximusPrime/PrimeDictate.git
Set-Location PrimeDictate
python -m venv .venv
.\.venv\Scripts\Activate.ps1
python -m pip install --upgrade pip
python -m pip install -r requirements.txt
python run.pyOnly one PrimeDictate instance runs per Windows user. Use the tray icon to reopen a running instance.
.\.venv\Scripts\python.exe -m compileall -q src run.py build.py
.\.venv\Scripts\python.exe build.py --check
.\.venv\Scripts\python.exe -m unittest discover -s tests -vThe regression suite covers backend dispatch and device validation, model catalogs, cloud contracts and consent, hotkey Toggle/Hold behavior, localization parity, operation coordination, file chunking/export, clipboard safety, diagnostics, navigation, overlay placement, and package resources. Windows CI runs the same compile, preflight, and test checks.
.\.venv\Scripts\python.exe build.pyExpected outputs:
dist\PrimeDictate-Portable.exe
dist\PrimeDictate\PrimeDictate.exe
dist\PrimeDictate-Setup.exe
The build uses the tracked PyInstaller specifications and validates locale catalogs, assets, metadata, and the bundled Vulkan integrity manifest before packaging. Inno Setup must be installed for PrimeDictate-Setup.exe.
PrimeDictate/
|-- run.py Controller and application state machine
|-- build.py Validation and Windows packaging
|-- installer.iss Inno Setup definition
|-- src/
| |-- audio/ Capture, resampling, adaptive VAD
| |-- engine/ STT, models, providers, file jobs
| |-- hotkey/ Global shortcut listener and validation
| |-- injector/ Focus-safe clipboard insertion
| |-- locales/ English and Turkish JSON catalogs
| `-- ui/ Window, pages, tray, overlay, styling
|-- runtime/whisper-vulkan/ Pinned CLI/server runtime, hashes, provenance
|-- tests/ Core and UI regression tests
`-- docs/ Guides, architecture, screenshots
See Architecture for component contracts and data flow.
Release validation uses the Windows paste compatibility matrix. Rows requiring Office, browsers, elevated applications, Remote Desktop or specific GPU hardware are physical release checks and cannot be proven by mocked CI alone.
- The application is Windows-only.
- Vulkan behavior depends on the installed driver and hardware despite runtime preflight checks. The persistent server binds only to loopback on a random port with a per-process unguessable request path.
- CUDA paths, catalogs, and failure handling are covered by automated tests, but this release was not physically benchmarked on an NVIDIA card by the maintainer producing these artifacts.
- Cloud behavior also depends on provider availability, account access, quota, and API changes.
- A successful build cannot replace validation on the target microphone, CPU/GPU, driver, and Windows configuration.
- PrimeDictate produces text; it does not execute spoken operating-system commands.
- Turkish User Guide
- Architecture
- Windows compatibility matrix
- Windows release checklist
- Contributing
- Security Policy
- Vulkan Runtime Provenance
- GPL-3.0 License
Project: PrimeDictate · Website: Maximus Prime Software · Email: maximusprimesoftware@gmail.com
Private by design. Built for productive Windows workflows.







