Local dictation
Speak a task, then review the words before sending.
Models are downloaded only on request. You can cancel a download, remove a model or choose another ready model in voice settings; downloading does not change your selection or start recording. Audio is processed locally through whisper.cpp. Temporary audio and recognition files are removed after processing or cancellation. No transcription API key or per-request transcription payment is needed. After you send it, the text becomes part of the task and follows your selected agent or provider connection.
Step by step
Use VECTA desktop on a Mac with Apple Silicon. Open a new task or an existing task conversation and select the microphone in the writing area.
On first use, review the model information and choose Download. Small is the default: 487,601,967 bytes, about 488 MB (465 MiB). Optional Base is 147,951,465 bytes, about 148 MB (141 MiB). Wait until the selected model is ready.
Choose your speech language or automatic detection. Voice settings let you choose a microphone and test its input level; the microphone test does not need a downloaded speech model.
Start recording when you are ready and allow microphone access if macOS asks. Watch the input level and timer. Stop to insert the recognised words at the saved cursor position, or cancel to discard the recording.
Read and edit the draft, especially names and numbers. Send only when you want the task to continue. If you edited the draft during recognition, insert the offered transcript explicitly to preserve your changes.