DictaFlow
DictaFlow Blog

Dot Agent Deck Voice Typing in 2026: Command Boundaries

A practical test for keeping spoken prompt text separate from voice actions.

October 04, 2026

Voice waveforms beside an editable prompt and a separate amber action boundary
A conceptual illustration of voice input connected to coding context, not a product screenshot.

You want to tell a coding agent, "Don't send it until the dry run passes." The voice interface also recognizes "send it" as an action. Before dictating a longer prompt, you need to know whether that sentence will become text or trigger a submission.

Dot Agent Deck's voice typing makes this distinction worth checking. An October 4, 2026 issue asks for more natural mode-switch phrases, including "talking on" and "talking off." The maintainer says the current phrase lists are too narrow. The request still requires matching the entire utterance, so a phrase inside a longer sentence must not switch modes. This is a pending request, not proof that your installed release supports those aliases. Read issue 1544.

This guide covers setup and acceptance testing from the documentation. It isn't a hands-on product review. The fictional examples below help distinguish transcription errors from mistakes at command boundaries. Use a disposable agent session with no production access.

Start with the documented controls

The project's voice guide covers desktop voice control, but the TUI has none. Settings keeps Speech and Commands separate. By default, speech recognition runs in a local container, while commands use an OpenAI-compatible service. Local audio recognition doesn't mean every command stays local.

In an agent pane, "type on" turns on typing mode, and "type off" turns it off. While typing mode is active, ordinary speech becomes prompt text. The guide says a pause alone won't send the prompt. A standalone send phrase can submit it, and some send phrases also work as a separate final sentence. Outside typing mode, a "type ..." request starts a five-second countdown before submission.

Read both paths carefully before turning on the microphone. They handle submissions differently. If you plan to inspect a long instruction, start in typing mode instead of assuming every voice entry path waits for approval.

Before testing, set up a throwaway task. Ask the agent to describe an empty sample directory without making edits or running shell commands. Keep the usual Stop typing and microphone controls visible. Leave enough space to view the prompt and status message at the same time.

Test quoted commands before real work

Handle one sentence at a time. First, dictate: "The button should say send it when the draft is ready." The desired result is the literal text, with the agent still waiting. Before doing anything else, compare the text with what you said.

Then try: "Explain why the end marker belongs after the final record." This tests a short command-like word in ordinary prose. Don't call an unexpected submission an accuracy problem. Accurate words paired with the wrong action are a separate failure.

For a deliberate submit test, use a harmless question and the exact documented send phrase. Watch the prompt field and the agent state. Record which utterance caused the submission. Reset the session before repeating the test so an earlier partial prompt doesn't affect the result.

Check the final sentence separately. Compare a sentence that contains the words "send it" with a short prompt followed by its own "Send it." sentence. The documentation gives special treatment to some send phrases at the end. If your dictation includes quoted interface text, don't end the spoken block with a phrase that also works as a submit command.

You can split up the task. Dictate the explanation, stop typing, then use the keyboard to enter the exact quoted label. That small manual step is better than repeatedly testing an unclear boundary during a live work request.

A stop phrase needs a matching start phrase

The new issue highlights a pairing problem: if you accept a start alias without its matching stop alias, the stop phrase may remain in the prompt. Check the documented pair for your installed version instead of choosing the phrase that sounds natural in conversation.

Keep a short note beside the screen listing the supported start, stop, and microphone-off controls. Keep them separate because stopping dictation and releasing the microphone serve different purposes. You may want to pause text entry but keep voice navigation available, or stop listening completely.

After you enter typing mode, say the supported stop phrase by itself. Check that the mode indicator changes. Then test a normal sentence with similar words, such as "We should stop typing duplicate headers into the export." It should enter the text without switching modes.

If the phrase appears as text instead of stopping, use the visible button and delete the unwanted words. Don't repeat it faster or louder while the agent is still ready to receive input. First reset it to a clean state, then check the transcript and the list of supported commands.

For an upstream report, include the app version, current screen, active mode, exact utterance, visible transcript, and expected outcome. Keep private directory names and real agent prompts out of public issues. A short reproducible example helps more than a recording of a confidential coding session.

Keep transcription, targeting, and permission separate

A useful acceptance test should check more than whether the words are spelled correctly. Put an invented instruction into a sandbox: "Describe the export-header mismatch. Do not edit files. Ask before running a command."

Inspect the transcript for the negative instruction. Then check the destination. Finally, confirm that submission was deliberate. Treat these as separate checks. A correct transcript in the wrong agent pane still isn't a successful dictation.

When testing navigation, stop dictating before opening another agent. Start a new voice block only after you can see the intended destination. Don't change the target and make a long request in the same test. You want a failure you can explain, not a chain of events you have to guess about.

The words "ask before running" don't replace tool permissions. Keep the agent's own approval settings enabled during tests. A dictation layer can capture your instruction, but it can't guarantee that a downstream coding agent will follow it.

When you only need text at the cursor

For voice navigation in Dot Agent Deck, use its native controls. If you already navigate with the keyboard and only want to dictate a prompt, a separate dictation tool may be simpler.

DictaFlow offers hold-to-talk input, app-aware formatting, custom vocabulary, and a Knowledge Base. It runs natively on Windows, Mac, iOS, and Android. Consumer Pro costs $7/month or $69/year. Local AI processing lets you transcribe offline after downloading the model, but it does not make the coding agent work offline.

Using this approach, focus the agent's text field, dictate a block, and review it before clicking the agent interface's send control. Treat it as a workflow to verify in your app, not a claim that every terminal surface inserts text the same way. DictaFlow doesn't replace Dot Agent Deck's navigation or agent-management commands.

Start with the getting-started guide. If you also need dictation in browser tickets, email, or remote desktops, read the comparison guide. Choose native voice control when you need to perform actions. Choose text entry when you want to edit the prompt before deciding whether to send it.