DictaFlowBlog
DictaFlow Blog

MAI-Transcribe-2-Streaming: The 100 ms Test in 2026

Early text is useful. Test the corrected sentence that reaches your app.

October 2, 2026

Unbranded microphone and abstract speech-processing panels
Conceptual speech-processing illustration. The panels do not show actual model output.

Microsoft says MAI-Transcribe-2-Streaming produces its first partial text in just over 100 milliseconds. That result is tentative, though. It doesn't say when a corrected, complete sentence will be ready to add to a document.[1] For everyday dictation, check how long it takes to respond after you stop speaking, then measure how long you need to fix the result.

The word "partial" should appear next to the number. On October 2, The Decoder reported on Microsoft's new streaming transcription model and two speech-generation models.[3] The speech-to-text announcement helps people compare voice input tools. But seeing the first word quickly doesn't answer every question for someone trying to finish an email.

Reading the announcement, I wanted to know how the timing would feel in an email composer after a spoken correction had changed a date. The launch figures do not answer that question.

What Microsoft's 100 ms figure measures

Microsoft announced MAI-Transcribe-2-Streaming on October 1, 2026. In its launch post, the company says the model can transcribe speech in 60 languages. It costs $0.54 per audio hour through the end of the year and returns its first partial results in just over 100 ms. The post also links to an Artificial Analysis accuracy ranking.[1] These figures come from Microsoft, not from a DictaFlow test.

Microsoft's launch includes a chart comparing final transcription accuracy with post-speech latency.[1] That latency isn't just the time needed to produce the first word. Microsoft's developer documentation says intermediate results update the transcript, while final results confirm each segment.[4] The model's output then has to appear in your chosen text field. Measure the time from releasing the dictation hotkey to seeing a corrected sentence in your email app. A model benchmark doesn't cover the whole task.

A partial result lets an interface display words before the speaker finishes. This can help you see whether the microphone picked up speech. A voice agent can also start working with those early words. The Decoder's report discusses this when explaining how agents respond during speech.[3] In an email field, you want a finished sentence you can leave in place.

Consider this made-up test sentence: "Send the proposal to Morgan on Thursday, sorry, Friday." The first words may appear immediately, but the final date depends on the correction. A good test waits until the full sentence is settled. It checks that Friday stays and Thursday disappears without manual cleanup.

Time the email you can actually send

Start with a simple task, such as a short reply that includes a person's name and a corrected date. Use made-up information. Keep the same microphone and computer for every trial, and read the same passage into each tool.

Note when you start speaking, when the first text appears, when you stop, and when the inserted sentence settles. A phone video of the screen is enough for a rough personal comparison. It won't measure latency like a lab test, but it can reveal a long wait hidden by a headline number.

Then edit the text until you'd feel comfortable sending it. Track that time separately. One wrong surname can take longer to fix than several punctuation errors. Check the spelling against the untouched result before judging the tool. A tool that quickly produces almost-right text still leaves you with work.

Include these details in the passage so the check can be repeated:

  • A name you type often, especially one with an unusual spelling.
  • A correction that changes a date or amount.
  • A sentence ending followed by a deliberate pause.
  • A second paragraph, if your usual workflow requires one.

First, dictate the passage into a plain note. Repeat it in your real email composer or document editor. Keep an untouched copy of both versions so you can see which mistakes came from transcription and which appeared when the text entered the app.

For example, a transcript can be accurate but still land in the wrong field. Record that failure too. The person using dictation still has to fix it, regardless of which part of the software caused the problem.

A speech model still needs an insertion workflow

Microsoft released a model, and the launch directs developers to Microsoft Foundry and its playground.[1] That doesn't mean every Microsoft app now has a dictation button. The announcement does not promise system-wide text insertion into every app.

An app built around a speech API still has to handle activation and text insertion at the cursor. The app also has to handle changing partial results. Edits inside a preview are easy enough. But once you move to the next field, later changes need more careful handling.

This is a good place to try hold-to-talk. Hold a key, speak a short block, release it, and see what appears at the cursor. You'll have a clear moment to judge the result. You can test a continuously streaming tool the same way. Use your normal pauses instead of assuming one style works for everyone.

DictaFlow lets you hold a hotkey and dictate text into other apps. I run the company behind it, so I have a commercial interest in this article. Microsoft's announcement doesn't show any integration with DictaFlow, and I don't claim that DictaFlow matches the model latency Microsoft published.

If you use a remote desktop or an app that blocks pasted text, try typing directly into it. DictaFlow's Citrix workflow sends output as typed text for these situations. Record insertion failures separately from recognition errors, including any retries in the remote session.

Keep the benchmark and your buying test separate

I'd keep Microsoft's streaming latency figure on the shortlist. Early feedback matters when a tool otherwise feels slow to respond. The release is worth watching, especially if you build voice interfaces. But first partial latency doesn't replace testing the completed writing.

You can learn a lot from one page of notes. Record the delay before the first text appears, the wait after you stop speaking, any corrections, and failed insertions. Note the exact app beside each result. A browser text box and a hosted work app may give very different results.

Save the test passage and reuse it when a vendor changes a model or an app adds streaming. Include the corrected date each time, then compare the final sentence in the field where you'd send it.

Sources

[1] Microsoft's streaming transcription announcement, October 1, 2026. Read from an archived copy captured October 1, 2026, at 18:29 UTC.[5]

[3] The Decoder's report on Microsoft's transcription and voice models, October 2, 2026.

[4] Microsoft Learn: MAI-Transcribe-2-Streaming overview, updated October 1, 2026, checked live October 2, 2026.