Skip to content
Productivity and professional use cases

Dictating Emails with AI: From Raw Audio to a Polished Message

Turn a messy voice note into a well-toned email, beyond your phone's built-in dictation. A step-by-step guide using AI transcription.

In this article

    Phone voice dictation has been around for years, and it still does the same single thing: turn sound into text, word for word, with commas scattered at random and no sense of what actually matters. If you dictate a long email with Siri or the Google keyboard, you get a literal transcript that you then have to reorder, trim, and proofread anyway. It doesn't draft, it doesn't prioritize, it doesn't adapt the tone.

    This article walks through a different workflow: record your raw ideas as if you were talking to a colleague, upload that audio to an AI transcription and analysis tool, and use the structure it hands back — summary, key points, decisions, action items — as the skeleton for writing the final email in the right tone. This isn't real-time dictation; it's an intermediate step that saves you the work of putting your own thoughts in order.

    The limits of your phone's built-in dictation

    The dictation built into iOS, Android, or any keyboard does exactly one thing: it converts speech into text literally. That works fine for a quick WhatsApp message, but it falls short when you need to write an email that covers several topics, includes specific requests, or has to sound professional.

    • It doesn't distinguish a main point from an offhand comment.
    • It doesn't notice if you mentioned a deadline or a person responsible for something.
    • It doesn't reorganize what you said: if you ramble while speaking, the text rambles right along with you.
    • It doesn't adjust register: whatever you say casually comes out just as casual in writing.

    The result is a draft that needs to be edited almost entirely by hand, which cancels out much of the time savings that dictating instead of typing was supposed to deliver.

    The real workflow: from raw voice note to structured email

    Instead of dictating word by word while carefully composing each sentence, record a voice note the way you'd talk on a call: explain the context, let the ideas come out in whatever order they occur to you, mention names, deadlines, and decisions without worrying about phrasing. You can use your phone's voice memo app or any recorder.

    That audio file (MP3, M4A, WAV, a WhatsApp voice note in OPUS, even an MP4 or MOV video) gets uploaded to VOCAP. Within a few minutes you get the full transcript as plain text, plus an AI analysis with:

    • A summary of what you dictated.
    • Key points broken out by topic.
    • Decisions and action items with owner and deadline, whenever you mentioned them.
    • Open questions you left hanging.
    • Direct quotes for specific phrases you want to keep exactly as said.

    That structured output is what you actually use to write the email: you're no longer starting from a chaotic block of text, but from an organized list of what you need to say.

    Adjusting the tone: from literal transcript to final draft

    VOCAP doesn't write the email for you, and it doesn't automatically apply a formal or casual tone — it transcribes and analyzes the content. Polishing the tone is still your job, but now you're doing it from an already organized base instead of an entire audio recording you've had to listen to twice.

    In practice, this usually means taking the key points and action items VOCAP returns and turning them into email-ready sentences: shorter, free of filler words, with whatever greeting and sign-off fit the person you're writing to. If you said something in the recording you specifically want to keep — a figure, a condition, an exact phrase you negotiated — the quotes section lets you copy it verbatim without digging through the whole transcript.

    For formal emails, it's worth checking that sentences dictated off the cuff don't come across as too casual; for quick internal replies, sometimes copying the key points almost as-is is enough. The difference from native dictation is that you're not starting from zero: you're starting from ideas that are already separated and prioritized.

    Common use cases

    This workflow performs especially well in situations where writing directly is slower than just talking it out:

    • Follow-up emails after a meeting: record a spoken recap the moment you step out, then use the decisions and action items with owner and deadline to write the recap email without losing any detail.
    • Long, explanatory replies: when you need to justify a decision or walk through a process, it's faster to say it out loud and then shape the summary into paragraphs afterward.
    • Dictating while doing something else: driving, walking, or switching between tasks — record the voice note and process it later instead of typing in the moment.
    • Delegating replies: if someone else drafts emails on your behalf, hand them the transcript and the list of action items with owner and deadline instead of a multi-minute audio file they'd have to listen to in full.

    What VOCAP doesn't do (and how to work around it)

    For this workflow to run without surprises, it helps to be clear on what VOCAP is and isn't. It's a transcription and analysis tool for audio or video files you've already recorded — not real-time dictation, and not an app that writes emails automatically.

    • It doesn't transcribe live while you speak: you have to upload the file once it's already recorded.
    • It doesn't separate who's speaking (no speaker diarization), which usually isn't an issue if you're the only one dictating.
    • It has no mobile app and doesn't integrate with your email client, calendar, or CRM: you copy or download the result as a Word or TXT file and paste it wherever you need it.
    • It doesn't keep the audio once processing is finished; the file is deleted from VOCAP's servers at the end (it's processed through OpenAI and Anthropic).
    • It doesn't guarantee 100% accuracy, so it's worth double-checking proper names or figures before sending an important email.

    If you want to try this workflow with your own voice notes, VOCAP offers free minutes to get started and one-time packs of hours when you need to process more volume, with no subscription required.

    FAQ

    Does VOCAP dictate or write the email for me in real time?

    No. VOCAP transcribes and analyzes audio or video files you've already recorded and uploaded; it doesn't work as live dictation and doesn't draft the email automatically, but it does give you the summary and key points to write it faster.

    What audio formats can I upload to dictate an email?

    VOCAP accepts MP3, M4A, WAV, WhatsApp voice notes in OPUS, and also MP4 or MOV video, among other formats, with a limit of 500 MB per file.

    Is the audio from my voice note kept after it's processed?

    No. The audio is deleted from VOCAP's servers once processing finishes; the file is handled through OpenAI and Anthropic during that process.

    Can I use this workflow to prepare emails after a meeting too?

    Yes. Uploading a meeting recording gets you a summary, key points, decisions, and action items with owner and deadline when mentioned — exactly the raw material for writing a follow-up email.

    What languages does the transcription support for dictating emails?

    VOCAP supports more than 50 languages, so you can dictate and transcribe voice notes in English or other languages before writing the final email.

    About the author

    Manuel Gregorio · Founder of VOCAP

    Founder of VOCAP. Since 2024 I help professionals — lawyers, doctors, journalists, podcasters and business teams — turn their recordings into searchable text with AI, GDPR-compliant and from EUR 1/hour.

    15 free minutes · then from €1/h, no subscription

    Start free