A meeting recording, a research interview, a voice memo, a podcast episode — at some point everyone needs an audio file turned into text. The usual route is uploading it to a cloud service that meters you by the minute and keeps a copy on its servers. Your Mac can do the whole job itself: Tota's Audio Drop transcribes audio files entirely on-device, labels who said what, and exports in five formats — with no upload, no account, and no minute limits.
Drag In a File, Get Back a Transcript
Drop an audio file into Tota's Audio Drop tab and it's transcribed with the same Whisper-class models Tota uses for dictation. Common formats work out of the box — MP3, WAV, AIFF, M4A, and more — and long recordings are handled in streaming chunks, so a multi-hour meeting doesn't need a multi-hour wait or a monster machine. A progress bar shows where you are, and you can cancel at any point.
Speaker Labels, Without the Cloud
The part that usually forces people into cloud tools is diarization — working out who said what. Tota runs speaker identification on-device too: speakers are detected automatically, each gets a colour, and clicking a label renames "Speaker 1" to a real name across the whole transcript and every export. If a recording's voices genuinely can't be separated, you get a clean single-speaker transcript rather than a garbled guess.
Timestamps are up to you: off, once per speaker turn (the default), or on every sentence.
Five Export Formats
Every transcript exports as plain text, Markdown, SRT subtitles, JSON, or CSV — each with copy and save options. That covers pasting minutes into a doc, subtitling a video, and feeding structured data into a pipeline, without a converter in between.
How It Compares
| Tota | Otter.ai | MacWhisper | |
|---|---|---|---|
| Processing | 100% on-device | Cloud | On-device |
| Speaker labels | Yes, automatic | Yes | Yes |
| Free-tier limits | 14-day trial, no quotas | 300 min/month, 3 lifetime file imports | Basic features free |
| Pricing | £19.99 once (includes dictation) | Subscription | One-time (file transcription focus) |
| Also does system-wide dictation | Yes | No | Limited |
The honest framing: if file transcription is your entire job, MacWhisper is a strong specialist tool — see our comparison in Tota vs MacWhisper. Tota's case is that you get file transcription with on-device speaker labels and full system-wide dictation, voice commands, and per-app glossaries in one £19.99 purchase. And if you're currently uploading recordings to Otter for minutes-metered cloud transcription, Tota vs Otter.ai covers that trade-off in detail.
Why On-Device Matters for Recordings
Dictation is ephemeral; recordings are evidence. Client calls, patient consultations, source interviews, board meetings — the recordings people most need transcribed are usually the ones that least belong on a third-party server. On-device transcription means there is no upload to retain, train on, or breach, and it works the same in a Faraday cage as on office WiFi. The privacy reasoning is laid out on our offline dictation page, and every network call Tota can make is enumerated on the security page.
Frequently asked questions
How do I transcribe an audio file on a Mac?
Open Tota's Audio Drop tab and drag the file in (or click to pick one). The transcript is produced entirely on your Mac — nothing is uploaded — with speakers detected and labelled automatically. Common formats like MP3, WAV, AIFF, and M4A are supported.
Can I transcribe an interview with speaker names?
Yes. Tota detects speakers automatically and colour-codes them. Click a label to rename "Speaker 1" to a real name and the whole transcript and every export update. If voices can't be separated, you still get a clean single-speaker transcript.
Is there a limit on audio length or monthly minutes?
No. Tota has no minute quotas or file limits — multi-hour recordings are processed in streaming chunks, and you can cancel at any point. Cloud services meter transcription (Otter's free plan is 300 minutes/month) because server time costs them money; local processing doesn't.
Which export formats are supported?
Plain text, Markdown, SRT subtitles, JSON, and CSV — each with copy and save options, including your speaker names and optional timestamps (off, per speaker turn, or every sentence).
Can I transcribe audio to text on a Mac without uploading it anywhere?
Yes — that's the point of Audio Drop. Transcription and speaker identification both run on-device, so confidential recordings (client calls, medical notes, unreleased interviews) never leave your Mac.

