Mountain landscape

How to Transcribe an Interview on a Mac — Offline, With Speaker Names

By Ryan Crabbe, developer of TotaLast verified against Tota on 31 July 2026

An hour of interview audio is an evening of typing — which is why interview transcription became a paid cloud industry with per-minute meters. But if you work on a Mac, the machine in front of you can do the whole job: transcribe the recording, work out who said what, and hand you a labelled transcript — without the audio ever leaving your disk. For journalists protecting sources and researchers bound by consent agreements, that last part isn't a nice-to-have.

TL;DR

Save your interview as an audio file, drop it into Tota's Audio Drop tab, rename "Speaker 1" and "Speaker 2" to real names, and export as text, Markdown, SRT, JSON, or CSV. Runs entirely on your Mac; no upload, no per-minute pricing.

Step 1: Get the Recording as a File

Audio Drop accepts the formats interviews actually arrive in — MP3, WAV, AIFF, M4A, and more:

  • Zoom / video calls: use the recording's audio file (Zoom saves one alongside the video), or export audio from your editor.
  • iPhone or voice recorder: AirDrop the file to your Mac, or drag it out of the Voice Memos app — the voice memos guide covers that in detail.
  • Multi-hour recordings: fine as they are. Tota processes long files in streaming chunks, so there's no need to split them first.

Step 2: Pick a Model Worthy of the Audio

Interviews are the hard case for speech recognition: two voices, different microphones and accents, café clatter in the background. This is exactly where the bigger Whisper tiers earn their download — before a batch of interviews, load Large v3 Turbo in Settings. The trade-offs between model sizes are covered in Whisper model sizes explained.

Step 3: Drop the File In

  1. Open Tota and go to the Audio Drop tab in the sidebar.
  2. Drag the recording in, or click to pick it.
  3. Watch the progress bar — you can cancel at any point. Everything runs locally; it works the same with WiFi off.

Step 4: Name the Speakers

This is the step that separates interview transcription from plain transcription. Tota runs speaker identification (diarization) on-device: each voice is detected automatically and colour-coded. Click a label, type the person's name, and "Speaker 2" becomes "Dr. Okafor" across the entire transcript and every export.

Timestamps are configurable — off, once per speaker turn (the default), or every sentence. For quoting, per-turn timestamps make it easy to jump back to the recording and check the exact wording.

Step 5: Export for Wherever the Work Happens

  • Markdown — into Obsidian, Notion, or your notes system, speaker names intact.
  • Plain text — for pasting quotes into a draft.
  • CSV / JSON — structured turns for qualitative analysis tools or your own scripts.
  • SRT — if the interview is destined for video with subtitles.

Why Offline Matters More for Interviews Than Anything Else

A dictated Slack message is ephemeral; an interview recording is a record of a real person saying real things, often under specific promises. Uploading it to a transcription service adds a third party to that promise — their retention policy, their training-data terms, their breach surface. On-device transcription keeps the chain of custody at one machine: yours. The full reasoning is on our transcribe audio files on Mac page, and every network call Tota can make is enumerated on the security page.

Honest note: if your newsroom or lab needs shared workspaces, collaborative highlighting, and meeting-bot integrations, that's what cloud tools like Otter.ai are actually for — Tota vs Otter.ai lays out that trade-off honestly. Tota's case is the individual doing the interviewing: unlimited local transcription with speaker names, for the price of two transcribed hours at a typical human-transcription rate.

Frequently asked questions

Can it tell the interviewer and interviewee apart?

Yes. Tota detects speakers automatically and labels each turn — click a label to rename "Speaker 1" to a real name, and the whole transcript plus every export updates. If a recording's voices genuinely can't be separated, you get a clean single-speaker transcript rather than a garbled guess.

Does it work with Zoom, phone, and voice-recorder files?

Yes. Anything you can save as a common audio file works: MP3, WAV, AIFF, M4A, and more. Export the audio from Zoom or your recorder app, then drop the file into Tota's Audio Drop tab.

How accurate is the transcription?

Tota uses OpenAI's Whisper models running locally, which is the same model family behind most modern transcription services. Accuracy depends mostly on your audio and the model tier you load — for interviews with accents, jargon, or room noise, load Large v3 Turbo before transcribing.

Is the recording uploaded anywhere?

No. Transcription and speaker identification both run on your Mac, and the process works with WiFi off. There is no server copy of the recording or the transcript — which matters when the recording involves a confidential source or was made under a specific consent agreement.

What does it cost per hour of audio?

Nothing per hour. Tota is a one-time £19.99 purchase with a 14-day free trial — no subscription, no per-minute metering, no monthly quota. Cloud services like Otter.ai meter free accounts at 300 minutes per month; a local app has no reason to count your minutes.