Real time transcription on a Mac, without the upload.
Coco writes down a meeting, a lecture, an interview or a podcast while it is still going on, using Apple's speech engine on your own machine. When you stop, you have a timestamped text file with each line marked Me or Others, sitting in a folder you control.
Why the wait is the worst part of transcription
Most transcription on a Mac still works the old way. You record the meeting, you find the file, you upload it, you wait, and some minutes later a web page shows you the words. The waiting is not the only cost. The audio has left your laptop, your minutes have been counted against a plan, and the finished transcript lives in someone else's account rather than in your documents folder. For a client interview or an internal review, that last part is often the one that stops people using these tools at all.
Real time transcription flips the order. The words appear as they are spoken, so the transcript is finished the moment the meeting is. There is nothing to upload because the recognition already happened, and nothing to wait for because you were reading the output the whole time. What you are left with is an artifact: a plain text file with a timestamp on every line, the speaker marked, searchable with Spotlight or grep like any other file on your disk. You can open it a year later without logging in anywhere.
Coco does this with Apple's own speech framework, the one behind dictation, running on the Mac whenever Apple supports on-device recognition for your Mac's language. That is what it takes to transcribe audio on a Mac without a single byte leaving it. If you mainly want to read along while audio plays rather than keep the file, Coco's live captions page at /features/live-captions covers the same engine from the reading-along side. This page is about what is left on disk afterwards: the transcript as a document, not as a strip of text that vanishes when the window closes.
What you get from a Coco transcription session
The words arrive while people are still talking
There is no record step and no process step. You choose Start, the window appears, and lines begin filling in a second or two behind the speech. Apple's engine closes each utterance when a speaker pauses, and Coco quietly opens a fresh request so the transcript keeps accumulating instead of rewriting one line over and over. That is the difference between a dictation box and a transcript. You can scroll back to the start of a two-hour meeting while the meeting is still running, and read what was said in the first ten minutes.
It hears the call and it hears you
System audio comes in through ScreenCaptureKit, the same route macOS uses for screen recording, so it covers any app: Zoom, Teams, Meet, a browser tab, a podcast player, a lecture recording. The microphone is captured separately through AVAudioEngine at the same time. Both streams land in one list, mic lines labelled Me and everything from your speakers labelled Others. That is two-channel labelling, not voice fingerprinting, so a room with four people on one microphone comes out as one speaker. For a call it is usually the split that matters.
No minute counter, no account, no upload
When Apple supports on-device recognition for your Mac's current language, Coco requires it, which means the audio never leaves the machine and a session has no time ceiling. A three-hour workshop costs the same as a five-minute standup, which is nothing, because Coco is a one-time purchase rather than a plan with minutes attached. There is no sign-up, no server of ours and no telemetry in this feature. Where Apple has no on-device model for your language, the framework falls back to Apple's servers exactly as dictation does, and needs a connection.
The output is a file you own
Stop the session and Coco writes a .txt with a timestamp on every line, plus two .m4a tracks, one for the microphone and one for system audio. They go to Live Captions/Recordings inside iCloud Drive/Coco when iCloud sync is on, and to ~/Library/Application Support/Coco when it is off. Keep Original Recordings is on by default; turn it off and a stopped session keeps only the text. A Recordings window lists past sessions so you can reread or copy one whole, and Copy Transcript puts the current session on the clipboard.
It follows your Mac's language, and says so when it cannot
Recognition uses SFSpeechRecognizer with punctuation on, in whatever language your Mac is set to in System Settings. There is no separate language picker inside Coco yet, so transcribing a Japanese interview on an English Mac means switching the system language first. That is a real limitation and we would rather write it down than let you find it during a meeting. If Apple has no recognizer for the language your Mac is set to, the window tells you instead of quietly producing nonsense.
Coco next to Otter.ai, macOS Live Captions and Raycast
Otter.ai is the tool most people mean by meeting transcription, and it is genuinely better than Coco at several things. It separates speakers by voice rather than by channel, syncs a meeting to your phone and your colleagues, writes a summary, and exports subtitle formats. It does that by taking your audio to a server, behind an account, with minutes counted against a plan. If you need a shared workspace and speaker names, that trade is worth making and Coco is not the tool.
Coco is the other trade. Nothing is uploaded, nothing is metered, there is no login, and the session ends as a text file and two audio tracks in a folder on your Mac. Apple's own Live Captions is free and on-device too, but it is a reading aid: it does not keep the session, mark who spoke, or leave you a file. Raycast is a launcher and does not transcribe a meeting at all; it is in the table because Coco is also a launcher for the rest of the day.
| Coco | Otter.ai | macOS Live Captions | Raycast | |
|---|---|---|---|---|
| Real-time (words appear while speaking) | ||||
| Runs on the device, works offline | ||||
| No account or sign-in | ||||
| No per-minute or monthly transcription limit | ||||
| One-time purchase, no subscription | ||||
| Transcribes system audio from any app | ||||
| Microphone and system audio at the same time | ||||
| Session saved as a plain .txt you own | ||||
| Original audio kept locally | ||||
| Your audio is never uploaded | ||||
| Tells apart speakers by voice | ||||
| AI summary of the meeting | ||||
| Shared team workspace | ||||
| Mobile app for recording away from the Mac | ||||
| Exports subtitle formats such as SRT | ||||
| Bundled with an app launcher |
Transcribing a meeting on your Mac with Coco
Start the session before the meeting does
Click Coco in the menu bar, open the Live Transcription submenu and choose Start. The window appears and begins listening straight away, so open it a minute early rather than during the introductions. Show brings it back if you hid it, and Stop ends the session and writes the files.
Grant the three permissions once
System audio needs Screen Recording, because that is where macOS keeps it. The microphone needs Microphone, and recognition needs Speech Recognition. macOS asks for each the first time. If you refused one before, enable it under Privacy & Security and relaunch Coco, because the screen-recording grant does not apply until you do.
Leave it running and read along if you want
The panel floats above your call without stealing focus, stays pinned across Spaces and over fullscreen apps, and can be moved, resized and set to one of three text sizes. Pause stops new lines without stopping the recording. Do not close the window mid-meeting, because closing it ends the session.
Collect the transcript afterwards
Stop the session, then choose Recordings from the same submenu to read past sessions and copy one whole. The .txt and the two .m4a tracks sit in Live Captions/Recordings inside iCloud Drive/Coco, or in ~/Library/Application Support/Coco with iCloud sync off. Clear All in settings deletes every stored session.
Questions about real time transcription on a Mac
- Can I transcribe a Zoom or Teams meeting without a bot joining it?
- Yes. Coco takes system audio through ScreenCaptureKit, the route macOS uses for screen recording, so it hears whatever any app is playing. That covers Zoom, Teams, Meet, a browser tab, a webinar player or a recorded lecture, with no plugin inside the meeting app and no assistant appearing in the participant list. Your microphone is captured at the same time, so the transcript shows the other side as Others and you as Me, in the order things were said. The only requirement is the Screen Recording permission, which macOS asks for once and which needs a relaunch of Coco before it takes effect.
- How long can one transcription session be?
- As long as the meeting is, when recognition is running on the device. Apple imposes no session ceiling in that mode, and Coco reissues a recognition request each time the engine closes an utterance, so a three-hour workshop accumulates into one continuous transcript rather than being cut into chunks. Nothing is metered, because there are no minutes to meter: Coco is a one-time purchase and there is no plan behind it. The practical limits are ordinary ones, disk space for the audio and battery if you are unplugged. If your Mac's language has no on-device model, recognition goes through Apple's servers instead and behaves like dictation.
- What exactly is in the file at the end?
- A plain .txt with one line per utterance, each carrying a timestamp and a speaker label of Me or Others, plus two .m4a tracks, one recorded from the microphone and one from system audio. Keep Original Recordings is on by default and turning it off leaves only the text. The files go to Live Captions/Recordings inside iCloud Drive/Coco when iCloud sync is on, and to ~/Library/Application Support/Coco when it is off. They are ordinary files: searchable with Spotlight, greppable, openable in any editor, and yours to move, back up or delete without asking anything for permission.
- Does it know who is speaking?
- Only in the two-channel sense. Lines from your microphone are labelled Me and everything coming out of your speakers is labelled Others, which is the split that matters on a call. Coco does not fingerprint voices, so it cannot tell you that the second speaker was Priya and the third was Tom, and a meeting room where four people share one microphone comes out as a single Me. Tools like Otter.ai do proper voice diarisation and the comparison table above says so plainly. If named speakers are what your work needs, that is the honest recommendation.
- How is this different from Coco's live captions?
- It is the same feature described from the other end. Coco's live captions page at /features/live-captions is written for reading along while audio plays, when a call is in a language you follow better on paper or the volume has to stay down. This page is about the file that exists afterwards. Both start from the same menu bar item, the same on-device Apple recognition and the same floating window, and both leave the same .txt and .m4a behind. Which page you read depends on whether you want the words during the meeting or a searchable record of it a month later.
- Can it translate the transcript, or export subtitles?
- Neither, and both are worth being clear about. Coco transcribes into the language being spoken and does not translate; translation is on the roadmap and the live captions page goes into more detail on that. It also writes plain text rather than SRT or VTT, so it is not a subtitling tool and there is no timed-caption export. What it gives you is a timestamped text file and the original audio, which is enough to feed whichever translator or subtitle editor you already use. Anything beyond that we would rather you got from a tool built for it.
Copy the prompt, or send it straight to the AI you already use — it drives Coco through the same commands you would type yourself.
Install the Coco skill (npx skills add butterflydream-ai/Coco --skill coco -g -y), then use coco to help me with Real time transcription on a Mac, without the upload..