
A Zoom recording has one structural problem no amount of editing skill fixes after the fact: by default, everyone's audio gets mixed into a single track in real time, compressed for a live call, not for post-production. If cross-talk, a ringing phone, or a dog barking in the background survives the mix-down, it is baked into every speaker's audio for good. The fix happens before you record, not in the edit: in Zoom's desktop app, under Settings > Recordings, turning on "Record a separate audio file for each participant" gives every speaker their own track, so you can clean, level, and denoise each one independently instead of fighting a single flattened mix. The setting only applies to local recordings on the desktop app, up to 80 separate tracks, and it does nothing for anyone who calls in by phone, since phone participants share one track regardless.
Zoom's gallery view resizes tiles based on who is talking, which looks fine live and terrible in a finished video, where a subject's frame should not jump size every time someone else unmutes. Lock speaker view during recording if the piece is a single-focus interview like a talking-head video, or plan to crop each participant's tile to a fixed frame in post if it is a multi-person panel, closer to editing a multicam interview. Either decision is easier to make before the call than after 90 minutes of footage exists.
Once the audio and framing problems are handled, a Zoom recording is functionally the same editing problem as any other interview: too much material, and a story buried somewhere inside it. Scrubbing through an hour of remote conversation to find the good parts is slow, and worse on Zoom footage specifically, because remote conversations run longer and looser than in-person ones. This is text-based editing: working from a transcript instead, reading the conversation and marking what is worth keeping, is the faster path regardless of where the footage came from. In ScriptCut, that transcript carries word-level timecodes, so a selected sentence maps to an exact frame in the recording, whether it is a single Zoom file or the separate per-participant tracks Zoom can output.
For a quick visual walkthrough of trimming, cropping, and captioning a Zoom recording before it goes anywhere near an NLE, this covers the basics:
Cross-talk is the most common complaint about remote interviews, and separate audio tracks are the actual fix, not a workaround. With one mixed track, two people talking over each other is one unusable moment. With separate tracks, you can duck one speaker under the other, choose which voice leads, or cut around the overlap entirely, none of which is possible once two voices are baked into a single file. If you did not enable separate tracks before recording, cross-talk moments in the finished mix usually have to be cut rather than salvaged.
Zoom is fine for a casual conversation, but purpose-built remote recording tools solve the audio problem structurally instead of through a settings toggle. Riverside.fm and similar tools record each participant locally on their own device in uncompressed audio and up to 4K video, then upload progressively in the background, so a bad internet connection during the call does not degrade the final file the way it can on Zoom's real-time stream.
| Tool | Recording method | Per-speaker tracks |
|---|---|---|
| Zoom | Real-time call, local or cloud recording | Only with "separate audio" enabled, desktop only, local recording only |
| Riverside.fm | Local recording on each device, progressive cloud upload | Yes, by default, uncompressed audio and up to 4K video |
| SquadCast | Local recording on each device, progressive cloud upload | Yes, by default |
| StreamYard | Browser-based live streaming with recording | Limited; built more for live production than raw multitrack capture |
If remote interviews or podcast episodes are a recurring part of the workflow, switching the recording tool solves the audio problem once, and it is worth doing before you are batch editing a full season of them. If it is a one-off, cleaning up a Zoom file the way described above is usually enough.
Enable separate audio tracks in Zoom's recording settings before the call. Lock or plan framing for how the finished video should look, not how gallery view behaves live. Record, then pull the transcript (Zoom generates one automatically for cloud recordings, or run the local file through a transcription tool). Read the transcript, mark the material worth keeping, and cut filler words and dead air inline. Arrange the selected moments into order, clean each speaker's audio track individually where cross-talk or noise survived, and export the sequence to an NLE for final mix and any graphics.
Recording without separate audio tracks and discovering the cross-talk problem only in the edit is the most common and the most preventable. Editing straight off the single mixed-down recording instead of pulling a transcript first wastes hours re-listening for quotes. And leaving Zoom's gallery-view resizing untouched in the final cut, so a speaker's frame visibly jumps size mid-sentence, is a small thing that reads as unpolished the moment a viewer notices it.
Related reading: What Is Text-Based Editing?, How to Cut a Documentary Interview, and What Is Timecode?.
Turn on separate per-participant audio tracks before recording, fix framing issues from gallery view, then work from a transcript to find and select the material worth keeping before final mix and export.
A separate track per speaker. Enable "Record a separate audio file for each participant" in Zoom's desktop recording settings before the call, so overlapping voices and background noise can be cleaned independently instead of fighting one mixed-down track.
Speaker view for a single-focus interview, since gallery view resizes tiles live and creates visible frame jumps in the finished edit. Plan a fixed crop per participant if the piece is a multi-person panel.
Separate audio tracks, enabled before recording, are the real fix, letting you duck or isolate one speaker under another. Without separate tracks, cross-talk baked into a single mixed file usually has to be cut rather than repaired.
Only for participants on the desktop app using local recording, up to 80 tracks. Anyone who joins by phone shares a single audio track regardless of the setting.
If remote interviews are a recurring workflow, yes. Riverside.fm and similar tools record each participant locally in uncompressed audio and up to 4K video, avoiding the real-time compression and dropped-connection risk built into a live Zoom call.