Skip to main content

Why Does the AI Frame the Wrong Person in Multi-Person Videos?

When your video has multiple speakers, Vizard's AI may zoom in on the wrong person. Here's how the face detection works and what you can do to control framing.

How Vizard's AI Decides Who to Frame

When you upload a video with multiple people on screen, Vizard's AI automatically detects faces and decides who to keep in frame for the vertical clip. The AI uses a combination of signals to make this decision:

  • Face size: The largest, most prominent face in the frame is usually treated as the primary subject.

  • Speech activity: The AI tries to detect who is speaking (based on lip movement and audio cues) and prioritizes keeping the active speaker in frame.

  • Screen position: Faces closer to the center of the frame are weighted more heavily than faces at the edges.

In a one-on-one interview or a solo speaker video, this works well. In panel discussions, group recordings, or videos where two people are roughly equal in size and distance from the camera, the AI can misidentify who the main speaker is — and lock focus on the wrong person, sometimes for an extended portion of the clip.

Why does the AI pick the wrong person?

The AI's face detection is designed for the most common use case: a single dominant speaker. When two or more faces are similar in size and the audio makes it hard to isolate which face is speaking, the AI makes a best-guess call — and it can be wrong. This is especially common in:

  • Panel discussions where speakers are side by side at similar distances from the camera

  • Interview formats where the host and guest are both prominent on screen

  • Videos where there's a speaker in the foreground and someone visible but not speaking in the background

  • Recordings where audio is mixed or unclear, making speaker detection less reliable

How to fix it: Manual crop in the editor

If the AI has framed the wrong person, you can manually adjust the crop in the clip editor:

  1. Open the clip in the editor by clicking Edit.

  2. Click on the video frame in the preview area. You'll see a crop/frame handle appear around the current focus area.

  3. Drag the frame to reposition it onto the correct speaker.

  4. If different speakers are active at different moments, you can split the scene on the timeline and apply a different crop to each segment.

  5. Once satisfied, export the clip.

How to prevent it: Recording tips for multi-person videos

The most effective fix is to make the AI's job easier at the recording stage:

  • Use separate camera angles: If possible, record each speaker with a dedicated camera. Upload the individual speaker feeds so each clip naturally has one dominant face.

  • Position the main speaker more prominently: In a shared frame, having the primary speaker slightly closer to or larger in frame gives the AI a clear dominant subject to track.

  • Avoid overlapping faces: Side-by-side layouts where both faces are the same size are the hardest case for AI framing. A slight difference in size or position helps.

  • For Zoom recordings specifically: Use side-by-side mode (not speaker-stacked mode) before uploading, and see our guide on Vizard and Zoom speaker detection for additional steps.

Can I set framing preferences before generating clips?

Yes — in Clip Preferences, you can choose a layout template. If you're working with a specific video format (e.g., screen share + speaker, two-person interview), selecting the closest matching template can improve how the AI crops the video from the start. Learn more in What are Clip Preferences?

Note that if the selected template doesn't match the actual video content (e.g., you pick a screen-share layout but the video has no screen share), the AI will fall back to its default framing behavior.

Did this answer your question?