Video Audio Sync Lab
A camera and a separate recorder often start at different times. Find the same clap or cue in both players, set their timestamps, and the tool shifts the audio relative to the video. Export keeps the picture duration, fills uncovered time with silence, trims audio beyond the edges, and replaces any sound that was originally in the video.
Key features
- Match a visible video cue to a separate audio cue and fine-tune by milliseconds
- Show the exact signed offset, leading silence and audio trimmed at each edge
- Preview the two sources from the first time both have content
- Replace original sound and export a playable synchronized MP4 or WebM
- Save a timing JSON to reproduce or audit your edit
- Try a generated video plus WAV with a known 600 ms mismatch
How to use
- Choose a short MP4/WebM video and a separate audio file, or create the clap example.
- Seek both players to the same clap or cue and capture their positions, or type exact times.
- Use the millisecond adjustment if needed, and review leading silence and trimming.
- Preview the aligned sources, then build the new video in this browser.
- Play the exported result and download the video and timing JSON.
Use cases
- Match a camera clip with sound from an external recorder
- Replace noisy camera audio with a cleaner microphone recording
- Correct an early or late interview soundtrack
- Document the exact timing offset for another editor
Frequently asked questions
Does it actually create a video with sound?
Yes. The browser decodes the video and separate audio, places sound at the calculated offset, and encodes a new playable MP4 with AVC/AAC or WebM with VP8/Opus according to available browser encoders.
What does a positive offset mean?
The separate audio starts later in the video. For example, a video clap at 1.25 s and audio clap at 0.65 s gives +0.60 s of leading silence before audio begins.
What if sound begins too early or ends too late?
A negative offset discards the beginning of the separate audio. Audio past the video end is trimmed. Missing beginning or ending sound is silence. The picture track keeps its duration; the audio codec may add a few milliseconds of container padding.
What happens to audio already in the video?
It is replaced, not mixed. Only the separate audio file appears in the result. If you need to preserve both tracks, this tool is not suitable.
Which inputs and output quality are supported?
Input video must be a browser-decodable MP4 or WebM up to 40 MiB and 20 seconds; audio must be browser Web Audio-decodable mono or stereo up to 20 MiB and 20 seconds. The result is re-encoded at up to 480 pixels on the long edge and about 12 frames per second. Original subtitles and metadata are not retained.
Does it automatically find the clap?
No. You mark the matching visual and audible times yourself. The example provides known markers. Check the exported preview because source codecs and re-encoding can add small timing differences.
Privacy
Files are decoded and re-encoded locally in this browser. This tool does not upload the files or keep them in browser storage. Exported timing JSON includes the local filenames you selected; clear the workspace to release previews.
Comments & questions