Audio Description Editor
Build a separate audio-description track for a short video. Mark dialogue times yourself, then record a scene description with microphone permission or import an existing recording. Place each cue inside an available gap, review overlaps and out-of-range warnings, and mix the cue with the original video audio. The WAV lasts exactly as long as the video; JSON/CSV record the placement. No speech recognition, automatic scene description or dialogue detection is used.
Key features
- Preview a local video and manually mark protected dialogue intervals
- Show the remaining gaps where a voice description can fit
- Record a microphone cue or upload a browser-decodable audio file instead
- Check cue length, video boundary, dialogue collision and cue-to-cue overlap before export
- Mute a narration cue, adjust its time and gain, and lower the source bed during active cues
- Mix actual samples into a fixed-length stereo WAV and export cue JSON/CSV
How to use
- Choose an MP4, MOV or WebM video up to 15 seconds or load the example.
- Mark each dialogue interval yourself and inspect the available gaps.
- Enter a scene note and start time, then record up to six seconds or upload a voice clip. Add more cues if needed.
- Move or mute cues until dialogue, cue-overlap and video-boundary warnings are gone.
- Render the full-length WAV, preview it with the muted video, and download WAV plus JSON/CSV cue records.
Use cases
- Add a brief description of a visual action between two spoken lines
- Check whether a recorded cue fits a marked silence before publishing
- Export cue times and source hashes for a human accessibility review
- Listen to the original audio bed underneath a separately recorded narration cue
Frequently asked questions
Does the tool write or recognize the scene description for me?
No. You write the scene note, mark dialogue times and supply the narration recording. It does not recognize speech, find dialogue automatically, identify scenes or synthesize a voice.
Can I use it without microphone permission?
Yes. Upload a previously recorded WAV, MP3, M4A, OGG or WebM audio file that your browser can decode. Microphone access is requested only when you press Record, and its tracks stop after recording, cancel or clear.
What happens if a cue overlaps dialogue or another cue?
Unmuted overlaps and cues past the video end are listed as warnings and block export. Dialogue windows are manually marked and are not verified against spoken content. Mute or move a cue, or correct the window.
Does the output include a video file?
No. The result is a stereo 48 kHz PCM WAV with original video audio when decodable plus narration, and cue JSON/CSV. Play the WAV beside the original video for review; export does not mux a new video or retain captions and video metadata.
How are original audio and narration combined?
The original video audio remains full level outside narration. Under active cues it fades toward the chosen bed gain over about 20 ms and returns afterward. The narration is added at its cue start frame. If peaks exceed full scale, the whole WAV is normalized to a 0.98 peak; the pre-peak and gain are reported.
What limits and timing accuracy apply?
Video is limited to 32 MiB and 15 seconds; each audio cue is up to 8 MiB and six decoded seconds, with at most eight cues and 16 dialogue windows. WAV start times round to 48 kHz samples. Browser codecs and MP4 audio timing can vary, and simultaneous preview may drift; review the downloaded WAV against the video.
Privacy
Video and voice files are decoded and mixed locally in your browser. This tool does not upload source media or cue text. The downloadable JSON contains source names and SHA-256 hashes, so review it before sharing. Clearing releases temporary URLs and microphone tracks.
Comments & questions