Audio Pitch Tempo Studio
Set a semitone shift and a tempo ratio separately. A phase-vocoder pass changes the time scale, then windowed-sinc resampling changes pitch while retaining the duration selected by the tempo control. A +12-semitone shift with tempo 1× aims for an octave higher at the same duration; tempo 1.5× with pitch zero aims for a shorter file at the same steady-note frequency. Play both versions before saving. This basic algorithm has no transient phase locking or formant preservation, so drums can smear and voices can sound thin or synthetic.
Key features
- Adjust pitch from -12 to +12 semitones and tempo from 0.5× to 2× independently.
- Use a 2048-point phase vocoder with a 512-sample analysis hop and per-bin phase propagation for time scaling.
- Apply a 16-sample-radius windowed-sinc resampler with a low-pass step when raising pitch.
- Preview unchanged source and rendered result at their original levels and see exact output frames, length and peak levels.
- Export 16-bit PCM WAV and a JSON manifest with settings, source SHA-256 and method limits; cancel processing in a worker.
How to use
- Choose browser-decodable audio up to 8 MiB and 0.5–20 seconds after decoding, or load the 440 Hz WAV example.
- Enter semitones (-12 to +12) and a tempo ratio (0.5 to 2); read the predicted output duration.
- Render, then A/B listen for pitch, timing, phase smear and changed vocal timbre.
- Save the rendered WAV and optional settings JSON; lower source level elsewhere if a clipping warning appears.
Use cases
- Raise a practice track by a few semitones while keeping roughly the same duration.
- Slow a steady instrumental phrase for study without lowering its note pitch.
- Create two reproducible versions of a short sound effect with different pitch and timing.
Frequently asked questions
Are pitch and tempo really independent?
The output frame count is rounded from input frames divided by tempo ratio. Semitone shift sets a separate pitch factor of 2^(semitones/12). The phase vocoder first stretches by pitchFactor/tempoRatio and sinc resampling then brings the length to the tempo target. Steady tones are tested at 440, 880 and 220 Hz, but complex audio can deviate slightly.
Does changing tempo keep speech and drums perfectly clear?
No. This is a basic per-bin phase vocoder without transient phase locking. Drum attacks may smear or sound flanged, and voices may have a phasey quality. Preview the output, especially at extreme ratios.
Does pitch shifting keep vocal formants?
No. Formants move with the pitch factor; this can make voices sound smaller or larger. The tool does not perform vocal formant correction, beat detection, or source separation.
Why is the saved audio different from the original at neutral settings?
At zero semitones and 1× tempo the decoded floating-point samples are copied exactly in memory. The download is still 16-bit PCM WAV, so lossy or higher-bit-depth input can change precision. Samples outside ±1 are clipped during WAV writing and counted.
What input and output limits apply?
The encoded file may be at most 8 MiB. Browser-decoded mono/stereo audio must be 0.5–20 seconds, resampled to 48 kHz. Tempo 0.5× can produce up to 40 seconds; the intermediate phase-vocoder buffer is capped at four million frames.
Is my file sent to a server?
No. Decode, phase processing, resampling and playback happen in browser memory. JSON contains the filename, byte size, SHA-256 and settings, but no audio samples. Saving WAV is an explicit local download.
Privacy
Audio is decoded, rendered and played locally in the browser. This tool sends no selected file or PCM arrays to a server API. The optional JSON includes a source SHA-256 and settings but no sound samples.
Comments & questions