You downloaded a 48-minute webinar as MP4 and only need the spoken track for your commute — not another gigabyte of video on your phone. Itqan's extract audio from video tool pulls the soundtrack from MP4, MOV, MKV, and eight other container formats, then saves MP3, WAV, AAC, or FLAC. Choose full audio, an approximate vocal stem, or an instrumental stem. Up to 200 MB per upload from your device, Google Drive, or Dropbox. No install, no account. This guide is written manually for platform users — it covers extraction modes, format choice, four real workflows, a pre-upload checklist, and what to do after the file lands on your desktop.
What the video audio extractor does
Every video file carries at least one audio track — narration, music, ambient sound, or a mix. The extractor demuxes that track and encodes it as a standalone audio file you can play, edit, or transcribe without opening a video editor.
This differs from the audio format converter, which converts existing audio files. Here the source is video and the result is audio — read directly from the embedded stream, not re-recorded through speakers.
When to extract audio from video
Not every project needs the picture. Audio alone is lighter, easier to scrub in a player, and compatible with tools that ignore video containers.
- Commute listening: Turn a recorded lecture or conference talk into an MP3 for offline playback.
- Podcast repurposing: Your show was filmed for YouTube — extract the track and publish the audio RSS feed without re-exporting from an NLE.
- Transcription pipeline: Strip audio first, then send the WAV or MP3 to audio to text for Whisper transcription.
- Post-production: Export WAV for editing in Premiere, DaVinci, or Audacity.
- Archiving: Store FLAC copies while deleting bulky 4K video masters.
Extraction modes explained
Itqan offers three modes. Full audio is the default and the most reliable. Vocal and instrumental modes apply source-separation heuristics — useful experiments, not studio-grade stems.
| Mode | What you get | Best for | Caveats |
|---|---|---|---|
| Full audio | Complete original soundtrack, unchanged | Lectures, interviews, screen captures, any workflow where you want exactly what is in the file | Includes background music and noise — clean later if needed |
| Vocals | Stem weighted toward human speech and singing | Interviews buried under light music, lecture clips with intro jingles you want to reduce | Leakage from drums and bass is common; not a forensic isolation tool |
| Instrumental | Stem weighted toward accompaniment | Karaoke practice, remix sketches | Vocal bleed on complex mixes |
Source separation depends on how the mix was built. Mono speech over silence splits poorly; stereo pop with centred vocals gives more noticeable — still approximate — results. When in doubt, use full audio first.
Input and output formats
Supported video containers
MP4, AVI, MOV, MKV, FLV, WMV, WEBM, M4V, and 3GP. Phones and screen recorders usually produce MP4 or MOV. Itqan extracts the primary audio stream in the file you upload.
Output audio codecs
MP3 (sharing), WAV (editing), AAC (Apple devices), FLAC (lossless archive). The tool re-encodes to your chosen format.
Size limit
Each upload may be up to 200 MB. A 45-minute 720p screen recording often fits; a long 1080p OBS capture may not. Trim the video locally or split it into chapters before uploading.
Choosing an output format
- MP3 — Default choice for sharing, messaging apps, and car stereo USB sticks. Smaller files, acceptable speech quality at standard bitrates.
- WAV — Pick when you will cut, normalise, or apply effects in Audacity or a DAW. Files are large but avoid generation loss during editing rounds.
- AAC — Good compromise for iPhone voice memos, iTunes libraries, and platforms that prefer .m4a containers.
- FLAC — Archive concerts, worship services, or legal depositions where you must preserve fidelity without keeping the video wrapper.
If the next step is transcription, MP3 or WAV both work in audio to text. WAV can help Whisper on very quiet recordings because it avoids additional lossy compression — at the cost of upload size.
Step-by-step walkthrough
- Open extract audio from video on desktop or mobile.
- Upload from your computer, Google Drive, or Dropbox, or drag a file onto the drop zone.
- Confirm the sidebar shows the correct file name, size, and format.
- Select an extraction mode: full audio, vocals, or instrumental.
- Choose an output format (MP3, WAV, AAC, or FLAC).
- Preview the embedded audio in the built-in player if you want to sanity-check levels before processing.
- Click Start extraction and wait for the progress indicator to finish.
- Download the audio file and verify playback start-to-end before deleting the source video from your workflow folder.
Webinar — MP4 to MP3
A coordinator saves a 52-minute demo as MP4 (118 MB). She needs offline listening on a flight. Settings: full audio, MP3. Result: a 48 MB file for any player — she notes timestamps and pulls quotes from the full video later on Wi-Fi.
Karaoke — instrumental attempt
A choir member practices harmonies against a cover video (personal use only). Settings: instrumental, WAV for pitch work in Audacity. Expect faint vocal ghosting — record her part separately rather than expecting a broadcast-clean stem.
Podcast — video episode to audio
An indie creator exports episode 12 as MOV (94 MB) but publishes audio-only to Spotify. Settings: full audio, MP3 or AAC. Run noise reduction if room noise was audible, then upload — no video NLE export needed.
Transcription prep — lecture clip
A student screen-recorded a 38-minute Zoom lecture as MP4 (86 MB). Settings: full audio, WAV. Pipeline: extract, optionally denoise, then audio to text with model Small.
Pre-upload checklist
- Confirm the video plays with audible sound — silent screen captures cannot be recovered.
- Check file size is under 200 MB; trim or re-encode locally if needed.
- Pick mode: full unless you have a specific reason to experiment with stems.
- Match output format to the next tool in your chain (MP3 for listening, WAV for editing, FLAC for archive).
- For transcription, consider WAV or high-quality MP3 and denoise first if HVAC hum dominates.
- Verify you have rights to extract the content.
After you download
Spot-check the first and last minute. Rename clearly (2026-03-webinar-audio.mp3). Normalise quiet levels in the audio editor before transcription. Need another codec? Use the audio converter instead of re-uploading video.
How we handle your files
Uploads travel over HTTPS for extraction only — not used to train models or sold to advertisers. Files are purged after processing. Download links expire; save locally. Review our privacy policy and security page for organisational policies.
Limitations and expectations
- Vocal and instrumental modes are approximate — not a replacement for professional stem separation or multitrack masters.
- 200 MB cap — split long recordings or lower bitrate before upload.
- Only the primary audio stream is processed; extra MKV tracks may be ignored.
- DRM-protected video will fail — use unencrypted files you legally own.
- Respect copyright — do not redistribute extracted commercial content.
Frequently asked questions
Is extract audio from video free?
Yes — use it in your browser with no account or subscription.
Which video formats are supported?
MP4, AVI, MOV, MKV, FLV, WMV, WEBM, M4V, and 3GP.
What is the maximum file size?
200 MB per upload.
Are my files stored permanently?
No. Files are processed and removed automatically after a short period. We do not share them with third parties.
Does it work on mobile?
Yes — responsive on Android and iPhone browsers. Upload clips directly from your camera roll or cloud apps.
Full audio vs vocals vs instrumental — which should I pick?
Full audio for almost every workflow. Vocals or instrumental only when you understand results will be imperfect stems.
Can I extract audio then transcribe it?
Yes — download MP3 or WAV here, then upload to audio to text for Whisper transcription.
Wrap-up
Upload your video to Itqan's extract audio from video tool, choose a mode and output format, and download a standalone soundtrack in minutes. Default to full audio and MP3 for listening, WAV when editing or transcribing, and treat vocal/instrumental modes as experiments — not mastering tools. From webinar commutes to podcast publishing pipelines, separating audio from video keeps your workflow lightweight without re-recording or heavyweight NLE exports.
Related tools
- Audio format converter — change codec after extraction without re-uploading video.
- Audio to text — transcribe the extracted track with Whisper.
- Noise reduction — clean noisy recordings before sharing or transcription.
- Audio editor — trim, normalise, and splice the extracted file.
Extraction pipeline for meetings and courses
Pull audio with the video audio extractor, reduce hiss in noise reduction, then transcribe with audio to text. Keep the original video file archived; extraction does not replace backup. If you only need a clip of the soundtrack, trim in the audio editor after extraction rather than re-encoding the whole movie.
Name files with date and speaker when several recordings land the same day. For bilingual lectures, extract once, then translate the transcript in the translator instead of re-recording. When the goal is podcast distribution, convert the extracted track to a platform-friendly bitrate via the audio converter before upload.
Compare extracted duration to the source movie after the video audio extractor finishes; a large gap usually means a muted track or an accidental trim upstream.