You recorded a 47-minute faculty lecture on your phone, and the deadline for typed notes is tomorrow morning. Manual transcription at roughly 4× real time would eat four hours you do not have. Itqan's audio to text tool runs OpenAI Whisper on your upload — MP3, WAV, M4A, AAC, OGG, FLAC, or video MP4/AVI/MOV — and returns editable text in 50+ languages. No software install, no account. This guide walks through Whisper model choice, four real workflows, an upload checklist, and how to polish the transcript before you publish or file it.
Speech-to-text basics
Audio to text converts spoken words into written text via OpenAI's Whisper model. Upload once, receive a draft in minutes — the inverse of text to speech. Paste results into any editor or check length with the word counter. The interface includes an audio preview, progress bar, and copy/download on the result page.
Why transcribe audio at all?
Text is searchable, quotable, and accessible. Transcripts help listeners skim podcasts, support students with hearing difficulties, and give search engines indexable content.
- Education: Turn a recorded seminar into study notes without replaying the full hour.
- Journalism: Pull exact quotes from a 35-minute source interview.
- Business: Draft meeting minutes from a Teams or Zoom export.
- Content marketing: Repurpose a 22-minute video monologue into a blog post.
- Legal and compliance: Create a searchable first draft of a recorded statement (always verify against the original).
- Accessibility: Publish captions or full transcripts alongside video.
Supported media and size limits
Audio formats
MP3, WAV, M4A, AAC, OGG, and FLAC. Phone voice memos usually arrive as M4A or AAC; professional recorders often export WAV or FLAC.
Video formats (audio extracted)
MP4, AVI, and MOV. You do not need to strip audio first — Whisper processes the soundtrack. Useful for webinar recordings, screen captures with narration, or field video where only the spoken track matters.
File size
Each upload may be up to 150 MB. Long WAV files can hit the cap sooner — convert with the audio converter if needed.
Choosing a Whisper model
Whisper ships in several sizes. Itqan offers three tuned for web use:
| Model | Speed | Accuracy | Best for |
|---|---|---|---|
| Base | Fastest | Good on clean speech | Quick drafts, short clips under 15 minutes, clear studio audio |
| Small | Medium | Better on accents and mild noise | Default choice for most interviews and lectures |
| Medium | Slower | Highest | Noisy rooms, multiple speakers, or text you will publish verbatim |
Pick auto-detect for mixed speech; lock to Arabic or English for monolingual recordings. Processing time scales with duration and model size.
Step-by-step walkthrough
- Open audio to text on desktop or mobile.
- Upload from your computer, Google Drive, or Dropbox, or drag a file onto the drop zone.
- Confirm the sidebar shows correct file name, size, and format.
- Select language (auto or explicit) and Whisper model.
- Preview audio in the built-in player — skip this only if you already know the file is valid.
- Click Start conversion and watch the progress bar.
- On the result page, copy the transcript or download it as a text file.
- Edit names, numbers, and technical terms before publishing.
Lecture — 90-minute recording
A biology student records a 90-minute midterm review on a laptop mic placed two metres from the lecturer. The file is 68 MB MP3. Goal: searchable notes before the exam in 48 hours.
Prep: Listen to the first 30 seconds — if HVAC hum dominates, run the file through noise reduction first, then upload the cleaned MP3.
Settings: Language set to English, model Small (Base might miss "mitochondria" at a distance; Medium is overkill unless the room was echoey).
Result: Roughly 12,400 words returned. The student fixes three chemical names and bolds definitions — 25 minutes of editing versus 6+ hours of manual typing.
Interview — journalist workflow
A reporter records a 38-minute phone interview saved as M4A (22 MB). The source discusses budget figures and proper nouns that must be exact.
Settings: Auto-detect language (English with occasional Spanish phrases), model Medium because quotes go to print.
Workflow: Highlight sentences with numbers and cross-check "$4.2 million" and "Q3 2025" against the audio. Denoise café noise first to save verification time.
Meeting — Zoom audio export
A project manager exports a 52-minute Zoom meeting as M4A (41 MB). Three people spoke; action items must reach the team by end of day.
Settings: English, model Small.
Post-processing: Whisper does not label speakers. The manager inserts names and bullets tasks. Screen-recorded MP4 uploads work directly — no separate audio export needed.
Podcast — show notes and SEO
An indie podcaster publishes weekly episodes averaging 28 minutes. Episode 47 (MP3, 26 MB) covers "remote hiring in 2026."
Settings: English, Base for a fast first draft (host reads from a loose outline, audio is studio-quality).
Repurposing: The 4,800-word transcript becomes a blog summary, timestamped show notes, and meta copy. Delete nonsense lines from ad reads and intro music.
Arabic transcription tips
Whisper handles Modern Standard Arabic and many dialects, but accuracy varies with code-switching, fast colloquial speech, and proper nouns transliterated from English.
- Select Arabic explicitly instead of auto when the recording is Arabic-only.
- Record closer to the mic — phone distance causes dropped short vowels in fast speech.
- For mixed Arabic–English meetings, auto-detect works, but expect English trade terms to need manual fixes.
- Review hamza and ta marbuta on names — Whisper may normalise spelling differently from your style guide.
- Pair with noise reduction when recording outdoors or in busy offices.
Pre-upload checklist
- Confirm the file plays start-to-finish — corrupted uploads fail silently until processing.
- Check size is under 150 MB; compress or trim if needed.
- Pick language: auto for mixed speech, explicit for monolingual Arabic or English.
- Choose model: Base for speed, Medium for publish-ready accuracy.
- Denoise if background hum or wind competes with speech.
- For video, verify the correct audio track exists (some screen captures have no mic input).
- Redact or avoid uploading recordings with legally protected or highly confidential content unless your policy allows cloud processing.
After the transcript arrives
Whisper output is a draft. Budget 10–20% of audio duration for editing. Fix names, verify numbers against audio, add paragraph breaks, and remove fillers for publication. Output is plain text — no SRT timecodes.
How we handle your files
Uploads travel over HTTPS to Itqan's servers solely for Whisper transcription. We do not use your recordings to train AI models, sell audio to advertisers, or retain files indefinitely. After processing completes, uploads and generated transcripts are purged on a short automatic schedule.
Session tokens on the result page expire — bookmark the downloaded text file if you need it later. For organisational policies on cloud transcription, review our privacy policy, cookie notice, and security page. When material is legally privileged or classified, consult your compliance team before uploading to any online service.
Limits and pitfalls
- Whisper does not diarise speakers — one continuous block of text, no "Speaker A" labels.
- Music, applause, and overlapping voices degrade accuracy; trim intros if possible.
- Very strong accents or rare languages may need Medium model plus heavy editing.
- 150 MB cap — split multi-hour WAV files or convert to efficient codecs first.
- Output is plain text — no built-in SRT/VTT export with timestamps.
- Auto-detect can mis-guess on short clips under 30 seconds — pick language manually.
- Server processing time grows with file length; start long jobs when you can wait.
Frequently asked questions
Is audio to text free?
Yes — use it in your browser with no account or subscription.
Which formats are supported?
MP3, WAV, M4A, AAC, OGG, FLAC, plus video MP4, AVI, and MOV for automatic audio extraction.
What is the maximum file size?
150 MB per upload.
Are my files stored permanently?
No. Files are processed for transcription and removed automatically after a short period. We do not share them with third parties.
Does it work on mobile?
Yes — responsive on Android and iPhone browsers. Upload voice memos directly from your phone.
Base vs Small vs Medium — which should I pick?
Base for quick drafts on clean audio. Small for everyday balance. Medium when accuracy matters more than speed.
Can it transcribe Arabic?
Yes — select Arabic from the language menu or use auto-detect on Arabic recordings.
Wrap-up
Upload your recording to Itqan's audio to text tool, choose language and Whisper model, and download a draft transcript in minutes instead of hours. Denoise noisy files first, pick Medium for publish-ready work, and always verify names and numbers against the original audio. From 90-minute lectures to 28-minute podcast episodes, speech-to-text turns listen-only content into text you can search, quote, and repurpose.
Related tools
- Text to speech — the reverse direction: written text to spoken audio.
- Noise reduction — clean recordings before transcription.
- Audio converter — change format or shrink large WAV files under 150 MB.
- Video audio extractor — pull audio from video if you need a standalone track first.
From transcript to publishable text
After speech-to-text in the audio to text tool, run the draft through the word counter for length targets, then the case converter if casing drifted. For voiceover of the cleaned script, open text to speech.