Formats, links, and upload limits
Which audio and video formats QuillAI accepts, how to paste a link, file size limits, and how to set audio language and speaker recognition.
Supported formats
QuillAI accepts most common audio and video formats:
- MP3
- MP4
- M4A
- MOV
- AAC
- WAV
- OGG
- OPUS
- MPEG
- WMA
- WMV
Links instead of a file
On the "Paste link" tab you can point to a video on YouTube, Vimeo, Facebook, Twitch, and similar services — the service downloads and processes the media itself, no local file needed.
File or link — which to pick
If the recording already lives on YouTube or another video host, use a link: it's faster and doesn't spend your bandwidth downloading it to your computer first. Uploading a file makes sense when the recording only exists locally or on your team's drive.
Limits
A regular upload is capped at 100 MB per file. Paid plans raise that: files up to 10 hours long and up to 5 GB, with up to 50 files processed at once. If your file is bigger, compress it or upload a smaller version.
Language and speaker recognition
In the upload settings — next to the "Start transcription" button — you can turn on "Recognize speakers": every section of the transcript then gets labeled with the speaker's name. The same panel lets you set "Audio language": leave it on auto-detect, or pick a language manually from the list (English, Spanish, French, German, Italian, Portuguese, Dutch, Japanese, Chinese, Russian, and more).
Tip
Picking the exact language manually usually gives a cleaner transcript than auto-detect — especially for short recordings or ones with a strong accent.
Several files at once
You can upload several files together — each one tracks its own progress and status, and an "Add more files" button lets you drop in another batch without starting over. A "Clear completed" button tidies finished uploads out of the list.
Drag it in or browse
You can drag a file straight into the upload area, or click "Browse files" and pick it from a file dialog — both work the same way. Audio clarity affects accuracy more than anything else: background noise, echo, and quiet speech hurt transcript quality more noticeably than the file's format or container.