Audio to Text

Convert audio or video files to text transcripts online. Supports multilingual detection and SRT/VTT subtitle export. Audio is extracted locally in your browser for privacy.

⬆️
Click to select, or drag audio / video here
Supports mp4 / mov / mp3 / wav / m4a, etc. · Up to 200MB per file · Audio is extracted from video locally; the original file isn't uploaded
Recognition mode

Tip: The first time you use a model, the server downloads it automatically (please wait a moment).

FAQ

What file formats are supported?

Supports common audio formats (mp3, wav, m4a, aac, etc.) and video formats (mp4, mov, mkv, etc.). Videos are processed locally in your browser to extract audio before upload.

Are my files uploaded to the server?

Videos are extracted into audio within your browser; the original video is never uploaded. Only the extracted audio is sent for recognition and deleted immediately after processing.

Which languages are supported?

Automatic language detection is enabled by default, with manual selection available for Chinese, English, Japanese, Korean, and more. Chinese automatically uses Mandarin and simplified Chinese optimization.

What's the difference between the three recognition modes?

Fast (small) mode is the quickest, ideal for quickly generating drafts from long audio. Standard (medium) offers balanced speed and accuracy—recommended for everyday use. High Precision (large-v3) delivers the highest accuracy, best for official subtitles, though it takes longer.

Explore more AI tools and products