Skip to content

Video Transcription for the SubtitleSwitcher Plugin

Daniel Neto edited this page Apr 29, 2023 · 7 revisions

Install Vosk-transcriber on Ubuntu

Vosk is an offline speech recognition toolkit that supports multiple languages. The Vosk-transcriber is a Python package that utilizes the Vosk library to transcribe audio files. By installing Vosk-transcriber in Ubuntu, you can enable the SubtitleSwitcher plugin to automatically transcribe videos and audio. This will allow you to generate accurate subtitles for your videos in various languages without the need for an internet connection. Follow this step-by-step guide to install Vosk-transcriber in Ubuntu and enhance the functionality of the subtitleSwitcher plugin with speech recognition capabilities.

Step 1: Update your system

Before installing new packages, it's a good practice to update your system. Open a terminal and run the following command:

sudo apt update && sudo apt upgrade

Step 2: Install Python and pip

Vosk-transcriber requires Python and pip (the Python package installer) to work. Install them by running:

sudo apt install python3 python3-pip

Step 3: Install FFmpeg (optional)

Vosk-transcriber uses FFmpeg to convert audio files to the required format. Install FFmpeg by running:

sudo apt install ffmpeg

Step 4: Install Vosk-transcriber

Now you can install the Vosk-transcriber Python package using pip:

pip3 install vosk-transcriber

Step 5: Download a Vosk language model

Vosk requires a language model to perform speech recognition. Download a pre-built language model from the Vosk website (https://alphacephei.com/vosk/models) or the GitHub repository (https://github.com/alphacep/vosk-api/blob/master/doc/models.md). For example, to download the English language model, run:

wget https://alphacephei.com/vosk/models/vosk-model-en-us-aspire-0.2.zip

Step 6: Extract the language model

Unzip the downloaded language model:

unzip vosk-model-en-us-aspire-0.2.zip

Clone this wiki locally