Skip to content

Video Transcription for the SubtitleSwitcher Plugin

Daniel Neto edited this page Apr 29, 2023 · 7 revisions

The transcriber

Vosk is an offline speech recognition toolkit that supports multiple languages. The Vosk-transcriber is a Python package that utilizes the Vosk library to transcribe audio files. By installing Vosk-transcriber in Ubuntu, you can enable the SubtitleSwitcher plugin to automatically transcribe videos and audio. This will allow you to generate accurate subtitles for your videos in various languages without the need for an internet connection. Follow this step-by-step guide to install Vosk-transcriber in Ubuntu and enhance the functionality of the subtitleSwitcher plugin with speech recognition capabilities.

Update your system

Before installing new packages, it's a good practice to update your system. Open a terminal and run the following command:

sudo apt update && sudo apt upgrade

Python installation from Pypi

The easiest way to install vosk API is with pip. You do not have to compile anything.

Make sure you have up-to-date pip and python3 versions:

  • Python version: 3.5-3.9
  • pip version: 20.3 and newer.

Upgrade Python and pip if needed. Then install vosk on Linux/Mac from pip:

pip3 install vosk

For more information check https://alphacephei.com/vosk/install

Download a Vosk language model

Vosk requires a language model to perform speech recognition. Download a pre-built language model from the Vosk website (https://alphacephei.com/vosk/models) or the GitHub repository (https://github.com/alphacep/vosk-api/blob/master/doc/models.md). For example, to download the English language model, run:

wget https://alphacephei.com/vosk/models/vosk-model-small-en-us-0.15.zip

Extract the language model

Unzip the downloaded language model:

unzip vosk-model-small-en-us-0.15.zip

Clone this wiki locally