Speaker diarization API
887
Audio speaker diarization and transcription API.
make build-prod
make run
make build-dev
make dev
We deploy continuously using Travis CI and a Docker Hub deploy bot.
You can also deploy the production and development images manually if you like.
make push-dev
make push-prod
You will need to set DOCKER_USERNAME and DOCKER_PASSWORD and be a member of the ubclaunchpad docker hub organization to deploy manually.
We have a pipeline that is designed to take YouTube videos with
transcripts and convert them into training data. pipeline.py
is a CLI that will attempt to download the transcript and audio
data for a given video, as well as prompt for some information
that the rest of the pipeline uses to create labelled data
(ie. what delimeters are used to identify speakers).
$ cd app/collector
$ ./pipeline.py <video_id>
You can push training data into the research environment on DigitalOcean.
You will need to collect the instance PEM file from your tech lead. Place
the PEM locally in ~/.ssh/id_minutes. Set the environment variable
MINUTES_RESEARCH_INSTANCE in your environment to the IP address of the
DigitalOcean instance (available on Slack or from your tech lead).
Then, if you wish to push the file bigdata.csv, use the following command:
make FILE=bigdata.csv push-data
It will appear in the data folder on the research platform.
Content type
Image
Digest
Size
462.8 MB
Last updated
almost 9 years ago
docker pull ubclaunchpaddeploybot/minutes-prod