Sign inSign up

ttlequals0/minuspod

By ttlequals0

•Updated 3 days ago

Removes ads from podcasts using Whisper + AI. Serves clean RSS feeds for any podcast app.

Image
Machine learning & AI
1

100K+

ttlequals0/minuspod repository overview

MinusPod

⁠MinusPod

Removes ads from podcasts using Whisper transcription and AI detection. Serves modified RSS feeds that work with any podcast app.

⁠How It Works

  1. Transcription -- Whisper converts audio to text with timestamps (local GPU or remote API)
  2. Ad Detection -- 3-stage pipeline: audio fingerprinting, learned text patterns, then LLM analysis (Claude, Ollama, or any OpenAI-compatible endpoint)
  3. Audio Processing -- FFmpeg removes detected ads with short audio markers
  4. Serving -- Flask serves modified RSS feeds and processed audio

Processing is on-demand. First play takes a few minutes; subsequent plays are instant (cached).

⁠Features

  • Web UI for feed management, ad review/correction, pattern management, and settings
  • Ad editor with boundary adjustment, confidence scores, and detection stage badges
  • Pattern learning system (podcast, network, and global scope) -- improves over time
  • Verification pass re-transcribes processed audio to catch missed ads
  • Audio analysis (volume anomalies, transition detection, DAI boundary enforcement)
  • Webhooks for notifications (Pushover, ntfy, or any HTTP endpoint)
  • OPML import/export, full-text search, processing history with stats
  • REST API with 71 documented endpoints (OpenAPI spec at /docs)

⁠Quick Start

cat > .env << EOF
ANTHROPIC_API_KEY=your-key-here
BASE_URL=http://localhost:8000
EOF

mkdir -p data
docker-compose up -d

Web UI at http://localhost:8000/ui/

⁠Requirements

  • NVIDIA GPU with Docker support (for local Whisper), or use a remote Whisper API
  • Anthropic API key, Ollama, or any OpenAI-compatible LLM endpoint

⁠LLM Providers

ProviderConfig
Anthropic (default)ANTHROPIC_API_KEY=your-key
Ollama (local, free)LLM_PROVIDER=ollama, OPENAI_BASE_URL=http://host.docker.internal:11434/v1, OPENAI_MODEL=qwen3:14b
OpenAI-compatibleLLM_PROVIDER=openai-compatible, OPENAI_BASE_URL=http://your-endpoint/v1

⁠Whisper Backends

BackendConfig
Local GPU (default)WHISPER_MODEL=small, WHISPER_DEVICE=cuda
Remote APIWHISPER_BACKEND=openai-api, WHISPER_API_BASE_URL=http://your-server:8765/v1
GroqWHISPER_BACKEND=openai-api, WHISPER_API_BASE_URL=https://api.groq.com/openai/v1

⁠Volumes

PathDescription
/app/dataDatabase, cached RSS, and processed audio
/app/assets(Optional) Custom replace.mp3 for ad break marker

⁠Environment Variables

VariableDefaultDescription
ANTHROPIC_API_KEY--Claude API key
LLM_PROVIDERanthropicanthropic, openai-compatible, or ollama
OPENAI_BASE_URL--Base URL for OpenAI-compatible/Ollama
OPENAI_API_KEYnot-neededAPI key for OpenAI-compatible endpoint
OPENAI_MODEL--Model name for non-Anthropic providers
BASE_URLhttp://localhost:8000Public URL for feed links
WHISPER_MODELsmallWhisper model size
WHISPER_DEVICEcudacuda or cpu
WHISPER_BACKENDlocallocal or openai-api
WHISPER_API_BASE_URL--Remote Whisper API URL
TUNNEL_TOKEN--Cloudflare tunnel token

⁠Disclaimer

This tool is for personal use only. Only use it with podcasts you have permission to modify or where such modification is permitted under applicable laws.

Tag summary

Content type

Image

Digest

sha256:246004689…

Size

4.4 GB

Last updated

3 days ago

docker pull ttlequals0/minuspod