Sign inSign up

dervish/unraidmonitorbot

By dervish

•Updated about 11 hours ago

Telegram bot for monitoring Unraid servers with AI diagnostics and container control

Image
Machine learning & AI
Monitoring & observability
0

6.9K

dervish/unraidmonitorbot repository overview

⁠UnraidMonitor

A Telegram bot for monitoring Docker containers and Unraid servers. Get real-time alerts, check container status, view logs, and control containers - all from Telegram.

scratch-it.co.uk/unraidmonitorbot⁠ has the illustrated tour. This README is the reference.

⁠Features

  • Interactive Setup Wizard - Guided first-run setup via Telegram with auto-classification of containers
  • Container Monitoring - Status, health checks, crash detection, and recovery notifications
  • Resource Alerts - CPU/memory usage with per-container thresholds, adjustable directly from alert buttons
  • Log Watching - Automatic alerts when errors appear in container logs
  • AI Diagnostics - LLM-powered log analysis and troubleshooting (Anthropic, OpenAI, or Ollama)
  • Smart Ignore Patterns - AI-generated patterns to filter known errors, with interactive toggle selection
  • Multi-Provider LLM - Switch between Anthropic Claude, OpenAI GPT, or local Ollama models at runtime
  • Container Control - Start, stop, restart, and pull containers with inline confirmation buttons
  • Image-Update Detection - Opt-in daily digest of containers with newer images available, with Pull buttons
  • Auto-Heal - Opt-in automatic restart of unhealthy containers (HEALTHCHECK failures) with storm guard
  • Unraid Server Monitoring - CPU/memory, temperatures, array health, and parity-operation progress
  • Unraid Notification Relay - Opt-in forwarding of Unraid's own notification feed (SMART, disk errors, share full) with an adjustable importance floor
  • UPS Monitoring (NUT) - Reads your UPS over the network from a NUT⁠ server: mains loss, low battery, overload and bypass alerts, plus /ups for battery and runtime
  • Memory Pressure Management - Automatic container priority handling during high memory
  • Mute System - Temporarily silence alerts per container, server, or array
  • Natural Language Chat - Ask questions naturally instead of using commands
  • Interactive Dashboard - /manage hub for status, resources, server, disks, ignores, mutes, and a Features panel to toggle optional monitors
  • Sectioned Help - /help with navigable category buttons instead of a text wall

⁠Screenshots

Setup wizard classifying containers into priority, protected, watched and killableLog error alert for plex with Ignore Similar, Mute, Logs and Diagnose buttons
Setup wizard. Scans your containers on first run and sorts them into priority, protected, watched and killable. Re-run any time with /setup.Log error alert. Errors in a watched container, with the latest line and one tap to ignore, mute, read the logs or diagnose.
AI diagnosis explaining that SABnzbd received SIGTERM and exited cleanlyResource alert showing plex at 527 percent CPU against a 400 percent threshold
AI diagnosis. /diagnose reads the logs and tells you what happened and why, instead of handing you a wall of text.Resource alert. CPU and memory against your thresholds, and you can change the threshold from the alert itself.
Digest listing sonarr and rreading-glasses-db with newer images and Pull buttonsThe bot answering a request for a status update written as a fairy tale
Image updates. An opt-in daily digest of containers running behind their registry, each with a Pull button.Natural language. Ask in plain English. It reads real server state, and it will humour you.
The /manage dashboard showing server CPU, RAM and uptime with buttons for Status, Resources, Server, Disks, Manage Ignores, Manage Mutes and FeaturesTelegram autocomplete listing the bot commands with a one-line description of each
The /manage hub. Server vitals at the top, then every panel one tap away, feature toggles included.Commands, if you want them. The menu is built from what your install actually has enabled, so it never offers something the bot cannot do.

⁠What's New in v0.22.3

  • Unraid notifications stop going missing - Ones with underscores in them, such as disk names and serial numbers, failed to send. Any alert whose formatting won't parse now goes out as plain text
  • No UPS, no fuss - The startup message only mentions UPS if you turned it off or it stopped answering

⁠What's New in v0.22.2

  • A startup message you can read at a glance - Anything that needs a look comes first, the rest is one line, and it names the AI model each feature uses. It also stopped claiming UPS monitoring was off when it was on

⁠What's New in v0.22.1

  • AI features use the newest Claude models again - Picking a model with /model used to freeze the bot on whatever was newest that day. sonnet and opus now always mean the latest, and existing choices upgrade themselves. A full model ID you pick on purpose stays pinned

⁠What's New in v0.22.0

  • Stopping a watched container no longer hammers Docker - The log watcher re-checked a stopped container hundreds of times a second until it came back. It now waits between checks
  • Crash alerts get through - An error alert no longer silences the crash alert that follows it, and "RESTART LOOP" now fires for containers crashing a minute or more apart
  • Memory handling can't get stuck - Four ways the memory monitor could go quiet for good are fixed, and container memory now matches docker stats on Unraid 7 instead of counting the disk cache
  • UPS and mutes - A power cut that starts during a mute is reported when the mute ends if you're still on battery, and /ups during a network blip no longer hides a real outage
  • Ignore Similar works on long errors - It used to save a pattern that could never match

⁠What's New in v0.21.2

  • Replying to an alert now picks the right container - /mute, /ignore and /diagnose read the container name off the alert you replied to. Replying to a restart-loop alert used to mute a container called "4" (the crash count) and tell you it had worked. Reply-to-alert /diagnose had never worked on anything but resource alerts
  • The 🔄 Restart button on an alert asks first - It restarted immediately on one tap, while /restart has always wanted a ✅. Alerts stay in your chat for days, so a stale one was a mis-tap away from bouncing a container
  • A full array no longer texts you every five minutes - The capacity warning now fires once per crossing and re-arms when usage drops back under the threshold
  • Fewer slow leaks - The Unraid client stopped leaking a connection on every network drop, and your runtime /model choice survives a crash mid-save

⁠What's New in v0.21.1

  • Memory was reported as ~98% when the real figure was ~55% - The percentage came from one Unraid API field and the gigabytes from another, and those two mean different things. /server, the Memory Critical alert body and the figures handed to the AI were all wrong together. Alert thresholds read the percentage, so no false alerts were firing
  • /server detailed now shows the total and the reclaimable disk cache - "55% used" next to "0.5 GB free" was baffling without them

⁠What's New in v0.21.0

  • UPS monitoring, over the network - The bot reads your UPS from a NUT⁠ server, so the UPS does not have to be plugged into the machine running the bot. Alerts on mains loss, low battery, overload and bypass
  • New /ups command - Model, status, battery percentage, runtime left, load and input voltage. /ups detailed dumps every variable
  • On by default, quiet by default - With no NUT server to talk to it logs one line and stays silent, rather than nagging the majority of installs that have no UPS. Turn it off in /manage -> ⚙️ Features
  • A UPS it cannot read says "unavailable", never "healthy" - A monitor that lost contact with upsd knows nothing, and reporting that as an OK would be worse than useless

Setup gotcha: upsd binds to 127.0.0.1 only by default. If the bot runs in a container you need LISTEN 0.0.0.0 3493 in upsd.conf before it can connect. See Configure NUT⁠.

⁠What's New in v0.20.0

  • No more false parity alarms - A parity sync or disk rebuild is reported as progress ("45% complete"), not as a disk problem, and you're told when it finishes. A genuinely failed disk still alerts during a sync
  • Unraid notifications in Telegram - Opt-in relay of Unraid's own notification feed (SMART, disk errors, share full, parity results) so everything lands in one place. Enable in /manage → ⚙️ Features
  • Control how chatty it is - The notification button cycles WARNING+ (default), ALERT only, or everything including INFO, and applies instantly

⁠What's New in v0.19.0

  • Four broken buttons fixed - Array threshold options no longer fail silently after saving the value; Stop buttons on memory alerts now work even with memory management disabled; "Re-mute 1h" means one hour rather than sixty
  • Command autocomplete - Type / in Telegram to see every command, built from the features your install actually has enabled
  • /manage panels have Back and Refresh - Status, Resources, Server and Disks are no longer dead ends
  • /pull keeps your GPU - nvidia device access, custom runtimes and supplementary groups now survive a container update
  • Failures are visible - A handler that crashes replies instead of going quiet, and no button can spin forever
  • Top memory users on pressure alerts - Memory warnings list the top 5 memory-consuming containers, largest first, with a one-tap 🔄 Restart for containers that just need a bounce (v0.18.0)

Note on UPS support: UPS monitoring reads from a NUT server, not from Unraid's API. Unraid exposes upsDevices, but that data comes from apcupsd over a local USB link, which is no use to a bot in a container and no use at all without the cable. NUT works over TCP, so it covers both cases. See the changelog⁠.

See the changelog⁠ for full details.


⁠Table of Contents


⁠Installation

The easiest way to install on Unraid.

  1. Install from Community Apps

    • Open the Unraid web UI
    • Go to Apps tab
    • Search for "Unraid Monitor Bot"
    • Click Install
  2. Configure the template

    • TELEGRAM_BOT_TOKEN - Your bot token (how to get one⁠)
    • TELEGRAM_ALLOWED_USERS - Your Telegram user ID (how to find it⁠)
    • ANTHROPIC_API_KEY (optional) - Enables AI features via Claude
    • OPENAI_API_KEY (optional) - Enables AI features via OpenAI
    • OLLAMA_HOST (optional) - Enables AI features via local Ollama (e.g., http://192.168.1.100:11434)
    • DEFAULT_MODEL (optional) - Override the default AI model (e.g., qwen2.5:7b, gpt-4o)
    • UNRAID_API_KEY (optional) - Enables server monitoring
  3. Start the container

  4. Message your bot on Telegram - send /start to begin the setup wizard

    • The wizard will guide you through connecting to your Unraid server
    • It auto-classifies your containers into categories (priority, protected, watched, killable, ignored)
    • When an Anthropic API key is configured, AI assists with classifying unknown containers
    • Review and adjust the categories, then confirm to save
    • The bot restarts automatically and begins monitoring
  5. Re-configure anytime (optional)

    • Send /setup to re-run the wizard (merges non-destructively with existing config)
    • Or edit /mnt/user/appdata/unraid-monitor/config/config.yaml directly and restart

⁠Docker on Unraid (Manual)

If not using Community Apps, you can set it up manually.

⁠Step 1: Create directories
mkdir -p /mnt/user/appdata/unraid-monitor/{config,data}
⁠Step 2: Create the environment file

Create /mnt/user/appdata/unraid-monitor/config/.env:

# Required
TELEGRAM_BOT_TOKEN=your_bot_token_here
TELEGRAM_ALLOWED_USERS=123456789

# Optional - AI features (configure at least one for /diagnose, NL chat, smart ignore)
ANTHROPIC_API_KEY=your_anthropic_api_key_here
OPENAI_API_KEY=your_openai_api_key_here
OLLAMA_HOST=http://localhost:11434

# Optional - override the default AI model (e.g. qwen2.5:7b, gpt-4o)
DEFAULT_MODEL=

# Optional - enables Unraid server monitoring
UNRAID_API_KEY=your_unraid_api_key_here

# Optional - only if your NUT server requires a login for reads
NUT_USERNAME=
NUT_PASSWORD=
⁠Step 3: Add the container in Unraid

Go to Docker → Add Container and configure:

FieldValue
Nameunraid-monitor-bot
Repositorydervish/unraidmonitorbot:latest
Network Typebridge or your preferred network

Add these paths:

Container PathHost PathAccess
/app/config/mnt/user/appdata/unraid-monitor/configRead/Write
/app/data/mnt/user/appdata/unraid-monitor/dataRead/Write
/var/run/docker.sock/var/run/docker.sockRead Only

Add these variables:

NameValue
TELEGRAM_BOT_TOKENYour bot token
TELEGRAM_ALLOWED_USERSYour user ID
ANTHROPIC_API_KEY(optional) Claude AI features
OPENAI_API_KEY(optional) OpenAI AI features
OLLAMA_HOST(optional) Ollama URL, e.g., http://192.168.1.100:11434
DEFAULT_MODEL(optional) Override default model, e.g., qwen2.5:7b
UNRAID_API_KEY(optional) Unraid server monitoring
NUT_USERNAME(optional) Only if your NUT server gates reads
NUT_PASSWORD(optional) Only if your NUT server gates reads
PUID(optional) Runtime user ID for file ownership (default: 99 — Unraid's nobody)
PGID(optional) Runtime group ID for file ownership (default: 100 — Unraid's users)
TZYour timezone (e.g., Europe/London)
⁠Step 4: Start and verify

Start the container and check the logs for any errors. Message your bot on Telegram with /start to begin the interactive setup wizard.


⁠Docker on Other Systems

For non-Unraid Docker hosts (Ubuntu, Debian, Synology, etc.), use docker-compose:

  1. Clone the repository and create your environment file:

    git clone https://github.com/dervish666/UnraidMonitor.git
    cd UnraidMonitor
    cp config/.env.example config/.env
    # Edit config/.env with your TELEGRAM_BOT_TOKEN, TELEGRAM_ALLOWED_USERS, etc.
    
  2. Adjust docker-compose.yml volume paths to suit your system (the defaults point to Unraid appdata paths). For example:

    volumes:
      - /var/run/docker.sock:/var/run/docker.sock:ro
      - ./config:/app/config
      - ./data:/app/data
    
  3. Check your Docker socket GID and set it if it differs from the default (281):

    ls -ln /var/run/docker.sock   # look at the 4th column
    echo "DOCKER_GID=999" >> .env  # adjust to match
    
  4. Build and start:

    docker-compose up -d
    
  5. Message your bot on Telegram with /start to begin the setup wizard.


⁠Prerequisites

⁠1. Create a Telegram Bot
  1. Open Telegram and message @BotFather⁠
  2. Send /newbot
  3. Follow the prompts to name your bot
  4. Copy the bot token (looks like 123456789:ABCdefGHIjklMNOpqrsTUVwxyz)
⁠2. Get Your Telegram User ID
  1. Message @userinfobot⁠ on Telegram
  2. It will reply with your numeric user ID (e.g., 123456789)

This ID is used to restrict who can control your bot. You can add multiple IDs separated by commas: 123456789,987654321

⁠3. Configure an LLM Provider (Optional)

At least one provider is needed for AI-powered features (/diagnose, smart ignore patterns, natural language chat). You can configure multiple providers and switch between them at runtime with /model.

Option A: Anthropic Claude (recommended)

  1. Sign up at console.anthropic.com⁠
  2. Go to API Keys and create a new key
  3. Add it as ANTHROPIC_API_KEY

Option B: OpenAI

  1. Sign up at platform.openai.com⁠
  2. Go to API Keys and create a new key
  3. Add it as OPENAI_API_KEY

Option C: Ollama (free, runs locally)

  1. Install Ollama from ollama.com⁠
  2. Pull a model: ollama pull llama3.1:8b
  3. Set OLLAMA_HOST to your Ollama URL (e.g., http://192.168.1.100:11434)

Models are auto-discovered from Ollama at startup. Note: some local models don't support tool calling, so NL chat actions (restart, etc.) may be limited.

⁠4. Get an Unraid API Key (Optional)

Required for Unraid server monitoring (CPU, memory, temps, array status).

  1. In Unraid web UI, go to Settings → Management Access
  2. Generate an API key
  3. Add it as UNRAID_API_KEY
⁠5. Configure NUT for UPS monitoring (optional)

Needed only if you have a UPS. NUT talks over the network, so the UPS can hang off any machine on the LAN, not necessarily the one running the bot.

  1. Install a NUT server. On Unraid, install the NUT plugin from Community Apps and point it at your UPS. On another Linux box, install nut and configure ups.conf for your model. Check it works locally first:

    upsc myups
    
  2. Let the bot reach it. This is the step people miss. upsd binds to 127.0.0.1 only by default, which a container cannot reach. Add this to upsd.conf and restart upsd:

    LISTEN 0.0.0.0 3493
    

    Then confirm from another machine: upsc myups@<your-nut-host>.

  3. Point the bot at it (optional). If your NUT server runs on the same box as Unraid, the bot uses your unraid.host automatically. Otherwise set nut.host in config.yaml.

  4. Credentials (optional). Most upsd setups allow anonymous reads, since upsd.users normally gates only SET and instant commands. If yours does not, set NUT_USERNAME and NUT_PASSWORD in config/.env.


⁠Configuration

Configuration is stored in config/config.yaml. On first run, the interactive setup wizard creates this file. You can also run /setup anytime to reconfigure.

Location:

  • Unraid: /mnt/user/appdata/unraid-monitor/config/config.yaml
  • Docker: ./config/config.yaml (relative to project root)
⁠Essential Settings
# Containers to watch for log errors
log_watching:
  containers:
    - plex
    - radarr
    - sonarr
    - lidarr
  error_patterns:
    - "error"
    - "exception"
    - "fatal"
    - "failed"
    - "critical"
  ignore_patterns:
    - "DeprecationWarning"
    - "DEBUG"
  cooldown_seconds: 900  # 15 min between alerts for same container

# Containers to hide from status reports
ignored_containers:
  - some-temp-container

# Containers that cannot be controlled via Telegram (safety)
protected_containers:
  - unraid-monitor-bot
  - mariadb
  - postgresql14
⁠Resource Monitoring

CPU is reported per-core on Linux, so multi-threaded apps can exceed 100% (e.g., 200% = 2 cores fully used). Set thresholds accordingly.

resource_monitoring:
  enabled: true
  poll_interval_seconds: 60
  sustained_threshold_seconds: 120  # Alert after 2 min exceeded

  defaults:
    cpu_percent: 80
    memory_percent: 85

  # Per-container overrides (also adjustable via Telegram)
  containers:
    plex:
      cpu_percent: 200   # Plex transcoding uses multiple cores
      memory_percent: 90
    handbrake:
      cpu_percent: 400   # Expected to max out all cores

Per-container thresholds can also be adjusted directly from Telegram: when a resource alert fires, tap ⚙️ Raise Limit to pick a new threshold. The change applies immediately and persists across restarts.

⁠Memory Pressure Management

Automatically kills low-priority containers when system memory is critical.

memory_management:
  enabled: false  # Disabled by default - enable with caution
  warning_threshold: 90      # Notify at this %
  critical_threshold: 95     # Start killing at this %
  safe_threshold: 80         # Offer restart when below this
  kill_delay_seconds: 60     # Warning before killing
  stabilization_wait: 180    # Wait between kills

  # Never kill these (highest priority)
  priority_containers:
    - plex
    - mariadb

  # Kill these in order during memory pressure (lowest priority first)
  killable_containers:
    - handbrake
    - tdarr

  # Offer a one-tap Restart button for these on pressure alerts — for
  # services that hog memory but recover after a bounce (classic Plex).
  # Pick them from Telegram via /manage → Features → Configure memory restarts.
  restart_containers:
    - plex

Memory warnings list the top 5 memory users and offer Restart/Stop buttons sorted largest-first, so the biggest win is always the top button.

⁠Unraid Server Monitoring
unraid:
  enabled: true
  host: "192.168.1.100"  # Your Unraid IP
  port: 443
  use_ssl: true
  verify_ssl: false  # Set true if using valid SSL cert

  polling:
    system: 30          # CPU/memory poll interval
    array: 300          # Array status poll interval
    notifications: 300  # Unraid notification feed poll interval
    # ups: 60    # IGNORED - UPS polling lives under the `nut:` section below

  thresholds:
    cpu_temp: 80         # Alert above this temp (C)
    cpu_usage: 95        # Alert above this %
    memory_usage: 90     # Alert above this %
    disk_temp: 50        # Alert above this temp (C)
    array_usage: 85      # Alert above this %
    # ups_battery: 30    # IGNORED - see `nut.thresholds.battery_charge` below

  notifications:
    enabled: false           # Relay Unraid's own notifications into Telegram
    min_importance: WARNING  # WARNING (default), ALERT, or INFO for everything

Unraid's notification feed is what sits behind the bell icon in the web UI: SMART warnings, disk errors, share-full warnings, parity results, plugin updates. Relaying it means one place to look instead of two. It is off by default and floored at WARNING, because the feed also carries routine INFO chatter (backup finished, parity-check tuning pausing and resuming).

Toggle it from Telegram with /manage → ⚙️ Features. Enabling or disabling restarts the bot; changing the importance floor applies immediately.

⁠UPS Monitoring (NUT)

Reads your UPS from a NUT⁠ server over TCP 3493. Enabled by default, but it does nothing until a host resolves, and it never alerts about a NUT server it has never reached.

nut:
  enabled: true          # Master switch (also toggled from /manage -> Features)
  host: ""               # Blank falls back to unraid.host
  port: 3493
  ups_name: ""           # Blank auto-picks when upsd serves exactly one UPS
  poll_seconds: 60

  thresholds:
    battery_charge: 50   # Warn below this %, but only while on battery
    load: 80             # Warn above this % of rated capacity

You get an alert when the mains drops (OB), when the battery gets low (LB), when the battery needs replacing (RB, nagged once a day rather than every poll), and on OVER, BYPASS, OFF, FSD and ALARM. Coming back to mains sends a recovery message with how long you ran on battery.

A runtime calibration (CAL) is not alerted on. It puts the UPS on battery deliberately, the same reason a parity sync is not reported as a failed disk.

If the bot cannot reach upsd, /ups and /health say unavailable and name the error. They never render a UPS it cannot read as healthy. Losing a server that was previously working sends an alert after three consecutive failed polls, so one dropped poll does not wake you up.

⁠Image-Update Detection

Checks once per day (configurable) whether a newer image is available for watched containers. Sends a single batched digest message with Pull buttons. Disabled by de

Tag summary

Content type

Image

Digest

sha256:4f43dc357…

Size

65.6 MB

Last updated

about 11 hours ago

docker pull dervish/unraidmonitorbot