Sign inSign up

repodb/telemetry-default-filedatasinker

By repodb

•Updated 3 months ago

A long-running worker that periodically reads telemetry data from PGSQL and archives them to disk.

Image
Integration & delivery
Developer tools
Monitoring & observability
0

304

repodb/telemetry-default-filedatasinker repository overview

⁠RepoDB Telemetry File Data Sinker Service

A long-running worker that periodically reads telemetry rows out of RepoDB PGSQL⁠ (DefaultTelemetry table) and archives them to Parquet files on disk (or a mounted volume), picking up exactly where the previous cycle left off.

⁠Running with Docker

Create the volume (to persist the status file and archived Parquet files across container restarts) and network (skip either if it already exists):

docker volume create repodb_telemetry_data
docker network create repodb

Run the image:

docker run -d --name repodb-telemetry-default-filedatasinker --cpus=0.25 \
  -e CONNECTION_STRING="postgresql://postgres:RepoDB2026@repodb-insights-postgres:5432/repodb_insights" \
  -e DIRECTORY_PATH="/data/telemetry/default" \
  -e FREQUENCY_IN_MINUTES=60 \
  -v repodb_telemetry_data:/data/telemetry/default \
  --network=repodb \
  repodb/telemetry-default-filedatasinker:latest

A volume mount for DIRECTORY_PATH (as shown above) is recommended so the status file and archived Parquet files persist across container restarts.

⁠How It Works

The service runs an immediate sink cycle on startup, then loops forever, sleeping between cycles:

  1. Load the checkpoint. The service reads a JSON status file (see "Status File" below). If it doesn't exist yet, the checkpoint is initialized to the earliest StartTime currently in the Telemetry table (or the current time if the table is empty).
  2. Determine the window. The window to archive runs from the checkpoint up to "now" (UTC, truncated to whole seconds).
  3. Split into daily chunks. The window is split at UTC day boundaries, so a multi-day gap since the last run is archived one calendar day at a time rather than in a single query.
  4. Archive each chunk. For each (start, end) chunk, rows where "StartTime" >= start AND "StartTime" < end are fetched and, if any exist, merged into that day's Parquet file.
  5. Advance the checkpoint. After a chunk with at least one row is written, the checkpoint moves to that chunk's end time and a run entry is appended to the status file's history. Chunks with zero rows are skipped without advancing the checkpoint, so an empty day is simply re-checked on the next cycle rather than recorded as "done."
  6. Handle failures quietly. Any exception during a cycle is caught and logged; the service does not crash and simply tries again on the next scheduled cycle.

⁠Configuration

The service is configured entirely via environment variables:

VariableRequiredDescription
CONNECTION_STRINGYesPostgreSQL connection string (e.g. postgresql://user:pass@host:5432/db). Read directly from the environment with no default, so the service fails to start if it's missing.
DIRECTORY_PATHNoRoot directory the status file and archived data are written under. Defaults to /tmp/repodb/telemetry. Backslashes are normalized to forward slashes.
FREQUENCY_IN_MINUTESNoMinutes between sink cycles. Defaults to 60. Also used to align each cycle's wake-up time to a fixed grid (see "Cycle Timing" below).
MAX_META_HISTORYNoMaximum number of run entries retained in the status file's history log. Defaults to 20000.
LOG_LEVELNoThe logging level (DEBUG, INFO, WARNING, etc.). Defaults to DEBUG.

The source table and timestamp column are fixed in code as the query used by the code (DefaultTelemetry table, filtered on StartTime field), matching the convention used by the Purger⁠ service. They are not configurable via environment variables.

Tag summary

Content type

Image

Digest

sha256:81868846b…

Size

138.5 MB

Last updated

3 months ago

docker pull repodb/telemetry-default-filedatasinker