Sign inSign up

xcy960815/data-middle-station

By xcy960815

•Updated 1 day ago

Image
0

9.0K

xcy960815/data-middle-station repository overview

⁠圭表

中文文档⁠

圭表 is a Nuxt 3 full-stack application for data visualization, analysis, and data processing workflows.

⁠Tech Stack

  • Frontend framework: Nuxt 3
  • UI framework: Element Plus
  • State management: Pinia
  • Data visualization: ECharts
  • Code editor: Monaco Editor
  • Styling: TailwindCSS + Less
  • Database: MySQL (platform metadata) and MySQL/PostgreSQL/ClickHouse/SQL Server/Oracle (analytics data sources)
  • Realtime communication: Socket.IO
  • AI integration: OpenAI-compatible Responses API, streamed NDJSON, AI SDK UI message types
  • Knowledge retrieval: RAG knowledge bases with structure-aware Markdown chunking, Qdrant vector search, an OpenAI-compatible Embedding endpoint, an optional cross-encoder rerank stage, and a local mock index provider
  • Monitoring: Prometheus, node-exporter, cAdvisor, prom-client
  • Scheduling: node-schedule, MySQL distributed execution locks
  • Logging: Winston
  • Testing: Vitest

⁠Features

  • Analysis: SQL-backed analysis with dimensions, measures, filters, sorting, drill state, configuration history, public access, and anonymous sharing.
  • Visualization: table, interval/bar, line, area, stacked, pie, scatter, funnel, combo, KPI, and other chart configurations.
  • AI assistance: analysis error diagnosis, dataset SQL error diagnosis, multi-turn analysis interpretation with optional server-side data tools, and an admin-only host-monitor assistant, optionally grounded by scoped knowledge base retrieval.
  • Knowledge base: admin-managed RAG bases with Markdown/TXT document upload, version history and per-scene scope control, structure-aware chunking (headings, SQL samples, tables, error cases), a calibrated cosine similarity floor with optional cross-encoder reranking, external index sync with reindexing, admin retrieval testing, and retrieval logs with retention cleanup.
  • Dashboards: separate view/edit workflows, drag and resize widgets, manual or scheduled refresh, version history, public access, and anonymous sharing.
  • Datasets: SQL modeling against the configured analytics database, Monaco schema completion, preview, field semantics and expressions, enable/disable state, and version switching.
  • Data sources: MySQL/PostgreSQL/ClickHouse/SQL Server/Oracle dedicated data sources with dialect-aware analysis queries, runtime connection mapping management, dataset source selection, and resource-level access control.
  • Permissions: view/edit/manage resource permissions for analyses, dashboards, datasets, and data sources, plus access applications, approval notifications, and public read-only fallback for content resources.
  • Reporting and alerts: immediate email reports, scheduled/recurring email tasks, server-rendered SVG chart snapshots, analysis alarm rules, and execution logs.
  • Operations: admin login/email/alarm log centers with on-demand retention cleanup per data domain, real-time system broadcasts, notification SSE, host and Nuxt monitoring, and an application update prompt based on the Nuxt build ID.
  • Accounts and reliability: self-service registration, profile avatar upload, JWT + Redis sessions with an online user list, distributed scheduled-job execution, stale-task recovery, and scheduled retention cleanup for operational data.
  • Optional demos: IndexedDB, Web Worker, EventSource, lazy loading, and canvas tables pages under /dev/*, guarded by ENABLE_DEMO_PAGES.

⁠Requirements

  • Node.js 22.x (Volta pins 22.22.2)
  • pnpm 9.x (packageManager pins [email protected]; corepack enable installs it)

Install with pnpm only: package.json carries pnpm.overrides that pin @nuxt/cli and @nuxt/schema, so npm or yarn resolve a different, untested dependency tree.

⁠Setup

  1. Clone the repository.
git clone [repository-url]
cd data-middle-station
  1. Install dependencies.
pnpm install
  1. Configure environment variables.

The application reads configuration through two chains, but the repository keeps only three files:

  • env/.env.daily, env/.env.pre: daily and pre-release application runtime configuration, loaded by pnpm dev / pnpm dev:pre / pnpm build:pre through --dotenv.
  • env/.env.prod: production configuration. One file serves both chains — Docker Compose deployment (docker compose --env-file env/.env.prod) and bare-metal PM2 / direct connection (pnpm build:prod, start:prod through --dotenv). Compose uses it to interpolate images, container names, host ports and bind-mount source paths, then injects NUXT_* variables so the app reaches the metadata MySQL, Redis, and the knowledge base embedding (Infinity) and qdrant services; postgres-data, clickhouse-data, sqlserver-data and oracle-data host the optional PostgreSQL, ClickHouse, SQL Server and Oracle demo business databases. Data source connections are registered at runtime on the Data Sources page instead of via environment variables. Within the same chain one setting can have two key names (for example MYSQL_ROOT_PASSWORD and SERVICE_DB_PASSWORD, APP_CONTAINER_NAME and MONITOR_NUXT_CONTAINER_NAME); both must carry the same value. Database addresses use the host LAN IP plus the published port, with MONITOR_PROMETHEUS_URL as the only exception — it stays a container name because Prometheus publishes no host port.

This repository is private and these environment files (with real values) are committed directly. When adding a new key, update the key set in the files below and keep each environment's value consistent with the actual deployment:

  • env/.env.daily, env/.env.pre, env/.env.prod

Example env/.env.daily:

NODE_ENV=daily
PORT=12581
# Application title
APP_NAME=圭表
APP_VERSION=daily-local
API_BASE=/api
ENABLE_DEMO_PAGES=false

# Database
SERVICE_DB_HOST=127.0.0.1
SERVICE_DB_PORT=3310
SERVICE_DB_USER=root
SERVICE_DB_PASSWORD=change_me
SERVICE_DB_NAME=data_middle_station
SERVICE_DB_TIMEZONE=+08:00
SERVICE_DB_DATE_STRINGS=true
SERVICE_DB_DECIMAL_NUMBERS=true
SERVICE_DB_CONNECTION_LIMIT=20
SERVICE_DB_CONNECT_TIMEOUT=10000
SERVICE_DB_IDLE_TIMEOUT=60000

# Encryption key for data source passwords at rest; 32-byte hex string (64 chars)
# Generate: node -e "console.log(require('crypto').randomBytes(32).toString('hex'))"
DATA_SOURCE_SECRET_KEY=change_me

# Redis
SERVICE_REDIS_HOST=127.0.0.1
SERVICE_REDIS_PORT=6383
SERVICE_REDIS_USERNAME=default
SERVICE_REDIS_PASSWORD=change_me
SERVICE_REDIS_DB=0
SERVICE_REDIS_BASE=dms-redis

# Logs
LOG_PATH=./logs
LOG_TIME_FORMAT=YYYY-MM-DD HH:mm:ss
LOG_LEVEL=info

# Auth and email
JWT_SECRET_KEY=replace_with_a_long_random_secret
JWT_EXPIRES_IN=2h
JWT_REFRESH_EXPIRES_IN=30d
AUTH_SESSION_FAIL_MODE=strict
SM2_PUBLIC_KEY=replace_with_sm2_public_key
SM2_PRIVATE_KEY=replace_with_sm2_private_key
# Session cookie SameSite policy: lax (default) / strict / none.
# Cross-origin development needs none; browsers then require Secure (set automatically),
# so cookies only work over HTTPS — plain-http cross-origin cannot carry cookies.
AUTH_COOKIE_SAMESITE=lax
# Comma-separated origins allowed to call the API cross-origin
CORS_ALLOWED_ORIGINS=
ENABLE_SOCKET_DEMO=false
SOCKET_ALLOWED_ORIGINS=
SMTP_HOST=smtp.example.com
SMTP_PORT=465
SMTP_SECURE=true
SMTP_REJECT_UNAUTHORIZED=true
[email protected]
SMTP_PASS=your_smtp_password
[email protected]

# LLM AI
LLM_PROVIDER=openai-compatible
LLM_API_KEY=
LLM_BASE_URL=
LLM_MODEL=

# Knowledge index
# Index provider: qdrant (native Qdrant) or mock (local stub); leave empty to disable
KNOWLEDGE_INDEX_PROVIDER=
# Qdrant HTTP base URL; use http://qdrant:6333 from the Compose app container
KNOWLEDGE_INDEX_BASE_URL=
# Qdrant API key
KNOWLEDGE_INDEX_API_KEY=
# Knowledge index request timeout in milliseconds
KNOWLEDGE_INDEX_TIMEOUT_MS=30000
# OpenAI-compatible Embedding base URL; /embeddings is appended automatically
KNOWLEDGE_EMBEDDING_BASE_URL=
# Embedding API key, independent from LLM_API_KEY
KNOWLEDGE_EMBEDDING_API_KEY=
# Embedding model and its native vector dimension
KNOWLEDGE_EMBEDDING_MODEL=
KNOWLEDGE_EMBEDDING_DIMENSIONS=
KNOWLEDGE_EMBEDDING_BATCH_SIZE=16
KNOWLEDGE_EMBEDDING_TIMEOUT_MS=30000
# Retrieval similarity floor (cosine): hits below it are dropped, 0 disables the filter.
# Calibrated for bge-small-zh-v1.5; re-calibrate whenever the embedding model changes.
KNOWLEDGE_RETRIEVAL_MIN_SCORE=0.5
# Optional per-path recall candidate cap (vector and keyword recall each fetch up to this many).
# Unset or blank falls back to the built-in default of 20; non-integer or out-of-range values (1-50) fail retrieval.
KNOWLEDGE_CANDIDATE_TOP_K=
# Character budget of one chunk, and the overlap used when splitting plain text.
# The overlap must stay smaller than the budget, and the budget well below the embedding context window.
KNOWLEDGE_CHUNK_MAX_CHARACTERS=400
KNOWLEDGE_CHUNK_OVERLAP=60
# Index sync lock TTL in milliseconds; also how long before a stale running task is reclaimed.
KNOWLEDGE_INDEX_LOCK_TTL_MS=600000
# Tasks claimed per worker tick, and the worker polling interval in milliseconds.
KNOWLEDGE_INDEX_WORKER_BATCH_SIZE=10
KNOWLEDGE_INDEX_WORKER_INTERVAL_MS=2000
# Retry back-off for failed index tasks, in minutes: min(base * 2^(n-1), max).
KNOWLEDGE_INDEX_RETRY_BASE_MINUTES=5
KNOWLEDGE_INDEX_RETRY_MAX_MINUTES=60
# Attempts allowed before a non-delete index task reaches a terminal state.
KNOWLEDGE_INDEX_MAX_ATTEMPTS=3
# Delay in minutes before retrying a task while no index provider is configured.
KNOWLEDGE_INDEX_PROVIDER_MISSING_DELAY_MINUTES=30
# Optional cross-encoder rerank service speaking the /rerank protocol (Jina, bge-reranker, ...).
# Leave the base URL empty to keep the plain vector score order.
KNOWLEDGE_RERANK_BASE_URL=
KNOWLEDGE_RERANK_API_KEY=
KNOWLEDGE_RERANK_MODEL=BAAI/bge-reranker-v2-m3
KNOWLEDGE_RERANK_TIMEOUT_MS=10000

# Monitor
MONITOR_PROMETHEUS_URL=http://localhost:9090
MONITOR_METRICS_TOKEN=replace_with_a_long_random_monitor_token
MONITOR_DISK_MOUNTPOINT=/
MONITOR_LOG_DISK_MOUNTPOINT=/
MONITOR_NUXT_CONTAINER_NAME=dms-app-container
MONITOR_CONTAINER_METRICS_ENABLED=false
MONITOR_REFRESH_INTERVAL_SECONDS=5
MONITOR_RETENTION_DAYS=30
MONITOR_MAX_HISTORY_POINTS=2000

# Log cleanup
SCHEDULED_EMAIL_LOG_RETENTION_DAYS=90
ANALYSIS_ALARM_LOG_RETENTION_DAYS=90
NOTIFICATION_READ_RETENTION_DAYS=90
SYSTEM_BROADCAST_ENDED_RETENTION_DAYS=365
ACCESS_APPLY_RESOLVED_RETENTION_DAYS=365
KNOWLEDGE_RETRIEVAL_LOG_RETENTION_DAYS=90
USER_LOGIN_LOG_RETENTION_DAYS=365

# Scheduled email recovery
SCHEDULED_EMAIL_RUNNING_TIMEOUT_MINUTES=10
SCHEDULED_EMAIL_RECOVERY_INTERVAL_MINUTES=5
SCHEDULED_EMAIL_ORPHANED_EXECUTIONS_THRESHOLD_MINUTES=30

# Distributed execution history cleanup
SCHEDULED_JOB_EXECUTION_COMPLETED_RETENTION_DAYS=30
SCHEDULED_JOB_EXECUTION_FAILED_RETENTION_DAYS=15
SCHEDULED_JOB_EXECUTION_CLEANUP_CRON=15 3 * * *
  1. Initialize the databases.

The schema and demo data are plain dumps under sql/, applied by hand — neither Compose nor the app imports them, and the app never migrates its own metadata schema:

  • sql/data_middle_station.sql: platform metadata schema (resources, permissions, datasets, schedules, knowledge bases, logs). Import it into the database named by SERVICE_DB_NAME; the dump has no CREATE DATABASE.
  • sql/mysql_business_center.sql: optional MySQL demo business database for analysis and dataset exercises. It creates business_center itself and must not be imported into the metadata database.
  • sql/postgresql_retail_center.sql: optional PostgreSQL demo business database matching the Compose postgres-data service. It targets retail_center and the public schema, which is the only schema the PostgreSQL data source driver reads.
  • sql/clickhouse_behavior_center.sql: optional ClickHouse demo business database matching the Compose clickhouse-data service. It targets behavior_center with ClickHouse native DDL (MergeTree tables, no transaction block), so it must be imported with clickhouse-client rather than a MySQL/PostgreSQL client.
  • sql/sqlserver_finance_center.sql: optional SQL Server demo business database matching the Compose sqlserver-data service (host port 3316). It targets finance_center with T-SQL DDL; the official image only ships linux/amd64 (Apple Silicon needs Rosetta emulation) and its container runs as UID 10001, so the host data directory must be chown -R 10001:0 before first start.
  • sql/oracle_logistics_center.sql: optional Oracle demo business database matching the Compose oracle-data service (host port 3317). Connect as the business user root to the service name FREEPDB1 and run the whole file in any client such as Navicat, DBeaver or sqlplus (it contains plain SQL only, no SQLPlus-specific commands). Table and column names are double-quoted lowercase identifiers, because Oracle folds unquoted identifiers to uppercase while the analysis queries quote lowercase column names. The image needs about 2 GB of memory and its container runs as UID 54321, so the host data directory must be chown -R 54321:54321 before first start.

Register the business databases you imported as data sources on the Data Sources page after the app is running.

⁠Development

Start the development server:

# Default environment, reads env/.env.daily
pnpm dev

# Against the pre or prod env file
pnpm dev:pre
pnpm dev:prod

The server port comes from PORT in the selected env file.

⁠Build and Deployment

Build the production version:

pnpm build        # env/.env.daily
pnpm build:pre    # env/.env.pre
pnpm build:prod   # env/.env.prod

To deploy with Docker Compose:

# First deployment or full-stack reconciliation
docker compose --env-file env/.env.prod -p dms-service -f docker-compose.yml \
  up -d --pull always

# Update only the application after a new image is published
docker compose --env-file env/.env.prod -p dms-service -f docker-compose.yml \
  up -d --pull always --force-recreate data-middle-station

DATA_MIDDLE_STATION_IMAGE may remain xcy960815/data-middle-station:latest. Nuxt exposes a single built-in build identifier to both the server and client; GitHub Actions injects a unique CI run identifier for each image publication so Docker cannot reuse an older Nuxt build layer. The deployment command explicitly pulls the remote latest digest and recreates the Nuxt container. After the replacement succeeds, already-open pages detect the changed build identifier and prompt the user to refresh. Restarting the same image keeps its original build identifier and does not produce a false update notification. Image publication is currently triggered by v* Git tags, while build identifier generation itself does not depend on the tag value.

Double-check the target host's real values in env/.env.prod before deployment. One up -d reconciles the whole stack: the application plus redis, mysql-main (platform metadata), mysql-data, postgres-data, clickhouse-data, sqlserver-data and oracle-data (demo business databases), embedding and qdrant (knowledge base), and prometheus, node-exporter and cadvisor (host monitoring). The Compose file exposes configurable host ports for the web service, Redis, the MySQL instances, PostgreSQL, ClickHouse, SQL Server, and Oracle; restrict the Redis/MySQL/PostgreSQL/ClickHouse/SQL Server/Oracle ports with the host firewall when remote access is not required. Prometheus, node-exporter, and cAdvisor remain internal to dms-monitor-network. The rerank service is not part of the stack and is reached over the LAN when KNOWLEDGE_RERANK_BASE_URL is set.

The host monitoring deployment adds Prometheus, node-exporter, and cAdvisor on the internal dms-monitor-network. Prometheus data is persisted under PROMETHEUS_DATA_DIR_PATH, and the Nuxt metrics endpoint is protected by MONITOR_METRICS_TOKEN. The Prometheus image runs as nobody (65534:65534), so the directory configured by PROMETHEUS_DATA_DIR_PATH must be writable by that user when it is pre-created, for example chown -R 65534:65534 /your/prometheus/data/path. cAdvisor runs with privileged host access and is only reachable from the internal monitor network. Set DOCKER_DATA_ROOT to the target host's actual Docker data root, and validate its /dev/kmsg device plus /var/run, /sys, and /dev/disk mounts. MONITOR_DISK_MOUNTPOINT selects the host filesystem card, while MONITOR_LOG_DISK_MOUNTPOINT selects the filesystem containing the Nuxt log bind mount.

⁠PM2 on a single host

Docker Compose is the deployment path for the NAS; PM2 runs the same build artifact directly on one machine for local or Windows hosts. The entry is ecosystem.config.js, which starts .output/server/index.mjs in fork mode, so build first:

pnpm build:prod && pnpm start:prod    # daily and pre have matching start:* scripts
pnpm restart:prod                    # or restart / restart:pre
pnpm stop && pnpm delete

start:* trigger a prestart* lifecycle that rebuilds .output/; keep pnpm dev stopped when you run them.

⁠Code Quality

The project uses the following tools:

  • ESLint - code linting
  • Prettier - code formatting
  • Husky - Git hooks
  • Commitlint - commit message linting
  • Conventional Changelog - changelog generation
  • Vitest - unit and service tests

Run the checks that match the change:

pnpm lint            # ESLint over the repository
pnpm lint:fix
pnpm format:check    # Prettier; run pnpm format to rewrite
pnpm test            # Vitest, tests live in test/
pnpm test:watch
pnpm exec nuxt typecheck   # vue-tsc over pages, stores and server types

⁠Project Structure

data-middle-station/
├── assets/              # Styles and compiled static assets
├── components/          # Shared Vue components
├── composables/         # Frontend composables
├── layouts/             # Layout components
├── middleware/          # Route middleware (auth.global.ts)
├── pages/               # File-based routes; pages/dev/ is the guarded demo area
├── plugins/             # Nuxt client plugins
├── public/              # Unprocessed public assets
├── shared/              # Contracts and domain rules consumed by both client and server
├── stores/              # Pinia stores
├── types/               # TypeScript declarations
├── utils/               # Frontend utilities
├── server/              # Nitro server: api -> service -> mapper, plus middleware, plugins, validation
├── sql/                 # Metadata schema and demo business database dumps
├── storage/             # Uploaded knowledge base documents (runtime data)
├── test/                # Vitest suites
├── env/                 # Application and Docker Compose env files (committed in this private repo)
├── docs/                # Temporary plans and archived material
└── .agents/             # Long-lived collaboration rules and per-module knowledge

The main request path is pages/components -> stores/composables -> server/api -> service -> mapper -> SQL.

⁠More Documentation

⁠Versioning and Changelog

A release is a vX.Y.Z Git tag. Pushing one runs .github/workflows/docker-publish.yml⁠, which publishes xcy960815/data-middle-station:latest plus the matching version label; package.json version is kept aligned with the newest tag. Commit messages follow Conventional Commits and are linted by Commitlint, which is what makes the changelog generatable.

CHANGELOG.md is generated from the tags, not written by hand:

pnpm changelog          # append the newest release range
pnpm changelog:version  # rebuild the whole file from every tag

After a full rebuild, fix up the top of the file by hand: the generator writes commits that are not covered by a tag into invented 0.0.N+1 headings, repeats the current version as an empty heading, and drops the # Changelog preamble. Fold those commits into one ## 未发布 (YYYY-MM-DD) block and restore the preamble. CHANGELOG.md is in .prettierignore, so formatting is not rewritten on commit.

⁠Contributing

  1. Fork the repository.
  2. Create a feature branch (git checkout -b feature/AmazingFeature).
  3. Commit your changes (git commit -m 'feat: add some amazing feature').
  4. Push the branch (git push origin feature/AmazingFeature).
  5. Open a pull request.

⁠License

MIT License. See the LICENSE file in the repository root.

⁠Contact

Please submit questions or suggestions through GitHub Issues.

Tag summary

Content type

Image

Digest

sha256:e8cd20e4a…

Size

90.7 MB

Last updated

1 day ago

docker pull xcy960815/data-middle-station