Sign inSign up

ssivart/ask-cloudscape

By ssivart

•Updated 9 months ago

An **offline-deployable** Cloudscape documentation retrieval service

Image
Developer tools
Web servers
0

273

ssivart/ask-cloudscape repository overview

⁠Cloudscape RAG Server — BM25 / Runtime-Offline

An offline-deployable Cloudscape documentation retrieval service:

Browsertrix crawl → Generate WACZ → (Optional) Generate TypeDoc Markdown with Node/Docker → Page-level SQLite FTS5 index (BM25) → MCP HTTP server for querying.

  • Runs fully offline (no model required at query time).
  • Results are grouped into four buckets: components.api, components.usage, patterns, typedoc.
  • Each bucket is ranked separately using FTS5 bm25(), then normalized to [0,1] (where 1 is best).

⁠Quick Start

⁠0) Requirements
  • Docker (for Browsertrix, TypeDoc, and server build)
  • Python 3.10+
⁠1) Crawl (requires internet access)
make crawl

Output: data/wacz/collections/cloudscape/<timestamp>.wacz (Generated with --text and --generateWACZ, containing pages/pages.jsonl and pages/extraPages.jsonl.)

To use a fixed filename, copy the latest WACZ to data/wacz/cloudscape.wacz after generation.

Converts npm packages (default: @cloudscape-design/components) into Markdown, included in the index:

make typedoc

Output: data/typedoc_md/**

⁠3) Build the Index (offline)
make index
  • Reads from data/wacz/.../cloudscape.wacz (or timestamped file) and data/typedoc_md/**
  • Produces build/index.db (SQLite + FTS5)
⁠4) Run the MCP HTTP Server (offline)
make build
make run
# Default: http://localhost:8000

⁠MCP Server Usage

This server implements the Model Context Protocol (MCP) using SSE (Server-Sent Events) transport. It provides two tools: search and page.

⁠Using with MCP Inspector

The easiest way to test and explore the server is using MCP Inspector:

npx @modelcontextprotocol/inspector http://localhost:8000/sse

This opens a web interface where you can:

  • Browse available tools (search and page)
  • Test tool calls interactively
  • View responses in real-time
⁠Using with MCP Clients
{
  "mcpServers": {
    "ask-cloudscape": {
      "url": "http://localhost:8000/sse"
    }
  }
}
⁠Available MCP Tools
⁠1. search Tool

Intelligent search for AWS Cloudscape documentation with component name detection and multi-bucket results.

Parameters:

  • q (string, required): Search query (supports component names, keywords, phrases)
  • k_components (integer, optional): Number of API/Usage results per bucket (default: 1)
  • k_patterns (integer, optional): Number of pattern results (default: 5)
  • k_typedoc (integer, optional): Number of TypeDoc results (default: 3)

Example calls:

Tool: search
Arguments: {"q": "Flashbar"}
→ Returns default results (1 API, 1 Usage, 5 Patterns, 3 TypeDoc)
Tool: search
Arguments: {"q": "Button", "k_components": 2, "k_patterns": 0, "k_typedoc": 0}
→ Returns only component results (2 API, 2 Usage)
Tool: search
Arguments: {"q": "\"status indicator\" AND color", "k_components": 1, "k_patterns": 3, "k_typedoc": 0}
→ Phrase search with FTS5 operators

Response structure:

{
  "query": "Flashbar",
  "detected_component": "flashbar",
  "components": {
    "api": [{"url": "...", "title": "...", "text_preview": "...", "text_len": 7784}],
    "usage": [{"url": "...", "title": "...", "text_preview": "...", "text_len": 9644}]
  },
  "patterns": [...],
  "typedoc": [...],
  "used_rag": true,
  "pack_id": "..."
}
⁠2. page Tool

Retrieve the full content of a specific documentation page.

Parameters:

  • url (string, required): Full URL of the page to retrieve

Example call:

Tool: page
Arguments: {"url": "https://cloudscape.design/components/flashbar/?tabId=api"}
→ Returns complete page content
⁠Running Tests
make test

This runs tests in Docker to avoid polluting the local environment. Tests verify that component queries return the correct API and Usage URLs at position 0.


⁠Common Makefile Targets

  • make crawl → Crawl with Browsertrix, output WACZ (pages.jsonl, extraPages.jsonl)
  • make typedoc → Generate TypeDoc Markdown via Node/Docker (data/typedoc_md/**)
  • make index → Build page-level SQLite FTS5 index from WACZ + Markdown
  • make build → Build Docker image with MCP server
  • make run → Run MCP HTTP server with build/index.db
  • make test → Run comprehensive search ranking tests in Docker

⁠Computation Principles & Indexing Rules

  1. Page-Level Indexing

    • Indexes whole pages from Cloudscape + TypeDoc Markdown (not fragmented by sentence/paragraph).
    • Agent-friendly: one hit = one page of context.
  2. URL Normalization & Buckets

    • components.api: .../components/<name>/?tabId=api

    • components.usage: .../components/<name>/?tabId=usage

    • patterns: .../patterns/** (no tab split)

    • typedoc: From npm packages like @cloudscape-design/components (typedoc://...)

    Noise pages (e.g., tabId=playground/testing or ?example=) are excluded.

  3. Content Cleaning (Denoise)

    • Removes cookie banners, footers, long sidebar menus.
    • Previews prioritize Properties / Usage / Guidelines.
  4. BM25 Ranking & Normalization

    • FTS5 bm25() (lower = better).
    • Min-max normalized per bucket to [0,1], exposed as score (higher = better).
  5. Previews

    • /search returns text_preview aligned to keywords/section headers (with highlights).
    • For full text, use /page?url=....

⁠Configuration

Optional .env:

SEEDS="https://cloudscape.design/get-started/ https://cloudscape.design/components/ https://cloudscape.design/patterns/"
COLLECTION=cloudscape
OUT_DIR=data/wacz
NODE_IMAGE=node:20-bookworm-slim   ; # Used by TypeDoc container

⁠Troubleshooting

  • Empty /search results → Ensure build/index.db exists with pages / pages_fts, and --text was used during crawl.
  • Previews still include noise → Re-run make index (latest indexer improves cleaning).
  • TypeDoc too noisy → Manually ignore very short .md files or reduce k_typedoc.

⁠Suggested Prompt for AI Agents

You are a Cloudscape Design System assistant. When the user asks about a component or pattern:

  1. Use the search MCP tool with {"q": "<user query>", "k_components": 1, "k_patterns": 3, "k_typedoc": 2}.
  2. Start with components.api[0] and components.usage[0] text_preview. Use the page tool to fetch full text if needed.
  3. For usage guidance, best practices, or UX guidelines, check the patterns bucket.
  4. For precise types/interfaces/events, check the typedoc bucket.
  5. When composing answers:
    • Cite sources using the returned url.
    • Integrate in the order API → Usage → Patterns → TypeDoc.
    • Use only retrieved content—no speculation.
    • Deduplicate redundant content, keeping the most specific/operational.
    • If no exact match, state "Not explicitly described in Cloudscape documentation" and provide closest reference.
  6. End with a "References" list of 1–4 URLs.

Queries may use keywords or phrases (e.g., "Flashbar", "\"status indicator\" AND color"). If results are too broad, start with components.api and usage; expand to patterns if more context is needed.

Tag summary

Content type

Image

Digest

sha256:c8de8b61b…

Size

70.7 MB

Last updated

9 months ago

docker pull ssivart/ask-cloudscape:sse-251228