An **offline-deployable** Cloudscape documentation retrieval service
273
An offline-deployable Cloudscape documentation retrieval service:
Browsertrix crawl → Generate WACZ → (Optional) Generate TypeDoc Markdown with Node/Docker → Page-level SQLite FTS5 index (BM25) → MCP HTTP server for querying.
components.api, components.usage, patterns, typedoc.bm25(), then normalized to [0,1] (where 1 is best).make crawl
Output: data/wacz/collections/cloudscape/<timestamp>.wacz
(Generated with --text and --generateWACZ, containing pages/pages.jsonl and pages/extraPages.jsonl.)
To use a fixed filename, copy the latest WACZ to
data/wacz/cloudscape.waczafter generation.
Converts npm packages (default: @cloudscape-design/components) into Markdown, included in the index:
make typedoc
Output: data/typedoc_md/**
make index
data/wacz/.../cloudscape.wacz (or timestamped file) and data/typedoc_md/**build/index.db (SQLite + FTS5)make build
make run
# Default: http://localhost:8000
This server implements the Model Context Protocol (MCP) using SSE (Server-Sent Events) transport. It provides two tools: search and page.
The easiest way to test and explore the server is using MCP Inspector:
npx @modelcontextprotocol/inspector http://localhost:8000/sse
This opens a web interface where you can:
search and page){
"mcpServers": {
"ask-cloudscape": {
"url": "http://localhost:8000/sse"
}
}
}
search ToolIntelligent search for AWS Cloudscape documentation with component name detection and multi-bucket results.
Parameters:
q (string, required): Search query (supports component names, keywords, phrases)k_components (integer, optional): Number of API/Usage results per bucket (default: 1)k_patterns (integer, optional): Number of pattern results (default: 5)k_typedoc (integer, optional): Number of TypeDoc results (default: 3)Example calls:
Tool: search
Arguments: {"q": "Flashbar"}
→ Returns default results (1 API, 1 Usage, 5 Patterns, 3 TypeDoc)
Tool: search
Arguments: {"q": "Button", "k_components": 2, "k_patterns": 0, "k_typedoc": 0}
→ Returns only component results (2 API, 2 Usage)
Tool: search
Arguments: {"q": "\"status indicator\" AND color", "k_components": 1, "k_patterns": 3, "k_typedoc": 0}
→ Phrase search with FTS5 operators
Response structure:
{
"query": "Flashbar",
"detected_component": "flashbar",
"components": {
"api": [{"url": "...", "title": "...", "text_preview": "...", "text_len": 7784}],
"usage": [{"url": "...", "title": "...", "text_preview": "...", "text_len": 9644}]
},
"patterns": [...],
"typedoc": [...],
"used_rag": true,
"pack_id": "..."
}
page ToolRetrieve the full content of a specific documentation page.
Parameters:
url (string, required): Full URL of the page to retrieveExample call:
Tool: page
Arguments: {"url": "https://cloudscape.design/components/flashbar/?tabId=api"}
→ Returns complete page content
make test
This runs tests in Docker to avoid polluting the local environment. Tests verify that component queries return the correct API and Usage URLs at position 0.
make crawl → Crawl with Browsertrix, output WACZ (pages.jsonl, extraPages.jsonl)make typedoc → Generate TypeDoc Markdown via Node/Docker (data/typedoc_md/**)make index → Build page-level SQLite FTS5 index from WACZ + Markdownmake build → Build Docker image with MCP servermake run → Run MCP HTTP server with build/index.dbmake test → Run comprehensive search ranking tests in DockerPage-Level Indexing
URL Normalization & Buckets
components.api: .../components/<name>/?tabId=api
components.usage: .../components/<name>/?tabId=usage
patterns: .../patterns/** (no tab split)
typedoc: From npm packages like @cloudscape-design/components (typedoc://...)
Noise pages (e.g.,
tabId=playground/testingor?example=) are excluded.
Content Cleaning (Denoise)
BM25 Ranking & Normalization
bm25() (lower = better).[0,1], exposed as score (higher = better).Previews
/search returns text_preview aligned to keywords/section headers (with highlights)./page?url=....Optional .env:
SEEDS="https://cloudscape.design/get-started/ https://cloudscape.design/components/ https://cloudscape.design/patterns/"
COLLECTION=cloudscape
OUT_DIR=data/wacz
NODE_IMAGE=node:20-bookworm-slim ; # Used by TypeDoc container
/search results → Ensure build/index.db exists with pages / pages_fts, and --text was used during crawl.make index (latest indexer improves cleaning)..md files or reduce k_typedoc.You are a Cloudscape Design System assistant. When the user asks about a component or pattern:
- Use the
searchMCP tool with{"q": "<user query>", "k_components": 1, "k_patterns": 3, "k_typedoc": 2}.- Start with
components.api[0]andcomponents.usage[0]text_preview. Use thepagetool to fetch full text if needed.- For usage guidance, best practices, or UX guidelines, check the
patternsbucket.- For precise types/interfaces/events, check the
typedocbucket.- When composing answers:
- Cite sources using the returned
url.- Integrate in the order API → Usage → Patterns → TypeDoc.
- Use only retrieved content—no speculation.
- Deduplicate redundant content, keeping the most specific/operational.
- If no exact match, state "Not explicitly described in Cloudscape documentation" and provide closest reference.
- End with a "References" list of 1–4 URLs.
Queries may use keywords or phrases (e.g.,
"Flashbar","\"status indicator\" AND color"). If results are too broad, start withcomponents.apiandusage; expand topatternsif more context is needed.
Content type
Image
Digest
sha256:c8de8b61b…
Size
70.7 MB
Last updated
9 months ago
docker pull ssivart/ask-cloudscape:sse-251228