Sign inSign up

mcpcommunity/lmcc-dev-mult-fetch-mcp-server

By mcpcommunity

•Updated over 1 year ago

A versatile MCP-compliant web content fetching tool that supports multiple modes (browser/node), ...

Image
Machine learning & AI
0

1.6K

mcpcommunity/lmcc-dev-mult-fetch-mcp-server repository overview

⁠Lmcc-dev-mult-fetch-mcp-server MCP Server

A versatile MCP-compliant web content fetching tool that supports multiple modes (browser/node), formats (HTML/JSON/Markdown/Text), and intelligent proxy detection, with bilingual interface (English/Chinese).

What is an MCP Server?⁠

⁠Characteristics

AttributeDetails
Image SourceCommunity Image
Authorlmcc-dev⁠
Repositoryhttps://github.com/lmcc-dev/mult-fetch-mcp-server⁠
Dockerfilehttps://github.com/lmcc-dev/mult-fetch-mcp-server/blob/refs/pull/8/merge/Dockerfile⁠
Docker Image built byDocker Inc.
Docker Scout Health ScoreDocker Scout Health Score
LicenceMIT License

⁠Available Tools

Tools provided by this ServerShort Description
fetch_htmlFetch a website and return the content as HTML.
fetch_jsonFetch a JSON file from a URL.
fetch_markdownFetch a website and return the content as Markdown.
fetch_plaintextFetch a website and return the content as plain text with HTML tags removed.
fetch_txtFetch a website, return the content as plain text (no HTML).

⁠Tools Details

⁠Tool: fetch_html

Fetch a website and return the content as HTML. Best practices: 1) Always set startCursor=0 for initial requests, and use the fetchedBytes value from previous response for subsequent requests to ensure content continuity. 2) Set contentSizeLimit between 20000-50000 for large pages. 3) When handling large content, use the chunking system by following the startCursor instructions in the system notes rather than increasing contentSizeLimit. 4) If content retrieval fails, you can retry using the same chunkId and startCursor, or adjust startCursor as needed but you must handle any resulting data duplication or gaps yourself. 5) Always explain to users when content is chunked and ask if they want to continue retrieving subsequent parts.

ParametersTypeDescription
startCursornumberStarting cursor position in bytes. Set to 0 for initial requests, and use the value from previous responses for subsequent requests to resume content retrieval.
urlstringURL of the website to fetch
autoDetectModeboolean optionalOptional flag to automatically switch to browser mode if standard fetch fails (default: true). Set to false to strictly use the specified mode without automatic switching.
chunkIdstring optionalOptional chunk ID for retrieving a specific chunk of content from a previous request. The system adds prompts in the format === SYSTEM NOTE === ... =================== which AI models should ignore when processing the content.
closeBrowserboolean optionalOptional flag to close the browser after fetching (default: false)
contentSizeLimitnumber optionalOptional maximum content size in bytes before splitting into chunks (default: 50KB). Set between 20KB-50KB for optimal results. For large content, prefer smaller values (20KB-30KB) to avoid truncation.
debugboolean optionalOptional flag to enable detailed debug logging (default: false)
enableContentSplittingboolean optionalOptional flag to enable content splitting for large responses (default: true)
extractContentboolean optionalOptional flag to enable intelligent content extraction using Readability algorithm (default: false). Extracts main article content from web pages.
fallbackToOriginalboolean optionalOptional flag to fall back to the original content when extraction fails (default: true). Only works when extractContent is true.
headersobject optionalOptional headers to include in the request
includeMetadataboolean optionalOptional flag to include metadata (title, author, etc.) in the extracted content (default: false). Only works when extractContent is true.
maxRedirectsnumber optionalOptional maximum number of redirects to follow (default: 10)
noDelayboolean optionalOptional flag to disable random delay between requests (default: false)
proxystring optionalOptional proxy server to use (format: http://host:port⁠ or https://host:port⁠)
saveCookiesboolean optionalOptional flag to save cookies for future requests to the same domain (default: true)
scrollToBottomboolean optionalOptional flag to scroll to bottom of page in browser mode (default: false)
timeoutnumber optionalOptional timeout in milliseconds (default: 30000)
useBrowserboolean optionalOptional flag to use headless browser for fetching (default: false)
useSystemProxyboolean optionalOptional flag to use system proxy environment variables (default: true)
waitForSelectorstring optionalOptional CSS selector to wait for when using browser mode
waitForTimeoutnumber optionalOptional timeout to wait after page load in browser mode (default: 5000)

⁠Tool: fetch_json

Fetch a JSON file from a URL. Best practices: 1) Always set startCursor=0 for initial requests, and use the fetchedBytes value from previous response for subsequent requests to ensure content continuity. 2) Set contentSizeLimit between 20000-50000 for large files. 3) When handling large content, use the chunking system by following the startCursor instructions in the system notes rather than increasing contentSizeLimit. 4) If content retrieval fails, you can retry using the same chunkId and startCursor, or adjust startCursor as needed but you must handle any resulting data duplication or gaps yourself. 5) Always explain to users when content is chunked and ask if they want to continue retrieving subsequent parts.

ParametersTypeDescription
startCursornumberStarting cursor position in bytes. Set to 0 for initial requests, and use the value from previous responses for subsequent requests to resume content retrieval.
urlstringURL of the website to fetch
autoDetectModeboolean optionalOptional flag to automatically switch to browser mode if standard fetch fails (default: true). Set to false to strictly use the specified mode without automatic switching.
chunkIdstring optionalOptional chunk ID for retrieving a specific chunk of content from a previous request. The system adds prompts in the format === SYSTEM NOTE === ... =================== which AI models should ignore when processing the content.
closeBrowserboolean optionalOptional flag to close the browser after fetching (default: false)
contentSizeLimitnumber optionalOptional maximum content size in bytes before splitting into chunks (default: 50KB). Set between 20KB-50KB for optimal results. For large content, prefer smaller values (20KB-30KB) to avoid truncation.
debugboolean optionalOptional flag to enable detailed debug logging (default: false)
enableContentSplittingboolean optionalOptional flag to enable content splitting for large responses (default: true)
headersobject optionalOptional headers to include in the request
maxRedirectsnumber optionalOptional maximum number of redirects to follow (default: 10)
noDelayboolean optionalOptional flag to disable random delay between requests (default: false)
proxystring optionalOptional proxy server to use (format: http://host:port⁠ or https://host:port⁠)
saveCookiesboolean optionalOptional flag to save cookies for future requests to the same domain (default: true)
scrollToBottomboolean optionalOptional flag to scroll to bottom of page in browser mode (default: false)
timeoutnumber optionalOptional timeout in milliseconds (default: 30000)
useBrowserboolean optionalOptional flag to use headless browser for fetching (default: false)
useSystemProxyboolean optionalOptional flag to use system proxy environment variables (default: true)
waitForSelectorstring optionalOptional CSS selector to wait for when using browser mode
waitForTimeoutnumber optionalOptional timeout to wait after page load in browser mode (default: 5000)

⁠Tool: fetch_markdown

Fetch a website and return the content as Markdown. Best practices: 1) Always set startCursor=0 for initial requests, and use the fetchedBytes value from previous response for subsequent requests to ensure content continuity. 2) Set contentSizeLimit between 20000-50000 for large pages. 3) When handling large content, use the chunking system by following the startCursor instructions in the system notes rather than increasing contentSizeLimit. 4) If content retrieval fails, you can retry using the same chunkId and startCursor, or adjust startCursor as needed but you must handle any resulting data duplication or gaps yourself. 5) Always explain to users when content is chunked and ask if they want to continue retrieving subsequent parts.

ParametersTypeDescription
startCursornumberStarting cursor position in bytes. Set to 0 for initial requests, and use the value from previous responses for subsequent requests to resume content retrieval.
urlstringURL of the website to fetch
autoDetectModeboolean optionalOptional flag to automatically switch to browser mode if standard fetch fails (default: true). Set to false to strictly use the specified mode without automatic switching.
chunkIdstring optionalOptional chunk ID for retrieving a specific chunk of content from a previous request. The system adds prompts in the format === SYSTEM NOTE === ... =================== which AI models should ignore when processing the content.
closeBrowserboolean optionalOptional flag to close the browser after fetching (default: false)
contentSizeLimitnumber optionalOptional maximum content size in bytes before splitting into chunks (default: 50KB). Set between 20KB-50KB for optimal results. For large content, prefer smaller values (20KB-30KB) to avoid truncation.
debugboolean optionalOptional flag to enable detailed debug logging (default: false)
enableContentSplittingboolean optionalOptional flag to enable content splitting for large responses (default: true)
extractContentboolean optionalOptional flag to enable intelligent content extraction using Readability algorithm (default: false). Extracts main article content from web pages.
fallbackToOriginalboolean optionalOptional flag to fall back to the original content when extraction fails (default: true). Only works when extractContent is true.
headersobject optionalOptional headers to include in the request
includeMetadataboolean optionalOptional flag to include metadata (title, author, etc.) in the extracted content (default: false). Only works when extractContent is true.
maxRedirectsnumber optionalOptional maximum number of redirects to follow (default: 10)
noDelayboolean optionalOptional flag to disable random delay between requests (default: false)
proxystring optionalOptional proxy server to use (format: http://host:port⁠ or https://host:port⁠)
saveCookiesboolean optionalOptional flag to save cookies for future requests to the same domain (default: true)
scrollToBottomboolean optionalOptional flag to scroll to bottom of page in browser mode (default: false)
timeoutnumber optionalOptional timeout in milliseconds (default: 30000)
useBrowserboolean optionalOptional flag to use headless browser for fetching (default: false)
useSystemProxyboolean optionalOptional flag to use system proxy environment variables (default: true)
waitForSelectorstring optionalOptional CSS selector to wait for when using browser mode
waitForTimeoutnumber optionalOptional timeout to wait after page load in browser mode (default: 5000)

⁠Tool: fetch_plaintext

Fetch a website and return the content as plain text with HTML tags removed. Best practices: 1) Always set startCursor=0 for initial requests, and use the fetchedBytes value from previous response for subsequent requests to ensure content continuity. 2) Set contentSizeLimit between 20000-50000 for large pages. 3) When handling large content, use the chunking system by following the startCursor instructions in the system notes rather than increasing contentSizeLimit. 4) If content retrieval fails, you can retry using the same chunkId and startCursor, or adjust startCursor as needed but you must handle any resulting data duplication or gaps yourself. 5) Always explain to users when content is chunked and ask if they want to continue retrieving subsequent parts.

ParametersTypeDescription
startCursornumberStarting cursor position in bytes. Set to 0 for initial requests, and use the value from previous responses for subsequent requests to resume content retrieval.
urlstringURL of the website to fetch
autoDetectModeboolean optionalOptional flag to automatically switch to browser mode if standard fetch fails (default: true). Set to false to strictly use the specified mode without automatic switching.
chunkIdstring optionalOptional chunk ID for retrieving a specific chunk of content from a previous request. The system adds prompts in the format === SYSTEM NOTE === ... =================== which AI models should ignore when processing the content.
closeBrowserboolean optionalOptional flag to close the browser after fetching (default: false)
contentSizeLimitnumber optionalOptional maximum content size in bytes before splitting into chunks (default: 50KB). Set between 20KB-50KB for optimal results. For large content, prefer smaller values (20KB-30KB) to avoid truncation.
debugboolean optionalOptional flag to enable detailed debug logging (default: false)
enableContentSplittingboolean optionalOptional flag to enable content splitting for large responses (default: true)
extractContentboolean optionalOptional flag to enable intelligent content extraction using Readability algorithm (default: false). Extracts main article content from web pages.
fallbackToOriginalboolean optionalOptional flag to fall back to the original content when extraction fails (default: true). Only works when extractContent is true.
headersobject optionalOptional headers to include in the request
includeMetadataboolean optionalOptional flag to include metadata (title, author, etc.) in the extracted content (default: false). Only works when extractContent is true.
maxRedirectsnumber optionalOptional maximum number of redirects to follow (default: 10)
noDelayboolean optionalOptional flag to disable random delay between requests (default: false)
proxystring optionalOptional proxy server to use (format: http://host:port⁠ or https://host:port⁠)
saveCookiesboolean optionalOptional flag to save cookies for future requests to the same domain (default: true)
scrollToBottomboolean optionalOptional flag to scroll to bottom of page in browser mode (default: false)
timeoutnumber optionalOptional timeout in milliseconds (default: 30000)
useBrowserboolean optionalOptional flag to use headless browser for fetching (default: false)
useSystemProxyboolean optionalOptional flag to use system proxy environment variables (default: true)
waitForSelectorstring optionalOptional CSS selector to wait for when using browser mode
waitForTimeoutnumber optionalOptional timeout to wait after page load in browser mode (default: 5000)

⁠Tool: fetch_txt

Fetch a website, return the content as plain text (no HTML). Best practices: 1) Always set startCursor=0 for initial requests, and use the fetchedBytes value from previous response for subsequent requests to ensure content continuity. 2) Set contentSizeLimit between 20000-50000 for large pages. 3) When handling large content, use the chunking system by following the startCursor instructions in the system notes rather than increasing contentSizeLimit. 4) If content retrieval fails, you can retry using the same chunkId and startCursor, or adjust startCursor as needed but you must handle any resulting data duplication or gaps yourself. 5) Always explain to users when content is chunked and ask if they want to continue retrieving subsequent parts.

ParametersTypeDescription
startCursornumberStarting cursor position in bytes. Set to 0 for initial requests, and use the value from previous responses for subsequent requests to resume content retrieval.
urlstringURL of the website to fetch
autoDetectModeboolean optionalOptional flag to automatically switch to browser mode if standard fetch fails (default: true). Set to false to strictly use the specified mode without automatic switching.
chunkIdstring optionalOptional chunk ID for retrieving a specific chunk of content from a previous request. The system adds prompts in the format === SYSTEM NOTE === ... =================== which AI models should ignore when processing the content.
closeBrowserboolean optionalOptional flag to close the browser after fetching (default: false)
contentSizeLimitnumber optionalOptional maximum content size in bytes before splitting into chunks (default: 50KB). Set between 20KB-50KB for optimal results. For large content, prefer smaller values (20KB-30KB) to avoid truncation.
debugboolean optionalOptional flag to enable detailed debug logging (default: false)
enableContentSplittingboolean optionalOptional flag to enable content splitting for large responses (default: true)
extractContentboolean optionalOptional flag to enable intelligent content extraction using Readability algorithm (default: false). Extracts main article content from web pages.
fallbackToOriginalboolean optionalOptional flag to fall back to the original content when extraction fails (default: true). Only works when extractContent is true.
headersobject optionalOptional headers to include in the request
includeMetadataboolean optionalOptional flag to include metadata (title, author, etc.) in the extracted content (default: false). Only works when extractContent is true.
maxRedirectsnumber optionalOptional maximum number of redirects to follow (default: 10)
noDelayboolean optionalOptional flag to disable random delay between requests (default: false)
proxystring optionalOptional proxy server to use (format: http://host:port⁠ or https://host:port⁠)
saveCookiesboolean optionalOptional flag to save cookies for future requests to the same domain (default: true)
scrollToBottomboolean optionalOptional flag to scroll to bottom of page in browser mode (default: false)
timeoutnumber optionalOptional timeout in milliseconds (default: 30000)
useBrowserboolean optionalOptional flag to use headless browser for fetching (default: false)
useSystemProxyboolean optionalOptional flag to use system proxy environment variables (default: true)
waitForSelectorstring optionalOptional CSS selector to wait for when using browser mode
waitForTimeoutnumber optionalOptional timeout to wait after page load in browser mode (default: 5000)

⁠Use this MCP Server

{
  "mcpServers": {
    "lmcc-dev-mult-fetch-mcp-server": {
      "command": "docker",
      "args": [
        "run",
        "-i",
        "--rm",
        "mcpcommunity/lmcc-dev-mult-fetch-mcp-server"
      ]
    }
  }
}

Why is it safer to run MCP Servers with Docker?⁠

Tag summary

Content type

Image

Digest

sha256:51ac31a25…

Size

98.5 MB

Last updated

over 1 year ago

docker pull mcpcommunity/lmcc-dev-mult-fetch-mcp-server