Sign inSign up

skumagai/semantics-text-extraction

By skumagai

•Updated over 8 years ago

text extraction

Image
1

405

skumagai/semantics-text-extraction repository overview

⁠Text extraction ('/text')

This endpoint performs text extraction given a URL.

In general, text extraction is performed with newspaper⁠. However, some websites require site-specific extraction in order to achieve accurate result. A list of such websites is maintained in code/text/site_specific_parsing.csv in CSV format with 3 columns. The columns are domain name, name of custom extraction module in code/text/site_plugins/, and language.

⁠Parameters

  • URL (required)
  • lang (optional; 2-letter code, e.g. nl, en)

⁠Return value (in JSON)

{
    "title": TITLE_TEXT,
    "description": DESCRIPTION_TEXT,
    "body": BODY_TEXT
}

⁠Custom site-specific modules

Before a control is passed on to a custom site-specific module for text extraction, a web page is rerquested from a remote server and the return status is 200. The resulted HTML page as a string is passed on to a function parse in the module. Therefore, parse function is mandatory for any site-specific module. This function returns a triplet of title, description, and body. Any of those fields can be an empty string. If text extraction fails, return a triplet of empty strins.

def parse(html):
    # parse the html.
    return (title, description, body)

The docker image already havs BeautifulSoup4 (bs4), lxml, and requests. See site_plugins/example.py and site_plugins/telegraaf.py for examples.

Tag summary

Content type

Image

Digest

Size

37.3 MB

Last updated

over 8 years ago

docker pull skumagai/semantics-text-extraction