This endpoint performs text extraction given a URL.
In general, text extraction is performed with newspaper.
However, some websites require site-specific extraction in order to achieve accurate result.
A list of such websites is maintained in code/text/site_specific_parsing.csv in CSV format with
3 columns.
The columns are domain name, name of custom extraction module in code/text/site_plugins/, and
language.
{
"title": TITLE_TEXT,
"description": DESCRIPTION_TEXT,
"body": BODY_TEXT
}
Before a control is passed on to a custom site-specific module for text extraction,
a web page is rerquested from a remote server and the return status is 200.
The resulted HTML page as a string is passed on to a function parse in
the module.
Therefore, parse function is mandatory for any site-specific module.
This function returns a triplet of title, description, and body.
Any of those fields can be an empty string.
If text extraction fails, return a triplet of empty strins.
def parse(html):
# parse the html.
return (title, description, body)
The docker image already havs BeautifulSoup4 (bs4), lxml, and requests.
See site_plugins/example.py and site_plugins/telegraaf.py for examples.
Content type
Image
Digest
Size
37.3 MB
Last updated
over 8 years ago
docker pull skumagai/semantics-text-extraction