Sign inSign up

skumagai/semantics-endpoints

By skumagai

•Updated over 9 years ago

Some Dutch NLP services

Image
0

1.3K

skumagai/semantics-endpoints repository overview

⁠Text extraction, NER, and sentiment analysis, trend matching in Python

⁠Prerequisite

  • docker
  • docker-compose
  • git (If clone this repository. Simply copying all contents also works.)

⁠Building services

To build the software locally:

docker-compose build

Unit tests are run during the build process. A new Docker image is not built if a test is failed.

A simple integration test is available in code/test. To perform the tests, run like so:

code/test/integration-runner

This brings all Docker containers up, perform the integration tests, and shutdown the containers.

⁠Installation

Once changes are pushed to the bitbucket repository, Docker images are automatically bulit at Docker hub⁠. This buliding process may take several hours. Once all images are build over Docker hub, those images can be pulled to deploy on any machine (see the next section).

⁠Deployment

After downloading docker-compose.yml from semantics-endpoints⁠, run:

docker-compose pull
docker-compose -d up

To stop the service run docker-compose down.

⁠Endpoints

GET request at an endpoint will return the usage of the endpoint. POST request at the same endpoint with mandatory and optional parameters will return actual results. If any endpoint takes a text string as a parameter or returns a text string as a part of a return value, the string should be in unicode/utf-8.

⁠Integrated service ('/')

This endpoint performs text extraction, named entity/keywords extraction, and sentiment analysis in a single step.

⁠Text extraction ('/text')

This endpoint performs text extraction given a URL.

In general, text extraction is performed with newspaper⁠. However, some websites require site-specific extraction in order to achieve accurate result. A list of such websites is maintained in code/text/site_specific_parsing.csv in CSV format with 3 columns. The columns are domain name, name of custom extraction module in code/text/site_plugins/, and language.

⁠Parameters
  • URL (required)
  • lang (optional; 2-letter code, e.g. nl, en)
⁠Return value (in JSON)
{
    "title": TITLE_TEXT,
    "description": DESCRIPTION_TEXT,
    "body": BODY_TEXT
}
⁠Custom module

Each custom module has to have an entry-point, a function named download_and_parse with one argument URL. Then, the function has to return a tuple of 3 elements: title, description, and body in this order. Some of these data may be an empty string.

def download_and_parse(url):
    # download html.
    # parse the html.
    return (title, description, body)

The docker image already include BeautifulSoup4 (bs4), lxml, and requests. See code/text/site_plugins/telegraaf.py for an example.

⁠Named entity/keyword recognition ('/ner')

This endpoint performs named entity/keyword recognition given text.

⁠Parameters
  • text (required; a string of text)
  • lang (required; 2-letter code)
  • extract (required; what to extract; one of ner, noun, proper noun; case-sensitive)
⁠Return value (JSON)
{
    "named_entities": [
        {
            "word": WORD,
            "lemma": LEMMA,
            "type": NER_TYPE (B_PER, B_LOC, O, etc),
            "pos": POS_TAG
            "count": INTEGER
        },...
    ]
}
⁠Sentiment analysis ('/sentiment')

This endpoint performs sentiment analysis given text.

⁠Parameters
  • text (required; a string of text)
  • lang (required; 2-letter code)
⁠Return value (JSON)
{
    "sentiment": REAL,
    "subjectivity": REAL
}
⁠Trend matching ('/trend_match')

This endpoint compares named entities/keywords against Twitter trends, and this is not a part of the integrated endpoint (/).

⁠Parameters
  • data (required: a JSON-serialized list of lists)
⁠Return value (JSON)
{
    "scores": {
        "KEYWORD1": [SCORE1, SCORE2, SCORE3],
        "KEYWORD2": [SCORE1, SCORE2, SCORE3],
        ...,
        "region_match": [SCORE1, SCORE2, SCORE3]
    }
}

disabled Need Google account to use.

Returns trending searches in a country in XML format (for now). Google doesn't provide regional- or city-level trending searches, neither is setting speficific languages. This module accesses an unofficial API, so it may break at any time. It is esitmated to have a rate limit of 200 calls / hour.

in the last step.

To enable this endpoint, two files need to be modified. Those files are:

  • code/web/nginx.conf
  • code/docker-compose.yml

The locations of modifications are clearly marked in both files. You can supply an account information when launching the entire service like so:

docker-compose build
GOOGLE_USER=USER_NAME GOOGLE_PASSWORD=PASSWORD docker-compose up -d

⁠Limitations and TODO

  • Only support Dutch for now. It should be straightforward to extend support for other languages except '/ner' endpoint.
  • I don't know how to deploy all of these images across multiple machines. Needs investigation.
  • Each endpooint is in a separate image. Some endpoints may be consolidated.

⁠Comments on the setup

  • /ner and /text use python 3, while /sentiment and /trend_match use python 2.
    • A library (pattern⁠ used in /sentiment only supports python 2.
    • A library (frog⁠) used in /ner supports both python 2 and python 3. However, the library is difficult to build. The latest frog compiled for python 3 is available in Arch Linux, which is used here.
    • /trend_match supports both versions. The reason this is currently deployed with python 2 is the base image (alpine:3.4) favors this version.
    • /text should work with python 2, but one of install libraries (newspaper⁠) is moving toward python 3 only in recent versions. There is no reason to use python 2 with an older version.

Tag summary

Content type

Image

Digest

Size

16 MB

Last updated

over 9 years ago

docker pull skumagai/semantics-endpoints:root-latest