HTML文書の本文部分を抽出します。
ブログやニュースなどのHTML文書は、メニュー・広告など本文以外の様々な情報が含まれます。これら余計な情報を除去し、本文部分のHTMLだけを抽出します。本文抽出処理はWebAPIとして提供するため、どの処理系でも使用することができます。
$ docker version
Client:
Version: 17.03.1-ce
API version: 1.27
Go version: go1.7.5
Git commit: c6d412e
Built: Tue Mar 28 00:40:02 2017
OS/Arch: windows/amd64
Server:
Version: 17.04.0-ce
API version: 1.28 (minimum version 1.12)
Go version: go1.7.5
Git commit: 4845c56
Built: Wed Apr 5 18:45:47 2017
OS/Arch: linux/amd64
Experimental: false
urlパラメータに、本文部分を抽出したいURLを指定します。
$ curl -v "http://localhost:5000/extract?url=https%3A%2F%2Ftechcrunch.com%2F2017%2F09%2F02%2Fthe-product-design-challenges-of-ar-on-smartphones%2F"
JSONが返ります。
< HTTP/1.0 200 OK
< Content-Type: application/json
< Content-Length: 280396
< Server: Werkzeug/0.12.2 Python/3.6.2
< Date: Tue, 05 Sep 2017 07:19:56 GMT
<
{
"content": "<html><body><div><div class=\"article-entry text l-featured-container\">...(中略)...</body></html>",
"full-content": "<!DOCTYPE html>\n<html xmlns=\"http://www.w3.org/1999/xhtml\" xmlns:og=\"http://opengraphprotocol.org/schema/\" xmlns:fb=\"http://www.facebook.com/2008/fbml\" lang=\"en\">...(中略)...</body>\n</html>\n<!--\n\tgenerated 268 seconds ago\n\tgenerated in 0.216 seconds\n\tserved from batcache in 0.002 seconds\n\texpires in 32 seconds\n-->\n",
"simplified-content": "<h1>extract-content</1>\n<p>HTML文書の本文部分を抽出します。...(中略)...<a href=\"https://github.com/u6k/extract-content/blob/master/LICENSE\">MIT License</a></p>",
"summary-list": [
"With the launch of ARKit, we are going to see augmented reality apps become available for about 500 million iPhones in the next 12 months, and at least triple that in the following 12 months \u2014 as we can now include the numbers of ARCore-supporting devices from Google.",
"Matt Miesnieks is a partner at [Super Ventures](http://www.superventures.com/).",
"How would I even think to get my phone out right now? This is a huge design problem for smartphone AR (and beacons before\u00a0this). (Image: [Estimote](https://estimote.com/))"
],
"title": "The product design challenges of AR on smartphones",
"url": "https://techcrunch.com/2017/09/02/the-product-design-challenges-of-ar-on-smartphones/"
}
実行用Dockerコンテナを起動します。
$ docker run \
-d \
--name extract-content \
-p 5000:5000 \
u6kapps/extract-content
開発用Dockerイメージをビルドします。
$ docker build -t extract-content-dev -f Dockerfile-dev .
開発用Dockerコンテナを起動します。
$ docker run \
--rm \
-it \
--name extract-content-dev \
-p 5000:5000 \
-v ${PWD}:/opt/extract-content \
extract-content-dev bash
アプリケーションを起動します。
$ python run.py
テストを実行するには、開発用Dockerコンテナを起動して、 py.test を実行します。
$ py.test --capture=no --junit-xml=build/results.xml
build/results.xml にXUnit形式のテスト結果が出力されます。
Swaggerドキュメントを参照するには、swagger-uiコンテナを起動します。
$ docker run \
--rm \
-p 8080:8080 \
-e "SWAGGER_JSON=/opt/extract-content/docs/swagger.yaml" \
-v ${PWD}:/opt/extract-content \
swaggerapi/swagger-ui
TODO: 2017/9/4時点のswagger-codegenは、openapi 3.0.0のSwaggerドキュメントを読み込めないもよう。対応されたら、ビルド・プロセスで静的ファイルを出力するようにします。
実行用Dockerイメージをビルドします。
$ docker build -t u6kapps/extract-content .
Content type
Image
Digest
Size
487.1 MB
Last updated
over 8 years ago
docker pull u6kapps/extract-content