Sign inSign up

binux/pyspider

By binux

•Updated over 8 years ago

Image
39

50K+

binux/pyspider repository overview

⁠pyspider Build Status Coverage Status Try

A Powerful Spider(Web Crawler) System in Python. TRY IT NOW!⁠

Tutorial: http://docs.pyspider.org/en/latest/tutorial/⁠
Documentation: http://docs.pyspider.org/⁠
Release notes: https://github.com/binux/pyspider/releases⁠

⁠Sample Code

from pyspider.libs.base_handler import *


class Handler(BaseHandler):
    crawl_config = {
    }

    @every(minutes=24 * 60)
    def on_start(self):
        self.crawl('http://scrapy.org/', callback=self.index_page)

    @config(age=10 * 24 * 60 * 60)
    def index_page(self, response):
        for each in response.doc('a[href^="http"]').items():
            self.crawl(each.attr.href, callback=self.detail_page)

    def detail_page(self, response):
        return {
            "url": response.url,
            "title": response.doc('title').text(),
        }

Demo

⁠Installation

WARNING: WebUI is open to the public by default, it can be used to execute any command which may harm your system. Please use it in an internal network or enable need-auth for webui⁠.

Quickstart: http://docs.pyspider.org/en/latest/Quickstart/⁠

⁠Contribute

⁠TODO

⁠v0.4.0

⁠License

Licensed under the Apache License, Version 2.0

Tag summary

Content type

Image

Digest

sha256:dde6b4431…

Size

334.5 MB

Last updated

over 8 years ago

docker pull binux/pyspider