Sign inSign up

deffyc/pyspider

By deffyc

•Updated about 9 years ago

daocloud.io测试

Image
0

380

deffyc/pyspider repository overview

⁠pyspider Build Status Coverage Status Try

A Powerful Spider(Web Crawler) System in Python. TRY IT NOW!⁠

Tutorial: http://docs.pyspider.org/en/latest/tutorial/⁠
Documentation: http://docs.pyspider.org/⁠
Release notes: https://github.com/binux/pyspider/releases⁠

⁠Sample Code

from pyspider.libs.base_handler import *


class Handler(BaseHandler):
    crawl_config = {
    }

    @every(minutes=24 * 60)
    def on_start(self):
        self.crawl('http://scrapy.org/', callback=self.index_page)

    @config(age=10 * 24 * 60 * 60)
    def index_page(self, response):
        for each in response.doc('a[href^="http"]').items():
            self.crawl(each.attr.href, callback=self.detail_page)

    def detail_page(self, response):
        return {
            "url": response.url,
            "title": response.doc('title').text(),
        }

Demo

⁠Installation

WARNING: WebUI is open to the public by default, it can be used to execute any command which may harm your system. Please use it in an internal network or enable need-auth for webui⁠.

Quickstart: http://docs.pyspider.org/en/latest/Quickstart/⁠

⁠Contribute

⁠TODO

⁠v0.4.0

⁠License

Licensed under the Apache License, Version 2.0

Tag summary

Content type

Image

Digest

Size

335.7 MB

Last updated

about 9 years ago

docker pull deffyc/pyspider