Sign inSign up

danger89/1337xscraper

By danger89

•Updated over 3 years ago

1337x web scraper using Scrapy

Image
0

2.0K

danger89/1337xscraper repository overview

⁠1337x Web Scraper

This script is web scraping the 1337x.to website and indexing data.

Ideal for creating your own application(s) on top of the 1337x meta-data (eg. creating your own search engine).

Currently used for web scraping specifically GNU/Linux games by johncena141. And scraping e-book by zakareya.
The data is stored into a MariaDB database.

We have two frontends:

⁠Usage

⁠Dependencies

Run-time dependency:

  • Python3 + pip (python3 python3-dev python3-pip)
  • MySQL database development files (libmysqlclient-dev)
  • Additional libs for Scrapy (libxml2-dev libxslt1-dev zlib1g-dev libffi-dev libssl-dev)

More packages will be downloaded via pip, see next section.

⁠Prepare

I advice you to use a Python virtual environment⁠, create & activate such an environment via:

python3 -m venv env
source env/bin/activate

Next, install the required packages via:

pip install -r requirements.txt

Rename mysqlconfig.sample.py⁠ to mysqlconfig.py inside the website1337xscraper folder. And update the MySQL configuration.

Import the new table structure into the database, see 1337x.sql file⁠.

⁠Run scraper

Execute scraper and start filling the database:

scrapy crawl 1337x

Or by running: ./start_spider.py

The goal of the scraper is to populate a (MySQL) database with all the meta-data, so a dedicated search engine can be build on top of this information.

Note: If we want to stored files, we will store them in the website sub-folder.

Optionally, execute scraper and output the meta-data to a "feed" file (eg. JSON file):

scrapy crawl 1337x -O 1337x.json

⁠Run 2nd scraper

Our 2nd web scraper is almost the same, but only scraps the first two pages:

scrapy crawl 1337x_homepage

Or by running: ./start_spider2.py

⁠Docker Image

The Docker image is available on DockerHub⁠.

Note: The Docker Image will start the scrawler using a cronjob, so the 1337x full spider runs automatically several times a week. And the 'homepage' spider runs even every 4 hours.

I provided a docker-compose file⁠ for convenience (be sure you add 1337xmysqlconfig.py file locally with the MySQL configuration. See example config⁠).

Building Docker image

Create a Docker image locally using:

docker build -t danger89/1337xscraper .

⁠Learn & Debug

You can use the Scrapy shell to help debugging or learn how to extract data when using scrapy:

scrapy shell 'https://1337x.to/user/johncena141/'

Check the response object for data, just an example:

response.css('.table-list-wrap td.name')[0].get()

More info:

Tag summary

Content type

Image

Digest

sha256:dc3c4143b…

Size

193.5 MB

Last updated

over 3 years ago

docker pull danger89/1337xscraper