This script is web scraping the 1337x.to website and indexing data.
Ideal for creating your own application(s) on top of the 1337x meta-data (eg. creating your own search engine).
Currently used for web scraping specifically GNU/Linux games by johncena141. And scraping e-book by zakareya.
The data is stored into a MariaDB database.
We have two frontends:
Run-time dependency:
python3 python3-dev python3-pip)libmysqlclient-dev)libxml2-dev libxslt1-dev zlib1g-dev libffi-dev libssl-dev)More packages will be downloaded via pip, see next section.
I advice you to use a Python virtual environment, create & activate such an environment via:
python3 -m venv env
source env/bin/activate
Next, install the required packages via:
pip install -r requirements.txt
Rename mysqlconfig.sample.py to mysqlconfig.py inside the website1337xscraper folder. And update the MySQL configuration.
Import the new table structure into the database, see 1337x.sql file.
Execute scraper and start filling the database:
scrapy crawl 1337x
Or by running: ./start_spider.py
The goal of the scraper is to populate a (MySQL) database with all the meta-data, so a dedicated search engine can be build on top of this information.
Note: If we want to stored files, we will store them in the website sub-folder.
Optionally, execute scraper and output the meta-data to a "feed" file (eg. JSON file):
scrapy crawl 1337x -O 1337x.json
Our 2nd web scraper is almost the same, but only scraps the first two pages:
scrapy crawl 1337x_homepage
Or by running: ./start_spider2.py
The Docker image is available on DockerHub.
Note: The Docker Image will start the scrawler using a cronjob, so the 1337x full spider runs automatically several times a week. And the 'homepage' spider runs even every 4 hours.
I provided a docker-compose file for convenience (be sure you add 1337xmysqlconfig.py file locally with the MySQL configuration. See example config).
Building Docker image
Create a Docker image locally using:
docker build -t danger89/1337xscraper .
You can use the Scrapy shell to help debugging or learn how to extract data when using scrapy:
scrapy shell 'https://1337x.to/user/johncena141/'
Check the response object for data, just an example:
response.css('.table-list-wrap td.name')[0].get()
More info:
Content type
Image
Digest
sha256:dc3c4143b…
Size
193.5 MB
Last updated
over 3 years ago
docker pull danger89/1337xscraper