Table of Contents
Abstract
Theoretical Framework
Methodology
Deployment
Technical Implementation
Installation
Usage
License
Abstract
soce_scraper is a web scraping tool designed to extract data from the Sistema Oficial de Contratación Pública del Ecuador (SOCE). It allows users to gather information on various procurement processes, including "Ínfima Cuantía", "Procesos", "Regimen Especial", and "Procedimiento especial". The extracted data is then exported to Excel files for further analysis.
Theoretical Framework
Web Scraping : The process of extracting data from websites using automated scripts.
Data Processing : Utilizing libraries like Pandas to manipulate and analyze the scraped data.
Excel Exporting : Using openpyxl to create Excel files with hyperlinks for easy access to detailed information.
Methodology
The soce_scraper follows a structured approach:
User Input : Users provide a date range and select the type of procurement process they wish to extract data for.
Data Extraction : The scraper sends requests to the SOCE website and retrieves the relevant data.
Data Processing : The scraped data is processed and formatted into a structured format.
Excel Export : The processed data is exported to Excel files, with appropriate formatting and hyperlinks.
Deployment
soce_scraper is built using Flask, allowing it to run as a web application. The application processes user requests and returns the extracted data in Excel format.
Technical Implementation
soce_scraper is developed in Python 3.9, utilizing:
Flask for the web interface.
Requests for making HTTP requests to the SOCE website.
Pandas for data manipulation and analysis.
Openpyxl for exporting data to Excel files.
Gemini for asking the llm model the data in the start resolution pdf
Installation
Prerequisites
Python 3.9 : Ensure that Python is installed and configured in your system.
or
Docker : Ensure to have Docker installed and running in the system.
Installation
Via Github
Clone the repository
git clone https://github.com/nava2105/soce_scraper.git
cd soce_scraper
Copy
Required Libraries : Install the necessary libraries using pip:
pip install -r requirements.txt
Copy
Run the application
Access the application: Open your web browser and navigate to http://localhost:5000
Via Dockerhub
Clone the docker image
docker pull na4va4/soce_scraper
Copy
Run the image in a container
docker run -p 5000:5000 na4va4/soce_scraper
Copy
Access the application: Open your web browser and navigate to http://localhost:5000
Usage
Once the application is up and running, you can interact with the system by:
Selecting the start and end dates for the data extraction.
Choosing the type of procurement process.
Submitting the form to generate and download the corresponding Excel file.
Submit an Excel file in witch one column must have the links to the procedures similar to the result of the procedure extraction.
Submitting the form to generate and download the corresponding Excel file.
Submit an Excel file in witch one column must have the links to the procedures similar to the result of the procedure extraction.
Submitting the form to generate and download the corresponding Excel file.
Submit a list of pdf´s files of the start resolutions of the processes.
Ask for the technical commission ang GEMINI AI shall search for the answer.
License
This project is licensed under the MIT License - see the LICENSE file for details.