A tool for analyzing the documentation of scientific datasets
231
Extract, in a structured manner, the general guidelines from the ML community about dataset documentation practices from its scientific documentation. Study and analyze scientific data published in peer-review journals such as: Nature's Scientific Data and Data-in-Brief.
Here you have a complete list of data journals suitable to be analyzed with this tool. Test the web UI of the tool in the following HuggingFace Space, and the API using our Docker image
The tools come with two UIs and a docker image. A web app built with Gradio intended to test the tool's capabilities and analyze a single document (you can try it in the HuggingFace Space). And a API built with FastAPI, suited to be integrated into any ML pipeline. The docker image is suited to deploy only the API.
First you need to install docker in your sistem. Then:
docker pull joangi/datadoc_analyzer
docker run --name apidataset -p 80:80 joangi/datadoc_analyzer
The API will be running in your localhost at port 80. (You can change the port in the command above)
The API imitates the behavior of the tabs of the web UI, but, in addition, you also have an endpoint to retrieve all the dimensions at the same time. The API's swagger documentation, which can be tested in situ, is published together along the API. The server will start at port 8000 by default (if not occupied by another app of your system). And the documentation will be found at http://127.0.0.1:8000/docs
Content type
Image
Digest
sha256:fab3616c9…
Size
893.6 MB
Last updated
over 3 years ago
docker pull joangi/datadoc_analyzer