eneric extraction recipes to extract schema.org entities for your software and data
354

This is a repository with example extractors and recipes intended to be used with schemaorg Python to help you to extract metadata from your datasets, software and other entities described in schema.org.
The following specifications have Dockerfiles (and associated Github actions) for you to use! See the subdirectories to get usage:
For both of the above, when you deploy to Github pages for the first time, you
need to switch Github Pages to deploy from master and then back to the gh-pages
branch on deploy. There is a known issue with Permissions if you deploy
to the brain without activating it (as an admin) from the respository first.
The following examples for entities (children of "Thing") defined in schema.org are also provided. These specifications don't yet have Docker containers or Github Action extractors.
For each of the above, the metadata shown is also embedded in the page as json-ld (when you "View Source.")
Each folder above includes an example python script to extract metadata (extract.py),
a recipe to follow (recipe.yml), and the specification in yaml format (in the
case of a specification not served by production schema.org).
For the Docker and Github Actions usage, see inside the ImageDefinition folder. For all other schema.org entities and local usage, details are provided here. Before running these examples, make sure you have installed the module (and note this module is under development, contributions are welcome!)
pip install schemaorg
To extract a recipe for a particular datatype, you can modify extract.py and the
recipe.yml for your particular needs, or use as is. Generally we:
The goal of the software is to provide enough structure to help the user (typically a developer) but not so much as to be annoying to use generally.
If I am a provider of a service and want my users to label their data for my service,
I need to tell them how to do this. I do this by way of a recipe file, in each
example folder there is a file called recipe.yml that is a simple listing of required fields defined for the entities that are needed. For example, the recipe.yml in the
"SoftwareSourceCode" folder tells the parser that we need to define
properties for "SoftwareSourceCode" and an Organization or Person. For example.
with the schemaorg Python module
I can learn that the "SoftwareSourceCode" definition has 121 properties,
but the recipe tells us that we only need a subset of those
properties for a valid extraction.
This is the code snippet that shows how you extract metadata and use the schemaorg Python module to generate the final template page. This file could be run in multiple places!
For the folders with associated containers, you will find a Dockerfile (and associated entrypoint.sh)! These containers will build the extractor into an image that can be used with Github Actions.
Content type
Image
Digest
Size
504.5 MB
Last updated
over 7 years ago
docker pull openschemas/extractors:ImageDefinition