Classifies texts in NOW-corpus format into 39 top-level Wikipedia domains (Sports, Politics, Science, Food & drink, …) using a fine-tuned German DistilBERT model, and writes the result as a stand-off metadata XML document — one ranked, scored list of domains per input text.
The output lives in the TEI namespace and reuses TEI's classification
vocabulary (taxonomy / category / catDesc / catRef), but is not
conformant TEI P5: the standOff, metadataLayer, source and textRef
containers are local extensions.
One text per record. A record starts with a line beginning @@, immediately
followed by the document id (the first whitespace-separated token) and the
text. Any following lines, up to the next @@, are appended to the same text.
Literal <p> paragraph markers are stripped.
@@EXAMPLE.00001 Die Heimmannschaft gewann das Fußballspiel am Samstagabend ...
Der Trainer lobte nach dem Spiel die Abwehr ... <p> Die Fans feierten den Sieg
@@EXAMPLE.00002 Das Parlament debattierte am Mittwoch über den Haushaltsentwurf ...
A <standOff> document containing a single <metadataLayer> with:
<source> — provenance: tool, tool version, model, run date, selection
(top<k>), threshold and scoreType.<taxonomy> — the full list of domains (<category xml:id="…">).<textRef target="DOCID"> per input text, holding up to --topk
<catRef> children. Each catRef points at a domain (target="#Domain"),
carries its 1-based rank (n) and the classifier confidence in
[0, 1] (cert).<textRef target="EXAMPLE.00001">
<catRef scheme="#wikitaxonomy" target="#Sports" n="1" cert="0.9527"/>
<catRef scheme="#wikitaxonomy" target="#History" n="2" cert="0.1688"/>
</textRef>
cat sample.now | docker run --rm -i korap/wiki-taxonomy > out.xml
Content type
Image
Digest
sha256:949f4e562…
Size
4.7 GB
Last updated
2 months ago
docker pull korap/wiki-taxonomy