Paper-dataset schema evidence: four categories, modular models, Mercury/Apptainer deployment.
43
English | 简体中文
Paper–dataset schema evidence pipeline: structure, encoding, value, syntax.
The paper workflow retains the full layout-aware text and freezes a shared classification index and task specification for three interchangeable local models and one frontier soft reference. The dataset workflow independently produces evidence through format parsers. Both paths meet in an evaluation packet, ending at an integrity and provenance gate. A soft reference is fallible; passing this gate does not establish semantic accuracy.
| Task | Documentation |
|---|---|
| Run the demo, prepare paper/dataset inputs, build a corpus, run jobs and verify packets | User guide |
| Deploy on Mercury, obtain a Docker/Apptainer image, select models and freeze a configuration | Deployment guide |
| Read the documentation in Chinese | 中文首页 · 使用指南 · 部署指南 |
plalelab/schema-study:0.1.0-cuda13 (Linux amd64)This Linux/Bash example needs no GPU, model weights or API key, and makes no model calls.
docker pull plalelab/schema-study:0.1.0-cuda13
mkdir -p "$PWD/schema-study-results"
docker run --rm \
--mount type=bind,src="$PWD/schema-study-results",dst=/outputs \
plalelab/schema-study:0.1.0-cuda13 \
offline-demo --output /outputs/demo-v1
Read schema-study-results/demo-v1/report.json after completion. The demo uses synthetic inputs and mock backends to exercise repeated local-model slots, a separate classifier role, a soft reference, resume behavior and packet verification. Use a new directory for another complete demo; resume an actual batch in its existing batch directory.
high_fidelity_schema_study/
four_category/ # Tasks, model adapters, datasets, batches, packets, Mercury
extractors/ # Format parser plugins
config/ # Shared specification and draft model configurations
templates/ # Exact prompts, taxonomy and output JSON schemas
container/
docker/ # OCI image recipe
apptainer/ # SIF recipe, build script and bind-mount launcher
docs/ # English guides and optional Chinese translations
tests/ # Synthetic and mock software tests
source-export-manifest.json # File-byte SHA-256 inventory for the current export
This repository distributes the runnable four-category workflow. Users manage original papers, datasets, model weights, API credentials and historical results in external directories. The export manifest covers only its listed files; CI and release metadata are not implicitly included. The original v0.1.0 inventory remains available in its tag and release assets.
git clone https://github.com/williamQ96/schema_study.git
cd schema_study
python3.12 -m venv .venv
source .venv/bin/activate
python -m pip install -r requirements_four_category_offline.txt
python -m pytest -q tests
python -m high_fidelity_schema_study.four_category.cli --help
The offline dependencies support parsing, mock demos and software tests. Real Transformers inference also requires the additional runtime libraries in the Linux CUDA image.
run verifies the model configuration, source code, SIF, checkpoint and hardware before dispatch. Sources and results retain their identities, and failures are isolated per task.Content type
Image
Digest
sha256:c6abbb7f8…
Size
7 GB
Last updated
15 days ago
docker pull plalelab/schema-study:0.1.0-cuda13