Software Requirements Specification - Analysis and feedback tooling
177
Tools for turning meeting notes and half-written specifications into a requirements document you can review — and for checking documents that already exist.
Ten commands in one image. No model and no GPU inside it. Only two of them talk to a model server: the one that reads your source text, and the one-command pipeline that calls it. Everything else is plain text processing and needs nothing but this image, so a single GPU host can serve everybody.
This page has three parts: using it, setting it up, and background. If somebody has already pointed you at a model server, the first part is all you need.
Two commands, from a raw Teams caption file to a reviewable draft:
docker run --rm nicholasyue/lacuna:0.4.0 lacuna-wrapper > lacuna
chmod +x lacuna
AGENTIC_BASE_URL=http://your-gpu-host:11434/v1 \
./lacuna pipeline meeting.vtt SPEC-100 "Whatever It Is"
The first pair writes a small script that supplies the mount, the user
mapping and the network flags, so you type filenames instead of eighty
characters of docker run. It is the same script for every command in the
image — ./lacuna agentic-conform spec.md works too.
The second turns captions into a transcript, extracts what can be quoted, lays it into the canonical sections, and checks the result.
One directory, named for the spec, holding everything that run produced:
SPEC-100-meeting/
README.md what each file is and who needs it
draft.md the draft specification — this is the document
draft.evidence.md every claim beside the words it came from
conformance.txt whether the draft is in the canonical format
draft.evidence.json the same evidence, for tooling
draft.provenance.json which source lines each requirement traces to
meeting.md your captions as a transcript, when the input was a .vtt
Read draft.md. The README.md beside it is written for that
particular run — it carries the counts, names each file and says who needs
it, so the whole directory can be forwarded to somebody who has never
heard of this tool.
Every claim is quoted verbatim from your source. A claim the model cannot ground in the actual text is dropped rather than written into the document, and the drop is recorded so "the tool lost it" and "the source never said it" stay distinguishable.
It never invents an acceptance criterion, a number, or a section. A
section nobody spoke about is marked TODO:, not asserted to be empty —
absence of evidence is not evidence of absence. Where a requirement has no
test, it emits a stub saying so rather than a plausible-looking scenario.
A clean conformance report is not a good document. conformance.txt
checks shape: sections present, identifiers well formed. It says nothing
about whether the content is right, and a draft can pass every structural
check while a requirement sits in the wrong section. The draft needs
reading by somebody who was in the room.
It does not replace an engineer. No quality score is produced, deliberately. A person remains the acceptance gate; these tools narrow what that person has to read.
The image ships six sample documents at /opt/lacuna/samples/, so you can
see real output before deciding whether this is useful to you. None of
these needs a model server.
# The same document after a hurried edit. Seventeen findings, deliberately.
docker run --rm nicholasyue/lacuna:0.4.0 \
agentic-conform /opt/lacuna/samples/broken-spec.md
# A specification in the canonical format. Reports nothing - it is clean.
docker run --rm nicholasyue/lacuna:0.4.0 \
agentic-conform /opt/lacuna/samples/conforming-spec.md
# What a review of a document looks like: which decisions are still open.
docker run --rm nicholasyue/lacuna:0.4.0 \
srs-eval /opt/lacuna/samples/conforming-spec.md
The first is the one worth running first. It reports 7 blockers, 8 concerns and 2 notes on a document that looks fine at a glance, and names the line and the reason for each.
conforming-spec.md | a specification in the canonical format |
broken-spec.md | the same document after a hurried edit — 17 findings |
meeting-transcript.md | a conversation; input for pipeline (needs a model) |
written-spec.md | a specification in somebody else's template; the other pipeline input |
evidence-table.json | what extraction produces; feeds render and merge with no model |
meeting-captions.vtt | the same meeting as Teams exported it |
To work on your own files, mount a directory at /work:
docker run --rm -v "$PWD:/work" nicholasyue/lacuna:0.4.0 \
agentic-conform /work/your-spec.md
Several of the words here are ordinary English with a narrower meaning. The image carries the full glossary:
docker run --rm nicholasyue/lacuna:0.4.0 glossary
The six worth knowing before you read a report:
| claim | one statement drawn from your source, with the exact words it came from |
| grounded | the quote was found in your document verbatim, so the claim survives |
| kind | what sort of statement it is — requirement, problem, goal — which decides its section |
| conformance | whether the draft is in the canonical format. Shape, not content |
| blocker / concern / note | the three severities. A report with only notes is a clean report |
TODO: | nobody said anything on that subject. A question for a person, not a fault |
| tool | what it does | needs a model server |
|---|---|---|
agentic-conform | Checks a document against the canonical specification format. Reports blockers, concerns and notes with line numbers. | no |
srs-eval | Reviews a requirements document against the SRS standard — what is unstated, unmeasurable or still undecided. Writes HTML too. | no (see note) |
agentic-render | Turns an evidence table into a specification document plus a provenance sidecar. Pure text; invents nothing. | no |
agentic-merge | Combines evidence tables drawn from several sources into one, keeping every claim attributed. | no |
agentic-score | Measures an extraction against a hand-written reference. For tuning and regression testing. | no |
vtt-to-transcript | Turns a Teams/WebVTT caption file into a transcript the rest of the pipeline can quote. pipeline calls it for you. | no |
glossary | Prints the glossary. | no |
lacuna-wrapper | Prints the host-side wrapper script that supplies the mount, user and network flags. | no |
agentic-extract | The one that needs a model. Reads source text and pulls out claims that can be quoted from it, producing an evidence table. | yes |
pipeline | The whole thing in one command: convert captions if needed, extract, render, conform. | yes, because it calls agentic-extract |
Note on srs-eval. It has an optional model-assisted judgment, but no
inference backend is wired into the command-line tool in this image, so it
always runs its deterministic checks only and reports the judgment as
Not run. It needs no model server, and --no-judge only silences the
notice. The judgment half is available to code that supplies its own
backend, not through this CLI.
So eight of the ten work fully offline, with no server, no network and no configuration. That is deliberate: reading and checking documents is text processing, and only extraction genuinely needs a model.
agentic-extract and pipeline talk to any OpenAI-compatible HTTP
endpoint — Ollama and
vLLM both work. The image contains no model
weights and needs no GPU itself; point it at a machine that has one.
localhost inside a container is the container, not the machine
running it. This is the one thing that costs an afternoon, because the
failure is a connection error against a URL that is obviously correct from
a shell on the same host.
Ollama binds 127.0.0.1 by default, so it is unreachable from any
other network namespace at any address. From inside a container,
localhost, host.containers.internal and the host's own LAN address all
fail for that single reason.
Model server on the same machine — share its network:
docker run --rm --network=host -v "$PWD:/work" \
-e AGENTIC_BASE_URL=http://localhost:11434/v1 \
nicholasyue/lacuna:0.4.0 \
pipeline /opt/lacuna/samples/meeting-transcript.md SPEC-014 "Playlist Handoff"
Model server on another machine — the normal deployment. That host
needs Ollama listening on something other than loopback
(OLLAMA_HOST=0.0.0.0 ollama serve), which exposes it to the network it is
on and is a decision worth making deliberately:
docker run --rm -v "$PWD:/work" \
-e AGENTIC_BASE_URL=http://gpu-host:11434/v1 \
nicholasyue/lacuna:0.4.0 pipeline /work/notes.md SPEC-100 "A Feature"
Reach Ollama by a name it recognises, not by a bare IP address. It
checks the Host header and answers 403 Forbidden otherwise. That 403 is
useful on sight: it means the networking is right and the URL is wrong,
which is the opposite diagnosis from a connection error.
14B is the floor, and a GPU with at least 12 GB. This is a measured conclusion, not a preference:
qwen2.5:14b and above — what the tool is tuned for. Recommended.Nothing enforces this. --model accepts anything, and a small model still
runs; agentic-extract prints a warning before starting, because a thin
result from an undersized model looks exactly like a thin result from a
quiet source document.
| Model | qwen2.5:14b (Q4_K_M, about 15 GB resident) |
| Server | Ollama, default settings |
| GPU | NVIDIA RTX A5000, 24 GB, running 100% GPU with no offload |
That is the reference configuration. If you match it, the numbers on this page are what you should see.
Throughput scales with the document. The work is chunked and the
default grouped strategy makes four model calls per chunk:
| on the reference configuration | ||
|---|---|---|
| a two-page document | 2 chunks | 34–40 seconds |
| a 45-minute Teams meeting, 644 cues, 36 KB | 69 chunks | 6 min 40 s |
Roughly six seconds a chunk: divide your transcript by the 600-character budget and you have an estimate before you start.
Nothing prints between the start of extraction and its summary, so a long run looks like a hang and is not.
| variable | default | |
|---|---|---|
AGENTIC_BASE_URL | http://localhost:11434/v1 | the model server |
AGENTIC_MODEL | qwen2.5:14b | the model to ask |
STRATEGY | grouped | grouped or single (pipeline only) |
BUDGET | 600 | characters of source per model call (pipeline only) |
OUT | SPEC-nnn-<input> | the output directory (pipeline only) |
The wrapper forwards all five, so STRATEGY=single ./lacuna pipeline ...
behaves the same way inside the container as outside it. Setting one on a
bare docker run needs an explicit -e STRATEGY=single, which is most of
why the wrapper exists.
STRATEGY=single for a document that is already a specification. It
files claims into the right sections considerably more accurately there;
grouped reaches more sections on a conversation and is the default for
that reason.
An explicit --base-url or --model beats either variable. Localhost is
left as the default even though it is almost certainly wrong inside a
container, so that a missing configuration fails by not connecting —
which is diagnosable — rather than by quietly reaching something
unintended.
Under rootless Podman, add --userns=keep-id --user "$(id -u):$(id -g)",
or output lands owned by a subordinate uid and the document you just
generated cannot be edited by you. Docker needs only the --user half.
The wrapper does this for you.
linux/amd64.openai, pinned exactly and used
purely as an HTTP client against the server you configure. No account, no
API key, no hosted service.AGENTIC_BASE_URL, and only agentic-extract makes it. The other tools
make no network calls at all.| tag | |
|---|---|
0.4.0 | current release, linux/amd64. One output directory per run, with a README in it |
0.3.0 | previous release. Output was loose files in the working directory |
0.2.0 | no caption support: a .vtt handed to agentic-extract loses about 42% of its quotes, silently |
Version tags are immutable in intent: a change means a new version rather
than a rebuild of an old one. latest is deliberately not published, so
you always pull a version you chose.
A caption file cannot be quoted, and the failure is silent. Teams cuts a cue every few seconds, mid-clause:
00:00:02.140 --> 00:00:05.980
<v Sam>We put together the playlist in editorial, then we</v>
00:00:05.980 --> 00:00:09.420
<v Sam>email a list of shot names over to the review room.</v>
That is one sentence. agentic-extract requires every claim to quote its
source verbatim — the model returns a quote and the tool finds it, which
is what stops a paraphrase arriving downstream looking like evidence. A blank
line, a GUID and a timestamp sit between the two halves, so the sentence is
not in the file, the claim is dropped, and you get a thin evidence table.
That is indistinguishable from a quiet meeting, which is why this matters more than it sounds. Measured on a real 644-cue Teams export:
| input | model calls | grounded claims | quotes rejected |
|---|---|---|---|
the .vtt as exported | 36 | 38 / 38 / 33 | 42% |
after vtt-to-transcript | 13 | 34 / 38 / 35 | 7% |
Three repeats each. Same number of usable claims for 2.8x the model calls, because timestamps and GUIDs make the same meeting 95 KB instead of 36 KB and the model reads all of it.
pipeline does this for you — hand it the .vtt and it converts first,
as step 0, then carries on. The transcript is kept in the output directory,
because it is the file the evidence table records a checksum of and the file
requirement line numbers point into. Converting again later with different
options moves every reference.
To convert on its own — worth doing if you want to read or edit the transcript before drafting from it:
./lacuna vtt-to-transcript meeting.vtt -o meeting.md
--title adds a heading, --max-block changes how much speech goes in one
block, --timestamps keeps the times (off by default, because anything on the
page is something the model can quote, and a requirement traced to [00:12:34]
helps nobody).
A Word transcript needs no conversion — Teams already groups it by speaker. Save it as plain text and use that.
Content type
Image
Digest
sha256:cf62b3ade…
Size
54.9 MB
Last updated
27 days ago
docker pull nicholasyue/lacuna:0.4.0