Repository hand over readiness reporter
207
Existing tools tell you if a repository is ready to publish. Babysitter tells you if it is ready to be someone else's problem — and prices that in engineering hours, not points.
Babysitter is a containerised acceptance gate that answers one question:
Could a competent engineer who did not write this code take responsibility for it — right now, from a clean checkout?
It runs like pytest or a linter: one command, a verdict, and an exit code your pipeline can gate on.
podman run --rm -v "$PWD:/repo:ro" nicholasyue/babysitter scan /repo
Runs rootless under Podman, with no daemon and no privileged mode. Docker works
identically — substitute docker for podman in any command here.
| Tag | Meaning |
|---|---|
0.1.1 | The current release. Pin an exact version like this in CI. |
0.1 | Latest patch of the 0.1 line — currently 0.1.1. |
latest | Latest release — currently 0.1.1. |
Every published patch version stays available, so an existing pin keeps resolving to the image it always did. Pin an exact version in CI: a scanner that changes underneath you changes your scores without anyone changing any code.
Platforms: linux/amd64 and linux/arm64. Each tag is a multi-architecture
manifest, so Apple Silicon, AWS Graviton and x86 hosts all pull the right image with no
extra flags.
Most quality tools read your source. Babysitter clones your repository into a clean room and runs it — installs the dependencies, builds it, executes the test suite, scans it, and reports what actually happened.
That distinction is the product. Cloning drops everything uncommitted, so a project that
builds only because of an untracked local_settings.py or a stray generated header fails
here — exactly as it will fail for the team receiving it.
Three outputs come back:
| Output | What it tells you |
|---|---|
| Readiness score | 0–100 across six weighted dimensions |
| Hard gates | Score-independent blockers. A repository cannot show 87/100 while nobody knows how to deploy it. |
| Babysitting Index | Estimated downstream cost in engineer-hours |
Repository Readiness: 39 / 100
NOT READY FOR HANDOFF
Hard gate: Cannot build from a clean checkout.
Hard gate: Critical tests do not exist.
Hard gate: Required production configuration undocumented.
Hard gate: License or compliance blocker.
Score 39 is below the fail threshold (50).
Dimensions
----------------------------------------------------------------
Build #######............. 33% (weight 15)
Tests .................... 0% (weight 25)
Architecture #################... 85% (weight 20)
Operational readiness ######.............. 30% (weight 15)
Security #################... 83% (weight 15)
Documentation .................... 0% (weight 10)
Babysitting Index: 72
Estimated downstream cost: 7.2-14.4 engineer-hours -> Remediate before handoff
Largest contributors:
+15 [TEST-001] No tests
+12 [BUILD-002] No installable dependency declaration
+10 [OPS-001] No deployment or run procedure
+9 [DOC-002] 1 environment variables are read but undocumented
+8 [DOC-001] README is unmodified project-template boilerplate
# Scan the current directory
podman run --rm -v "$PWD:/repo:ro" nicholasyue/babysitter scan /repo
# Scan a remote repository, no checkout needed
podman run --rm nicholasyue/babysitter scan https://example.com/acme/service.git
# Machine-readable output
podman run --rm -v "$PWD:/repo:ro" nicholasyue/babysitter scan /repo --json
# Rank a whole portfolio, worst first
podman run --rm -v "$PWD/repos:/repos:ro" nicholasyue/babysitter \
compare /repos/alpha /repos/beta /repos/gamma
# List every check, with its dimension and scope
podman run --rm nicholasyue/babysitter checks
# Full option reference
podman run --rm nicholasyue/babysitter scan --help
| Code | Meaning |
|---|---|
0 | READY — no gates tripped, no high-severity findings, score above threshold |
1 | CONDITIONAL — usable, but remediation is required before handoff |
2 | NOT READY — a hard gate tripped, or the score is below the fail threshold |
3 | ERROR — babysitter itself failed; the repository was not assessed |
Code 3 is deliberately distinct. A broken gate must never look like a failed repository.
Babysitter is built for rootless use: it runs as a non-root user inside the container,
needs no daemon, no --privileged, and no capabilities. It reads your repository through
a read-only mount and never writes to it.
Two things are worth knowing when running rootless.
Restrictive file permissions. Rootless Podman maps your host user to root inside the
container, and other identities to a subordinate UID range. Because Babysitter runs as an
unprivileged user in the container, a repository whose files are not group- or
world-readable (mode 0600, say) will be unreadable and the scan will fail. Map your own
UID straight through:
podman run --rm --userns=keep-id -v "$PWD:/repo:ro" nicholasyue/babysitter scan /repo
SELinux. On RHEL, Rocky, AlmaLinux and similar with SELinux enforcing, add a relabelling suffix to the mount so the container may read it:
podman run --rm -v "$PWD:/repo:ro,z" nicholasyue/babysitter scan /repo
Use :ro,z (shared label) rather than :ro,Z (private label) — Z relabels the
directory exclusively for one container, which is not what you want on a working
repository other processes also use.
Babysitter reads commit history for change-frequency ranking, contributor analysis, and
scanning past commits for secrets. Give it a full clone. In CI that usually means setting
fetch depth to unlimited — GIT_DEPTH: 0 on GitLab, fetch-depth: 0 on GitHub Actions.
A shallow clone still scans, but history-based findings will undercount.
31 high-signal checks rather than hundreds of style rules:
| Dimension | Weight | Question |
|---|---|---|
| Buildability | 15 | Can someone else clone it and reliably build it? |
| Correctness | 25 | Is there credible evidence that it works? |
| Maintainability | 20 | Can another engineer change it without fear? |
| Operational readiness | 15 | Can it be run, deployed, diagnosed — and does anyone still know how? |
| Security & dependencies | 15 | Secrets, CVEs, dependency pinning, licence |
| Documentation | 10 | Does it explain what it is supposed to do? |
Build and test verification: Python and C++/CMake. Every other repository still gets the language-agnostic checks — documentation, ownership, secrets, licence, complexity, deployment — and the report states plainly that build and test evidence is absent.
Coverage weighted by what matters. A repository with 90% coverage of trivial getters should not outscore one with 60% coverage of the code that decides things. Files are ranked by change frequency × complexity × how much depends on them, and coverage is reported against that.
Bus factor, not just documentation. A repository can be documented, tested, owned on paper and score 100 while exactly one person has ever written a line of it. Babysitter reads git history to flag critical paths with a single author, and critical paths that no currently active contributor has ever touched — the shape that turns one resignation into an outage.
It tells you what it could not check. A missing scanner is reported as "not scanned", never as "nothing found". A build needing a vendor SDK the image cannot carry is reported as "not verified", not as a failure. An unassessed dimension is excluded from the score rather than counted as zero — and then blocks a READY verdict, because "we never built it" must not round up to "ready for handoff".
Optional .babysitter.toml at your repository root. Every field shown is optional:
[babysitter]
ready_at = 80 # score at or above which a gate-free repo is READY
fail_under = 50 # score below which the verdict is NOT READY
timeout_seconds = 900 # per-check timeout for build and test commands
[weights] # dimension weights; defaults shown
buildability = 15
correctness = 25
maintainability = 20
operational = 15
security = 15
documentation = 10
[checks]
exclude = ["ARCH-005", "python.dead_code"] # finding ids or check ids
[paths]
exclude = ["generated/**"]
[build]
# System dependencies the clean-room image cannot be expected to carry — a licensed
# vendor SDK, say. A build failure attributable to one of these is reported as
# NOT VERIFIED rather than failed, and does not trip a hard gate.
system_requires = ["FooSDK>=7.0"]
[index.points]
# Recalibrate the Babysitting Index against your own observed handoffs.
"BUILD-001" = 24
[ownership]
path = "docs/OWNERS"
Configuration is read from the clean-room checkout, not your working tree — configuration that is not committed does not reach the receiving team either.
# GitLab CI
handoff-readiness:
stage: test
image: nicholasyue/babysitter:0.1.1
variables:
GIT_DEPTH: 0
script:
- babysitter scan . --format engineer
- babysitter scan . --json --output readiness.json
artifacts:
when: always
paths: [readiness.json]
allow_failure: true # remove once your baseline is green
# GitHub Actions
- uses: actions/checkout@v4
with:
fetch-depth: 0
- name: Handoff readiness
run: |
podman run --rm -v "$PWD:/repo:ro" \
nicholasyue/babysitter:0.1.1 scan /repo
Start with the gate permissive, measure your baseline, then enforce. A gate switched on before anyone knows their starting score is a gate everyone learns to override.
Scanning executes the target repository's own build and test code. That is not
incidental — it is how the tool establishes that a clean checkout reproduces. Running it
on a repository you do not trust is equivalent to running that project's make install on
your machine.
The container runs as a non-root user, scrubs the host environment before every subprocess so your credentials and toolchain roots never reach the code under test, and applies per-check timeouts. It is a boundary, not a guarantee. For untrusted code, harden the invocation:
podman run --rm \
--network=none \
--read-only --tmpfs /tmp:rw,nosuid \
--pids-limit 512 --memory 4g \
--cap-drop ALL --security-opt no-new-privileges \
-v "$PWD:/repo:ro" \
nicholasyue/babysitter scan /repo
--cpus is deliberately absent: the cpu cgroup controller is not delegated to
unprivileged users by default, so a CPU limit makes the command fail outright under
rootless Podman. --memory and --pids-limit need only the memory and pids
controllers, which usually are delegated. Per-check timeouts bound runtime regardless.
With --network=none dependency installation cannot reach a package index, so the build
and test phases are reported as NOT VERIFIED — a stated gap in the scan, never a
finding against the repository — and you keep every language-agnostic check. If you need
build evidence for untrusted code, run the scan in a disposable VM or an ephemeral CI
runner rather than relaxing the container.
Babysitter sends nothing anywhere. No telemetry, no phone-home, no analytics, no LLM calls. The only network access is dependency resolution for the repository being scanned, against whatever package index that repository declares.
Note that a report naming a leaked credential contains evidence of that credential's
location. Treat --json output with the same care as the repository it describes.
Two scans of the same commit produce byte-identical output. Every scanner and toolchain in the image is pinned by version and checksum, so a given image tag always returns the same verdict for a given commit — which is what makes it usable as a gate rather than as a suggestion. Upgrading the image is therefore a deliberate act that may move your scores.
v0.1.1. Honest about what it is not yet:
.mailmap is honoured and obvious automation is
excluded, but treat contributor counts as a strong hint, not a roster.The image is licensed under Apache-2.0. Use it freely, including commercially and
inside proprietary pipelines — there is no per-seat cost, no licence server, and no
obligation to publish anything you scan with it or build around it. The full licence text
is inside the image at /usr/share/doc/babysitter/LICENSE:
podman run --rm --entrypoint cat \
nicholasyue/babysitter /usr/share/doc/babysitter/LICENSE
The build source for the image is proprietary and not publicly available. Apache-2.0 places no source-disclosure obligation on the distributor, so this is simply how the project is published — it does not restrict what you may do with the image.
Third-party components — gitleaks, osv-scanner, cppcheck, ruff, vulture, pytest, coverage, lizard, and the Debian and Python base layers — remain under their own licences, and each ships its licence text in the image alongside its package metadata.
For support, defect reports, or questions about deploying this in your environment, contact [email protected].
Content type
Image
Digest
sha256:4d564d357…
Size
273.1 MB
Last updated
24 days ago
docker pull nicholasyue/babysitter