Sign inSign up

nicholasyue/babysitter

By nicholasyue

•Updated 24 days ago

Repository hand over readiness reporter

Image
0

207

nicholasyue/babysitter repository overview

⁠Babysitter

Existing tools tell you if a repository is ready to publish. Babysitter tells you if it is ready to be someone else's problem — and prices that in engineering hours, not points.

Babysitter is a containerised acceptance gate that answers one question:

Could a competent engineer who did not write this code take responsibility for it — right now, from a clean checkout?

It runs like pytest or a linter: one command, a verdict, and an exit code your pipeline can gate on.

podman run --rm -v "$PWD:/repo:ro" nicholasyue/babysitter scan /repo

Runs rootless under Podman, with no daemon and no privileged mode. Docker works identically — substitute docker for podman in any command here.


⁠Supported tags

TagMeaning
0.1.1The current release. Pin an exact version like this in CI.
0.1Latest patch of the 0.1 line — currently 0.1.1.
latestLatest release — currently 0.1.1.

Every published patch version stays available, so an existing pin keeps resolving to the image it always did. Pin an exact version in CI: a scanner that changes underneath you changes your scores without anyone changing any code.

Platforms: linux/amd64 and linux/arm64. Each tag is a multi-architecture manifest, so Apple Silicon, AWS Graviton and x86 hosts all pull the right image with no extra flags.


⁠What it actually does

Most quality tools read your source. Babysitter clones your repository into a clean room and runs it — installs the dependencies, builds it, executes the test suite, scans it, and reports what actually happened.

That distinction is the product. Cloning drops everything uncommitted, so a project that builds only because of an untracked local_settings.py or a stray generated header fails here — exactly as it will fail for the team receiving it.

Three outputs come back:

OutputWhat it tells you
Readiness score0–100 across six weighted dimensions
Hard gatesScore-independent blockers. A repository cannot show 87/100 while nobody knows how to deploy it.
Babysitting IndexEstimated downstream cost in engineer-hours
⁠Example
Repository Readiness: 39 / 100
NOT READY FOR HANDOFF
  Hard gate: Cannot build from a clean checkout.
  Hard gate: Critical tests do not exist.
  Hard gate: Required production configuration undocumented.
  Hard gate: License or compliance blocker.
  Score 39 is below the fail threshold (50).

Dimensions
----------------------------------------------------------------
  Build                  #######.............   33% (weight 15)
  Tests                  ....................    0% (weight 25)
  Architecture           #################...   85% (weight 20)
  Operational readiness  ######..............   30% (weight 15)
  Security               #################...   83% (weight 15)
  Documentation          ....................    0% (weight 10)

Babysitting Index: 72
  Estimated downstream cost: 7.2-14.4 engineer-hours  ->  Remediate before handoff
  Largest contributors:
    +15  [TEST-001] No tests
    +12  [BUILD-002] No installable dependency declaration
    +10  [OPS-001] No deployment or run procedure
    +9   [DOC-002] 1 environment variables are read but undocumented
    +8   [DOC-001] README is unmodified project-template boilerplate

⁠Usage

# Scan the current directory
podman run --rm -v "$PWD:/repo:ro" nicholasyue/babysitter scan /repo

# Scan a remote repository, no checkout needed
podman run --rm nicholasyue/babysitter scan https://example.com/acme/service.git

# Machine-readable output
podman run --rm -v "$PWD:/repo:ro" nicholasyue/babysitter scan /repo --json

# Rank a whole portfolio, worst first
podman run --rm -v "$PWD/repos:/repos:ro" nicholasyue/babysitter \
  compare /repos/alpha /repos/beta /repos/gamma

# List every check, with its dimension and scope
podman run --rm nicholasyue/babysitter checks

# Full option reference
podman run --rm nicholasyue/babysitter scan --help
⁠Exit codes
CodeMeaning
0READY — no gates tripped, no high-severity findings, score above threshold
1CONDITIONAL — usable, but remediation is required before handoff
2NOT READY — a hard gate tripped, or the score is below the fail threshold
3ERROR — babysitter itself failed; the repository was not assessed

Code 3 is deliberately distinct. A broken gate must never look like a failed repository.


⁠Running rootless

Babysitter is built for rootless use: it runs as a non-root user inside the container, needs no daemon, no --privileged, and no capabilities. It reads your repository through a read-only mount and never writes to it.

Two things are worth knowing when running rootless.

Restrictive file permissions. Rootless Podman maps your host user to root inside the container, and other identities to a subordinate UID range. Because Babysitter runs as an unprivileged user in the container, a repository whose files are not group- or world-readable (mode 0600, say) will be unreadable and the scan will fail. Map your own UID straight through:

podman run --rm --userns=keep-id -v "$PWD:/repo:ro" nicholasyue/babysitter scan /repo

SELinux. On RHEL, Rocky, AlmaLinux and similar with SELinux enforcing, add a relabelling suffix to the mount so the container may read it:

podman run --rm -v "$PWD:/repo:ro,z" nicholasyue/babysitter scan /repo

Use :ro,z (shared label) rather than :ro,Z (private label) — Z relabels the directory exclusively for one container, which is not what you want on a working repository other processes also use.

⁠Git history

Babysitter reads commit history for change-frequency ranking, contributor analysis, and scanning past commits for secrets. Give it a full clone. In CI that usually means setting fetch depth to unlimited — GIT_DEPTH: 0 on GitLab, fetch-depth: 0 on GitHub Actions. A shallow clone still scans, but history-based findings will undercount.


⁠What it checks

31 high-signal checks rather than hundreds of style rules:

DimensionWeightQuestion
Buildability15Can someone else clone it and reliably build it?
Correctness25Is there credible evidence that it works?
Maintainability20Can another engineer change it without fear?
Operational readiness15Can it be run, deployed, diagnosed — and does anyone still know how?
Security & dependencies15Secrets, CVEs, dependency pinning, licence
Documentation10Does it explain what it is supposed to do?

Build and test verification: Python and C++/CMake. Every other repository still gets the language-agnostic checks — documentation, ownership, secrets, licence, complexity, deployment — and the report states plainly that build and test evidence is absent.

⁠Three things that make it different

Coverage weighted by what matters. A repository with 90% coverage of trivial getters should not outscore one with 60% coverage of the code that decides things. Files are ranked by change frequency × complexity × how much depends on them, and coverage is reported against that.

Bus factor, not just documentation. A repository can be documented, tested, owned on paper and score 100 while exactly one person has ever written a line of it. Babysitter reads git history to flag critical paths with a single author, and critical paths that no currently active contributor has ever touched — the shape that turns one resignation into an outage.

It tells you what it could not check. A missing scanner is reported as "not scanned", never as "nothing found". A build needing a vendor SDK the image cannot carry is reported as "not verified", not as a failure. An unassessed dimension is excluded from the score rather than counted as zero — and then blocks a READY verdict, because "we never built it" must not round up to "ready for handoff".


⁠Configuration

Optional .babysitter.toml at your repository root. Every field shown is optional:

[babysitter]
ready_at = 80             # score at or above which a gate-free repo is READY
fail_under = 50           # score below which the verdict is NOT READY
timeout_seconds = 900     # per-check timeout for build and test commands

[weights]                 # dimension weights; defaults shown
buildability = 15
correctness = 25
maintainability = 20
operational = 15
security = 15
documentation = 10

[checks]
exclude = ["ARCH-005", "python.dead_code"]   # finding ids or check ids

[paths]
exclude = ["generated/**"]

[build]
# System dependencies the clean-room image cannot be expected to carry — a licensed
# vendor SDK, say. A build failure attributable to one of these is reported as
# NOT VERIFIED rather than failed, and does not trip a hard gate.
system_requires = ["FooSDK>=7.0"]

[index.points]
# Recalibrate the Babysitting Index against your own observed handoffs.
"BUILD-001" = 24

[ownership]
path = "docs/OWNERS"

Configuration is read from the clean-room checkout, not your working tree — configuration that is not committed does not reach the receiving team either.


⁠CI integration

# GitLab CI
handoff-readiness:
  stage: test
  image: nicholasyue/babysitter:0.1.1
  variables:
    GIT_DEPTH: 0
  script:
    - babysitter scan . --format engineer
    - babysitter scan . --json --output readiness.json
  artifacts:
    when: always
    paths: [readiness.json]
  allow_failure: true      # remove once your baseline is green
# GitHub Actions
- uses: actions/checkout@v4
  with:
    fetch-depth: 0
- name: Handoff readiness
  run: |
    podman run --rm -v "$PWD:/repo:ro" \
      nicholasyue/babysitter:0.1.1 scan /repo

Start with the gate permissive, measure your baseline, then enforce. A gate switched on before anyone knows their starting score is a gate everyone learns to override.


⁠Security

Scanning executes the target repository's own build and test code. That is not incidental — it is how the tool establishes that a clean checkout reproduces. Running it on a repository you do not trust is equivalent to running that project's make install on your machine.

The container runs as a non-root user, scrubs the host environment before every subprocess so your credentials and toolchain roots never reach the code under test, and applies per-check timeouts. It is a boundary, not a guarantee. For untrusted code, harden the invocation:

podman run --rm \
  --network=none \
  --read-only --tmpfs /tmp:rw,nosuid \
  --pids-limit 512 --memory 4g \
  --cap-drop ALL --security-opt no-new-privileges \
  -v "$PWD:/repo:ro" \
  nicholasyue/babysitter scan /repo

--cpus is deliberately absent: the cpu cgroup controller is not delegated to unprivileged users by default, so a CPU limit makes the command fail outright under rootless Podman. --memory and --pids-limit need only the memory and pids controllers, which usually are delegated. Per-check timeouts bound runtime regardless.

With --network=none dependency installation cannot reach a package index, so the build and test phases are reported as NOT VERIFIED — a stated gap in the scan, never a finding against the repository — and you keep every language-agnostic check. If you need build evidence for untrusted code, run the scan in a disposable VM or an ephemeral CI runner rather than relaxing the container.

Babysitter sends nothing anywhere. No telemetry, no phone-home, no analytics, no LLM calls. The only network access is dependency resolution for the repository being scanned, against whatever package index that repository declares.

Note that a report naming a leaked credential contains evidence of that credential's location. Treat --json output with the same care as the repository it describes.


⁠Reproducibility

Two scans of the same commit produce byte-identical output. Every scanner and toolchain in the image is pinned by version and checksum, so a given image tag always returns the same verdict for a given commit — which is what makes it usable as a gate rather than as a suggestion. Upgrading the image is therefore a deliberate act that may move your scores.


⁠Status

v0.1.1. Honest about what it is not yet:

  • The Babysitting Index weights are placeholders. They encode a plausible ordering of cost, not measured truth. The ranking is trustworthy well before the hours are; recalibrate the point values against your own observed handoffs.
  • Semantic checks are not implemented. Comment-versus-code drift and duplicate concept detection need language understanding; this release is fully deterministic and offline by design.
  • Two ecosystems for build verification: Python and C++/CMake.
  • The arm64 build is produced under emulation and has had less real-world exercise than amd64. Report anything that behaves differently.
  • Bus factor is an approximation. Git identity is unreliable — one person commits under three addresses, bots commit as people. .mailmap is honoured and obvious automation is excluded, but treat contributor counts as a strong hint, not a roster.

⁠Licence

The image is licensed under Apache-2.0. Use it freely, including commercially and inside proprietary pipelines — there is no per-seat cost, no licence server, and no obligation to publish anything you scan with it or build around it. The full licence text is inside the image at /usr/share/doc/babysitter/LICENSE:

podman run --rm --entrypoint cat \
  nicholasyue/babysitter /usr/share/doc/babysitter/LICENSE

The build source for the image is proprietary and not publicly available. Apache-2.0 places no source-disclosure obligation on the distributor, so this is simply how the project is published — it does not restrict what you may do with the image.

Third-party components — gitleaks, osv-scanner, cppcheck, ruff, vulture, pytest, coverage, lizard, and the Debian and Python base layers — remain under their own licences, and each ships its licence text in the image alongside its package metadata.

⁠Support

For support, defect reports, or questions about deploying this in your environment, contact [email protected]⁠.

Tag summary

Content type

Image

Digest

sha256:4d564d357…

Size

273.1 MB

Last updated

24 days ago

docker pull nicholasyue/babysitter