Management UI and REST API on every SEAPATH node: inventory, cluster join, Ansible runs
10K+
A management UI and REST API running on every SEAPATH node, so that a machine installed from the ISO is usable from a browser, several machines can be joined into a cluster, and Ceph can be deployed on top, with no separate Ansible control machine.
The configuration of a machine is edited here as an inventory and applied by the SEAPATH playbooks. SEAPATH is a function converting an inventory into a running infrastructure, and this service is a front end onto that function. What it writes itself is that inventory, its own trust material, and the Pacemaker metadata a guest carries on its disk image, which D31 records and bounds.
The Ansible control machine a site usually sets up beside a cluster has two jobs, holding the desired state and running the playbooks, and both move into the cluster itself.
Concretely, the service does five things:
ansible-runner and turns their event
stream into a readable progress view;No SEAPATH role is rewritten, and no configuration file is rendered twice. What the UI runs is what the CI tests.

Node describes what the machine is. The hostname and the distribution, the
isolated and housekeeping CPUs as the kernel command line and sysfs report
them, the disks under the stable by-path name Ceph wants, and the interfaces.
It reads, and offers two acts. Reboot this machine is an administrator's: a
run of one task that schedules the reboot a few seconds after the run ends, the
way an update reboots this machine, so the run records its status before the
machine goes down, and the confirmation says where the guests go. The console
button opens a shell on the ansible account for the times a page is not
enough, on this machine or on any other machine or guest of the inventory a run
reaches over SSH. That account may run /bin/sh as root with no password, the
rule Ansible's escalation uses, so sudo sh in a console is root on the
machine, and opening one asks for an administrator. Other sudo commands ask
for a password, sudo -s included when the login shell is bash.

Inventory is the desired state, edited as the folder of files it is. The
left column lists what the repository carries, meaning the inventory and every
quadlet, rule and template it names, with the history of who changed what. The
editor parses, checks the rules and asks ansible-inventory about the result
before committing anything.
The box is a text area, and a file whose indentation carries meaning needs more
of one than a browser gives: Tab and Shift+Tab shift the lines the
selection touches, Enter carries the indentation of the line it leaves, one
level in under a key that opens a block and the dash repeated in a list, and
Ctrl+/ comments the block out.
Beside them is a switch. An inventory is YAML with no schema, so
cephadm_netwrok is a name the file accepts, ansible-inventory parses and
every rule passes, and the answer arrives three minutes into a convergence from
a role that read a variable nobody set. With the assistant on, a name being
typed is completed from what may be written at that point in the file, each
candidate carrying the role that reads it and what goes wrong when it is wrong,
so a guest entry is offered vm_disk and a hypervisor is not. What is already
written is read back the other way: a name nothing reads, with the one that was
probably meant, and a variable written where nothing will read it. Nothing reads
it means none of the three readers an inventory has, the collection this node
runs, Ansible itself, and the file, which reads its own variables through every
{{ }} in it. None of this refuses a commit, because a variable of a site's own
is a legitimate name this service has never read and is written exactly the way
a misspelling is.
Under it sits the copy each of the other machines holds. Every node clones the repository, and bringing the clones together is an act: the panel asks each machine the inventory declares which commit it is on, and one button pushes this node's branch to all of them over the connection a run already makes. Git accepts a fast forward, so a machine carrying commits this node has never seen is named and left alone, and forcing the push says in the panel what that discards. See D32.

Deployment is where a machine actually changes. Commissioning runs the full
convergence; the picker beside it runs a single playbook when a single thing
was edited, and every entry says what it plays, what it will restart, and why
this node may not be allowed to run it. A playbook that needs a value for one
run only, such as the machine cluster_remove_machine takes out of the
cluster, says so on its card: the launch asks for it, and it is kept with the
run rather than in the inventory. Under them sit two panels, shut until
they are needed. Reaching the other machines is the SSH trust: the site key
this node holds, and the host keys it has accepted, both undone in one click.
The code this node runs is the pair that decides what an apply executes: the
seapath.ansible collection, which arrives as a file when the fix is upstream
and the image is not out yet, and this service itself, which is
seapath_webui_image in the inventory and changes the way every other change
to a machine does, by an apply.
The arrow in the top bar of every page asks the registry for a newer seapath-webui. It asks once on its own when an administrator signs in, and again on each click; a tick means the inventory already names the newest version the registry holds. When a newer version exists, the button becomes Upgrade to that version: one click pins it in the inventory as one commit and launches the playbook that deploys it on every machine naming an image. That run restarts this service and no guest, and the run window follows it over the page.



A run plays what its playbook names, less the guests of the VMs group: a
site's guests are appliances and vendor images as often as SEAPATH VMs, and the
first one refusing a connection would end the convergence of the machines the
operator came for. Choose machines, beside Apply, narrows a run before it
is launched. The window lists the groups, the machines and the guests of the
inventory as checkboxes, because --limit takes a union and converging two
hypervisors should be one launch, and the line under them names the machines
the current boxes resolve to. Nothing checked is the playbook's own scope. A
machine this node cannot reach is marked there, and narrowing to machines that
answer is how a node whose neighbour is down still converges itself. See D39 in
docs/decisions.md.

VMs is the guests, and it exists because a guest is one object whose parts
sit on three other pages. Its definition is an entry of the VMs group in the
inventory, the disk image and the libvirt XML it names are in the two stores
around that file, and what it is doing right now is one line of the Pacemaker
resource table. The page puts them on one row: the guest, which of the two
deployments creates it, whether it is running and where, and what can be done
to it. A Creation column joins them while a guest still carries its recipe or
force: whether a deployment would find the two files the entry names, and
the warning that matters most on the page, since the roles destroy and
recreate a guest that carries force. A converged inventory carries neither,
and the column stays away. The node it runs on says in its
colour what holds it there: blue where nothing does and the cluster placed it,
green where a constraint holds it exactly where its inventory entry declares,
amber where the cluster and the inventory disagree, with the whole sentence on
hover. A Specs column gives its size, the vCPUs, the memory and the disks
summed, each disk on hover, as libvirt-exporter reports the running domain,
so a guest that is shut off shows no disk. The address a run reaches it at
sits beside them, with whether the entry carries a cloud-init seed and a
Console button for a shell inside the guest at that address. Serial
console is for the guest that no longer answers there: it runs
vm-mgr console as root on a hypervisor, and vm_manager finds the machine
running the guest and attaches to its serial port. A switch above the table
shows all the guests, the cluster ones or the standalone ones, with a count on
each. Beside the table sit the guests Pacemaker runs and the inventory
describes nowhere, with the same buttons: a convergence leaves them alone, and
a guest added under one of their names would collide with it.
Where an inventory describes a cluster and a standalone machine at once, each
guest says which of the two creates it, through cluster_VMs and
standalone_VMs as children of VMs. Everything follows from that one fact:
the playbook that deploys it, the module that starts and stops it, the options
its entry may carry, and whether it has an RBD image to hold metadata. A file
with one flat group says nothing and its guests take the file's own mode, which
is every inventory written before those groups existed.

Adding a VM is the act it performs whole. Name the guest, give it a disk image
and a libvirt XML, and the page commits the entry and launches the deployment.
Each of the two files is either uploaded for this guest or picked from what
this node already holds, since one qcow2 built with cloud-init and SEAPATH's
guest.xml.j2 serve every guest of a site once what makes each guest different
lives in its entry.
That difference is the network section. It writes three variables with three
readers: bridges, the interface the template renders; cloud_init, the
mapping the upstream cloud_init_seed role builds the guest's NoCloud seed
from, with its address, gateway, resolvers, hostname and the packages it
installs on its first boot; and ansible_host, where a later run reaches inside
the guest. The address is typed once and written twice. The MAC the seed matches
the interface by is generated in the QEMU range when none is given. A network
that could not work is refused before it is written: an address with no prefix,
a gateway outside the guest's network, an address or a MAC another host of the
file already holds. Two boxes, checked by default, make the guest reachable by
the runs that follow: this node's public key, and the site key where one is
held, go into the seed for the ansible account, and the entry accepts the
guest's host key on the first connection. The seed is built by the role inside
this container, like on any control machine, which is why the image carries
cloud-localds. See D48 in docs/decisions.md.
Three more things sit under the network. Ping beside the address sends
three echo requests from this node, because the file can refuse an address one
of its own hosts holds and has no view of the rest of the network. A box gives
the account the seed creates passwordless sudo, for a generic cloud image
whose ansible account otherwise stops every task a run becomes root for. And
a root password for the serial console, for the guest whose network did not
come up: it travels with the run that creates the guest, is spliced into that
run's copy of the inventory and wiped when the run ends, so no form of it is
ever committed. See D52.
A guest whose domain XML has a VNC <graphics> also gets a Graphic console,
whatever created it: its screen in the browser, from the VNC server QEMU runs on
the hypervisor's loopback, with a button for Ctrl+Alt+Del, a toggle between
fitting the panel and scrolling at actual size, and full screen. It is how a
Windows guest whose network is down is reached. See D62.
Folded under all of that is what cluster_vm create is given: placement,
priority, live migration and its timeouts, colocation, disk bus, the pinning
profile. They are asked there because each is written once into the guest's
image metadata, and changing one afterwards costs an outage. Live migration is
on unless it is unchecked: without it, every move Pacemaker makes for the guest,
a failover or the standby before a reboot, is a stop on one machine and a boot
on the other. Underneath, those are the writes this service has always made and
the upstream playbook it has always run: the image to the store git does not
carry, the XML committed with the inventory, the guest a splice into the file
checked like every other write, and a whole playbook of the collection. The
operator is spared the trip through two pages and a group name they have no
reason to know.

The run that follows opens in a window over the page, here the VMs page the guest was added from, with the playbook, its state, the task being played and the same task stream the Runs page draws, so the operator reads the result of their own action on the page they launched it from. Closing the window stops this browser watching and nothing else: the run belongs to the service, keeps going and lands in the history. A run that ends under an open window has the panels of the page read again, since what the operator stayed for is what the run changed in the table underneath. Every action of this UI that is a playbook opens the same window. See D43 in docs/decisions.md.
Editing a guest's metadata is there too. A guest's Pacemaker configuration lives
as metadata on its RBD image, vm_manager writes those keys at creation and
never again, and the only upstream way to change one is to recreate the guest
from its seed image and lose its disk. So the page asks Ceph directly, with rbd image-meta, the way the cluster view asks the exporters: a window lists what
the image carries, another edits one value, wide enough for the libvirt domain
that lives in there under xml, and removing a key is asked before it happens
because nothing here puts back what it took away. Every write reads the image
before and after and says what moved. Applying a change stops the guest and
rebuilds its Pacemaker resource, so it is a second button that names the outage
and appears only when something did move.
A standalone guest has no RBD image and so no metadata. Its row offers its
libvirt domain and its pinning profile instead. The profile is
seapath_alloc on its inventory entry, committed like any edit and put on
the machine by a run of the two tasks deploy_vms_standalone writes it with,
and a profile naming an isolation seapath-alloc does not know is refused
before it is committed, since the hook would quietly leave those vCPUs on the
housekeeping cores. The domain has no path through the inventory, since the role
skips a guest libvirt already has: the window reads virsh dumpxml --inactive
over the SSH path a run takes, checks the edit keeps the name and UUID libvirt
holds, and defines it with a run of community.libvirt.virt. Both are read when
the guest starts from shut off, so saving either asks once whether to shut the
guest down and start it, and one run does all of it. See D66.
Starting and stopping a guest are there too, one button per row, offered as whichever of the two would change something. Each is a run: one task calling the upstream module, over the SSH path a convergence uses, under the same lock so a start cannot slip in under a convergence. The confirmation names the guest and says what stopping it does, because on these machines what a guest serves is a substation function.
A cluster guest can also be taken out of the cluster and put back. Disable
is cluster_vm disable, run the same way: the guest is stopped and its
Pacemaker resource removed, while its RBD image, its metadata and its inventory
entry stay, and the row then says it is out of the cluster and that a
deployment leaves it alone. Enable builds the resource back from the image.
A disabled guest can be deleted for good, by an administrator: the entry is
committed out of the inventory first, so no deployment creates the guest again,
then cluster_vm remove deletes its RBD group, images and metadata. The disk
image and the XML stay, since another guest may be made from them.
Once a guest exists, the lines that created it are read by nobody: the roles
read vm_disk, vm_template, cloud_init and their neighbours only for a
guest the hypervisor does not have yet. So the deployment run that creates a
guest takes them out of its entry when it ends, as one commit authored by the
operator who launched the run and naming it, and leaves the variables later
runs read. Each guest is judged on what the machines report, so a guest the run
did not manage to create keeps its recipe for the next one. The Creation column
names the files while the entry still carries them. Afterwards a struck out
sheet of paper beside the guest's name offers to delete them: the image and the
XML the guest was made from, deleted from this node when no other entry names
them. See D49 and D50 in docs/decisions.md.

Containers is the same idea one layer down, and almost all of it already
existed. A container in SEAPATH is a quadlet: upload_extra_files copies a
.container file to /etc/containers/systemd, podman's generator turns it
into a systemd unit at the next daemon-reload. On a cluster, a container
Pacemaker runs is a workload of cluster_containers, which
deploy_containers_cluster deploys on every hypervisor with its images, its
RBD image and its resource. Variables the upstream roles already read, which
the page reads back and joins to what the machines publish: the unit state comes from the systemd
collector of the node_exporter every node runs, out of the same exposition
the CPU pool is read from, so the reading costs no new request anywhere.
Who owns a container decides what the row offers. Pacemaker holds a resource for it: one row, one act, and the cluster chooses the node. Nothing holds one: a row per machine, because the same quadlet is a unit on each machine the inventory sends it to and stopping it on one says nothing about the others. Declaring one writes the same entries a site would have written by hand, at the scope the operator picks, and the entry lands where those machines already read the list from, because Ansible replaces a variable rather than merging it and an entry in the wrong place silently stops the site's other uploads. What makes it real is a run, named rather than launched: the playbook that uploads a quadlet is the prerequisites one, which reconfigures a great deal more than a container.

The Quadlet column names the file podman reads, and clicking it shows that
file, since its dozen lines answer most of what the row raises: which image,
which ports, and whether an [Install] section is about to start the container
behind Pacemaker's back. It is read from the inventory folder and from nowhere
else on this machine, and editing it stays on the Inventory page, where a write
is a commit. A container the cluster holds says where it runs in the colour of
its node name, like a guest, and carries Move and Return, the two acts
the Cluster page offers on the same Pacemaker resource. Their confirmation says
what a move costs a container: podman has no live migration, so the unit stops
on one member and starts on the other.
A workload of cluster_containers also carries Remove: the entry is
marked state: absent, deploy_containers_cluster takes it off every node,
its RBD image too when the operator checks it, and once that run has
succeeded the entry and its files leave the inventory.


Usage says how busy each machine is and which guest or container is making
it so. A card per machine of the inventory carries its CPU, memory and what it
runs, and the machine selected below it gets five minutes of charts: CPU and
memory stacked by workload, with the isolated and housekeeping CPUs summed
apart, then disk and physical port traffic. Under them sit a table of what
each guest and container consumes, the file systems, the disks and the
interfaces. It is read from node_exporter, libvirt-exporter and
prometheus-podman-exporter, which the monitoring playbook already installs on
every machine. The service reads them every five seconds and keeps the last
five minutes in memory, and only while somebody signed in has used it in the
last fifteen minutes, so a page opened late starts full and two operators cost
the machines one reading. History and alerting stay Prometheus's. See D67.

Cluster is what the machines are doing right now, which is the one question the other pages cannot answer: which node that VM is on, whether the cluster has quorum, whether an OSD went down last night. Membership leads with quorum, because a cluster without it moves nothing. Resources is the table with the failure in it, and in SEAPATH those are mostly VMs, one Pacemaker resource per guest. Storage is Ceph: health with the c
Content type
Image
Digest
sha256:6ed7265b7…
Size
207.9 MB
Last updated
1 day ago
docker pull insatomcat/seapath-webui