Intel PMEM-CSI is a storage driver for container orchestrators like Kubernetes. It makes local persistent memory (PMEM) available as a filesystem volume to container applications.
It can currently utilize non-volatile memory devices that can be controlled via the libndctl utility library. In this readme, we use persistent memory to refer to a non-volatile dual in-line memory module (NVDIMM).
The PMEM-CSI driver follows the CSI specification by listening for API requests and provisioning volumes accordingly.
The PMEM-CSI driver can operate in two different DeviceModes: LVM and Direct
The following diagram illustrates the operation in DeviceMode:LVM:

In DeviceMode:LVM PMEM-CSI driver uses LVM for Logical Volumes Management to avoid the risk of fragmentation. The LVM logical volumes are served to satisfy API requests. There is one Volume Group created per Region, ensuring the region-affinity of served volumes.
The driver consists of three separate binaries that form two initialization stages and a third API-serving stage.
During startup, the driver scans persistent memory for regions and namespaces, and tries to create more namespaces using all or part (selectable via option) of the remaining available space. The namespace size can be specified as a driver parameter and defaults to 32 GB. This first stage is performed by a separate entity pmem-ns-init.
The second stage of initialization arranges physical volumes provided by namespaces into LVM volume groups. This is performed by a separate binary pmem-vgm.
After two initialization stages, the third binary pmem-csi-driver starts serving CSI API requests.
The PMEM-CSI driver can pre-create Namespaces in two modes, forming
corresponding LVM Volume groups, to serve volumes based on fsdax or
sector (alias safe) mode Namespaces. The amount of space to be
used is determined using two options -useforfsdax and
-useforsector given to pmem-ns-init. These options specify an
integer presenting limit as percentage, which is applied separately in
each Region. The default values are useforfsdax=100 and
useforsector=0. A CSI request for volume can specify the Namespace
mode using the driver-specific argument nsmode which has a value of
either "fsdax" (default) or "sector". A volume provisioned in fsdax
mode will have the dax option added to mount options.
The PMEM-CSI driver can leave space on devices for others, and
recognize "own" namespaces. Leaving space for others can be achieved
by specifying lower-than-100 values to -useforfsdax and/or
-useforsector options. The distinction "own" vs. "foreign" is
implemented by setting the Name field in Namespace to a static
string "pmem-csi" during Namespace creation. When adding Physical
Volumes to Volume Groups, only Physical Volumes that are based on
Namespaces with the name "pmem-csi" are considered.
The following diagram illustrates the operation in DeviceMode:Direct:

In DeviceMode:Direct PMEM-CSI driver allocates Namespaces directly from the storage device. This creates device space fragmentation risk, but reduces complexity and run-time overhead by avoiding additional device mapping layer. Direct mode also ensures the region-affinity of served volumes, because provisioned volume can belong to one Region only.
In Direct mode, the two preparation stages used in LVM mode, are not needed.
The PMEM-CSI driver creates a Namespace directly in the mode which is asked by Volume creation request, thus bypassing the complexity of pre-allocated pools that are used in DeviceMode:LVM.
In DeviceMode:Direct, the driver does not attempt to limit space use. It also does not mark "own" namespaces. The Name field of a Namespace gets value of the VolumeID.
The PMEM-CSI driver supports running in different modes, which can be controlled by passing one of the below options to the driver's '-mode' command line option. In each mode, it starts a different set of open source Remote Procedure Call (gRPC) servers on given driver endpoint(s).
Controller should run as a single instance in cluster level. When the driver is running in Controller mode, it forwards the pmem volume create/delete requests to the registered node controller servers running on the worker node. In this mode, the driver starts the following gRPC servers:
One Node instance should run on each worker node that has persistent memory devices installed. When the driver starts in such mode, it registers with the Controller driver running on a given -registryEndpoint. In this mode, the driver starts the following servers:
This gRPC server operates on a given endpoint in all driver modes and implements the CSI Identity interface.
When the PMEM-CSI driver runs in Controller mode, it starts a gRPC server on a given endpoint(-registryEndpoint) and serves the RegistryServer interface. The driver(s) running in Node mode can register themselves with node specific information such as node id, NodeControllerServer endpoint, and their available persistent memory capacity.
This gRPC server is started by the PMEM-CSI driver running in Controller mode and serves the Controller interface defined by the CSI specification. The server responds to CreateVolume(), DeleteVolume(), ControllerPublishVolume(), ControllerUnpublishVolume(), and ListVolumes() calls coming from external-provisioner and external-attacher sidecars. It forwards the publish and unpublish volume requests to the appropriate Node controller server running on a worker node that was registered with the driver.
This gRPC server is started by the PMEM-CSI driver running in Node mode and implements the ControllerPublishVolume and ControllerUnpublishVolume methods of the Controller service interface defined by the CSI specification. It serves the ControllerPublishVolume() and ControllerUnpublish() requests coming from the Master controller server and creates/deletes persistent memory devices.
This gRPC server is started by the driver running in Node mode and implements the Node service interface defined in the CSI specification. It serves the NodeStageVolume(), NodeUnstageVolume(), NodePublishVolume(), and NodeUnpublishVolume() requests coming from the Container Orchestrator (CO).
The following diagram illustrates the communication channels between driver components:

All PMEM-CSI specific communication shown in above section between Master Controller(RegistryServer, MasterControllerServer) and NodeControllers(NodeControllerServer) is protected by mutual TLS. Both client and server must identify themselves and the certificate they present must be trusted. The common name in each certificate is used to identify the different components. The following common names have a special meaning:
pmem-registry is used by the RegistryServer.pmem-node-controller is used by NodeControllerServersThe test/setup-ca-kubernetes.sh
script shows how to generate certificates signed by Kubernetes cluster
root Certificate Authority. And the provided deployment
files shows how to use the generated
certificates to setup the driver. The test cluster is setup using
certificates created by that script. The
test/setup-ca.sh script also shows how to
generate self signed certificates. These are just examples,
administrators of a cluster must ensure that they choose key lengths
and algorithms of sufficient strength for their purposes and manage
certificate distribution.
A production deployment can improve upon that by using some other key delivery mechanism, like for example Vault.
In a typical CSI deployment, volumes are provided by a storage backend that is independent of a particular node. When a node goes offline, the volume can be mounted elsewhere. But PMEM volumes are local to node and thus can only be used on the node where they were created. This means the applications using PMEM volume cannot freely move between nodes. This limitation needs to be considered when designing and deploying applications that are to use local storage.
Below are the volume persistency models considered for implementation in PMEM-CSI to serve different application use cases:
Persistent Volumes
A volume gets created independently of the application, on some node
where there is enough free space. Applications using such a volume are
then forced to run on that node and cannot run when the node is
down. Data is retained until the volume gets deleted.
Ephemeral Volumes
Each time an application starts to run on a node, a new volume is
created for it on that node. When the application stops, the volume is
deleted. The volume cannot be shared with other applications. Data on
this volume is retained only while the application runs.
Cache Volumes
Volumes are pre-created on a certain set of nodes, each with its own
local data. Applications are started on those nodes and then get to
use the volume on their node. Data persists across application
restarts. This is useful when the data is only cached information that
can be discarded and reconstructed at any time and the application
can reuse existing local data when restarting.
| Volume | Kubernetes | PMEM-CSI | Limitations |
|---|---|---|---|
| Persistent | supported | supported | topology aware scheduling1 |
| Ephemeral | in design | in design | topology aware scheduling1, resource constraints2 |
| Cache | supported | supported | topology aware scheduling1 |
1 Topology aware scheduling ensures that an application runs on a node where the volume was created. For CSI-based drivers like PMEM-CSI, Kubernetes >= 1.13 is needed. On older Kubernetes releases, pods must be scheduled manually onto the right node(s).
2 The upstream design for ephemeral volumes currently does not take resource constraints into account. If an application gets scheduled onto a node and then creating the ephemeral volume on that node fails, the application on the node cannot start until resources become available.
Kubernetes cluster administrators can expose above mentioned volume
persistency types to applications using
StorageClass Parameters. An
optional persistencyModel parameter differentiates how the
provisioned volume can be used.
if no persistencyModel parameter specified in StorageClass then
it is treated as normal Kubernetes persistent volume. In this case
PMEM-CSI creates PMEM volume on a node and the application that
claims to use this volume is supposed to be scheduled onto this node
by Kubernetes. Choosing of node is depend on StorageClass
volumeBindingMode. In case of volumeBindingMode: Immediate
PMEM-CSI chooses a node randomly, and in case of volumeBindingMode: WaitForFirstConsumer Kubernetes first chooses a node for scheduling
the application, and PMEM-CSI creates the volume on that
node. Applications which claim a normal persistent volume has to use
ReadOnlyOnce access mode in its accessModes list. This
diagram
illustrates how a normal persistent volume gets provisioned in
Kubernetes using PMEM-CSI driver.
persistencyModel: cache
Volumes of this type shall be used in combination with
volumeBindingMode: Immediate. In this case, PMEM-CSI creates a set
of PMEM volumes each volume on different node. The number of PMEM
volumes to create can be specified by cacheSize StorageClass
parameter. Applications which claim a cache volume can use
ReadWriteMany in its accessModes list. Check with provided cache
StorageClass
example. This
diagram
illustrates how a cache volume gets provisioned in Kubernetes using
PMEM-CSI driver.
NOTE: Cache volumes are local to node not Pod. If two Pods using the same cache volume runs on the same node, will not get their own local volume, instead they endup sharing the same PMEM volume. Applications has to consider this and use available Kubernetes mechanisms like node anti-affinity while deploying. Check with provided cache application example.
Building of Docker images has been verified using Docker-ce: version 18.06.1
Persistent memory device(s) are required for operation. However, some development and testing can be done using QEMU-emulated persistent memory devices, see the "QEMU and Kubernetes" section for the commands that create such a virtual test cluster.
The driver does not create persistent memory Regions, but expects Regions to exist when the driver starts. The utility ipmctl can be used to create Regions.
PMEM-CSI driver implements CSI specification version 1.0.0, which only supported by Kubernetes versions >= v1.13. The driver deployment in Kubernetes cluster has been verified on:
| Branch | Kubernetes branch/version | Required alpha feature gates |
|---|---|---|
| devel | Kubernetes 1.13 | CSINodeInfo, CSIDriverRegistry |
| devel | Kubernetes 1.14 |
Use these commands:
mkdir -p $GOPATH/src/github.com/intel
git clone https://github.com/intel/pmem-csi $GOPATH/src/github.com/intel/pmem-csi
Use make build-image to produce Docker container image.
Use make push-image to build and push Docker container image to a Docker images registry. The
default is to push to a local Docker registry.
See the Makefile for additional make targets and possible make variables.
This section assumes that a Kubernetes cluster is already available with at least one node that has persistent memory device(s). For development or testing, it is also possible to use a cluster that runs on QEMU virtual machines, see the "QEMU and Kubernetes" section below.
The method to configure alpha feature gates may vary, depending on the Kubernetes deployment.
$ kubectl label node <your node> storage=pmem
If you are not using the test cluster described in Starting and stopping a test cluster where CRDs are installed automatically, you must install those manually. Kubernetes 1.14 and higher have those APIs built in and thus don't need these CRDs.
$ kubectl create -f https://raw.githubusercontent.com/kubernetes/kubernetes/release-1.13/cluster/addons/storage-crds/csidriver.yaml
$ kubectl create -f https://raw.githubusercontent.com/kubernetes/kubernetes/release-1.13/cluster/addons/storage-crds/csinodeinfo.yaml
Certificates are required as explained in Security.
If you are not using the test cluster described in
Starting and stopping a test cluster
where certificates are created automatically, you must set up certificates manually.
This can be done by running the ./test/setup-ca-kubernetes.sh script for your cluster.
This script requires "cfssl" tools which can be downloaded.
These are the steps for manual set-up of certificates:
$ curl -L https://pkg.cfssl.org/R1.2/cfssl_linux-amd64 -o _work/bin/cfssl --create-dirs
$ curl -L https://pkg.cfssl.org/R1.2/cfssljson_linux-amd64 -o _work/bin/cfssljson --create-dirs
$ chmod a+x _work/bin/cfssl _work/bin/cfssljson
$ KUBCONFIG="<<your cluster kubeconfig path>> PATH="$PATH:_work/bin" ./test/setup-ca-kubernetes.sh
$ sed -e 's/192.168.8.1:5000/<your registry>/' deploy/kubernetes-<kubernetes version>/pmem-csi-lvm.yaml | kubectl create -f -
$ sed -e 's/192.168.8.1:5000/<your registry>/' deploy/kubernetes-<kubernetes version>/pmem-csi-direct.yaml | kubectl create -f -
The deployment yaml file uses the registry address for the QEMU test cluster
setup (see below). When deploying on a real cluster, some registry
that can be accessed by that cluster has to be used.
If the Docker registry runs on the local development
host, then the sed command which replaces the Docker registry is not needed.
The deploy directory contains one directory or symlink for each
tested Kubernetes release. The most recent one might also work on
future, currently untested releases.
$ kubectl get pods
NAME READY STATUS RESTARTS AGE
pmem-csi-8kmxf 2/2 Running 0 3m15s
pmem-csi-bvx7m 2/2 Running 2 3m15s
pmem-csi-controller-0 4/4 Running 1 3m15s
pmem-csi-fbmpg 2/2 Running 2 3m15s
$ kubectl get nodes --show-labels
The command output must indicate that every node with PMEM has these two labels:
pmem-csi.intel.com/node=<NODE-NAME>,storage=pmem
If storage=pmem is missing, label manually as described above. If pmem-csi.intel.com/node is missing, then double-check that the alpha feature gates are enabled and the CSI driver is running.
$ kubectl create -f deploy/kubernetes-<kubernetes version>/pmem-storageclass-ext4.yaml
$ kubectl create -f deploy/kubernetes-<kubernetes version>/pmem-storageclass-xfs.yaml
$ kubectl create -f deploy/kubernetes-<kubernetes version>/pmem-pvc.yaml
$ kubectl get pvc
NAME STATUS VOLUME CAPACITY ACCESS MODES STORAGECLASS AGE
pmem-csi-pvc-ext4 Bound pvc-f70f7b36-6b36-11e9-bf09-deadbeef0100 4Gi RWO pmem-csi-sc-ext4 16s
pmem-csi-pvc-xfs Bound pvc-f7101fd2-6b36-11e9-bf09-deadbeef0100 4Gi RWO pmem-csi-sc-xfs 16s
$ kubectl create -f deploy/kubernetes-<kubernetes version>/pmem-app.yaml
These applications use storage: pmem in the nodeSelector list to ensure scheduling to a node supporting pmem device, and each requests a mount of a volume, one with ext4-format and another with xfs-format file system.
$ kubectl get po my-csi-app-1 my-csi-app-2
NAME READY STATUS RESTARTS AGE
my-csi-app-1 1/1 Running 0 6m5s
NAME READY STATUS RESTARTS AGE
my-csi-app-2 1/1 Running 0 6m1s
$ kubectl exec my-csi-app-1 -- df /data
Filesystem 1K-blocks Used Available Use% Mounted on
/dev/ndbus0region0fsdax/5ccaa889-551d-11e9-a584-928299ac4b17
4062912 16376 3820440 0% /data
$ kubectl exec my-csi-app-2 -- df /data
Filesystem 1K-blocks Used Available Use% Mounted on
/dev/ndbus0region0fsdax/5cc9b19e-551d-11e9-a584-928299ac4b17
4184064 37264 4146800 1% /data
$ kubectl exec my-csi-app-1 -- mount |grep /data
/dev/ndbus0region0fsdax/5ccaa889-551d-11e9-a584-928299ac4b17 on /data type ext4 (rw,relatime,dax)
$ kubectl exec my-csi-app-2 -- mount |grep /data
/dev/ndbus0region0fsdax/5cc9b19e-551d-11e9-a584-928299ac4b17 on /data type xfs (rw,relatime,attr2,dax,inode64,noquota)
Content type
Image
Digest
Size
193.1 MB
Last updated
over 7 years ago
docker pull jascott1org/pmem-csi