Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions docs/guides/index.md
Original file line number Diff line number Diff line change
Expand Up @@ -18,6 +18,7 @@
* [Directive Breakdown](directive-breakdown/readme.md)
* [User Interactions](user-interactions/readme.md)
* [System Storage](system-storage/readme.md)
* [Redundant Allocations](redundant-allocations/readme.md)

## NNF User Containers

Expand Down
38 changes: 38 additions & 0 deletions docs/guides/redundant-allocations/readme.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,38 @@
---
authors: Matt Richerson <matt.richerson@hpe.com>
categories: provisioning
---

# Redundant Allocations

## Background

Lustre, XFS, and Raw allocations can be created using a redundant configuration either through zpools or LVM. This guide documents how these configurations might be used.

## Creating a Redundant Allocation

An allocation may choose to use a redundant configuration to allow access to data even when a drive has failed. In addition, rebuild commands can be specified to allow a replacement drive to be added to the RAID.

The [Storage Profiles](../storage-profiles/readme.md) section details the command lines needed to create an allocation using a RAID device.

## Rebuilding a Redundant Allocation

Choosing whether to allow a redundant allocation to be rebuilt with a new drive depends on the file system and allocation type.

### LVM

For XFS and Raw allocations using LVM, the RAID device can only be rebuilt when the Rabbit node has the LV activated. For `jobdw` allocations, this means that the RAID cannot be rebuilt while the compute nodes have the LV activated between `PreRun` and `PostRun`. For this reason, the most common way to use RAID configurations in this situation will be without rebuilding enabled.

If an XFS or Raw allocation is made with `create_persistent`, then the rebuild commands should be specified to allow the RAID to be rebuilt when no compute nodes have the LV activated.

### Zpool

For Lustre allocations using zpool, there are no restrictions on when rebuild commands can be used. The RAID can be rebuilt for job and persistent instances regardless of whether the compute nodes have the file system mounted. This is possible because the Lustre targets are only mounted on the Rabbits.

### Replacing a Drive

If a drive fails and must be replaced, the Rabbit-p and Rabbit-s should be powered off to replace the drive. After replacing the drive and powering the Rabbit back on, the `nnf-node-manager` pod restarts on the Rabbit. On initialization, Rabbit software locates the new drive and adds an NVMe namespace for each allocation. If rebuild commands have been specified, the Rabbit software will run them to add the new NVMe namespace into the RAID device and restore the correct data.

## Allocation Status

If the allocation is a `PersistentStorageInstance`, the status of the RAID device is available in the `Status.State` field. If all the RAIDs in the persistent storage are healthy, `Status.State=Active`. If any of the RAIDs are degraded, `Status.State=Degraded`.
73 changes: 73 additions & 0 deletions docs/guides/storage-profiles/readme.md
Original file line number Diff line number Diff line change
Expand Up @@ -286,6 +286,79 @@ The `count` field may be useful when creating a persistent file system since the

In general, `scale` gives a simple way for users to get a filesystem that has performance consistent with their job size. `count` is useful for times when a user wants full control of the file system layout.

### RAID Configurations

Allocations can be set up to use a RAID device to provide continued access in the event of a drive failure. Optionally, commands can be specified to rebuild the RAID device after a replacement drive has been added. The storage profile parameters differ depending on whether the allocation is using LVM or zpool.

#### Zpool

To create a Lustre confiuration with a redundant zpool, the `raidz` option is required in `zpoolCreate` command. To allow the zpool to be rebuilt with a new drive, the `zpoolReplace` command is required.

The example below shows a RAID configuration for the OST, however, the same options can be specified for any of the Lustre targets.

```yaml
apiVersion: nnf.cray.hpe.com/v1alpha8
kind: NnfStorageProfile
metadata:
name: lustre-raid-example
namespace: nnf-system
data:
[...]
lustreStorage:
[...]
ostCommandlines:
mkfs: --ost --backfstype=$BACKFS --fsname=$FS_NAME --mgsnode=$MGS_NID --index=$INDEX
--mkfsoptions="nnf:jobid=$JOBID" $ZVOL_NAME
mountTarget: $ZVOL_NAME $MOUNT_PATH
postActivate:
- mountpoint $MOUNT_PATH
zpoolCreate: -O canmount=off -o cachefile=none $POOL_NAME raidz $DEVICE_LIST
zpoolReplace: $POOL_NAME $OLD_DEVICE $NEW_DEVICE
```

#### LVM

A RAID logical volume can be used with XFS and Raw allocations.
NOTE: gfs2 allocations cannot use RAID logical volumes because the LV is shared.

To create a redundant LV, `--type raid[x]` and `--nosync` should be specified in the `lvcreate` command. Also, the `--stripes` parameter should be adjusted accordingly to specify the number of data stripes. For `raid5`, `$DEVICE_NUM-1` is used.

To allow the LV to rebuild after a drive is replaced, `vgExtend`, `lvRepair`, and `vgReduce` should be specified in the `lvmRebuild` section.

```yaml
apiVersion: nnf.cray.hpe.com/v1alpha8
kind: NnfStorageProfile
metadata:
name: xfs-raid-example
namespace: nnf-system
data:
[...]
xfsStorage:
commandlines:
lvChange:
activate: --activate y $VG_NAME/$LV_NAME
deactivate: --activate n $VG_NAME/$LV_NAME
lvmRebuild:
vgExtend: $VG_NAME $DEVICE
vgReduce: --removemissing $VG_NAME
lvRepair: $VG_NAME/$LV_NAME
lvCreate: --activate n --zero n --nosync --type raid5 --extents $PERCENT_VG --stripes $DEVICE_NUM-1
--stripesize=32KiB --name $LV_NAME $VG_NAME
lvRemove: $VG_NAME/$LV_NAME
mkfs: $DEVICE
mountCompute: $DEVICE $MOUNT_PATH
mountRabbit: $DEVICE $MOUNT_PATH
postMount:
- chown $USERID:$GROUPID $MOUNT_PATH
pvCreate: $DEVICE
pvRemove: $DEVICE
sharedVg: true
vgChange:
lockStart: --lock-start $VG_NAME
lockStop: --lock-stop $VG_NAME
vgCreate: --shared --addtag $JOBID $VG_NAME $DEVICE_LIST
vgRemove: $VG_NAME

## Command Line Variables

### global
Expand Down
3 changes: 2 additions & 1 deletion mkdocs.yml
Original file line number Diff line number Diff line change
@@ -1,4 +1,4 @@
# Release: 0.1.22
# Release: 0.1.23
site_name: NNF
site_description: 'Near Node Flash'
docs_dir: docs/
Expand Down Expand Up @@ -28,6 +28,7 @@ nav:
- 'Switch a Node From Worker to Master': 'guides/node-management/worker-to-master.md'
- 'Directive Breakdown': 'guides/directive-breakdown/readme.md'
- 'System Storage': 'guides/system-storage/readme.md'
- 'Redundant Allocations': 'guides/redundant-allocations/readme.md'
- 'Repo Guides':
- 'Releasing NNF Software': 'repo-guides/release-nnf-sw/release-all.md'
- 'CRD Version Bumper': 'repo-guides/crd-bumper/readme.md'
Expand Down