Oracle RAC 05 - Storage and ASM disks
Oracle RAC Solution · Previous: Operating system preparation · Next: Grid Infrastructure installation
A RAC database is one set of files opened by two instances at once, so everything the database stores has to be on disks that both nodes see, under names and permissions that are the same on both and survive a reboot. This Article covers the storage of the nodes from the bottom up: one local volume per node for the software, five shared iSCSI LUNs, the udev rules that give them stable names and the right owner, and ASMLib, which stamps them as ASM disks. The disk groups themselves are created later by the installers. Everything here ran as root.
What each node sees
| Name | Capacity | Mount point | Purpose | Count |
|---|---|---|---|---|
ORADB-INSTALL | 107.4 GB | /data | Oracle install destination | one per server, two in all |
ORADB-CRS | 1073 MB | none | Cluster Registry (CRS) | one, shared by both servers |
ORADB-DATA01 | 107.4 GB | none | Oracle data disk 1 | one, shared |
ORADB-DATA02 | 107.4 GB | none | Oracle data disk 2 | one, shared |
ORADB-FRA | 214.7 GB | none | Fast Recovery Area | one, shared |
ORADB-ONTRLG | 21.5 GB | none | Online transaction logs | one, shared |
On both nodes the kernel presents them identically. The LUN names at the end of each line are my annotation in the notes, not part of the output, and the fence leaves out what ls -ld prints in front of each path: permissions, link count, owner, group, size and date.
$ ls -ld /sys/block/sd*/device
output 7 lines
/sys/block/sda/device -> ../../../2:0:0:0 /sys/block/sdb/device -> ../../../3:0:1:0 -> ORADB-INSTALL /sys/block/sdc/device -> ../../../5:0:0:0 -> ORADB-CRS /sys/block/sdd/device -> ../../../5:0:0:1 -> ORADB-DATA01 /sys/block/sde/device -> ../../../5:0:0:2 -> ORADB-DATA02 /sys/block/sdf/device -> ../../../5:0:0:3 -> ORADB-FRA /sys/block/sdg/device -> ../../../5:0:0:4 -> ORADB-ONTRLG
sda is the system disk, sdb the install disk on a SCSI host of its own, and the five shared LUNs sit on SCSI host 5 as LUN 0 to 4. They arrive over the iSCSI network on eth4 (see Network and DNS plan). How they got there is not in my notes: there is no iSCSI initiator configuration, no target configuration and nothing about the Hyper-V side. The output for oradb02 in the notes is identical to that of oradb01 down to the timestamps, so one of the two is a copy.
The local volume /data
The software homes, the inventory and the unpacked installation media live under /data, a volume that each node has for itself. The whole of sdb became an LVM physical volume, without a partition table, in a volume group named DATA.
$ pvcreate /dev/sdb $ vgcreate DATA /dev/sdb $ vgdisplay -v DATA $ lvcreate -L 99G -n data DATA $ mkfs.ext3 /dev/DATA/data $ mkdir -p /data $ mount -a
Between mkfs.ext3 and mkdir the mount was added to /etc/fstab:
# Oracle Install Destination /dev/mapper/DATA-data /data ext3 nodev,acl,user_xattr 0 0
The logical volume is 99 GB on a 107.4 GB disk; the notes do not say what the rest was kept for. The file system is ext3 like every other one on these SLES 11 systems, and the last two fields are zero, so /data is neither dumped nor checked at boot, unlike the system volumes. Do not confuse the volume group DATA with the ASM disk group DATA further down: the first is local LVM, the second is shared and belongs to ASM. The whole file is a Config document: fstab.
From LUN to disk group
flowchart LR subgraph lun["iSCSI LUN"] l1["ORADB-CRS, 1073 MB"] l2["ORADB-DATA01, 107.4 GB"] l3["ORADB-DATA02, 107.4 GB"] l4["ORADB-FRA, 214.7 GB"] l5["ORADB-ONTRLG, 21.5 GB"] end subgraph blk["Kernel device, partition 1"] s1["sdc1"] s2["sdd1"] s3["sde1"] s4["sdf1"] s5["sdg1"] end subgraph ud["udev symlinks, grid:asmadmin 0660"] u1["/dev/ORADB-CRS and /dev/ASMDISK1"] u2["/dev/ORADB-DATA01 and /dev/ASMDISK2"] u3["/dev/ORADB-DATA02 and /dev/ASMDISK3"] u4["/dev/ORADB-FRA and /dev/ASMDISK4"] u5["/dev/ORADB-ONTRLG and /dev/ASMDISK5"] end subgraph al["ASMLib disk"] a1["ORCL:ASMDISK1"] a2["ORCL:ASMDISK2"] a3["ORCL:ASMDISK3"] a4["ORCL:ASMDISK4"] a5["ORCL:ASMDISK5"] end subgraph dg["ASM disk group"] g1["OCR"] g2["DATA"] g3["FRA"] g4["ONTRLG"] end l1 --> s1 --> u1 --> a1 --> g1 l2 --> s2 --> u2 --> a2 --> g2 l3 --> s3 --> u3 --> a3 --> g2 l4 --> s4 --> u4 --> a4 --> g3 l5 --> s5 --> u5 --> a5 --> g4
Everything up to the ASMLib disk is built in this Article. The disk groups are not: OCR is created by the Grid Infrastructure installer and receives the cluster registry and the voting files (Grid Infrastructure installation), and DATA, FRA and ONTRLG are created with ASMCA before the database (Database software, listener and disk groups). All four have external redundancy: ASM mirrors nothing and trusts the storage underneath.
Partitions
Each shared device got one primary partition of type 83 over its whole size. The notes show the fdisk dialogue for sdc and say "the same" for the other four.
$ fdisk /dev/sdcoutput 5 lines
Command (m for help): n
Command action (e=extended, p=primary): p
Partition number (1-4, default 1): 1
First sector (2048-2095103, default 2048): ENTER
Last sector, +sectors or +size{K,M,G} (2048-2095103, default 2095103): ENTERThe dialogue stops there in the notes; the w that writes the table is not shown. Partitioning happens once, on one node, because the disks are shared. The other node still holds the old, empty partition tables in memory, so both nodes were told to read them again:
$ partprobeASM itself does not need a partition table. The notes give no reason for creating one, but the udev rules below match on the partition, so from here on the partition is the ASM disk.
Stable names and ownership with udev
sdc is a name the kernel gives in the order of discovery. Nothing guarantees that the cluster registry disk is sdc on both nodes or after the next boot, and a device node that udev creates belongs to root:disk, which the grid account cannot open. Both problems are solved by udev rules keyed on the one thing that does not change, the SCSI identifier of the LUN.
$ /lib/udev/scsi_id -g -u /dev/sdcoutput 1 line
14f504e46494c45526469736b30312d303030302d30303031
The command was run for sdc to sdg on both nodes, and the five identifiers were the same on both, which is the proof that the two nodes really look at the same five LUNs. All five begin with 14f504e46494c4552; after the leading 1, that is the ASCII string OPNFILER in hexadecimal. It is the only thing my notes say about the storage system, and they say it only by accident.
The rules file is /etc/udev/rules.d/59-persistent-disk.rules, with two rules per disk:
# The first partition of the block device with ID "14f504e46494c45526469736b30312d303030302d30303031" will be called "ORADB-CRS" and "ASMDISK1". KERNEL=="sd?1", SUBSYSTEM=="block", PROGRAM=="/lib/udev/scsi_id -g -u -d /dev/$parent", RESULT=="14f504e46494c45526469736b30312d303030302d30303031", SYMLINK+="ASMDISK1", OWNER="grid", GROUP="asmadmin", MODE="0660" KERNEL=="sd?1", SUBSYSTEM=="block", PROGRAM=="/lib/udev/scsi_id -g -u -d /dev/$parent", RESULT=="14f504e46494c45526469736b30312d303030302d30303031", SYMLINK+="ORADB-CRS", OWNER="grid", GROUP="asmadmin", MODE="0660"
| Part of the rule | Meaning |
|---|---|
KERNEL=="sd?1", SUBSYSTEM=="block" | applies to the first partition of any SCSI disk |
PROGRAM=="/lib/udev/scsi_id -g -u -d /dev/$parent" | asks the parent disk, not the partition, for its identifier |
RESULT=="…" | the rule goes on only if the identifier is this one |
SYMLINK+="ASMDISK1", SYMLINK+="ORADB-CRS" | two additional names in /dev, one per rule |
OWNER="grid", GROUP="asmadmin", MODE="0660" | the device node itself changes owner, group and mode |
The result is that /dev/ORADB-CRS and /dev/ASMDISK1 both point to whatever sdX1 carries that LUN today, and that the node behind them is readable and writable for grid and for the group asmadmin, the owner and group Oracle ASM requires. The file is identical on both nodes because the identifiers are. It is a Config document: udev rules for the ASM disks. The notes do not show how the rules were loaded (a udevadm trigger or a reboot), nor a listing of /dev that proves the result.
The multipath variant that was not used
The nodes reach the storage over one iSCSI interface, so there is one path per LUN and no multipath layer. The notes nevertheless keep the rules "were an MPIO layer in use": with DM multipath the stable name is the name of the multipath map, and the rule only has to set the owner.
ENV{DM_NAME}=="ORADB-CRS", OWNER:="grid", GROUP:="asmadmin", MODE:="660"That variant is a Config document of its own, udev rules for DM multipath, not used, with the warnings that belong to it: its device names do not match this system (it has four data disks named ORADB-DATA1 to ORADB-DATA4), and it was not used here.
ASMLib
ASMLib is Oracle's kernel driver and library for marking disks as ASM disks and finding them again by label. With it, ASM is given the discovery string ORCL:* and does not care about device names at all. The packages were installed in Operating system preparation; the commands with their complete output are a Config document: ASMLib commands.
| Step | Command | Where | What it does |
|---|---|---|---|
| 1 | oracleasm configure -i | both nodes | driver interface owned by grid:asmadmin, start and scan at boot |
| 2 | oracleasm init | both nodes | loads the module, mounts /dev/oracleasm |
| 3 | oracleasm createdisk ASMDISKn /dev/ORADB-… | one node only | writes the ASMLib label to the partition |
| 4 | oracleasm scandisks, oracleasm listdisks | both nodes | finds the labelled disks, lists them |
| 5 | oracleasm-discover 'ORCL:*' | both nodes | shows what ASM will see through the library |
The labels are written once, on one node, through the names udev provides:
$ oracleasm createdisk ASMDISK1 /dev/ORADB-CRS $ oracleasm createdisk ASMDISK2 /dev/ORADB-DATA01 $ oracleasm createdisk ASMDISK3 /dev/ORADB-DATA02 $ oracleasm createdisk ASMDISK4 /dev/ORADB-FRA $ oracleasm createdisk ASMDISK5 /dev/ORADB-ONTRLG
The other node finds them with scandisks, and listdisks must then print the same five names, ASMDISK1 to ASMDISK5, on both. The last check runs the discovery the way the Grid Infrastructure installer will:
$ oracleasm-discover 'ORCL:*'
output 7 lines
Using ASMLib from /opt/oracle/extapi/64/asm/orcl/1/libasm.so [ASM Library - Generic Linux, version 2.0.4 (KABI_V2)] Discovered disk: ORCL:ASMDISK1 [2095104 blocks (1072693248 bytes), maxio 128] Discovered disk: ORCL:ASMDISK2 [419428352 blocks (214747316224 bytes), maxio 128] Discovered disk: ORCL:ASMDISK3 [209713152 blocks (107373133824 bytes), maxio 128] Discovered disk: ORCL:ASMDISK4 [209713152 blocks (107373133824 bytes), maxio 128] Discovered disk: ORCL:ASMDISK5 [41940992 blocks (21473787904 bytes), maxio 128]
The names ASMDISK1 to ASMDISK5 therefore exist twice: as udev symlinks in /dev and as ASMLib labels. ASM uses only the second; the ASMDISKn symlinks are not referenced anywhere else in the notes.
Where the notes disagree with themselves
The order of the chapters. The notes describe ASMLib in chapter 5 and udev in chapter 6, but createdisk in chapter 5 writes to /dev/ORADB-CRS and its siblings, names that exist only once the rules of chapter 6 are in place. The udev chapter was done first. The rules in turn match partitions, so the order that works is: partition the disks, read the identifiers and write the rules, then ASMLib. The partitioning step for the other four disks also says "the same as in step [6.1]", where the fdisk step is [5.1]; the chapters were renumbered at some point.
The sizes. The table of block devices and ASMCA agree with each other; the oracleasm-discover output does not agree with them.
| ASMLib disk | LUN by the table | Size in the table | oracleasm-discover | ASMCA, later |
|---|---|---|---|---|
ASMDISK1 | ORADB-CRS | 1073 MB | 1072693248 bytes | used by the installer for OCR |
ASMDISK2 | ORADB-DATA01 | 107.4 GB | 214747316224 bytes | 102399 MB, in DATA |
ASMDISK3 | ORADB-DATA02 | 107.4 GB | 107373133824 bytes | 102399 MB, in DATA |
ASMDISK4 | ORADB-FRA | 214.7 GB | 107373133824 bytes | 204799 MB, in FRA |
ASMDISK5 | ORADB-ONTRLG | 21.5 GB | 21473787904 bytes | 20479 MB, in ONTRLG |
By the discovery output ASMDISK2 is the 214 GB disk and ASMDISK4 one of the 107 GB disks; by everything else it is the other way round. The notes do not explain it, and I do not correct either side. A smaller oddity is in the same output: ASMDISK1 has 2095104 blocks, which is exactly the number of sectors of the whole device in the fdisk dialogue, not of a partition that starts at sector 2048.
What I would do differently
- One mechanism, not two. Oracle documents device persistence with ASMLib and device persistence with udev rules as alternatives. With ASMLib the access rights come from the driver interface (
grid:asmadmin, set byoracleasm configure), so the owner, group and mode in my udev rules were redundant; the rules served only to givecreatediskstable names, and theASMDISKnsymlinks served nothing. - Write down the checks. A listing of
/dev/ORADB-*on both nodes and the loading of the rules are missing from the notes, and they are what would have settled the size question above. - A larger disk for
OCR. A 1 GB disk was above the 11.2 minimum of 300 MB each for a voting file and an OCR. It is below every current minimum: Oracle AI Database 26ai asks for 2 GB with external redundancy. - External redundancy means trusting one box. Oracle recommends external redundancy for storage that protects the data itself. All four disk groups, including the single voting file, therefore depended on the one storage system behind the LUNs, and my notes say nothing about how that system was protected.
ASMLib itself is still current: version 3 uses io_uring and needs no kernel module, and the discovery string is still ORCL:*. The rules file would work today with one change, the path /usr/lib/udev/scsi_id. Details are in the "Checked against" sections of the Config documents.