Oracle RAC 02 - Network and DNS plan
Oracle RAC Solution · Previous: Overview and design · Next: DNS server: Knot
A RAC cluster is unforgiving about names and networks. The Grid Infrastructure installer wants a public name and a virtual name per node, a SCAN name that resolves to three addresses, and a private network between the nodes that it can use for the cluster heartbeat. All of that has to be decided, and resolvable, before the installer starts. This Article is the plan: which interface does what, every name with its address, how the names are resolved, and why one of the five interfaces ended up without a job.
Five interfaces, five roles, one interface unused
Each node has the interfaces eth0 to eth4. My notes describe the layout for oradb01; the second node follows it with the host part .12.
| Interface | Role | Address on oradb01 | Name |
|---|---|---|---|
eth0 | Backup network | 10.30.30.11 | oradb01-bck.example.net |
eth1 | Production network, the public network of the cluster | 10.30.10.11 | oradb01.example.net |
eth2 | Unused, because of the Hyper-V switches | none | none |
eth3 | Management network | 10.30.40.11 | oradb01-mng.example.net |
eth3:INT | Oracle RAC interconnect, an alias on eth3 | 10.30.20.11 | oradb01-int.example.net |
eth4 | iSCSI SAN | 10.30.50.11 | oradb01-iscsi.example.net |
flowchart LR subgraph n1["oradb01"] e0["eth0, 10.30.30.11"] e1["eth1, 10.30.10.11"] e2["eth2, no address"] e3["eth3, 10.30.40.11"] e3i["eth3 alias INT, 10.30.20.11"] e4["eth4, 10.30.50.11"] e3 --- e3i end bck(["Backup 10.30.30.x"]) pub(["Public 10.30.10.x"]) cloud(["Hyper-V switch on the Cloud Root Network"]) mng(["Management 10.30.40.x"]) int(["Interconnect 10.30.20.x"]) san(["iSCSI SAN 10.30.50.x"]) e0 --- bck e1 --- pub e2 -. "no multicast, no broadcast, not used" .- cloud e3 --- mng e3i --- int e4 --- san bck --- bsrv["mng-backupsrv01-bck, 10.30.30.14"] mng --- dns["dns1, 10.30.40.13"] mng --- bsrvm["mng-backupsrv01, 10.30.40.14"] pub --- vip["VIPs .21 and .22, SCAN .31 to .33"] int --- peer["oradb02-int, 10.30.20.12"]
The notes never state a network mask. The reverse zones are cut at the third octet, and the installer answers below name the subnets as 10.30.10.0 and so on.
Why eth2 is unused
The line for eth2 in my notes reads: unused because of the Hyper-V switches; Hyper-V switches bound to the so-called Cloud Root Network support neither multicast nor broadcast.
eth2 was the interface meant for the private interconnect. The Oracle cluster uses multicast and broadcast for the communication between its nodes. The 11.2 Grid Infrastructure guide says so: from release 11.2.0.2 on, "multicasting is required on the private interconnect", on the ranges 224.0.0.0/24 and 230.0.1.0/24, and its troubleshooting chapter has an entry of its own for root.sh failing on the second node because of multicast, with a test tool, mcasttest.pl. I did not know either at the time. On a virtual switch that passes neither, the first node comes up alone and the second cannot join it: root.sh on the second node failed, and so did the attempt to bring the node into the cluster at all. Before I found the cause I applied two cumulative Grid patches in the belief that they would fix it. They did not. That detour is the subject of Grid patches and the multicast problem.
The outcome is what the table shows. The interconnect subnet 10.30.20.x was put on the management interface as the alias eth3:INT, and eth2 was left without an address. It follows that the switch behind eth3 did pass the traffic the cluster needs; the notes do not say how that switch differed. They also do not describe the move itself: there is no interface configuration file in them and no command that created the alias. Two lines in the kernel parameters are the only trace on the operating system side:
# Here the adapter "eth3" (eth3:INT) is used as the Oracle interconnect net.ipv4.conf.eth3.rp_filter = 0 net.ipv4.icmp_echo_ignore_broadcasts = 0
The first line turns off reverse-path filtering on the interface that now carries two subnets, the second lets the node answer broadcast pings. The whole file is a Config document of Operating system preparation: sysctl.conf, Oracle block.
The price of the workaround is plain: the private traffic of the cluster shares one virtual interface with the management traffic, and with DNS and NTP, because dns1 is reached over the same interface.
Names and addresses
The public network carries three kinds of addresses. The host addresses are configured on eth1. The VIPs and the SCAN addresses are not configured by me anywhere on the nodes: Clusterware brings them up on the public interface of whichever node holds them.
| Kind | Name | Address | FQDN |
|---|---|---|---|
| Public | oradb01 | 10.30.10.11 | oradb01.example.net |
| Public | oradb02 | 10.30.10.12 | oradb02.example.net |
| Virtual (VIP) | oradb01-vip | 10.30.10.21 | oradb01-vip.example.net |
| Virtual (VIP) | oradb02-vip | 10.30.10.22 | oradb02-vip.example.net |
| SCAN | oradb-scan | 10.30.10.31 | oradb-scan.example.net |
| SCAN | oradb-scan | 10.30.10.32 | oradb-scan.example.net |
| SCAN | oradb-scan | 10.30.10.33 | oradb-scan.example.net |
The other three networks of the nodes have one name per node, built from the host name and a suffix for the role.
| Network | Name | Address | FQDN |
|---|---|---|---|
| Interconnect | oradb01-int | 10.30.20.11 | oradb01-int.example.net |
| Interconnect | oradb02-int | 10.30.20.12 | oradb02-int.example.net |
| Backup | oradb01-bck | 10.30.30.11 | oradb01-bck.example.net |
| Backup | oradb02-bck | 10.30.30.12 | oradb02-bck.example.net |
| Management | oradb01-mng | 10.30.40.11 | oradb01-mng.example.net |
| Management | oradb02-mng | 10.30.40.12 | oradb02-mng.example.net |
| iSCSI SAN | oradb01-iscsi | 10.30.50.11 | oradb01-iscsi.example.net |
The iSCSI row has no partner. The notes give an iSCSI address only for oradb01, in the interface layout and in the DNS zone; the second node sees the same shared disks, so it must have had one, but it is written nowhere and so it is not in this table either.
The servers around the cluster:
| Server | Network | Name | Address |
|---|---|---|---|
| Backup Exec server | Management | mng-backupsrv01.example.net | 10.30.40.14 |
| Backup Exec server | Backup | mng-backupsrv01-bck.example.net | 10.30.30.14 |
| DNS and NTP server | Management | dns1.example.net | 10.30.40.13 |
One more name was added to DNS later, for the backup: rac-clsopdb-1234567890, with two address records pointing at the backup addresses of both nodes. It is the virtual RAC node of Backup Exec, explained in Backup Exec server and RAC backup.
What the Grid installer was told
On its "Network Interface Usage" screen the Grid Infrastructure installer lists every subnet it finds on the node and asks for one of three types. These were the answers.
| Interface | Subnet | Interface type |
|---|---|---|
eth0 | 10.30.30.0 | Do Not Use |
eth1 | 10.30.10.0 | Public |
eth2 | none | Do Not Use |
eth3 | 10.30.40.0 | Do Not Use |
eth3 | 10.30.20.0 | Private |
eth4 | 10.30.50.0 | Do Not Use |
The installer shows eth3 twice, once per subnet, which is how an alias appears to it; only the subnet 10.30.20.0 is Private. "Do Not Use" means only that Clusterware does not manage the network. Backup, management and iSCSI keep working as ordinary operating system interfaces. On the screens before it the installer was given the SCAN name oradb-scan with port 1521, no GNS, and the pairs of public and virtual host names from the table above. All answers are in Grid Infrastructure installation.
Name resolution
Every name is in DNS, in the zone example.net on dns1, and the nodes use that server and nothing else:
search example.net nameserver 10.30.40.13
That is the whole /etc/resolv.conf of a node. The search line is what makes the short names work: the tests of password-less SSH in the notes call the other node as ORADB02, ORADB02-int and ORADB02.example.net, and the first two resolve only through the search domain. There is one name server and no second one. dns1 is also an authoritative-only server, so it answers for example.net and the four reverse zones and for nothing outside them; the notes do not say how the nodes resolved other names, for example when the ASMLib packages were downloaded from Oracle's servers earlier in the preparation.
flowchart LR node["oradb01 or oradb02, resolv.conf"] -->|"query to 10.30.40.13"| dns["dns1, zone example.net"] dns --> scan["oradb-scan"] dns --> v1["oradb01-vip"] dns --> v2["oradb02-vip"] scan --> s1["10.30.10.31"] scan --> s2["10.30.10.32"] scan --> s3["10.30.10.33"] v1 --> a1["10.30.10.21"] v2 --> a2["10.30.10.22"] s1 --> cl["SCAN VIPs and SCAN listeners 1 to 3, port 1521, placed by Clusterware"] s2 --> cl s3 --> cl cl --> p1["two of them on oradb01, one on oradb02 at the check"] a1 --> h1["ora.oradb01.vip, online on oradb01"] a2 --> h2["ora.oradb02.vip, online on oradb02"]
The SCAN is one name with three address records. Without GNS, which was not configured, the three records have to exist in DNS before the installation. After it, Clusterware runs three SCAN VIPs with a SCAN listener each and spreads them over the nodes: at the status check after the installation ora.scan1.vip was on oradb02, ora.scan2.vip and ora.scan3.vip on oradb01. The output does not say which of the three addresses belongs to which resource. Each node VIP was online on its own node. The notes hold no client configuration and no failover test, so how clients connected and what happened to a VIP when a node went down is not recorded.
The zone file is a Config document: Zone example.net. The server that holds it is the subject of DNS server: Knot.
The hosts file that was written and not used
The preparation notes also contain a complete /etc/hosts with the same names, and above it the sentence that it is not in use at present because every record is in DNS. It is kept as a Config document: /etc/hosts of the nodes. Its first lines carry a warning that is independent of the rest:
# The loopback must not carry the host name. The Oracle cluster services are unhappy about it # and the cluster will not come up, at least not on the second and later nodes. 127.0.0.1 localhost.localdomain localhost
If the host name of the node stands on the loopback line, the node resolves its own name to 127.0.0.1. The remark is my own experience from this build and not a quotation from Oracle's documentation; the notes do not record the error that led to it.
The file also shows an idea that was dropped. It has commented-out entries oradb01-bck-vip and oradb02-bck-vip (10.30.30.21, 10.30.30.22) and oradb01-mng-vip and oradb02-mng-vip (10.30.40.21, 10.30.40.22), with the comment "in case the Oracle listener also listens on these addresses". They were never activated and are not in DNS.
Had the file been used, its three SCAN lines would have been a mistake. Oracle's guide, then and now, advises against SCAN addresses in the hosts file because a name resolved from that file yields one address only.
One name, two answers
There is one place where a name deliberately resolves differently depending on who asks. In DNS, oradb01.example.net is the public address 10.30.10.11. The notes give the reason as a condition: if the production network is in effect a back-end network for the application servers and not directly reachable from the backup server, the backup server needs addresses it can reach. So the hosts file of the Windows backup server maps oradb01.example.net and oradb02.example.net to the backup addresses 10.30.30.11 and 10.30.30.12. The file and the reason for it are in Backup Exec server and RAC backup.
What I would do differently
The notes themselves point at the weak spots of this plan.
- The interconnect shares an interface with management, DNS and NTP. It was the way out of a platform limit, not a design. A private network of its own, on a virtual switch that passes multicast and broadcast, is what
eth2was for. - One DNS server is a single point of failure for a cluster whose every name is in DNS only. A hosts file on the nodes would have been a fallback; the notes say only that the file they contain is not currently used because every record is in DNS. Oracle's current network checklist wants the public node names in DNS and in
/etc/hosts, both. - Test the private network for multicast before the installation, with Oracle's
mcasttest.pl, and not find out from a failingroot.sh. - The iSCSI network is documented for one node only, and has neither a second forward record nor a reverse zone.