LINUXOR.SK ... open source notes ...

Balabit - Manual NTP synchronisation

category: solutionz · date: 2018-12-31 · updated: 2026-10-03 · author: LALA

Balabit SCB Solution · Config document · referenced from High availability and Operations, upgrades and troubleshooting

noteThe procedure stops the NTP daemon and lets ntpd -g set the clock in one step, however large the offset. On an appliance that timestamps audit trails, a jump of the clock shows in every trail recorded around it; do it when no sessions are running. The notes keep the commands only, no output, and do not record whether, on which node or on which date they were run.

The vendor's support sent this procedure to bring the clock of a node back in line. The web interface has a banner for the situation it addresses, "Slave is out of sync with the master" under Basic Settings > Date & Time, and I kept a screenshot of it, but the notes do not say that the two belong together. The procedure measures the offset between the two nodes, looks at the NTP peers, stops ntpd, synchronises once by hand, starts ntpd again and measures once more.

ItemValue
WhereSSH console of the SCB as root, boot shell ("SSH -> BOOT shell")
Which nodeNot recorded; scb-other is the other node of the HA pair, reached over SSH from the boot shell
Which clusterNot recorded; the banner screenshot is kept with my operation notes, the commands with the troubleshooting notes
NTP serversThe Infoblox pair of site 1, 10.11.16.145 and 10.11.18.145: chapter 7 of the design for site 1, the config.xml for site 2
SourceTroubleshooting notes, the support's answer "Manual NTP synchronization."

The commands

The commands as the support gave them, to be run as root in the boot shell of a node, in the support's order. The comments that the support put before each group of lines were: "Actual time offset.", "NTP queries.", "Stop ntpd", "Execute a manual synchronisation.", "Start ntpd", "A new time offset measurement."

bash
$ date -R; ssh scb-other date -R;date -R
$ watch -n 1 'date -R; ssh scb-other date -R;date -R'
$ ntpq -c peers; ssh scb-other ntpq -c peers
$ ntpq -c sysinfo
$ systemctl stop ntp
$ ntpd -q -g -l /tmp/ntp.manual_connect.log
$ systemctl start ntp
$ date -R; ssh scb-other date -R;date -R
$ watch -n 1 'date -R; ssh scb-other date -R;date -R'

The options are explained from my general knowledge of the tools; the notes give the commands only.

LineWhat it does
date -R; ssh scb-other date -R;date -RLocal time, the other node's time, local time again. The two local readings bracket the remote one, so the offset can be read without a time server; -R prints RFC 2822 format with seconds and zone
watch -n 1 '…'The same every second, to see whether the offset stays or drifts
ntpq -c peersThe peer table of ntpd on this node and, over SSH, on the other. A peer with reach 0 and refid .INIT. has never answered
ntpq -c sysinfoThe daemon's own state: stratum, reference, offset
systemctl stop ntpThe service is called ntp on the boot firmware; ntpd must not run while it is started once more by hand
ntpd -q -g -l /tmp/ntp.manual_connect.log-q sets the clock and quits, -g allows the first correction to exceed the panic threshold of 1000 seconds, -l writes the log to the file
systemctl start ntpBack to normal operation

What the support bundle of the site 2 cluster showed on 2018-09-17 fits the symptom: the master synchronised from both Infoblox servers of site 1, while on the slave both of them were .INIT. with reach 0, and the master, listed there as scb1, had reach 0 as well. The slave had no source it could reach. The listing is in HA state of the site 2 cluster. The dates do not show that the banner, the support's procedure and that bundle belong to the same event.

Checked against One Identity Safeguard for Privileged Sessions 9.0

As builtToday
Master and slave nodePrimary and secondary node ("previously also referred to as the master node and the slave node")
Basic Settings > Date & Time with NTP serversThe page remains, and the alert xcbTimeSyncLost is still in the alert list
← solutionz