High availability recommendations

These recommendations make
FactoryTalk® Optix™
High Availability clusters easier to deploy, validate, and recover.

Cluster foundation

Follow the supported cluster layout and network design.
  • Use exactly three dedicated nodes on separate machines, one reserved Virtual IP address, and one shared Layer 2 segment or broadcast domain.
  • Ensure all three nodes meet the documented system requirements.
  • Keep all nodes on the same dedicated VLAN or other isolated Layer 2 segment, and keep Keepalived VRRP traffic on that segment.
  • Keep node time synchronized. Provisioning requires node clocks to differ by no more than 1 second, and a difference of 5 seconds or more is not recommended.

Client access and data continuity

Keep one stable client endpoint, and store runtime data where failover preserves it.
  • Use the Virtual IP as the only endpoint for browser clients, OPC UA clients, MQTT clients, and
    FactoryTalk Optix
    Studio deployments.
  • Keep project ports aligned with cluster ports. For example, the web port in the cluster must match the
    WebPresentationEngine
    port in the project.
    Preconfigured ports
    Port name
    Default port
    Web presentation engine (http)
    80
    Web presentation engine (https)
    443
    OPC UA server (opcua)
    59100
    MQTT broker (mqtt)
    8883
  • Configure retentivity storage for runtime changes that must persist. When you add the HighAvailability module, retentivity storage uses the shared cluster database automatically.
  • Store other failover-critical data outside the application files folder. Runtime-generated files written there do not synchronize across nodes, embedded database synchronization is not supported, and logging to an external database is the documented design.
  • During failover, the runtime application stops running temporarily. Operations such as data logging also stop. During the node transition, temporary connection loss between the web client and the runtime application occurs. The failover time varies depending on failure type and application size.
  • Keep OPC UA identity and trust stable across failover. Supply a server certificate and private key in the project archive, and pre-populate the trusted store with expected client certificates.

Security and administration

Harden nodes before cluster installation, and limit exposure to trusted networks.
  • Enable disk encryption on every node before cluster installation. Without disk encryption, persistent volumes under
    /var/lib/rancher/k3s/storage
    remain plaintext.
  • Use SSH key-based authentication only, limit SSH access to the management network, keep automatic security updates enabled, disable unnecessary services and packages, keep the audit daemon running, protect K3s data and agent credentials, and apply operating system and K3s security patches on each node.
  • Restrict firewall rules to required ports from trusted CIDRs. Block the NodePort range from external sources, and block other inbound traffic by default.
  • Use the documented secure settings on exposed services. Configure the Web Presentation Engine for HTTPS, use TLS when MQTT is used, and enforce an OPC UA security mode on the Runtime OPC UA Server.

Diagnostics and recovery

Confirm the real cluster state before you change the system, and protect data before disruptive actions.
  • Confirm cluster problems with diagnostics and exported logs before you make corrective changes.
  • Back up persistent data before you uninstall or rebuild the cluster. Uninstalling the cluster deletes persistent storage, project data, and database data.
Provide Feedback
Have questions or feedback about this documentation? Please submit your feedback here.
Normal