High Availability cluster troubleshooting

This table lists common cluster symptoms and resolutions.
Symptom
Resolution
SSH connectivity fails between the Windows workstation and an Ubuntu node.
Correct the SSH credentials or network configuration so the workstation can reach each node.
Web UI access is blocked because the cluster and the deployed project use different ports.
Align the project WebPresentationEngine port with the cluster configured web port, default 80 for HTTP or 443 for HTTPS.
The database stops accepting writes.
This issue occurs when two of the three nodes fail, and the database loses quorum.
Restore at least one failed node to establish quorum and resume writes.
Persistent data is missing after uninstalling the cluster.
Uninstalling the cluster removes persistent storage.
Recover data only from a backup created before uninstall.
The cluster dashboard displays a status that seems outdated or unexpected.
Refresh the dashboard.
If the status still looks wrong, check Diagnostics and exported logs to confirm the actual pod and node status.
Deploying a large project that exceeds the runtime's memory limit causes the deployment to report success in Studio, but the runtime immediately shuts down with little information logged.
  1. To back up current Helm values and revision history, run:
    helm get values ftoptix -n ftoptix > pre-change-values-backup.yaml && helm history ftoptix -n ftoptix
  2. To confirm node RAM can support the new memory limit, run:
    kubectl describe node <node-name> | grep -A2 "Capacity:"
  3. To apply the override, run:
    helm upgrade ftoptix <chart-path> -n ftoptix --reuse-values --set resources.runtime.requests.memory=<value> --set resources.runtime.limits.memory=<value>
  4. To verify the pods picked up the new limit and are healthy, run:
    kubectl rollout status statefulset/ftoptix-core-ha-runtime -n ftoptix kubectl get pods -n ftoptix kubectl describe pod ftoptix-core-ha-runtime-0 -n ftoptix | grep -A2 "Limits:"
  5. To confirm the change, deploy the project again, confirm the runtime reaches a running or active state, and confirm the Web UI is reachable.
Kubernetes returns 401 Unauthorized, runtime identity retrieval fails, the runtime enters demo mode, Update Server authorization fails intermittently, or runtime pods do not recover as expected.
These symptoms commonly result from clock skew between cluster nodes.
Configure all nodes to use the same approved NTP sources, restore time synchronization, confirm clocks differ by no more than 1 second, and retry the failed operation.
Do not continue provisioning if any node is unsynchronized or differs by 5 seconds or more.
Provide Feedback
Have questions or feedback about this documentation? Please submit your feedback here.
Normal