Disaster Recovery Plan
This section describes the disaster recovery options available for AppViewX deployments. It explains different recovery approaches to protect data and restore services during outage scenarios. Customers can choose the approach that best aligns with availability, data protection, and infrastructure requirements.
Approach 1 - VM Snapshot Based Recovery
- As part of this disaster recovery approach, periodic snapshots of the virtual machines hosting AppViewX nodes are taken at a configurable interval, with a recommended default of every 2 hours.
- The snapshot interval can be customized by the customer based on business requirements, infrastructure capacity, and acceptable data loss tolerance.
- In a cluster failure or complete outage, the most recent successful snapshot can be used to restore the system to a previously stable state. VM snapshots must be taken for all nodes and restored together during disaster recovery.
- Based on the configured snapshot interval, data loss can occur up to the duration of the last backup interval (for example, up to 2 hours when snapshots are taken every 2 hours).
- This approach provides fast recovery with minimal turnaround time, enabling the AppViewX cluster to return online quickly.
- The backup and restore mechanism is implemented at virtual machine level. The customer infrastructure team is responsible for snapshot scheduling, retention, storage, monitoring, and restore operations.
Approach 2 - VM Snapshot with Application-Level Backup Restore
Prerequisites
- Regular application-level backups of MongoDB and Vault must be configured using scheduled cron jobs, with frequency based on customer requirements.
- MongoDB and Vault backup files must be securely transferred and stored in a customer-managed remote location.
- The remote location must be reachable during recovery and must provide sufficient retention and storage capacity.
Recovery Approach
- Virtual machine snapshots are taken at longer intervals (for example weekly, monthly, or after patching or upgrade), or a single known-good snapshot is retained.
- In a cluster failure or complete outage, the VM snapshot restores the cluster to a working but potentially outdated state.
- To recover the latest data, stored MongoDB and Vault backups are restored on top of the recovered cluster, bringing it closer to the most recent pre-disaster state.
- VM snapshot backup and remote storage are managed by the customer infrastructure team, while MongoDB and Vault restoration is done using AppViewX utilities. VM snapshots must be taken for all nodes and restored together.
Approach 3 - New Cluster Rebuild with Data Restore
Prerequisites
- Regular application-level backups of MongoDB and Vault must be configured using scheduled cron jobs, with frequency based on customer requirements.
- MongoDB and Vault backup files must be securely transferred and stored in a customer-managed remote location.
- The remote location must be reachable during recovery and must provide sufficient retention and storage capacity.
Recovery Approach
- If the existing cluster is unavailable or unrecoverable, install a new AppViewX cluster as described in the installation guide.
- After deployment and validation of the new cluster, restore MongoDB and Vault backups from remote storage.
- Reapply all customer-specific configurations and customizations manually, including ConfigMaps, environment variables, integrations, certificates, and other environment-specific changes.
- This approach involves full rebuild and reconfiguration, so it has the highest turnaround time among disaster recovery options.
- Use this approach for IP migration or infrastructure changes where AppViewX runs on nodes with new IP addresses, making snapshot-based recovery unsuitable.
Approach 4 - Active-Standby Cluster with MongoSync utility
Prerequisites
- Deploy active-standby AppViewX clusters across two separate data centers.
- Provision a primary active cluster and a secondary standby cluster with network connectivity between them.
- Install and configure MongoSync to synchronize data from active MongoDB to standby MongoDB.
- Open required network ports and firewall rules between both data centers.
Recovery Approach
- AppViewX runs in active-standby mode, with production traffic served by active cluster and standby kept ready.
- MongoSync continuously replicates active MongoDB data to standby MongoDB.
- Vault is configured once and key configuration remains unchanged across both data centers, with no key rotation or reinitialization required during failover.
- If active data center is impacted, traffic is switched to standby cluster using synchronized MongoDB data and same Vault configuration.
- This approach provides minimal data loss and faster recovery, depending on MongoSync synchronization lag.
- Because the standby cluster is already deployed, recovery time is significantly lower than rebuild or backup-restore approaches.
