Availability Sets vs. Availability Zones, SLA tiers, backup strategies, and disaster recovery architectures across Azure, AWS, and GCP — the decisions that define your SAP uptime guarantee.
All three hyperscalers use the same SAP-certified HA pattern: two HANA nodes running synchronous system replication, managed by a Pacemaker cluster with automatic failover via a virtual IP.
The difference between deployment models is measured in minutes of annual downtime — and dramatic differences in cost and complexity.
Each cloud implements the same SAP-certified patterns differently — with distinct services, SLAs, and DR mechanisms that materially affect your RPO/RTO.
Standalone HANA instance with Azure auto-restart (service healing). No redundancy — relies on platform fabric controller to detect host failure and redeploy. Kernel panic requires kernel.panic=20 sysctl for automatic restart.
Two HANA VMs deployed across fault domains within the same datacenter, running synchronous HSR with Pacemaker. Protects against rack-level failure (power, networking, disk) but both nodes share the same datacenter — a facility-wide event takes both down.
Two HANA VMs in separate physical datacenters within the same region, each with independent power, cooling, and networking. Synchronous HSR + Pacemaker with Standard Load Balancer health checks for VIP failover. The recommended production deployment.
HANA System Replication across Availability Zones within the same region. Automatic failover via Pacemaker. Best for metro-area resilience but does not protect against region-wide disasters.
HANA System Replication (async) to a paired Azure region + Azure Site Recovery for app-tier VMs. ASR provides continuous replication for non-DB VMs — databases must use native HSR (ASR cannot guarantee DB consistency).
Azure Backup for VM snapshots + HANA Backint to Azure Blob Storage. Lowest cost DR option. Rebuild compute from backups in DR region. Cross-region backup support with GRS vaults.
Standalone HANA on a single EC2 instance with CloudWatch StatusCheckFailed_System alarm for automatic recovery. Instance retains its ID, private IPs, elastic IPs, and metadata on recovery. Limited to same-AZ recovery.
Two HANA instances across separate Availability Zones with synchronous system replication and Pacemaker-managed automatic failover. Virtual IP routed via AWS Transit Gateway or Network Load Balancer (overlay IP routing). AWS Launch Wizard automates the full cluster setup.
Same as HA — synchronous HANA SR across AZs with Pacemaker. AWS skips the Availability Set concept entirely — you go directly to Multi-AZ with placement groups for low-latency inter-node communication.
Asynchronous HANA SR to a secondary region. Supports multi-target replication (primary replicates to both local AZ and remote region simultaneously) and multi-tier (chained) topologies. Manual cross-region failover with automated intra-region.
AWS Backint Agent for HANA backups directly to S3 with cross-region replication. EBS snapshots auto-replicate across AZs and can be copied cross-region. Rebuild via CloudFormation templates + Backint restore.
Standalone HANA on a Compute Engine VM with live migration (transparent maintenance) and automatic restart on host failure. GCP's live migration is unique — it transparently moves the running VM to a healthy host without downtime during maintenance.
Two HANA instances in separate zones, running synchronous system replication with Pacemaker. VIP implemented via Internal Passthrough Network Load Balancer with health-check probes (socat on ports 49152–65535). The ILB itself carries a 99.99% SLA.
Synchronous HANA SR across zones within a region. Like AWS, GCP has no Availability Set equivalent — you go directly to multi-zone. Pacemaker with Corosync (token timeout: 20,000 ms for cloud latency tolerance).
Asynchronous HANA System Replication to a secondary region. Google Cloud's Agent for SAP supports Backint for direct backup to Cloud Storage. Cross-region persistent disk snapshots provide additional protection.
Google Cloud's centralized Backup and DR Service for SAP HANA. Supports application-consistent backups, Backint integration with Cloud Storage, and Persistent Disk snapshots for compute-level recovery.
| Criterion | Azure | AWS | GCP |
|---|---|---|---|
| SLA & Availability | |||
| Single VM SLA | 99.9% (Premium SSD) | 99.5% | 99.5% |
| Same-DC Redundancy | 99.95% (Availability Set / Flex Scale Set) | N/A — no equivalent | N/A — no equivalent |
| Cross-Zone HA SLA | 99.99% (Availability Zones) | 99.99% (Multi-AZ) | 99.99% (Multi-Zone) |
| Planned Maintenance | Reboot-based (Live Migration for some VMs) | Reboot-based | Live Migration (no reboot) |
| HA Mechanism | |||
| Cluster Framework | Pacemaker (SLES/RHEL) | Pacemaker (SLES/RHEL) | Pacemaker (SLES/RHEL) |
| VIP Implementation | Standard Load Balancer health check | Overlay IP via Transit Gateway or NLB | Internal Passthrough NLB with socat health check |
| HANA Replication | Sync HSR (intra-region), Async HSR (cross-region) | Sync HSR (intra-region), Async HSR (cross-region), Multi-target | Sync HSR (intra-region), Async HSR (cross-region), Active/Active read |
| Automated Deployment | ARM Templates + Azure Center for SAP | AWS Launch Wizard for SAP | Deployment Manager / Terraform modules |
| Disaster Recovery | |||
| App-Tier DR Service | Azure Site Recovery (ASR) | Elastic Disaster Recovery (DRS) | No native equivalent — PD snapshots + templates |
| App-Tier DR RPO | Minutes (continuous replication) | Sub-second (block-level replication) | Snapshot interval |
| DR Region Model | Fixed region pairs | Any region (flexible) | Any region (flexible) |
| Cross-Region Failover | Manual (with ASR orchestration) | Manual (with CloudFormation automation) | Manual |
| Backup & Recovery | |||
| HANA Backup Method | Azure Backup + Backint to Blob | AWS Backint Agent → S3 | Google Agent for SAP (Backint → Cloud Storage) |
| Backup Throughput | Varies by storage tier | Up to 16.8 GB/s (scale-out verified) | Varies by PD throughput |
| Snapshot Technology | Managed Disk snapshots (ZRS/GRS) | EBS Snapshots (cross-AZ/region, Fast Restore) | Persistent Disk Snapshots (multi-regional) |
| NFS/Shared Storage DR | Azure NetApp Files cross-region replication | EFS replication / FSx | Filestore or NetApp CVS |
| Cost Factors | |||
| Standby Instance Cost | Full price (AvZone requires identical VM) | Full price (Multi-AZ requires identical instance) | Full price (Multi-Zone requires identical VM) |
| Cross-Zone Data Transfer | Charged (varies by region) | Charged ($0.01/GB inter-AZ) | Free within region |
| DR Capacity Reservation | On-demand capacity reservation (ASR integrated) | On-demand capacity reservation | Reservations available |
Every hyperscaler supports the SAP Backint interface for native HANA backup integration. The key difference is the target storage service and restore throughput.
Synchronous log replay to standby node via HANA System Replication. Zero data loss within the HA cluster.
Weekly/daily full backup via Backint to object storage (Blob/S3/Cloud Storage). Baseline for point-in-time recovery.
Hourly differential backups capture only changed data blocks. Dramatically reduces backup window and storage consumption.
Crash-consistent disk snapshots for rapid recovery. EBS/Managed Disk/PD snapshots replicate across zones automatically.
Backups replicated to DR region via storage-level replication (GRS/S3 CRR/multi- regional Cloud Storage) for geographic protection.
Match your deployment model to your business requirements — not every workload needs four nines.
Single VM, backup-based recovery. Tolerate hours of downtime. Minimize cost by using smaller instances and infrequent backup schedules.
SLA: 99.5–99.9% · RTO: 2–6 hrs · RPO: Hours
Azure: Single VM + Premium SSD (99.9%)Cross-zone HA with synchronous HANA SR and Pacemaker. Backup-based DR to a secondary region. Balances cost against sub-5-minute intra-region recovery.
SLA: 99.99% · RTO: < 5 min (HA) / 2–4 hrs (DR) · RPO: 0 (HA) / Hours (DR)
All clouds: Multi-AZ/Zone HA + Backup DRCross-zone HA + cross-region async HSR with live standby. App-tier replicated via ASR/DRS. Full redundancy with manual regional failover in under 30 minutes.
SLA: 99.99% · RTO: < 30 min (regional) · RPO: Seconds–Minutes
Azure: AZ HA + ASR + Async HSRAWS: Multi-AZ + DRS + Multi-target HSRFull redundancy in both regions: 2×AZ HA in primary + 2×AZ HA in secondary with multi-target HSR. Withstands failure of 3 AZs across two regions. Highest cost, lowest risk.
SLA: 99.999% target · RTO: < 15 min · RPO: ≈ 0
AWS Pattern 7: Dual-region, dual-AZ, multi-target HSR| Provision | What to Negotiate | Why It Matters |
|---|---|---|
| RTO/RPO Guarantees | Explicit RTO/RPO SLAs in the RISE contract, not just infrastructure uptime. Require SAP to define application-level recovery targets, not just VM availability. | Infrastructure SLA ≠ application SLA. A 99.99% VM SLA means nothing if HANA takes 30 minutes to restart and 2 hours to recover logs. |
| DR Region Choice | Right to specify the DR region and hyperscaler zone configuration. Avoid being locked into SAP's default region selection. | Data sovereignty, latency requirements, and compliance mandates may require specific geographic placement. |
| DR Testing Rights | Quarterly DR testing at no additional cost, with SAP providing documentation of successful failover/failback and measured RTO/RPO. | Untested DR is no DR. Many RISE customers discover their DR doesn't work only during an actual incident. |
| Backup Retention | Define retention periods (e.g., 90 days online, 7 years archive) and require SAP to provide backup verification reports. | Compliance requirements (SOX, HIPAA) demand verifiable backup retention that outlasts the RISE contract itself. |
| Cross-Region Egress | Cap cross-region data transfer costs for DR replication at a fixed monthly amount, or negotiate inclusion in RISE base pricing. | Continuous async HSR generates significant cross-region traffic. Without a cap, DR costs can exceed the standby compute cost. |
| SLA Credits | Require SLA credits calculated against application availability (not just infrastructure), with meaningful credit rates (10–30% of monthly fees per breach). | Standard cloud credits are 10% for missing a 99.99% target — inadequate for SAP workloads where an hour of downtime costs $100K+. |
| Availability Zone Mandate | Require deployment across Availability Zones (not just Availability Sets or single-zone) for all production HANA instances in the RISE contract. | Some RISE deployments default to single-zone or Availability Set placement. The difference is 99.95% vs 99.99% — 4 hours vs 52 minutes of annual downtime. |
Estimate the business impact of different HA/DR configurations based on your organization's revenue profile.
Skynome's SAP governance framework audits your HA/DR posture, validates SLA coverage, and ensures your RISE contract includes the recovery guarantees your business demands.
Assess Your HA/DR Readiness →