Datrium Elevate NetApp Snap & Replicate to New Levels with Reverse Snaps

· 4 min read


Multisite Replication Topology

In my article (Disaster Recovery Seeding is a Pain!) I write about Datrium universal fingerprinting coupled with over-the-wire deduplication to drastically reduce the amount of data transferred between sites when replicating. Data is always globally fingerprinted, deduplicated, compressed and erasure coded, and data blocks are uniquely identified and logically aggregated. The system holds a full understanding of existing and missing blocks of all sites. However, depending on the data change rate, link and site availability there could be replication issues. Let's explore.

The Replication Lag Obstacle

The replication issues go something like this:

  1. IT organisations are unable to meet SLAs because there is replication lag.
  2. There is a lag because the network was down for a day, and all systems are trying to catch up and run forever behind.
  3. Because of the lag, the local systems start to get filled up, which causes a separate set of headaches.
  4. The replication lag further poses hurdles for being ready for DR.

A Short Background

Netapp introduced elegant snap-and-replicate techniques in the 90s (at least their academic papers are from that era). Their designs made it easy to create, retain and replicate multiple snapshots. This was revolutionary at that time.

Many decades later every storage and HCI system has some snap-and-replicate implementation with small improvements here and there, but the core implementation remains the same.

The basic design is you set-up a schedule at your primary site storage system that snaps every N minutes and retains them for a period of M days. You may also choose to replicate snaps to another site, and in case of a disaster, the latest snap is used on the DR site to recover workloads.

However, If there are any WAN or system unavailability issues the replication schedule falls behind, and in the meantime, snaps on the primary system are tied down and cannot be replicated to the DR site. Those snaps are designed to replicate in sequential time order, and hence will start replicating in order once the WAN or target site are restored.

Long story, but because snaps are designed to replicate in order, the latest snap will have to wait until all the older snaps have been replicated to the target site, and this causes even further delays.

In the meantime, SLAs are not met because the last snap is going to take time to catch up on the secondary site, and if there is a disaster now the secondary site has very stale data to recover from. Not a good situation to be in, but this is a common outcome.

Datrium elevates Snap & Replicate to new levels and solves the Replication Lag Obstacle.

The Datrium founding team were the original engineers behind Data Domain, VMware ESX, and NetApp SnapVault, and armed with the knowledge of issues, and they decided to handle that a bit differently.

  1. You already know, de-dupe over WAN to make replication faster.
  2. Always replicate the Latest snap first.

It takes excellent thinking to come up with such a simple and effective solution, but when WAN or site connectivity is restored after an issue, Datrium always replicates the Latest snap first. Subsequently, it replicates the second latest snap, and so on.

To demonstrate this behavior I kicked off IO generation (FIO) with a heavy bias towards writes (about 80K IOP/s), and then I setup snapshots and replication for every 5 minutes. The screenshot below demonstrates how Datrium DVX intelligently skipped a few snapshots and moved to the most recent at the time. When the replication if finished the system would then start replicating the latest snapshot again, and overtime as time and bandwidth allows the Skipped snapshots would be replicated.

This approach allows the target site always to have the latest snap first, and this is a critical element for DR readiness. Moreover, by going back and replicating older snaps in reverse order, all the compliance rules are honored.

Thanks, Sazzala Reddy (@sazzala) for content and reviews.

This article was first published by Andre Leibovici (@andreleibovici) at myvirtualcloud.net

storagevirtualization

DRDatriumNetApp