Datrium is gaining market share, and FUD escalates to new levels. Here I technically dismiss legacy vendor FUD against Datrium.

No matter what, when up and coming technology vendors start to disrupt legacy players, the amount of FUD (fear, uncertainty, and doubt) broadcasted by such vendors increase. Datrium being the major disruptor to the legacy storage and HCI business make competitors starting to spread ‘misguided’ information about our technology.
I respect my readers and for this reason, I will not create a battle card with X and Y vendors; that’s just not my style. Rather, I am going to educate on what is real and what is not real in regards to Datrium disruptive architecture. I want to take the opportunity to technically dismiss some of this FUD.
FUD 1: Datrium only support VMware hypervisor and vVols replace their VM-level granularity advantages.
Datrium supports KVM and Docker containers running bare-metal as well as VMware, so we're not VMware exclusive. However, yes, the Datrium integrate with VMware is unique. VMware is critical to the environment of the vast majority of enterprises and where they run their business-critical applications.
VMWare introduced vVols as a mechanism for storage vendors to be able to group different capabilities and functions into logical elements (i.e., vVols) that can be more granular than SAN LUNs. For example, a "Platinum" service level that is tied to a flash pool, a "Gold" service level that can be tiered between flash and SAS drives, or a single tier SAS with RAID 10.
In all vVols and their inherent service levels are another form of storage management that requires system administrators to have to determine if an application should be granted best performance (e.g. flash or tiered storage), best data availability (e.g. single or dual fault tolerance), or best capacity efficiency (e.g. RAID 10, RAID 5, or RAID 6).
Datrium has architected a system that always delivers the best performance, best data availability and best capacity efficiency with no storage management and its associated considerations and trade-offs.
In the world of private and hybrid clouds, self-service portals, and higher level orchestration services, it does not make sense to require users to identify, and in most cases make assumptions, about applications and data behavior. Users should not bother if application data is de-dupable, or RF2 vs. RF3 vs. Erasure Coding, or if compression delay should be 30min or 60min, or if checksumming, erasure coding and compression are to remain enabled for a given application. We don’t worry about all that when using the public cloud, so why should we when using private clouds?

Most of the user experience today is driven by the consumer products, and enterprise software vendors must incorporate Design Thinking into the development process. Having made the initial architecture efforts necessary to implement features that just work, and avoiding knobs as much as possible, Datrium has paved the way to an uncomplicated system that is ready to manage applications and workloads at scale.
FUD 2: Datrium datanode is a simple, dual controller JBOD with active/passive failover and NVRAM for a write cache (read cache is in each compute server).
This statement is extremely accurate. Datrium is a software-based platform that uses commodity hardware, both on the compute server as well as the data node storage appliance. We believe platforms should be simple to the administrator, and complex hardware designs drive complicated management.
FUD 3: Datrium is not true scale-out. Their scalability only match VMware's cluster size, 32 nodes right now.
1. The maximum size of a vSphere cluster as of 6.7 is 64 nodes, not 32. Perhaps this vendor is referencing the ESX host maximum that was first introduced in 3.5 and was the limit up to the 6.0 release?
2. One Datrium system can support up to 128 compute nodes, and this is a QA limit, not an architectural one. As such, each Datrium system can support two fully maxed out vSphere clusters under the same management domain.
3. Datrium scales out compute (performance) and data (durable capacity) independently. In fact storage analyst Chris Mellor recently reviewed our scale-out capability in terms of performance in his "Blocks & Files" site in a post entitled: "Datrium out-guns VMAX, XtremIO, and FlashArray all-flash boxes".
His summary says it all: "It beat Dell EMC's VMAX 950F and XtremIO X2 arrays and the Pure Storage Flash Array//m70 and //x70 in a range of mixed random read and write workloads and bandwidth tests."
I recently published “Datrium Scalability – Up to 138 Nodes” where I talk about Datrium disaggregated architecture and provided some impressive performance numbers achieved white protecting with triple mirroring (erasure coding), and executing deduplication, compression and encryption.
FUD 4: Datrium built a part HCI and part Array solution to differentiate themselves.
Datrium does not hope to differentiate ourselves this way, we do. Arrays and HCI have pros as well as cons, by taking the pros and removing the cons Datrium has built a platform with none of their trade-offs.
As noted by a judge when Storage Magazine and SearchStorage awarded Datrium the storage product of the year "…An innovative architecture that has clear advantages relative to both shared storage arrays and hyperconverged architectures".
A couple of specific comparison points:
Scale-out:
Arrays have none or limited scale-out performance potential with expensive and complicated upgrades. HCI likes to claim scale-out however the reality is that the scaling is very prescriptive and typically requires homogenous nodes.
Datrium performance is delivered via compute nodes which can scale to 128 delivering up to 18M IOPs. Performance scaling is in your hands with the server of your choosing, any form factor, make or model, with the flash of your choosing. The only requirement is that the server you select is on the hypervisor HCL.
Performance Contention:
Arrays must transmit all IO over the data network fabric, every read and write, there is no optimization of this path. This induces latency to all IO operations, even if flash is on the array the applications can never take full advantage of their performance. All applications must, therefore, share the performance of the array controllers. One demanding workload (e.g., VDI) can be impacted if other high IO workloads (e.g., SQL) are also running, this is known as the IO blender effect from the North <-> South traffic.
HCI presents a tightly coupled architecture which differs from arrays and presents new challenges this being the East <-> West traffic, as each node is also dependent on other nodes for its storage. All writes must first be protected on adjacent nodes before an application can continue. Some HCI doe not enforce data locality and reads from a VM may also be coming from another node in the cluster, in fact, the larger the cluster, the more likely it is that data will be served remotely as it needs to be balanced.
Datrium's unique performance profile mitigates common performance bottlenecks, and this is fundamental to why we can scale performance in a linear fashion that readily beats other approaches. Reads are architected to come locally from flash in the compute node. Each compute node has its flash-based, and software-based storage controller that's independent of every other compute node, this model is profoundly different than SANs or HCI. Writes are written to RAM and twice to NVRAMs in the Data Node in a compressed format and acknowledged back to the application. That's all the application waits for: reads from local flash and compressed writes to NVRAM. Datrium can optimally eliminate all reads from the network, and we optimize the writes.
FUD 5: Datrium places the software in the hypervisor. This splits the costs between hosts and the datanodes – hiding part of the cost in simple comparisons.
Yes, Datrium is a software platform that uses commodity hardware. Two of our founders created ESX with Dianne Green back in the day, and Datrium uniquely integrates with the ESX hypervisor via a VIB (vSphere Installation Bundle). To my knowledge, Datrium is the only commercially available platform that integrates with ESX this way, which provides a distinct advantage as deployments and upgrades are non-disruptive to the compute node.
Arrays have their own CPU, memory, and storage resources that are required to deliver performance. Consider the $ per CPU, memory, and flash on an array as compared to that of a commodity x86 server, and it's easy to appreciate the cost model advantages.
Datrium strives to be transparent, and customer trust is paramount. True to this all of our pricing models are all-inclusive of the costs to implement Datrium. We are incredibly confident that Datrium can provide a superior TCO no matter the comparison model, be it HCI or SAN.
FUD 6: Allowing any server to be used creates support issues, so Datrium is adding their own compute nodes to compensate.
This statement is patently false.
Datrium provides branded compute nodes as a convenience for customers who prefer a single support model. These compute nodes are off of the shelf Dell R640 servers, and there is nothing special or unique about them. The vast majority of deployments are using non-Datrium compute nodes with UCS blades being one of the most common compute platforms in use. Datrium has one simple requirement for compute nodes and that is that they are listed on the vSphere HCL, pick any server, new or old, on that comprehensive list, and we'll support it.
FUD 7: Datrium only offers snapshots and their own version of DIA via erasure coding.
Let's unpack this claim starting at the end first.
DIA (Data Invulnerability Architecture) is an end-to-end verification of data integrity that was first pioneered by Data Domain, that would be 3 of Datrium's Founders (Brian Biles, Hugo Patterson, and Sazzala Reddy). It is strictly not data protection against drive failures which is what erasure coding provides, so the claim is inherently incorrect.
Here is Sazzala Reddy's blog about his experiences with developing DIA as the CTO of Data Domain and how he applied the same approach to Datrium's data integrity.
Datrium's application protection, recovery, and replication solution is meant to operate in a completely seamless fashion inside of what's already being used in the vast majority of enterprises to manage the majority of applications, i.e., vCenter.
Datrium is a platform that natively provides end-to-end encryption with full data reduction, application mobility between private and public clouds, tier 1 scalable primary storage, backup, and both the retention and DR policy engines required to orchestrate the infrastructure.
Other than a Windows VSS agent for consistent application snapshots for the likes of SQL, no software needs to be installed directly by enterprise administrators for these capabilities… none, zip, nada. Please contemplate that fact for a moment and its operational benefit, it is a new paradigm.

Infrastructure approaches change as new ways are developed to address the inefficiencies of the past. That's what VMware did with virtualization on commodity x86 servers replacing physical ones, that's what Data Domain did with replacing tape as the backup medium, and this is what Datrium is doing today for infrastructures.
Consider that in April 2017 Datrium released our integrated backup capability to become self-protecting storage. Later that year Gartner released their analysis of Data Center Backup and Recovery Solutions, their guidance was that by 2022 20% of storage systems will be self-protecting obviating the need for backup – Datrium is already there!
FUD 8: Datrium is yet another thing to manage... it's own island.
Datrium can seamlessly integrate into existing VM environments and run alongside them for as long as required. Customers do not need to rip and replace any existing products, and you can continue to manage your environment as you do today and where it makes sense use Datrium to augment.
Calling Datrium an island is a misperception.
Differently to most HCI offerings, as an integrated platform the only point of management is the upgrade of the environment. So yes, from time-to-time you will need to push the upgrade button for Datrium in vCenter, but that's it. Datrium is an integrated platform, not a heterogeneous layered approach.

FUD 9: Datrium is still a Startup
This is one of those FUD that come from competitive decks where whoever put the deck together has no idea what they are writing.
Datrium is 6 years old and has been shipping product for 3 years. As of today Datrium has hundreds (closer to 1000) customers worldwide with multiple thousands software licenses shipped. Datrium customers host business critical applications, VDI, Test/Dev and other workloads.
Datrium secured $165M through Round D funding, and our investors see how we have built a repeatable and scalable business model.
That said, Datrium is a new supplier, and as such we're not tied to legacy approaches and we had the opportunity to provide a clean sheet architecture that allows IT to redo the delivery of IT resources. Gartner has even called out how new suppliers are disrupting the norm, and their recommendations fit Datrium like a glove.
I would like to acknowledge Dave Zenz work who pretty much did the heavy lifting for this blog post.
This article was first published by Andre Leibovici (@andreleibovici) at myvirtualcloud.net
2 comments
Karl
loving Datrium for a real and useful technology!
Edouard
loving Datrium for a real and useful technology!
Comments are preserved from the original site and are closed.