Showing posts with label Rich Brumpton. Show all posts
Showing posts with label Rich Brumpton. Show all posts

Friday, September 18, 2009

How to use PVS without having to wait for file copies

By:Rich Brumpton

As I have mentioned before I love PVS. The entire concept of live booting a large number of XenApp servers off of a single image simply makes sense. Works great for VDI or Desktop OS streaming too and one of the facts of life on all three of these scenarios is that you are going to have to do updates to your images. No matter how good a job you do a separating your apps from your image using application virtualization tools, you will still have to deal with OS patches and other required changes.

Typically, this means that you are versioning anywhere from once a month up to several times a week depending on the life-cycle stage and stability of the requirements for the image. In this post I will outline an approach that I have started using in my lab and show how to make these updates simple and without requiring a wait for file copies.

The general principle behind this update method is the same as what Citrix uses themselves with XenServer. "If you have to copy something big, you might as well ask the storage system to do it for you." This way the data never has to leave the SAN to be copied, and if you have a SAN that supports it you can even do thin cloning for instant writable copies of data.

Because that is what I wanted, the ability to update a vDisk image without having to copy the file. In my lab environment I often have to version a vdisk several times in one day while working through several generations of a build and the file copy process was limiting how many versions I could iterate through in a day. I also have in my lab right now a NetApp SAN that provides me with thin cloning, and one day during a demo I hit on the way to make this all easier.

During a presentation that included a demo the NetApp SnapDrive component (more on this later) I was asked by a member of the audience "How hard is it to take a snapshot and mount it?" I was demoing SnapDrive on my PVS server so I said "easy, like this" and demoed taking and mounting a snapshot about as fast as I could run through the wizard. After this demo I got to thinking, I had just created a copy of 12 different PVS images in seconds, not the ~1hour copy time I usually faced for a full versioning of all of my vDisks. This in turn led me to investigate how I could make this work in production, and the result is:

How to use NetApp and PVS to make instant version changes
The Key to making this work for me is the ability of the NetApp FAS to take a snapshot and make it writable. There are two mechanisms that can do this for an iSCSI LUN (which is where I will focus) hosted on NetApp storage, FlexClone and LUN Clone. While it's possible to do what I'm about to describe with LUN Clone, it will take several more steps or more extensive scripting to get the same ease of use as FlexClone coupled with SnapDrive, which is the solution I'll focus on here.

FlexCloneas mentioned above has the ability to take a snapshot of point in time data and make a writable version of it that only grows by the deltas between the original and modified version. The other piece to the solution from NetApp is a piece of software that you install on your servers that connect to the SAN called SnapDrive. This software component snaps into computer management and allows you to manage the storage attached to the server from the server itself. This is a very easy way for server administrators to take advantage of features like storage snapshots, provisioning, and cloning without having to understand the underlying SAN operations.

With these components in place you can manage versions of images on provisioning server like so:

Preparation
We have to build the environment properly to allow for this update method, so we begin by setting up a new store. Note that in this model you may have a large number of LUNs per server which is why we are using folder mount points instead of drive letters.
  1. Create a thinly provisioned FlexVol on the NetApp Array for use by PVS.
  2. Create a folder on the PVS server to server as a container for multiple vDisk stores (i.e. c:\stores)
  3. Create an iSCSI LUN and attach it to your PVS Server using SnapDrive and mount the LUN in a subfolder of your stores folder (i.e. c:\stores\WindowsXP.001)
  4. Use the PVS Console to create a new store pointed to this new LUN
  5. Create a new vDisk or copy an existing one to this location
  6. Assign this initial version to devices and use normally
Versioning
The actual versioning process is used for each version of the vDisk throughout it's lifecycle. The longest part of this process will be the actual updating process in the VM.
  1. Use SnapDrive to clone this LUN and mount under a new subfolder (i.e. c:\stores\WindowsXP.002)
  2. Create a new PVS store and import existing vDisks
  3. Change vDisk properties to Private mode, change version number (if using automatic updates) and assign to an update VM.
  4. Use update VM to apply updates
  5. Change vDisk back to standard mode and then assign to dev, test or prod machines as appropriate
Retirement
Because the child clone relies on it's parent, we do have to take steps to remove old versions cleanly.
  1. "Split" the child clone from the parent
  2. Ensure that all vDisks in the store are unused and unlocked
  3. Delete the old store
  4. Once this split is complete, delete the parent LUN and mount folder.

As you can see this update method, is not too unlike a traditional file copy update method in the number of steps required, but is much faster. The other really cool thing is that many of these steps can be scripted. Both PVS and SnapDrive provide scripting interfaces that can be used to perform many of these commands either on demand or on a regular basis. In fact I'm working on a variation of this for a customer who is looking for a way to script the behavior of PVS to allow scheduled updates trough SCCM. This process is to be completely automated up until it is tested and released into production by the administrators. The process for them would be very simular except instead of cycling ever upward through version number for their folders, they would have a series of folders that are re-used on a monthly basis (i.e. vDisks.1 to vDisks.4 for a weekly update schedule.)

The other cool thing about this solution is that you don't have to sacrifice performance in order to save space and time. Because the NetApp system is aware that the 2+ (virtual) copies of a given block in different thin clones are really the same physical block, it can simply cache it once, and avoid having to go to the disk for new instances of the same vDisk. This means that not only are you not slowing down the SAN by making a big file copy, your new Disk image is already in the cache of the storage array!

This kind of true synergy between vendors really gets be excited. To be able to take these features and tie them together ourselves without relying on them to do the integration for us by using a product like WorkFlow Studio is simply amazing.

More information on Provisioning Server



Read more!

Sunday, September 13, 2009

HA Provisioning Server without SAN or NAS!

By:Rich Brumpton

I've loved Provisioning Server for a long time now, and Citrix has been doing a great job of moving forward with additional features and functionality. They have also provided a very thorough job of documenting the many possible configurations of PVS in an HA configuration. One that has been possible for a while, but has not made it into the white papers is local disk HA. (Explained in this article and also detailed by Jarian Gibson on his blog here).

This local disk HA option is a perfect example of the trade-off we often have to make between a low upfront cost and easy operations in the future. I've run into many great PVS prospect customers who do not own a SAN, but still require high-availability from PVS because so many of their users rely on it for XenApp server or VDI machines. It does however increase operating expenses to deploy it this way because you have to copy multi-gigabyte files around multiple times throughout a single versioning of a vDisk. Compare this to using a shared LUN ($) (with a cluster file system ($)) where you only have to copy the file once.

So when choosing a HA method for PVS, be sure to take into account where the best place to make that investment is. Lower initial cost or easier operations, which one is worth more?

In a future blog post I will explore an undocumented method for managing vDisks that can be even quicker and easier than using a regular shared LUN or folder, on the right kind of storage.

More Recomended Reading


More information on Provisioning Server



Read more!

Monday, July 20, 2009

Why NetApp is a great choice for VDI

By:Rich Brumpton
While working with a customer on a large VDI architecture recently we were comparing the required storage across several vendors and after looking at the proposed solutions from several I was asked the question:
In regard to the configs I’m really surprised at the low number of spindles relative to the IOPS req[uirement]s. Can you please help me understand the PAM a little more?"

The short answer is that the PAM can greatly increase performance in an environment that is heavy on small random reads like VDI, but that is not the only technology that NetApp uses to help optimize the storage of VDI.


There are a few things working in NetApp’s favor to keep the spindle count low. To begin with on the NetApp system one or more large pools of disks (Aggregates) are created that allow thinly provisioned volumes to be created that are striped across the entire aggregate. These volumes then contain one or more LUNs, more on this part later.

The Raid Groups that make up these aggregates use RAID-DP which offers double disk failure protection like RAID6, but because the NetApp storage system always writes full stripes and never has to do the read, read, write, write operation that gives RAID5 it’s 4:1 overhead and a 6:1 overhead for RAID6. In fact since NetApp can write it’s metadata anywhere in the file system the RAID write overhead is 1:1, in fact the only place I have to calculate overhead on RAID-DP is for IOPS (I=P(N-2), where I is total raid group IOPS, P is single disk IOPS and N is the number of disks in the array) to account for parity disks.

Other technologies on the NetApp storage system combine to reduce the physical size of the working set including thin cloning and primary storage deduplication.

The first of these, thin cloning, allows a snapshot of a single master copy of a volume (FlexClone) or LUN (LUN clone) to be presented read-write to hosts. This appears to hosts as a separate full copy of the data, but in fact only the deltas between the old and new blocks are written to disk, all the common OS components that make up a good portion of the working set for VDI actually remain in the same, single location. This technology can be used using the NetApp Rapid Cloning Utility (RCU) for VMware View or when using Citrix XenServer as the host for a XenDesktop machine. In either case this allows the working set to be decreased from N*W to W+((N-1)*D)W (where N is the number of clones, W is the working set size, and D is the delta percent of change from the master.)

Data Deduplication also plays a role when more than one VM is stored in the same volume. This feature looks through a volume for duplicate data blocks and removes all but a master copy and places metadata pointers back to this copy for each other copy of the block. Like thin cloning this feature is available because of NetApp’s ability to store metadata anywhere within the file system and creating pointers is nothing unusual given the structure of the WAFL file system. In a VDI scenario Data Dedupe helps contain the size of the deltas between VM’s by removing duplicate blocks created by OS or software updates, but this is a batch process so short lived data structures may not benefit.

So now that we have reduced the size of the working set, let’s talk about the PAM card which is 16GB of DRAM on a PCI-E card coupled with FlexScale software to act as an intelligent read cache. For folks from the server world the PAM operates like an L2 cache on a processor as an accelerator between the controller RAM and data on disks. There are 2 modes that are interesting to us in this discussion, default mode and metadata-only mode. In default mode the PAM card caches ONLY small random reads and metadata. This allows a large majority of the pointers used for deduplication and thin cloning to be stored very close to RAM which will already be used to cache the MOST frequently used data and metadata. If thin cloning and deduplication are used intelligently with an optimized configuration this mode can be used to retain a large number of the random read blocks in the PAM, greatly reducing the amount of time that users have to wait for blocks to come all the way from disk. This is the mode that I would use with VMware View or persistent XenDesktop VM’s and this is the mode that helps the most with events like boot-storms in View and persistent VDI scenarios. Metadata-only mode is used when there is a large working set and there is no way that it can fit enough of it in the PAM to avoid simply churning through the cached data. Metadata is cached in the PAM while data blocks are not allowing instant access to metadata blocks and a shorter access time for data stored on disk. This mode is the one I would use with XenDesktop in which each VM can be configured to store its own Provisioning Services write cache on the SAN, but this cache will be unique to each VM.


RAID-DP: http://media.netapp.com/documents/wp_3298.pdf
VMware RCU: http://blogs.netapp.com/virtualization/2009/03/netapp-and-vmware-view-vdi-best-practices-for-solution-architecture-deployment-and-management-part-8.html
Deduplication: http://media.netapp.com/documents/tr-3505.pdf
XenDesktop on NetApp: http://www.citrix.com/site/resources/dynamic/partnerDocs/CitrixXD2.0withNetAppStoragePilotDeploymentOverview.pdf
PAM: http://blogs.netapp.com/storage_nuts_n_bolts/2008/08/performance-acc.html


Read more!
Microsoft Virtualization, Citrix, XENServer, Storage, iscsi, Exchange, Virtual Desktops, XENDesktop, APPSense, Netscaler, Virtual Storage, VM, Unified Comminications, Cisco, Server Virtualization, Thin client, Server Based Computing, SBC, Application Delivery controllers, System Center, SCCM, SCVMM, SCOM, VMware, VSphere, Virtual Storage, Cloud Computing, Provisioning Server, Hypervisor, Client Hypervisor.