Showing posts with label MCITP. Show all posts
Showing posts with label MCITP. Show all posts

Monday, 14 October 2013

WYCRMS Part 5. Nobody Ever Runs a Server That Long

In 1997, a HP 9000 engineer wouldn't blink telling about a server that had been running continuously for over five years.I found this remarkable at the time, and couldn't imagine a Windows server lasting that long. I have moved on, and frankly expect my Windows servers to survive that long today. Very few share this position, and I'm trying to find out why it's so lonely on this side of the fence.

5. Nobody Ever Runs a Server That Long

Uptime

IT Engineers can take this stuff really seriously. I was quite proud of my own server at home that ran, uninterrupted, for one year and three months, during which time I upgraded the RAID array without any filesystem disruption, hosted at least twenty different types of VM (from Windows 3.11 to Windows 2008), served media files and MySQL databases, shared photos with Apache and personal files with Samba. Only the fast-moving BTRFS in-kernel driver broke me from that little obsession, but you don't need to run a Unix-variant to get that kind of availability.

Windows admins are simply used to "bouncing" computers to correct a fault. Hey, it works at home right? It's a complacency, a quick-fix, and often in response to an urgent need to restore service - troubleshooting can and often does take a back seat when management is standing over your shoulder.

Since Windows was a command launched at a DOS prompt, the restart has been a panacea, and often a very real one that unfortunately actually works. It almost always adds no insight into the fault's cause. Perhaps there's a locked file you're not aware of preventing a service from restarting, or a routing entry updated somewhere in the past that isn't reapplied when the system starts up again; there are myriad ways that a freshly started server can have a different configuration from the one you shut down, allowing service to resume.

Once a server is built, it is configured with applications and data, then generally tested (implicitly or explicitly) for fitness of purpose. Once that is done, it processes a workload and no more attention is required assuming all tasks such as backups, defragmentation etc are scheduled and succeed. Windows isn't exactly a finite-state machine (in the practical sense), but it is nonetheless a closed system that can only perform a limited set of tasks and failure modes should be easy to predict

Servers are passive things. They serve, and only perform actions when commanded to. Insert the OS installation DVD, run a standard installation, plug in a network cable, and the system is ready for action. Only it's not configured to take action just yet - it's waiting. In this state, I think most engineers would expect it to keep waiting for quite some time - perhaps a year or more. But add a workload - say a database instance or a website, and attitudes change.

I've had frequent discussions with engineers who will tell me things like "this server has been up for over a year, it's time to reboot it". Somewhere between an empty server and a billion-row OLTP webshop is a place where the server goes from idle and calm to something that genuinely scares engineers - just for running a certain amount of time.

When pressed for exactly which component is broken (or likely will) that this mystic reboot is supposed to fix, I never get anything specific, just a vague "it's best practice".

Windows Updates are frequently cited as a reason to reboot servers, and thanks to the specifics of how Windows implements file locking yes the reboot there is unavoidable. This leads to the unfortunate tendency to accept reboots as a normal part of Windows Server operation, but tend to see the reboot as the point (with an update thrown in since the server is rebooting anyway) instead of an unfortunate side-effect. I realise the need to keep servers patched, but again, when pressed for a description of which known defects (that have actually or could probably affected service) a particular update application - with associated downtime - will fix, the response comes in: "Um, best practise?".

In the absence of an actual security threat, known defect fix or imminent power failure, I am rarely convinced to shut a server down. I first included "vendor recommendation" in that list, but realised I've yet to see one. Ever.

Even at three in the morning when no sane customer could be relying on a system, during a once-quarterly change window when all services are nominated unavailable so service providers can make radical changes, even then: No, you can't reboot my server.

If engineers took the time to think about where in the continuum from empty server to complex beast the point of fear arrives, they can figure out which bit is scaring them and make sure those are well-understood, properly configured and maintained.Unfortunately, that takes time, effort and sometimes a bit of theory and modelling. Rebooting is so much less effort.

Windows Server can, and should, be expected to remain ready for service for as long as the hardware can last. With the advent of virtualisation and VMotion, even that obstacle is gone, and the limits are practically nowhere to be found. Applictions are another story, and if the developer/support specialist think they need restarting, that's fine, but they have zero authority to suggest this for Windows.

I've heard the phrase "excessive uptime" identified as the root cause of outages. I doubt Microsoft would like to know the engineers they certify are - and I don't say this lightly - genuinely afraid of their product doing its job, as designed, for years. If that doesn't happen and a genuine OS fault occurs that only a reboot can solve, it is quite shocking how many engineers will actually report this problem to the vendor and tolerate workarounds, design hacks and cludgy scripts.

In the same way that one can learn a procedure for changing the spark plugs on a specific model of engine, while completely missing the black smoke ejected from the exhaust thanks to a chronic ignition timing failure, so too engineers who have not yet attained a mediocre grasp of computing theory can continue to diagnose and treat only symptoms.

A server failing to continue to do its core function of staying up is not a mild symptom that a reboot can fix. It is a fundamental failure of the product, and failing to do the hard thing of actually understanding why and demanding the vendor improves their product does nobody any service.

In fact, it's negligent.

Previous: Part 4. Windows Updates and File Locking

Tuesday, 8 October 2013

WYCRMS Part 2. Windows Just Isn't That Stable

In 1997, a HP 9000 engineer wouldn't blink telling about a server that had been running continuously for over five years.I found this remarkable at the time, and couldn't imagine a Windows server lasting that long. I have moved on, and frankly expect my Windows servers to survive that long today. Very few share this position, and I'm trying to find out why it's so lonely on this side of the fence.

2. Windows Just Isn't That Stable

Ah BlueScreen of Death, how I've missed you. Actually, I haven't, since finding out what caused them was a nightmare, and recovering without a remote console solution is not conducive to a predictable social life (or sleep schedule). That said, they were so common we even had joke screen savers mimicking them for our own geekish amusement. Since Microsoft acquired Sysinternals they're even available to download directly from Microsoft. Imagine your in-car entertainment system being configured to show you fake warnings of a failed brake line, or a cracked cylinder head. "Would you like the free video package of Ford vehicles endangering passengers' lives with your new Focus sir?". IT people are weird.

I've analysed my Windows 7 x64 installation, and in the last three years I've had six bluescreens. Once was my graphics card (pretty unique), all the others were my Bluetooth headphones putting my cheapo-Bluetooth dongle in a spin. I blame the dongle, not Windows.

OK, that's not fair to the dongle maker: I blame Windows, but only the Bluetooth stack since it's never been something I expect Windows to do well - multiple dongle-headphone combinations have yet to produce a pleasant experience (three dongles, two headphone models). The network card, storage stack, print drivers, memory management, process scheduler (NUMA-aware these days apparently): These all work so well I haven't notice them doing their job, and I am very familiar with what a complex job they have.

I expect roughly once a month to see a BSoD on public transport, or at stations, or many airports, or billboards. The layout of the BSoD has changed over the years, with each version of Windows getting a little tweak so that you can spot the version even if the error itself is gibberish, and I conclude from viewing these blue non-advertisements: These systems tend to be A) old, B) written in languages and coding styles that aren't that good, and C) interface with devices with terrible drivers.

This is not typical of modern Windows servers.

I would never dream of subjecting a server to the amount of change my hard-working personal workstation endures. AMD updates my video drivers multiple times a year, I attach and detach USB/phone/iSCSI devices more often than I refill my car's tank, and run code from pretty much anywhere as long as it promises me utility or entertainment. A server is different, running things I trust to go on processing without attendance, cleaning up after itself, and basically staying up. If I do make changes, it's controlled, tested and left the hell alone.

Windows Server is solid, and every iteration gets more solid. It's expanding to 64-bit spaces, handling multipath-iSCSI with ease, more cores than I have fingers in byzantine NUMA layouts, hosting server instances in their own right with Hyper-V and pushing gigabytes around through network cards and storage interfaces, crunching data and most importantly providing services.

Yet the very people who spend time and money proving they are skilled in designing and administering these systems so that they can adorn their signatures and office receptions with impressive Microsoft-approved decals are the first to tell you not to trust a given server (without even knowing the workload or configuration) to remain available. They express surprise and concern on viewing a server continuously running for over a year.

I'm surprised and, yes, concerned that they react this way. Isn't this what your sales folks promised me in the first place?

Previous: Part 1: But I Have to Reboot My Own Windows System All the Time!

Tuesday, 22 December 2009

Putting an MCITP in its place

I have noticed that the new raft of credentials from Microsoft don’t necessarily make sense to folks, especially those that are already familiar with the old set of MCSE-type credentials. I mentioned to some friends that I've got the “new MCSE”, a lot of them got it, but it dawned on me that this is, in fact, a field of some confusion. A quick search on Google came up with one fault, and that is how this new credential relates to the ones we (certainly I) already know.

The point of any credential, be it Cisco, VMware, Microsoft or embroidery is to show to an external party that you are qualified in a particular field of endeavour. This was quite plain with Microsoft’s old regime, the Microsoft Certified Professional (MCP), and the Microsoft Certified Systems Engineer (MCSE), as well as the Microsoft Certified Database Administrator (MCDBA). However, as Microsoft branch out into new fields and offer solid, integrated products in fields not entirely related to Windows Server or SQL Server, the approach of cobbling together a new acronym for a new product or role is unwieldy – imagine the Microsoft Certified System Centre Engineer – MCSCE??

So, what is the transition?

MCP –> MCTS

The first Microsoft certification I got in 1997 was an MCP: Windows 95. This showed that, according to Microsoft, I was competent in installing, administering and troubleshooting Windows 95. I remember just how proud I was that day.

The problem though is that the term, “Certified Professional”, encompasses both the specific credential and the entire field of Microsoft certified persons, so is not entirely appropriate. “Technology Specialist” on the other hand, clearly shows what the candidate is trying to demonstrate, that he knows his stuff on a particular product. This bit is key, a specialist in SQL Server configuration is not necessarily a specialist in database development or administration, and in larger organisations the roles are very clearly separate. The MCTS credential clearly segregates say, an application server specialist who can administer web applications, from the server network specialist who will hook it up to the various internal and external parties accessing it.

MCDST/MCSA/MCSE/MCDBA –> MCITP

A big failing of the old MCSE credential was the elective system. While Microsoft may introduce the idea in the future, I sincerely hope not as it adds doubt and confusion to the mix.

I hold an MCSE on Windows NT 4 (incorporating the MCP on Windows 95 I mentioned above). It included two “elective” exams from a list of many more, specifically TCP/IP Networking and Exchange Server 5.5. This means I need to explain to anyone asking just what kind of MCSE I’ve achieved. This was partially remedied in the 2003 track to include an MCSE: Messaging credential, but no such moniker exists for a SQL Server specialist.

The phrase “Systems Engineer” was especially limiting, since it implies an ability to design and implement server infrastructure centred on Windows Server. While that is indeed my own focus, it is of little use to someone specialising in monitoring and management systems, or even the venerable desktop support guru. While the DBA and the Desktop Support guy had their own acronym (MCDBA and MCDST respectively), I certainly don’t want to have to memorise the ever-growing list as a hiring or support manager.

By asserting that someone is an IT Professional in a named field, it indicates a proficiency in a technology set rather than one product. It also narrows the competency; while an Enterprise Administrator demonstrates competency in designing and implementing infrastructure from SANs and Terminal Services down to the desktop, the Server Administrator credential is more focused on those with competency in Windows Server itself.

These credentials are not easy to come by, and are especially hard if the individual has no relevant experience in the real world.

While the plethora of MCITP credentials may seem like a dilution of the fairly focused MCSE, it offers the opportunity for many more product specialist to demonstrate their competency in their field, with a credential on a par with the more established Systems Engineer we’ve come to know.

MCM/MCA

Now we get to the good stuff. The Microsoft Certified Master and Architect credentials are not for the faint of heart or newbies. The intensive certifications are for those with five or more years experience leading complex design, implementation and migration projects and a demonstrated history as a technology leader and expert. Standing up in front of a panel of recognised experts purporting to know your stuff is a daunting proposition, probably even for a few of the members of the very panel you’d be standing before.

For anyone claiming to be hot stuff on the range of Microsoft products, services and solutions, this is where you should be aiming. If you’re already that good, convincing your company to stump up for the three weeks training in Redmond for the MCM should be no effort, and I look forward to getting to that level myself.