Sunday, June 11, 2006
An exceptionally busy week
I spent two days last weekend at the IBM disaster recovery site. We got further than the last test in April, but we were told that the CEO was not impressed after we got back. Why should the CEO care about the results of a disaster recovery test-doesn't he have bigger things to worry about? The Windows sysadmin who said they wouldn't leave me high and dry with MQSI, a product I am not familiar with, walked out with a password I needed. This was after midnight Saturday. I got home at three in the morning, grabbed a few hours sleep and went back at it. All in all, it seemed like a typical test to me. We ran into a few roadblocks, but that's exactly why you test. It did seem a bit more disorganized though.
Friday before I left I was warned by the server team leader that we would be required to set up two new acceptance servers on a rush basis Monday. Monday morning came and went with no word. I found out that the cloning process Sunday night had not worked, and it had been rescheduled for that night (people involved in the DR test were also involved in the cloning). Tuesday morning came and went without word one. Once again, the cloning process has been rescheduled. Wednesday morning, all hell breaks loose. The original server has been corrupted by the cloning process. Only a fragment of the MQ queue manager is left on the machine. The most recent source code is datestamped 2003. The sysadmin tells me that he had to restore the c: drive where MQ is installed on d:. It must have impacted the Windows registry.
I am leary, but we decide to delete and recreate the queue manager. It's just like the disaster recovery exercise. There are six people milling around my desk. I tell them that I can't type with people watching. In the middle of this, my senior VP cruises by. Meanwhile, another acceptance queue manager has gone down, a deployment to acceptance for another application isn't working and the clones aren't working either. All of these applications are waiting to test!
I proceed to knock them off one by one. I recreate the corrupted queue manager but forget to reset the channel sequence numbers which is an automatic for a DR exercise. The one queue that uses JMS isn't working because of authorization but the sysadmin figures out there is a userid missing as a result of how the server was rebuilt (darn VMWare). I call IBM about the queue manager that won't start and wait for a callback. We (well mostly the application architect) figure out that the deployment isn't working because of missing transmission queue switch to the cluster on the mainframe. Repeated attempts to work on the clones are unsuccessful. IBM calls back late in the day and tells me to trying deleting a checkpoint file. I've never had to do this before but it works. I go home that night guilt free.
The next morning, the sysadmin tells me to report the problem with the clones to the IBM Support Centre. I go to collect some fresh diagnostics, and everything works OK. I tell the sysadmin to do to the other server whatever he did to that one. He had added it to the Windows domain. After four days of waiting then struggling, it takes 45 minutes to complete the setups. We then hit the brick wall that is the very obtuse mainframe programmer who claims he hasn't received his assignment yet.
After working 11 days in a row, I take Friday off but spend the day close to the phone (stupid, stupid, stupid).
D.
Friday before I left I was warned by the server team leader that we would be required to set up two new acceptance servers on a rush basis Monday. Monday morning came and went with no word. I found out that the cloning process Sunday night had not worked, and it had been rescheduled for that night (people involved in the DR test were also involved in the cloning). Tuesday morning came and went without word one. Once again, the cloning process has been rescheduled. Wednesday morning, all hell breaks loose. The original server has been corrupted by the cloning process. Only a fragment of the MQ queue manager is left on the machine. The most recent source code is datestamped 2003. The sysadmin tells me that he had to restore the c: drive where MQ is installed on d:. It must have impacted the Windows registry.
I am leary, but we decide to delete and recreate the queue manager. It's just like the disaster recovery exercise. There are six people milling around my desk. I tell them that I can't type with people watching. In the middle of this, my senior VP cruises by. Meanwhile, another acceptance queue manager has gone down, a deployment to acceptance for another application isn't working and the clones aren't working either. All of these applications are waiting to test!
I proceed to knock them off one by one. I recreate the corrupted queue manager but forget to reset the channel sequence numbers which is an automatic for a DR exercise. The one queue that uses JMS isn't working because of authorization but the sysadmin figures out there is a userid missing as a result of how the server was rebuilt (darn VMWare). I call IBM about the queue manager that won't start and wait for a callback. We (well mostly the application architect) figure out that the deployment isn't working because of missing transmission queue switch to the cluster on the mainframe. Repeated attempts to work on the clones are unsuccessful. IBM calls back late in the day and tells me to trying deleting a checkpoint file. I've never had to do this before but it works. I go home that night guilt free.
The next morning, the sysadmin tells me to report the problem with the clones to the IBM Support Centre. I go to collect some fresh diagnostics, and everything works OK. I tell the sysadmin to do to the other server whatever he did to that one. He had added it to the Windows domain. After four days of waiting then struggling, it takes 45 minutes to complete the setups. We then hit the brick wall that is the very obtuse mainframe programmer who claims he hasn't received his assignment yet.
After working 11 days in a row, I take Friday off but spend the day close to the phone (stupid, stupid, stupid).
D.