Sunday, April 20, 2008

 

Back to the drawing board

I have been working on a cluster upgrade of WebSphere MQ from version 5.3 to version 6.0.

I used a cutover approach (aka big-bang) for development and QA. We only had one problem in QA-a queue manager with a lot of client connections wouldn't stop. We got a local fix from IBM for that one.

The project team decided to use a migration approach for production. We've only migrated two applications, but we've had a major problem emerged with each one. For the first application, the queue manager was stuck in a STARTING status. IBM recommended a fixpack for that one. Then we had a problem where the queue manager didn't start at all but just the once in production. This one remains outstanding. After migrating the second application, its queue manager didn't start completely. IBM doesn't have a fix for this one and has asked for a trace.

In addition, the heartbeat mechanism was set up using the network which has its own set of problems. This causes the loss the shared disk where the queue manager data resides. The shared disk uses GFS (Global File System). I discovered this week that this allows simultaneous access to the data from both sides of the cluster. Each queue manager has its own unique data so GFS is overkill. In the virtual QA environment, GFS has corrupted the data several times-the last time disasterously. The sysadmin tried to change the heartbeat mechanism and ended up spending the whole night getting back to a working configuration on just the one server.

You shouldn't have this many problems going live in production. I think there are serious issues with the design that must be addressed before we migrate any more applications. I'm also wondering if our mysterious problems will disappear when these are corrected. This means a major delay for the final migrations. I've been working on this project since last May. WMQ v5.3 went off support last September.

D.

Comments: Post a Comment

<< Home

This page is powered by Blogger. Isn't yours?