Saturday, May 05, 2007

 

This week at "work"

Here's how the week went:
1) Monday: I caused an MQ QA cluster outage doing some testing.
2) Tuesday: we went out to lunch.
3) Wednesday: I re-investigated a known problem.
4) Thursday: we went out the lunch.
5) Friday: I broke two applications in development.
I couldn't be more productive if I tried.

I've been firefighting on the QA cluster for the last few months. Every two or three weeks, it would go "split brain" and fail over unexpectedly. We blamed it on the VMWare full image backup at first so that was stopped, but the failovers continued and even became more frequent. I finally did what the errors indicated and increased deadtime for the heartbeat (only a circumvention in my opinion). The cluster was stable the following weekend so I decided to make a script LSB compliant that heartbeat had been complaining about. When I tested failover, one of the filesystems wouldn't remount. I rebooted the server but still no joy. I got my new cluster sysadmin involved. He rebooted both servers and the mount worked this time. He also suggested that the heartbeat config might be out of synch which I dismissed (he and I don't always see eye-to-eye). Back at my desk, I had second thoughts so I compared the two haresources files and found a slight discrepancy! This had happened when I had made some changes to try to diagnose the original failover problem. I may have been the author of my own misfortune here.

While out at the first lunch, a bright yellow Porsche Carrera GT parked beside the restaurant. The same sysadmin said it was worth a quarter of million dollars. He was almost drooling over this car. Every guy who went past it stopped to gawk. Some even took pictures with their cellphones. My boss wagged his finger at someone who was looking at it giving the impression it was his. Most women, on the other hand, walked by without even giving it a glance.

Friday's adventures also involved the cluster, but in the development environment. I am trying to migrate one application from its own dedicated queue manager. At first when it didn't work, we discovered that a queue was missing on the mainframe or so we thought. I was dealing with the backup admin who put it back to humour me, and the app worked. When the primary returned from vacation, he informed us that the app was set up for dynamic reply queues on the mainframe side so the queue shouldn't exist. This was news to me, both technically and architecturally. Now, we'll have to figure out how to get it working in a clustered environment. This will probably involve reconfiguring the JBoss Message Driven Bean to which I am just getting my first real exposure.

D.

Comments: Post a Comment

<< Home

This page is powered by Blogger. Isn't yours?