Saturday, December 17, 2005

 

What's wrong with this picture?

Once I got access to a production server (a couple of weeks ago), I started reviewing the MQ error logs. I see a TCP/IP communications error. This is appearing in all four queue managers used by the premier web application here.

I report the problem to the Service Desk. I don't know it at the time but external business partners also report the problem. The trouble ticket get assigned to the sysadmin for this application. He consults the mainframe MQ guy who says that the occasional TCP/IP error is normal and tolerable. In four years of supporting MQ on distributed platforms, I have never seen this error except during scheduled network maintenance or as a result of a real problem.

The network guy says that he needs to be involved at the time of the problem and that he doesn't know how to use the sniffer. The network has just been outsourced to Bell, and the rules of engagement have yet to be defined. In fact, there seems to be a question whether Bell is responsible for the internal switches and routers. If they don't, these devices are without support.

I do some more digging. The problem is occurring at 35 minutes after the hour, only during even numbered hours but not consistently. The problem reoccurs a week to the minute after the previous report. I try to engage someone (anyone!) to take ownership of the problem. The sysadmin is over his head and really shouldn't be the primary. The new IT problem manager has "bigger" fish to fry as does the network manager. These people keep referring me to other people who can't help with the problem.

The sysadmin does tell me that the problem has also manifested itself between the application server and the database server. I do a traceroute and there is only one hop between two. The problem must be on the device that they are both plugged into. I decide to bounce my theory off the security operations manager and my boss. I go to the former because the firewall is the default gateway of the application server where MQ also resides. My boss know a lot about the network. They both reply with the same info.

The next day, the security manager comes back to me with a suspect, but he wants to confirm with his staff. A day or so after that, his firewall guy comes to talk to me, and yes, it is a known problem. There is a multicast (IGMP?) on the core switch that is causing the connectivity break if I understood him correctly.

His manager gives me more background. This problem has been going on for over two years. When the problem was first diagnosed, the resolution was judged too painful to pursue. Actually, the firewall guy did rebuild his cluster as part of this. It was one of the first things he did after joining the company.

This is just unacceptable to me. Everything on the backbone is a victim of this problem. The high profile web application supporting 10,000 brokers and external business partners (who we white label for) is being impacted. I haven't received an answer from the network manager about the devices involved. He appears not to know. I think his VP is going to be in over her head as well with this one (she's more a call center/voice person). I hope I'm pleasantly surprised.

I asked the security manager what the next step was. I told him I was here to fix new problems not find old ones. He is going to escalate it to his manager. I sent an email to my boss asking him to do the same. I expect he'll try to escalate technically rather than politically because that's the way he is. We'll see if us newbies can make anything happen here.

D.

Comments: Post a Comment

<< Home

This page is powered by Blogger. Isn't yours?