Saturday, November 18, 2006
Best practice bites
An MQ administrator should monitor the dead letter queue for undelivered messages. This is a sign of major system problems.
My pager started going off at the beginning of October. I tried to see what was on the dead letter queue, but the message was too big for the sample programs I normally use. Then the message expired. I let the usual suspects know there was a problem.
While I was on WMB training, I got exposure to a utility, rfhutil(c), that allows access to queues. This problem gave me the impetus to set it up on my laptop. Late one afternoon, I was able to capture one of the DLQ messages and the application developer. Turns out the data is too long to get to the backend and it's an application problem. This should be impossible for a well designed application. Someone hasn't done their testing properly. We only have a SIT team, a UAT team and a QA team, but this got by all of them. I ask for program correction, but it's a piece of the application that's been outsourced so there's no telling when this problem will be fixed.
While I'm investigating this problem, I see that messages are also hitting the mainframe DLQ. I know that my mainframe counterpart ignores this because of another design issue-messages are arriving at the backend without a valid userid (TAM and RACF aren't synchronized). I start sending out alerts when these problems occur, but no one seems to know what they're about or what to do about them, and they keep trying to involve me in the problem resolution.
Maybe I should increase the maximum message length and just let them flow through to the mainframe DLQ and not have to worry about it. One of my teammates is setting up a Tivoli alert for the support areas, but it's taken forever to get through the change management process and I'm not even sure it is working yet.
D.
My pager started going off at the beginning of October. I tried to see what was on the dead letter queue, but the message was too big for the sample programs I normally use. Then the message expired. I let the usual suspects know there was a problem.
While I was on WMB training, I got exposure to a utility, rfhutil(c), that allows access to queues. This problem gave me the impetus to set it up on my laptop. Late one afternoon, I was able to capture one of the DLQ messages and the application developer. Turns out the data is too long to get to the backend and it's an application problem. This should be impossible for a well designed application. Someone hasn't done their testing properly. We only have a SIT team, a UAT team and a QA team, but this got by all of them. I ask for program correction, but it's a piece of the application that's been outsourced so there's no telling when this problem will be fixed.
While I'm investigating this problem, I see that messages are also hitting the mainframe DLQ. I know that my mainframe counterpart ignores this because of another design issue-messages are arriving at the backend without a valid userid (TAM and RACF aren't synchronized). I start sending out alerts when these problems occur, but no one seems to know what they're about or what to do about them, and they keep trying to involve me in the problem resolution.
Maybe I should increase the maximum message length and just let them flow through to the mainframe DLQ and not have to worry about it. One of my teammates is setting up a Tivoli alert for the support areas, but it's taken forever to get through the change management process and I'm not even sure it is working yet.
D.