Thanks. We wanted to be able to blame the guy who wrote iw.MasterScript, but we couldn't. Just kidding.
So the mystery remains as to why both TS servers run into exactly the same errors when we try and start things after a core.
.......TS on that box is faster than any of our other environments. We haven't had a core dump since. It's only been 3 days, but the difference has been pretty remarkable.
Why it took us this long to figure out that we had a bad filer is a mystery we'd like to solve internally, but that appears to be the cause of most, if not all of our issues.
Good to hear it's back up! Say, just out of curiosity, what is a 'filer'?
Good to hear that the problem has been resolved. Then what Rare scenario Interwoven was mentioning where they see the bug in Teamsite itself. Can that scenario come into picture for some other clients?
Engineering response [...] this is an issue that may occur in rare instances under specific circumstances...
They probably filtered out part of a response that stated "So far only Smitty's Team managed to achieve these circumstances....That's actually how we knew that these instances/circumstances exist in the first place..."
[AIWOV]"... this is an issue that may occur in rare instances under specific circumstances. There is no current workaround, but a fix is being developed for a future software update."
----------There is some retry logic where if we attempt to send a request to iwauthend and we get a BAD_REQUEST failure, the connection is attempted a 2nd time. This logic was refactored a few years back and a new issue was introduced (after the initial request fails, we clear out a certain pointer, which is then null on the retry – hence a crash occurs). The fact that this rather obvious issue has been in the field for a few years now and we haven't had a report of it, shows how rare it is for the initial attempt to fail. In fact, in verifying the fix, we had to hard-code a failure into the code to simulate the problem. As to the cause of the bad request, there could be reasons for a HOPI connection failure. The only instances he found in authend itself where it would return a bad request is if the timestamp in the request handle is earlier than authend's recorded start time (which maybe could happen if authend is manually restarted or if the clock drifts a lot and is reset, maybe?) or where an external database (not to be confused with LDAP) is used and the URL used for authentication returns BAD_REQUEST to authend.----------So, to summarize, on an authentication request, the first attempt fails and the second retry attempt crashes the server because of a null variable. These initial failures are rare and could be due to several factors, including an incorrect timestamp. I think that a part of your problem on this particular server is that nscd was disabled and so the authentication attempts to the OS were not to memory cache, but to the file system (/etc/passwd and shadow files) instead. Enabling nscd seems to have ameliorated the issue quite a bit and moving the file systems to a new filer has increased the performance dramatically.I'm wondering now if the manual startup scripts that were put in place might also be tied into this. I don't know why that was done in the first place, but it might be a good idea to look into the history of that decision.
[...]They also seem to be concerned about our special startup script that a certain famous DevNet poster wrote - I'm not sure why they are concerned about this as it does nothing more than run all the appropriate startup scripts instead of having to run them all separately.