Can someone tell me what the clickable "!" on the active tasks does? Is it to move the priority of a task up or down if there are more then one task?Trying to find information on how to retry/restart tasks.
Thanks - we will have developers look at working on the WFT. The issue I have now is that even after killing jobs that are stalled and starting a new workflow the new jobs stall. Other workflows and deployments are working fine. This is specific two workflows that deploy to the servers that those jobs that stalled do.Not sure why new jobs continue to fail - all services up and healthy - servers are up and healthy. Looking at logs to try and find something ...
OK - so is the solution to end the job and redeploy? It seems that even after ending jobs new jobs using the same workflow stall on the actual deploy task.
You need to determine if the deploy task script / program is failing to run (and hence the task is hanging because there's no timeout handler) or whether the deployment is being initiated and is, itself, hanging (which again could be caught by a timeout handler).If this is a custom script / program - you should make sure that it logs diagnostics somewhere.If this is an OOTB script / program - you need to check all the possible log files - add debugging if you can, or replace the OOTB script/program with your own custom one and then ... (see above).
Seth's company has done waht way too many people do & that is copy the configurable_default_submit WF to their on and customized it to work in their environment. That is a bad idea (IMHO) a custom one can be written without much trouble and fit the needs much better.
Cavalier locking
Yes - true and noted. :-) Well - I could see this as an issue if we were seeing this on all/more deployments - it occurs mainly when a large number of files or large files are deployed. I'm thinking it is more of a time out issue. So - I think we need to increase what we log and try and find some threshold for these timeout issues (maybe find out why they are occuring) and see if we can provide some workflow task that the users can kick off to either (or both) remove some files and restart the task.
If it's a large number of files (e.g.: > 300) - it sounds like a case where iw.cfg wasn't "fixed" to not include files on the command line for external task scripts.If it's fewer ( e.g.: < 100 files) larger files - then it may just be a timeout issue, though I would suggest checking the network bandwidth between sender and receiver and look at the various options being used in the deployment that might cause performance issues.