I'm using a Quarkus REST application that connects to Documentum via DFC. Everything worked fine when running on an Amazon Linux image. However, after moving to a Chainguard Java image, we started seeing random socket I/O exceptions from various DFC threads, including executor threads, license threads, and others.
After investigating, it appears that the underlying RPC sockets are getting disconnected after roughly 5 minutes. This behavior was never observed on the Amazon Linux image.
As a mitigation, we added a health check that runs every 3 minutes to keep the service session active. This has helped somewhat, but user sessions still fail at random points. Sometimes the failure occurs while a user is in the middle of filling in a license thread, sometimes during session creation, and sometimes on the very first request handled by an executor thread.
For each request, I obtain a session from the session manager and release it when the request completes. My requests all complete within 5 minutes, so my assumption was that user sessions would not encounter socket timeout issues. However, that doesn't seem to be the case. Even though I request a "new" session, DFC appears to be returning an existing pooled session whose underlying socket may already have been disconnected. That stale session then fails in the middle of request processing.
I have also tried configuring the session connection timeout, but that has not been particularly helpful because my code already acquires and releases sessions on a per-request basis. From my understanding, the issue appears to be at the socket/RPC connection layer rather than the session lifetime itself.
One interesting observation is that when these socket I/O exceptions occur, the number of active sessions on the Documentum server starts increasing significantly. The failures seem to prevent sessions from being cleaned up properly. In many cases DFC eventually retries and functionality is preserved, but the accumulating sessions can eventually cause broader stability issues.
Typical errors look like:
Plain Text12026-07-13 16:52:06.591 [LicenseThread-3] WARN2c.d.f.c.i.c.d.DocbaseConnection:176 -3Socket IO Exception, so reopening the socket4 5[DM_SESSION_E_RPC_ERROR]6error: "Server communication failure"Show more lines
Has anyone successfully run DFC on a Chainguard image? Are there any known compatibility issues with the minimal Chainguard Java images?
I'm wondering whether DFC relies on OS-level networking behavior, socket keepalive settings, native libraries, DNS configuration, or other components that are available in Amazon Linux but absent or configured differently in Chainguard. Any suggestions, workarounds, or recommendations would be greatly appreciated.