[SEV3] Impact on DocV and Watchlist monitoring events

Incident Report for Socure

Postmortem

Root Cause Analysis

Incident: Service Interruption Affecting Account Data Access
Date Range: September 8, 2026 (10:04 AM to 10:31 AM EDT)
Impact: DocV and Watchlist Monitors
Status: Resolved

1. Summary

An unexpected depletion of backend database capacity temporarily prevented our platform from retrieving account information for a subset of client requests.

During the incident window (10:04 AM to 10:31 AM EDT of 8-Sep-2026), some of the client requests experienced elevated error rates when attempting to query account details.

Following a routine maintenance update, concurrent connections from backup environments exceeded normal thresholds, saturating available database capacity.

Engineering teams safely disconnected the redundant environments and refreshed active connections, restoring full system responsiveness and stability.

2. Timeline

Date / Time (ET) Event
First error  10:04 AM EST, 8 Sep 2026 
Incident reported 10:17 AM EST, 8 Sep 2026 
Redundant environment connection pool isolated 10:30 AM EST, 8 Sep 2026 
Incident resolved 10:31 AM EST, 8 Sep 2026 

3. Root Cause

  • Primary Root Cause: During a scheduled system update, elevated connection retention between deployment environments led to database connection pool saturation. Under peak traffic, this caused processing delays and request timeouts for dependent platform services.
  • Contributing Factors:

    • Rigid connection timeout parameters delayed the timely release of idle database allocations by passive environment.
    • Inadequate database node headroom hindered seamless processing of extra connections generated by the standby environment.
  • Non-contributing Factors:

    • None

4. Resolution

Full service integrity was restored through the following immediate actions:

  1. Isolating non-active environment connections to immediately free database pool capacity.
  2. Performing a rolling reset of active service nodes to re-establish clear database connection baselines.

5. Corrective and Preventive Actions

Action Description ETA / Status
Improvements to data access patterns Improve caching strategy and data access patterns in internal services that rely on accounts data. 15 Oct 2026
Automated Connection Management Enhance connection lifecycle management and automated cleanup routines for inactive background tasks. 15 Oct 2026

6. Lessons Learned

  • Enhancing Deployment Isolation: We are updating our deployment procedures to ensure complete separation of resources across active and passive environments, eliminating cross-environment resource contention during upgrades.
  • Optimizing Database Resource Allocation: To ensure peak performance, we are rebalancing our workload and increasing overall system capacity and efficiency.

7. Next Steps & Ongoing Commitment

We are committed to:

  • Resolve the identified root causes by the timelines shared
  • Implement proactive safety guards and monitoring thresholds while permanent application and infrastructure enhancements are finalized.
Posted Sep 17, 2026 - 15:52 EDT

Resolved

The issue impacting DocV and Watchlist monitoring services on September 8 has been resolved. No customer impact has been observed since 10:35 AM EST. We apologize for any disruption this may have caused.
Posted Sep 08, 2026 - 11:41 EDT

Investigating

We are currently investigating an issue that created an impact on DocV and Watchlist monitoring services where errors were observed today September 8th, from 09:05 AM EST to 09:30 AM EST.

System has recovered, we are working on determining a root cause at this time.
Posted Sep 08, 2026 - 10:54 EDT
This incident affected: Global Watchlist (Global Watchlist Screening with Monitoring), Socure ID+ Platform (Predictive DocV), and RiskOS Platform (Predictive DocV).