- Platform: WD MyCloud PR2100, WD OS5 5.33.102, x86_64
- PMS: 1.43.4.10903-e5521bd8c, but note clearly that the problem has occurred across multiple earlier releases too.
- Symptom: PMS becomes completely unresponsive while the process remains alive. Even
http://127.0.0.1:32400/identityhangs. - Healthy baseline: 22 threads, 89 open FDs, ~73 MB RSS, localhost
/identityresponds immediately. - Hung state: 78–79 threads, 184 FDs, ~78 MB RSS, no swap use.
- Critical evidence: 50 threads are named
PMS ReqHandler, and all 50 were in syscall 202 (futex). Of those, 46 were waiting on exactly the same futex address. - Sockets: very large number of port 32400 sockets in
CLOSE_WAIT, mostly from the client/router address, plus the localhost diagnostic requests. - Database: both Plex SQLite databases pass
PRAGMA integrity_check; current hang had only 891 main DB WAL frames and zero blobs WAL frames, so this particular freeze was not caused by a huge WAL backlog. - Prior distinct issue: an older scheduled DB-backup deadlock was identified and database backup was disabled. This newer hang still occurs without that condition.
- No evidence of: OOM kill, Plex crash, database corruption, whole-NAS failure, or excessive PMS memory consumption.
- Recovery: restarting Plex immediately restores service.
- Frequency: the latest instance ran about 27 hours before becoming wedged.
- Client environment: Plex Web in Chrome was open for testing. Sony Google TV and NVIDIA Shield are also Plex clients, behind the Google Home network/NAT.
During the latest hang, the same PMS process was still alive, but localhost /identity would not respond. PMS had grown to 78-79 threads and 184 open FDs, while memory remained normal at ~78 MB RSS with no swap use.
The key finding was 50 threads named PMS ReqHandler. All 50 were blocked in x86-64 syscall 202 (futex), and 46 of those 50 were waiting on the exact same futex address:
46 0x7fd21e6d4880
3 0x7fd2184d2df0
1 0x7fd220ba3eb0
netstat also showed a large number of Plex port 32400 connections in CLOSE_WAIT.
SQLite does not appear to be the cause of this hang. Both Plex databases pass integrity checks, and during the hang the main DB WAL had only 891 frames while the blobs DB had 0.
The PMS log stops abruptly around 17:14 while it was still successfully answering requests. The last sequence included successful /media/providers and /status/sessions requests, a /:/prefs request that does not appear to complete, and a successful /updater/status. When checked later, PMS was still running but completely unresponsive.
No OOM kill, process crash, database corruption, or whole-NAS failure was found. Restarting PMS immediately restores normal operation.
This appears to be an internal PMS synchronization/deadlock issue where request-handler threads accumulate behind the same futex. I can provide full logs and additional /proc diagnostics if useful.
Note: ChatGPT helped me run through all the commands, which was helpful and identified the hung threads, which is beyond my art and knowledge.
Short term solution is a cron job that checks if it’s hung and restarts automatically.