Summary
On Apple Silicon (M-series, arm64) hosts running Docker Desktop, every signify-py client.boot() against a KERIA cloud-agent deterministically fails with lmdb.BadRslotError: mdb_txn_renew: MDB_BAD_RSLOT: Invalid reuse of reader locktable slot. The except Exception: handler in keri.app.habbing.SignifyHab.processEvent masks this as Improper Habitat event type=icp, which is what most users will see — making the underlying LMDB issue invisible.
I'm using this for a master's thesis at HSLU (Switzerland), spent a day diagnosing it, and have a working patch. Filing this both as a bug report and a documentation pointer for other Apple Silicon users hitting the same wall.
Environment
|
|
| Host |
Apple M2 Pro, 10 cores, macOS 26.4.1 (arm64) |
| Docker |
25.0.3 / engine 29.4.1, Docker Desktop with arm64 backend |
| KERIA images tested |
weboftrust/keria:0.4.0, weboftrust/keria:latest, gleif/keria:0.3.0, custom build from main HEAD |
| Container internal |
Linux 6.12.76-linuxkit aarch64, Python 3.12.8, lmdb==1.6.2 |
| signify-py |
0.4.0 and 0.4.1 both reproduce |
| Witness pool |
gleif/keri:1.2.9 (3 production witnesses, kli witness start, TOAD=2) |
All KERIA images are confirmed arm64-native (docker image inspect shows Arch: arm64). This is not an emulation issue.
Reproduction
- Start KERIA on arm64 host
- Generate a fresh client AID with signify-py and call
client.boot() followed by client.connect()
- KERIA crashes during the inception event processing
Minimum repro:
from signify.app.clienting import SignifyClient
c = SignifyClient(url='http://org-keria:3901', passcode='0A...', tier='low')
c.boot() # 200 OK, but...
c.connect() # KeyError: 'controller' (because the agent crashed)
KERIA container exits with a KeyboardInterrupt cascade from the HIO doist; the actual triggering exception is buried.
Real root cause (after un-masking)
keri/app/habbing.py:2455 swallows the original exception:
def processEvent(self, serder, sigers):
try:
self.kvy.processEvent(serder=serder, sigers=sigers)
except Exception:
raise kering.ConfigurationError(f\"Improper Habitat event type={serder.ked['t']} for \"
f\"pre={self.pre}.\")
Patching this to except Exception as ex: raise ... from ex reveals the real cause:
lmdb.BadRslotError: mdb_txn_renew: MDB_BAD_RSLOT: Invalid reuse of reader locktable slot
File \".../keri/db/dbing.py\", line 552, in getVal
with self.env.begin(db=db, write=False, buffers=True) as txn:
This is py-lmdb's thread-local reader-slot table racing under HIO's coroutine scheduling on Apple Silicon's Docker backend (linuxkit hypervisor). Reader slots assigned by one coroutine are not properly released before another coroutine attempts to renew them. Same problem hits both iss/ixn/icp event paths and the partial-witness escrow loop.
Not reproducible on linux/amd64 in CI — explains why this hasn't surfaced upstream.
Working patch
Open every LMDB environment with lock=False. KERIA is single-process per container, so the cross-process locking that the lock-table provides is dead weight; disabling it eliminates the reader-slot race entirely.
Patch in keri/db/dbing.py:419:
# before
self.env = lmdb.open(self.path, max_dbs=self.MaxNamedDBs, map_size=self.MapSize,
mode=self.perm, readonly=self.readonly)
# after
self.env = lmdb.open(self.path, max_dbs=self.MaxNamedDBs, map_size=self.MapSize,
mode=self.perm, readonly=self.readonly, lock=False)
try:
self.env.reader_check() # belt-and-suspenders cleanup of stale slots
except Exception:
pass
With this patch applied on top of gleif/keria:0.3.0, client.boot()/connect() succeeds and full delegated-AID bootstraps run end-to-end on Apple Silicon. Verified with two independent organisations, fresh state each time.
(Tried notls=True first — that kwarg is not exposed in py-lmdb 1.6.2.)
Side-finding: Suber.rem(key=...) cleanup-path bug
While debugging, keria/src/keria/app/agenting.py:378 (and equivalent line on main) still calls Suber.rem(key=...). Since at least keri 1.2.6, this method signature is rem(keys=...) — see keri/db/subing.py:348-356. The call only fires on the post-failure cleanup path, so it's not the root cause, but it makes the cascade noisier:
TypeError: Suber.rem() got an unexpected keyword argument 'key'
Trivial fix: rename to keys=.
Suggested upstream actions
- Stop masking exceptions in
SignifyHab.processEvent (and the other except Exception: block at habbing.py:2561) — at minimum, chain the cause (raise ... from ex).
- Add
lock=False to the LMDBer.reopen() open call (keri/db/dbing.py:419) — safe because KERIA assumes single-writer-per-DB anyway.
- Fix
Suber.rem(key=...) → rem(keys=...) in the agency-delete cleanup.
- Document Apple Silicon as a supported development target if you'd like to widen the contributor base — the issue surfaces immediately for anyone trying to develop against KERIA from an M-series Mac.
Happy to open a PR for any of these if helpful — the patch script I'm using locally is in my fork.
References
- LMDB
MDB_NOTLS / lock=False semantics: https://lmdb.readthedocs.io/
- This issue was diagnosed during HSLU master's thesis work on a KERI/vLEI IoT platform.
Summary
On Apple Silicon (M-series, arm64) hosts running Docker Desktop, every
signify-pyclient.boot()against a KERIA cloud-agent deterministically fails withlmdb.BadRslotError: mdb_txn_renew: MDB_BAD_RSLOT: Invalid reuse of reader locktable slot. Theexcept Exception:handler inkeri.app.habbing.SignifyHab.processEventmasks this asImproper Habitat event type=icp, which is what most users will see — making the underlying LMDB issue invisible.I'm using this for a master's thesis at HSLU (Switzerland), spent a day diagnosing it, and have a working patch. Filing this both as a bug report and a documentation pointer for other Apple Silicon users hitting the same wall.
Environment
weboftrust/keria:0.4.0,weboftrust/keria:latest,gleif/keria:0.3.0, custom build frommainHEADlmdb==1.6.20.4.0and0.4.1both reproducegleif/keri:1.2.9(3 production witnesses,kli witness start, TOAD=2)All KERIA images are confirmed arm64-native (
docker image inspectshowsArch: arm64). This is not an emulation issue.Reproduction
client.boot()followed byclient.connect()Minimum repro:
KERIA container exits with a
KeyboardInterruptcascade from the HIO doist; the actual triggering exception is buried.Real root cause (after un-masking)
keri/app/habbing.py:2455swallows the original exception:Patching this to
except Exception as ex: raise ... from exreveals the real cause:This is
py-lmdb's thread-local reader-slot table racing under HIO's coroutine scheduling on Apple Silicon's Docker backend (linuxkit hypervisor). Reader slots assigned by one coroutine are not properly released before another coroutine attempts to renew them. Same problem hits bothiss/ixn/icpevent paths and the partial-witness escrow loop.Not reproducible on linux/amd64 in CI — explains why this hasn't surfaced upstream.
Working patch
Open every LMDB environment with
lock=False. KERIA is single-process per container, so the cross-process locking that the lock-table provides is dead weight; disabling it eliminates the reader-slot race entirely.Patch in
keri/db/dbing.py:419:With this patch applied on top of
gleif/keria:0.3.0,client.boot()/connect()succeeds and full delegated-AID bootstraps run end-to-end on Apple Silicon. Verified with two independent organisations, fresh state each time.(Tried
notls=Truefirst — that kwarg is not exposed inpy-lmdb1.6.2.)Side-finding:
Suber.rem(key=...)cleanup-path bugWhile debugging,
keria/src/keria/app/agenting.py:378(and equivalent line onmain) still callsSuber.rem(key=...). Since at leastkeri 1.2.6, this method signature isrem(keys=...)— seekeri/db/subing.py:348-356. The call only fires on the post-failure cleanup path, so it's not the root cause, but it makes the cascade noisier:Trivial fix: rename to
keys=.Suggested upstream actions
SignifyHab.processEvent(and the otherexcept Exception:block at habbing.py:2561) — at minimum, chain the cause (raise ... from ex).lock=Falseto theLMDBer.reopen()open call (keri/db/dbing.py:419) — safe because KERIA assumes single-writer-per-DB anyway.Suber.rem(key=...)→rem(keys=...)in the agency-delete cleanup.Happy to open a PR for any of these if helpful — the patch script I'm using locally is in my fork.
References
MDB_NOTLS/lock=Falsesemantics: https://lmdb.readthedocs.io/