Skip to content

lmdb.BadRslotError on Apple Silicon (arm64) blocks all signify-py bootstraps #436

Description

Summary

On Apple Silicon (M-series, arm64) hosts running Docker Desktop, every signify-py client.boot() against a KERIA cloud-agent deterministically fails with lmdb.BadRslotError: mdb_txn_renew: MDB_BAD_RSLOT: Invalid reuse of reader locktable slot. The except Exception: handler in keri.app.habbing.SignifyHab.processEvent masks this as Improper Habitat event type=icp, which is what most users will see — making the underlying LMDB issue invisible.

I'm using this for a master's thesis at HSLU (Switzerland), spent a day diagnosing it, and have a working patch. Filing this both as a bug report and a documentation pointer for other Apple Silicon users hitting the same wall.

Environment

Host Apple M2 Pro, 10 cores, macOS 26.4.1 (arm64)
Docker 25.0.3 / engine 29.4.1, Docker Desktop with arm64 backend
KERIA images tested weboftrust/keria:0.4.0, weboftrust/keria:latest, gleif/keria:0.3.0, custom build from main HEAD
Container internal Linux 6.12.76-linuxkit aarch64, Python 3.12.8, lmdb==1.6.2
signify-py 0.4.0 and 0.4.1 both reproduce
Witness pool gleif/keri:1.2.9 (3 production witnesses, kli witness start, TOAD=2)

All KERIA images are confirmed arm64-native (docker image inspect shows Arch: arm64). This is not an emulation issue.

Reproduction

  1. Start KERIA on arm64 host
  2. Generate a fresh client AID with signify-py and call client.boot() followed by client.connect()
  3. KERIA crashes during the inception event processing

Minimum repro:

from signify.app.clienting import SignifyClient
c = SignifyClient(url='http://org-keria:3901', passcode='0A...', tier='low')
c.boot()       # 200 OK, but...
c.connect()    # KeyError: 'controller' (because the agent crashed)

KERIA container exits with a KeyboardInterrupt cascade from the HIO doist; the actual triggering exception is buried.

Real root cause (after un-masking)

keri/app/habbing.py:2455 swallows the original exception:

def processEvent(self, serder, sigers):
    try:
        self.kvy.processEvent(serder=serder, sigers=sigers)
    except Exception:
        raise kering.ConfigurationError(f\"Improper Habitat event type={serder.ked['t']} for \"
                                        f\"pre={self.pre}.\")

Patching this to except Exception as ex: raise ... from ex reveals the real cause:

lmdb.BadRslotError: mdb_txn_renew: MDB_BAD_RSLOT: Invalid reuse of reader locktable slot

  File \".../keri/db/dbing.py\", line 552, in getVal
    with self.env.begin(db=db, write=False, buffers=True) as txn:

This is py-lmdb's thread-local reader-slot table racing under HIO's coroutine scheduling on Apple Silicon's Docker backend (linuxkit hypervisor). Reader slots assigned by one coroutine are not properly released before another coroutine attempts to renew them. Same problem hits both iss/ixn/icp event paths and the partial-witness escrow loop.

Not reproducible on linux/amd64 in CI — explains why this hasn't surfaced upstream.

Working patch

Open every LMDB environment with lock=False. KERIA is single-process per container, so the cross-process locking that the lock-table provides is dead weight; disabling it eliminates the reader-slot race entirely.

Patch in keri/db/dbing.py:419:

# before
self.env = lmdb.open(self.path, max_dbs=self.MaxNamedDBs, map_size=self.MapSize,
                     mode=self.perm, readonly=self.readonly)

# after
self.env = lmdb.open(self.path, max_dbs=self.MaxNamedDBs, map_size=self.MapSize,
                     mode=self.perm, readonly=self.readonly, lock=False)
try:
    self.env.reader_check()  # belt-and-suspenders cleanup of stale slots
except Exception:
    pass

With this patch applied on top of gleif/keria:0.3.0, client.boot()/connect() succeeds and full delegated-AID bootstraps run end-to-end on Apple Silicon. Verified with two independent organisations, fresh state each time.

(Tried notls=True first — that kwarg is not exposed in py-lmdb 1.6.2.)

Side-finding: Suber.rem(key=...) cleanup-path bug

While debugging, keria/src/keria/app/agenting.py:378 (and equivalent line on main) still calls Suber.rem(key=...). Since at least keri 1.2.6, this method signature is rem(keys=...) — see keri/db/subing.py:348-356. The call only fires on the post-failure cleanup path, so it's not the root cause, but it makes the cascade noisier:

TypeError: Suber.rem() got an unexpected keyword argument 'key'

Trivial fix: rename to keys=.

Suggested upstream actions

  1. Stop masking exceptions in SignifyHab.processEvent (and the other except Exception: block at habbing.py:2561) — at minimum, chain the cause (raise ... from ex).
  2. Add lock=False to the LMDBer.reopen() open call (keri/db/dbing.py:419) — safe because KERIA assumes single-writer-per-DB anyway.
  3. Fix Suber.rem(key=...) → rem(keys=...) in the agency-delete cleanup.
  4. Document Apple Silicon as a supported development target if you'd like to widen the contributor base — the issue surfaces immediately for anyone trying to develop against KERIA from an M-series Mac.

Happy to open a PR for any of these if helpful — the patch script I'm using locally is in my fork.

References

  • LMDB MDB_NOTLS / lock=False semantics: https://lmdb.readthedocs.io/
  • This issue was diagnosed during HSLU master's thesis work on a KERI/vLEI IoT platform.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions