Skip to content

[Bug] add_episode prompts are not order-stable: node labels, dedup candidates and previous episodes come in hash or database order #1931

Description

@nvinnikov

Before you file

  • I searched existing issues and did not find a duplicate.
  • I read the contributing guide and included a minimal reproduction.

Affected component

graphiti-core

Bug description

The prompts that add_episode builds are not order-stable. The same episode, ingested twice from identical model answers, produces prompts that differ only in the order of their lists, so the model is asked a slightly different question on each run. There are four sources:

  1. Labels of a freshly extracted node. _create_entity_nodes in utils/maintenance/node_operations.py (L313 on main) builds them as list({'Entity', str(entity_type_name)}). The iteration order of a set of strings depends on string hashes, which are salted per process (PYTHONHASHSEED), so the same node gets ["Place", "Entity"] in one process and ["Entity", "Place"] in another. These labels reach the prompts of the same episode as "entity_types": [...].
  2. Labels of a node read back from the database. get_entity_node_from_record in nodes.py (L1063) takes labels in the order Neo4j returns them from labels(n), which is not guaranteed. We observed both orders for the same node across runs.
  3. Dedup candidates. resolve_extracted_edge in utils/maintenance/edge_operations.py (L700) numbers related_edges and existing_edges in search-result order ({'idx': i, 'fact': ...}). Ties and embedding noise reorder the candidates, and the model's answer refers back to idx.
  4. Previous episodes. add_episode(previous_episode_uuids=[...]) loads them with EpisodicNode.get_by_uuids (graphiti.py L1153), which returns them in database order rather than the caller's order, so <PREVIOUS_MESSAGES> changes order between runs.

All four are still present in v0.30.2 and on main (checked 2026-09-27).

Minimal reproduction

Source 1 needs no database. The set-to-list conversion used for node labels gives different orders under different hash seeds:

PYTHONHASHSEED=0 python -c "print(list({'Entity', 'Place'}))"   # ['Place', 'Entity']
PYTHONHASHSEED=1 python -c "print(list({'Entity', 'Place'}))"   # ['Entity', 'Place']

Since Python randomises the seed per process by default, two runs of the same ingestion render different entity_types lists into the extraction, dedupe and summary prompts for the same node.

For sources 2–4: ingest the same episode twice into an empty Neo4j graph with an LLM client that returns recorded answers keyed by a hash of the prompt. The graphs are identical, but some prompts miss the cache because only the ordering of entity_types, the idx candidate list or <PREVIOUS_MESSAGES> differs.

Expected behavior

Given the same inputs and the same model answers, add_episode sends byte-identical prompts, so recorded answers can be replayed by prompt hash and a live model sees the same question each time.

Actual behavior

With identical answers the resulting graph is identical (we compared exports byte for byte), but the prompts are not, so replaying recorded answers by prompt hash misses, and live runs present differently ordered prompts.

Environment

graphiti-core 0.30.1 (also reproduced by reading the v0.30.2 and main sources), Neo4j 5.26 community, Python 3.12, macOS.

Logs or traceback

No error is raised; the symptom is prompt drift. Example: the same node rendered as "entity_types": ["Place", "Entity"] in one run and ["Entity", "Place"] in the next.

Proposed fix

Five one-line changes, which we run locally as a patch:

  • nodes.py: labels = sorted(record.get('labels', []))
  • node_operations.py: labels: list[str] = sorted({'Entity', str(entity_type_name)})
  • edge_operations.py: related_edges = sorted(related_edges, key=lambda e: (e.fact, e.uuid)), and the same for existing_edges, before the idx context is built (later index lookups use the same local names)
  • graphiti.py: after get_by_uuids, reorder previous_episodes to follow previous_episode_uuids

Happy to open a PR with these if the approach is acceptable.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions