Skip to content

[TEST] Neo4j conformance in CI: ratcheted three-route differential against pinned Neo4j, allocation gate for hot paths, and Neo4j TestKit for driver-level conformance #754

Description

@mm0nst3r

Summary

Divergences from Neo4j keep being found by manual differential runs after they reach main, not by CI. Examples from the last two days:

Each was found by comparing NornicDB with the pinned Neo4j over several routes, or by comparing allocations against the previous main. None of these runs is in CI.

This issue proposes three CI additions so that a divergence from Neo4j, or a hot-path allocation regression, fails the PR that introduces it.

What CI has today

.github/workflows/cypher-conformance.yml:

  • make cypher-conformance: the openCypher TCK over Bolt, with a ratchet.
  • make cypher-differential (scripts/cypher-tck/run-differential.sh): starts the pinned neo4j:5.26.30-community image and checks that a fixed corpus returns the same results as Neo4j (TestFixedDifferentialCorpusMatchesPinnedNeo4j).

The differential covers only statements that already match, and only Bolt auto-commit.

Proposal

1. A ratcheted, three-route differential

2. An allocation gate for hot paths

Allocation counts at a fixed iteration count (-benchtime=Nx -benchmem) are deterministic, unlike timings on shared runners. A job builds the base and the head, runs the cypher and storage benchmarks at a fixed count, and fails when any benchmark allocates more per operation than on the base, unless the PR lists it with a reason. Timing comparisons stay out of this job (they need a dedicated runner).

3. Neo4j TestKit (driver-level conformance)

neo4j-drivers/testkit is Neo4j's own harness that runs each official driver (Go, Python, Java, JavaScript, .NET) against a server. Its server-backed test suites, run against NornicDB, exercise Bolt-level behaviour through real drivers: value types, transactions, bookmarks, error classification and retry behaviour. Today these are checked only indirectly through one driver.

  • Run it as a separate workflow, nightly and on demand, with Python and Go to start, against NornicDB and the pinned Neo4j side by side.
  • Record a baseline of expected failures (tests needing cluster routing or Enterprise features) as a ratchet, the same way as the TCK. Fail only when a test that passed starts failing.

Acceptance

  • A PR that changes a statement's result, error code or written graph state relative to Neo4j (for a statement that matched before) fails CI, on whichever route it differs.
  • A PR that adds allocations per operation to any cypher or storage benchmark fails CI unless the increase is declared.
  • TestKit runs nightly with a committed baseline, and its report lists newly failing tests.

Consequence (of not having this)

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions