This guide explains how to build, test, and extend Didi (godot-mcp-native).
Two Python floors, and they are different numbers. The C++ build needs 3.9 or
newer, which is what a stock macOS ships through the Xcode Command Line Tools;
CMakeLists.txt states it, so a generator step that slips past it fails at
configure with a version rather than inside the generator. The Python test
suite needs 3.10 or newer, because requirements-dev.txt pins
jsonschema==4.26.0 and that release declares requires-python >= 3.10. On 3.9
the C++ build and its native suite are fine and pip install -r requirements-dev.txt fails with a resolution error that never names a Python
version.
Install requirements-dev.txt into the Python environment used for tests.
If it differs from the default interpreter, pass its absolute executable to
CMake with -DPython3_EXECUTABLE=... so generators and CTest use that environment.
- Visual Studio 2022 / Build Tools with C++20 support
- CMake 3.20+
- Python 3.9+ to build, 3.10+ to run the Python suite
- Godot 4.5+
- Windows PowerShell 5.1 or newer for the live integration harness (PowerShell 7 also works)
# Generate CMake solution
cmake -B build -S .
# Compile Release binaries
cmake --build build --config Release
# Run both registered suites
ctest --test-dir build -C Release --output-on-failurecmake -B build -S . -DCMAKE_BUILD_TYPE=Release
cmake --build build -j$(nproc)
ctest --test-dir build --output-on-failure| Option | Default | Effect |
|---|---|---|
DIDI_ELASTIC_INGRESS |
OFF |
Enables the explicitly requested safe-v1 profile on three read-only tools; strict calls stay strict. See Elastic Ingress for the contract and paired-build verification. |
DIDI_BUILD_TESTS |
ON |
Builds the native suite, the Phase 7 signal bridge fixture and its probe, and registers both CTest entries. OFF produces only didi and didi_extension, and the configure fails if a test-only target is defined anyway. |
DIDI_ENABLE_SANITIZERS |
OFF |
ASan and UBSan on GCC and Clang. MSVC has no supported combination here, so the configure refuses rather than building without the runtime. CI runs this on Ubuntu. |
ctest runs the native suite and the Python suite, the latter discovered rather
than listed, so the two cannot drift. Running didi_tests by path still works
and is what the --list and --filter flags below are for.
didi/
├── include/didi/
│ ├── common/ # Result<T>, Error, Logger, Base64/PNG, JSON, STB, IPC channels
│ ├── mcp/ # JSON-RPC 2.0, MCP server, tool/resource/prompt registries
│ ├── offline/ # GDScript diagnostics, resource indexer, test runner
│ └── gdextension/ # GDExtension interface, editor queue, Godot bridge, viewport renderer
├── src/
│ ├── common/ # Platform IPC (Win32 Named Pipes, POSIX sockets)
│ ├── mcp/ # MCP protocol handlers
│ ├── offline/ # AST analysis, file indexing, headless subprocess runner
│ ├── tools/ # Public tool handlers for the canonical and legacy surfaces
│ ├── gdextension/ # In-engine GDExtension module & renderer
│ └── standalone/ # main.cpp entry point for didi.exe
├── tests/ # Native suite plus real Godot smoke fixture/harness
├── addons/didi/ # Godot addon manifest, including the editor console's GDScript
│ # and brand marks; the built addon is staged in build/addons/didi
└── demo/ # Reference Godot 4 test project
Test Inventory carries the current totals, derived from the suites themselves and regenerated by python tools/test_inventory.py; CI runs --check after the build and fails when the page and the suites disagree. The native runner's reported total is authoritative as the suite evolves. Coverage is organized by contract rather than duplicated here as a brittle test-name inventory:
- JSON-RPC/MCP lifecycle, registration counts, capability metadata, output redaction, and structured errors.
- IPC framing, bounded transport waits, editor-hook state transitions, and single-response ownership.
- Descriptor validation, PID/start identity, opened-handle TOCTOU defenses, host publication, retirement, and cleanup races.
- Native checkpoint-store tests cover snapshot bounds, manifests, corruption, unsafe paths/links, and restore staging; managed-process tests cover owned-child identity, launch, exit, and termination behavior.
- Transactional attach, deterministic auto-selection, fresh identity handshakes, route supersession, kind-aware availability, deadlines, and quarantine.
- Runtime log cursor/gap/filter behavior, UTF-8 and payload bounds, runtime-tree bounds, and expression-sandbox policy.
- Tool/resource live and offline provenance, viewport/image encoding, GDScript diagnostics/patching/reflection, and resource indexing.
- Blackboard path rejection, atomic patching, expiry, bounds, and the cross-process file lock; task claim exclusivity under concurrent claimers, lease expiry and reclaim, dependency readiness, cycle refusal, and lease ownership on update and complete.
kServerInstructions in src/mcp/mcp_server.cpp is the shared compiled guide
for initialize and server/discover; docs/LLM_INSTRUCTIONS.md is the expanded
reference, not a file loaded at runtime. When a documented route or boundary
changes, update both and the API contract.
Keep the guide compact and do not embed project paths, client data or live
availability in it.
McpServer.InitializeInstructions checks the serialized field and resolves
named routes against the registry. It also holds the guide to what discovery
says about them: a tool the guide tells hosts not to call must still be
unimplemented, every other tool it names must be implemented, and every
parameter it spells as tool(param, ...) must be in that tool's inputSchema.
Implementing or retiring a tool the guide names, or renaming a parameter it
cites, fails the test until the guide is updated to match.
tests/test_initialize_instructions.py
drives the actual binary through malformed-handshake recovery, protocol
fallback, pipelined/batched requests, concurrent clients and the guide's
read workflow. Run it with python -m unittest tests.test_initialize_instructions -v.
The normalization wire suite adds boundary and per-request opt-in tests; run
tests.test_elastic_ingress with both compile-time profiles, setting
DIDI_TEST_ELASTIC_INGRESS to 0 or 1 to verify the binary being tested.
New Python modules must participate in discovery and an explicit invocation
in the existing CI workflow. Support both tests.<module> imports and discovery
with -t tests. For simple test-owned stdio children, use
tests/stdio_process.py to kill if needed, drain/close pipes and wait with a
finite timeout. Managed editor fixtures keep their process-tree ownership
cleanup. Avoid shared project/session state in new real-wire fixtures.
See the 2026-09-25 exploratory report for measured latency, cross-version live coverage, harness fixes and outstanding issues. Its local results do not substitute for cross-platform CI.
tests/contract_snapshots/ holds what a client is shown, so a change to it is a
diff a reviewer reads rather than a surprise a client finds
(Q3).
offline.json:initialize,tools/list,resources/list,resources/templates/listandprompts/listfrom a server with no engine. Entries are keyed by name, because the wire order is a hash map's and differs between standard libraries.offline-core.json: what--tools corechanges about those listings. A tool the profile leaves out is one<absent>line.live-<line>.json, one for each engine line in CI's matrix: what attaching an editor changes in those listings, and the answers to the calls incalls.json. The calls run against a fresh copy oftests/contract_fixture/in a headless editor with its own session directory and editor settings.calls.json: the read-only calls, and every implemented read-only tool that is not called, with the reason. A call that should be refused carries"expect_error": true; any other error fails the recording.
Session ids, pids, the pipe endpoint, the build id, the server version,
temporary paths, durations and timestamps become placeholders. The session id,
endpoint, build id, version and paths are replaced by value, taken from the
session Didi reports, so they are caught wherever they appear. The pid and the
start time are short numbers that could match an unrelated value, so they are
replaced by key, as durations are, and timestamps inside strings by pattern.
A value none of these rules covers fails the two-recording check below; add
the rule to tools/contract_snapshots.py. A text block that repeats
structuredContent is stored as a marker, and any other JSON text, which
includes every error, is stored parsed under text_json.
When a change to a schema or an answer is intended, regenerate in the same pull request and read the diff:
python tools/contract_snapshots.py --godot <Godot 4.5 exe> --godot <Godot 4.6 exe> --godot <Godot 4.7 exe>Regenerating records everything twice, each time from a fresh editor, and
writes nothing when the two disagree. The error names the JSON path that moved:
a value the recorder does not yet normalise, which belongs in
tools/contract_snapshots.py rather than in a snapshot. --check records once
and prints a diff, and without --godot only offline.json is recorded.
--binary names the server to record (the newest build by default, as the
Python suites pick it), --out <dir> writes what a run recorded even when it
failed, and --keep-work keeps each temporary project with its editor.log.
The editor is given its own DIDI_SESSION_DIR, APPDATA and GODOT_BIN, so
neither another editor on the machine nor your editor settings reach an answer.
CI checks offline.json in tests.test_contract_snapshots on all three build
platforms, and each live-<line>.json in the Godot job for that line. A failed
live check uploads what it recorded with the job's logs, so a contributor
without Godot can read it or adopt it. A change to the fixture changes every
live snapshot, so regenerate all three lines.
didi --tools full|core chooses at startup which tools a session lists (Q4).
tests/tool_list_budgets.json holds a byte budget for each profile's
tools/list, measured by tests/test_tool_profiles.py from a server with no
engine attached, and CI fails a listing over its budget. A new tool, or longer
prose, can exceed one. Raising a budget takes its own pull request that says
why, merged before the change that needs it, so growing what every session
pays for is a decision rather than a side effect.
core is the union of the tools recorded field trials reached, in
tools/field-trial/reached_tools.json, and every implemented tool the handshake
guide names. The names are written out as coreProfileTools() in
src/mcp/tool_registry.cpp, and the same test fails when that list drifts from
either source: after a trial, or after the guide starts naming another tool.
offline-core.json in the contract snapshots records what the profile changes
about the full listing.
A client that declares the didi/responseEconomy extension can decline the
text copy of structuredContent and a session descriptor it already holds
(Q5, API specification). The whole
of it is economizeToolResult in src/mcp/response_economy.cpp, applied in
one place: every tools/call answer in McpServer::handleRequest is encoded
through its encode lambda, so a new return path there must use it too.
Handlers never see the declaration and never need to. --session-descriptor once is the operator's way to the descriptor half for a client that declared
nothing; it enters in McpServer::responseEconomyFor, beside the declarations.
Every refusal names what fixes it: retry_with, field, next_call,
restart_with or retry_after_ms in error.data, or no_remedy saying why
nothing can (Q6). Say it at the site when the site knows. Otherwise the error
floor fills it from src/mcp/refusal_remedies.cpp, which is keyed by
data.code and may look at the tool. A new code needs an entry there, or one in
refusalsWithoutRemedy with the reason: tests/test_refusal_remedies.py scans
the source for every code and fails on one the table does not cover, and the
live harness fails a refusal that reaches a caller without a fix. Answer a
failure with CallToolResult::errorJson, fromError or notConnected, never
plain text, which has no error.data to carry a remedy. The constructor that
makes plain text is private, so a new one does not compile, and the harness
fails a plain-text failure it sees. RepeatedFailures, in
src/mcp/repeated_failures.cpp, marks the second identical failure of a call on
its way out of McpServer::handleRequest; it keys on the arguments as sent, so
nothing a handler does changes what counts as the same call.
A mutation that leaves work undone names it under follow_up (Q6).
src/mcp/follow_ups.cpp derives each step from a fact the answer already
carries, such as scene_saved: false or requires_editor_restart: true, so
publish the fact and the step follows. tests/follow_ups.json needs an entry
for every mutating tool: the work it can leave and the fact behind it, why it
leaves none, or an exemption with its issue. tests/test_follow_ups.py fails a
mutating tool with no entry and a declared kind of work no rule produces, and
the live harness fails a step a tool did not declare.
Every read-only tool has an entry in tests/bounded_reads.json: bounded,
naming what can cut its answer short, or unbounded, saying why nothing can
(Q5). A new read with neither fails tests/test_bounded_reads.py on every
build and the live harness's coverage check. A bounded tool puts a boolean
truncated on every successful answer, on every path, true when a bound was
reached; a limit that refuses the call instead makes a tool unbounded. The
harness records every bounded answer it makes through Invoke-Didi and fails
on one without the flag, and the Python test drives the offline paths and
forces a bound to bite.
fields comes from ToolDefinition::sections: list the top-level keys a
caller may choose between, never one the output schema requires, and
registerTool publishes the argument and omitted_fields, while
dispatchTool narrows the answer. It is for a large read made of separable
sections; a read that is one long list takes a bound instead.
A client that declares nothing must get the same bytes. The contract snapshots are the check: they record as a client that declared nothing, so any change this makes to that client's answers shows as a diff. The live harness runs #776's seven-call arc as both kinds of client on every engine line, and fails unless the declared arc costs under half the bytes.
The opt-in Python recovery suites launch real Godot through MCP stdio and cover owned-child crashes, saved-file persistence, the single restart, no replay, restore, source preservation, previews, and corrupt snapshots. Set DIDI_TEST_BINARY to the built host and DIDI_RECOVERY_GODOT to an absolute Godot executable:
$env:DIDI_TEST_BINARY = "D:/didi/build/Release/didi.exe"
$env:DIDI_RECOVERY_GODOT = "C:/Godot/Godot_v4.5.1-stable_win64.exe"
python -m unittest discover -s tests -t tests -p "test_managed_recovery*.py"See Managed Recovery verification for the live and adversarial suite scope. These are opt-in engine tests; a skipped test is not live recovery evidence.
tests/test_editor_startup_live.py launches a fresh editor on a project with a main scene, opens another scene as soon as the session is listed, and reads the edited scene for ten seconds (#1069). Set DIDI_STARTUP_GODOT to a Godot executable beside DIDI_TEST_BINARY; CI runs it on every engine line after the recovery suites. It retries an open the route deadline cut short, and it skips, rather than fails, only when the editor died on the documentation-thread crash #285 tracks. It keeps the editor's output and prints it on any other failure.
$env:DIDI_STARTUP_GODOT = "C:/Godot/Godot_v4.6.2-stable_win64_console.exe"
python -m unittest tests.test_editor_startup_live -vThe Windows live integration harness copies the tracked fixture into build/ and starts real Godot processes. It preserves the Phase 1/2 sequence, adds Phase 3 concurrent editor/game routing, and now exercises Phase 4 bounded search, SVG reimport, reversible isolation, capture IDs, mutation diffs, exact undo restoration, and cleanup. The Phase 8 import block runs a menu scene whose music stops, as the control, then sets the loop options on an OGG and a WAV import with asset_configure_import after a dry run, checks seven refusals, and runs the scene again to see the music still playing. The earlier coverage still checks scripts, groups, autoloads, nested settings, InputEvent forms, persistence rollback, scene lifecycle, resource ownership, unsafe paths, and honest errors:
.\tests\run_godot_integration.ps1 `
-GodotExecutable C:\Godot\Godot_v4.5.1-stable_win64_console.exeThe harness runs to completion on Windows PowerShell 5.1, including the persistence-rollback case that denies write rights through icacls. That case previously aborted the run: it selects its platform branch with $IsWindows, an automatic variable introduced in PowerShell 6, which is undefined on 5.1 and so took the POSIX branch and called chmod.
The Phase 1 substrate has also been run against Godot 4.6.2 and 4.7.2. The compatibility floor remains Godot 4.5.1; the live CI matrix runs the complete integration harness on 4.5.1, 4.6.2 and 4.7.2, which is every line the supported range covers, and bridge method hashes must remain valid on all three versions.
addons/didi/ carries the plugin that draws Didi's main screen tab. It is
GDScript, it compiles nothing, and it is the one part of Didi that is not behind
the native test boundary, so it is worked on differently from the rest.
- Adding a file to the addon means adding it to
DIDI_ADDON_MANIFEST_FILESinCMakeLists.txt, to theexpectedlist in the staged-addon step of.github/workflows/ci.yml(which compares againstLC_ALL=C sortorder), and todemo/addons/didi/.tests/test_editor_console.pyfails the build when any of those disagree with the addon directory, so the list cannot go stale silently — but it will not add the file for you. Four more places list the files and no suite checks them: theexpectedstrings intools/localci/lane.shandtools/localci/macos.sh, and the install trees indocs/QUICKSTART.mdanddocs/INTEGRATION_GUIDE.md. A new.gdships with the.uidsidecar an engine mints for it;tools/vibe/addon_script_engines.pyprints that uid and checks the script loads on every engine. A script the extension itself loads, asdidi_await.gdanddidi_import_watch.gdare, is also copied into the harness fixture bytests/run_godot_integration.ps1, because the fixture's addon is not the repository's. - The console must not become an MCP client. It reads the session descriptors Didi publishes and reports them. It calls no tool, speaks no part of the IPC protocol, and never reads the token out of a descriptor; a test enforces the last of those.
- Verify it in a real editor, on the floor and the ceiling. Copy the addon
into a scratch project and open it headless on 4.5.1 and on the newest engine
in CI:
Godot --headless --editor --path <project> --quit-after 6000. A GDScript parse error, a missing API on the 4.5 floor, and a tab that fails to build all show up there, and none of them show up in the native suite. - Brand marks are copies, not forks. The three SVGs in the addon are
byte-identical to
docs/brand/svg, held so by a test. Change the geometry indocs/brand/build.pyand copy the regenerated source across.
Do not register a success stub. A new name must either have a tested execution path or be classified as unimplemented and rejected.
A new name needs an accepted entry in Surface Amendments before it is registered. New capabilities normally arrive as an item in the Build Queue, whose section says what the tool has to do and when it is done, and whose Design Principles are the rules it has to keep.
To add a new tool (e.g. export_mesh_glb):
ToolDefinition export_tool;
export_tool.name = "export_mesh_glb";
export_tool.description = "Exports a target 3D node to a GLB file.";
export_tool.inputSchema = {
{"type", "object"},
{"properties", {
{"node_path", {{"type", "string"}, {"description", "Path to node"}}},
{"output_path", {{"type", "string"}, {"description", "Target .glb path"}}}
}},
{"required", {"node_path", "output_path"}}
};
export_tool.handler = [this](const json& args) {
return handleExportMesh(args, m_ipcClient);
};
registerTool(std::move(export_tool));tools/call checks arguments against this schema before it dispatches, so what
you declare here is enforced, not documentation. Declare required for every
argument the handler cannot do without. A tool's arguments are closed by
default, so a mistyped property name is a message naming it and listing what
the tool does take; you do not have to remember "additionalProperties": false,
and a new tool is covered on arrival. Declare "additionalProperties": true
only if the tool genuinely takes arguments it does not publish. Nested objects
are the other way round, because several of them are deliberately free-form
maps: close one explicitly when it has a fixed set of keys.
type, enum, required, minimum, maximum, minLength, maxLength,
minItems, maxItems, pattern and uniqueItems are all enforced. Do not
publish a keyword outside that list expecting it to hold; a schema keyword the
server ignores is worse than one it never advertised. Keep the handler's own
checks for anything the schema cannot express.
Update capabilityForTool in src/mcp/tool_registry.cpp. Choose only modes backed by tests: live, offline_fallback, both, or unimplemented.
If the handshake guide (kServerInstructions in src/mcp/mcp_server.cpp) names the tool, moving it between implemented and unimplemented also means updating the guide; see Protocol guidance.
CallToolResult handleExportMesh(const json& args, std::shared_ptr<ipc::IIpcClient> ipc) {
if (ipc && ipc->isConnected()) {
auto res = ipc->sendRequest("mesh.export", args);
if (res.isOk()) {
return CallToolResult::successJson(res.value());
}
return CallToolResult::error(res.error().message);
}
return CallToolResult::error("Godot Editor is offline.");
}Route the method from EditorHook::executeOnMainThread into a bounded implementation that performs real Godot calls on the main thread. Return a structured error whenever a required object, method bind, or engine operation is unavailable. Never return success metadata before the operation completes.
- Add native tests for capability metadata, offline behavior, and error propagation.
- Add a real Godot integration case for live behavior and UndoRedo where relevant.
- If the tool mutates, which is what gives it
dry_run, give it an entry intests/observed_post_state.json. Either list the answer fields it reads back after the write and add a case totests/observed_post_state.ps1that reads the same state throughtests/godot_smoke/observed_witness.gd, or exempt it with the reason and the issue that tracks the gap. A field copied from the request or read before the write is not observed, whatever its name.tests/test_observed_post_state.pyfails the build until the entry exists, and the live harness fails when an answer disagrees with the engine. - Regenerate the contract snapshots, because a new tool
changes
tools/list. It also grows the full profile's bytes, so it may need a budget raised first, in a pull request of its own. If the tool only reads, also give it a call intests/contract_snapshots/calls.json, or an exclusion with the reason it cannot be recorded;tests/test_contract_snapshots.pyfails the build until it has one. - Update Current Capability Matrix and Tool Reference.
src/runtime/checkpoint_store.cpp: bounded saved-file inventory, manifest/hash validation, checkpoint publication, retention, and restore staging.src/runtime/managed_process.cpp: owned-child launch, process identity, exit observation, logs, and termination.src/runtime/managed_recovery.cpp: copied-project lifecycle, readiness, pre/post checkpoints, reconciliation latch, one automatic restart, and preserved-project restore.src/runtime/session_client.cpp: descriptor discovery, opened-handle validation, cross-platform PID/process-start identity, transactional handshake on a finite deadline, token insertion, and local route state.src/gdextension/session_host.cpp: bind-before-publish editor/game endpoint lifecycle, private descriptor generation, authentication stripping, and safe no-replace descriptor retirement.src/gdextension/runtime_log.cpp: bounded 2,000-record ring, UTF-8-safe 16 KiB messages, 64 KiB details, cursor gaps, filtering, and logger sink mirroring.src/gdextension/runtime_bridge.cpp: SceneTree resolution, UTF-8 field limits, 10,000-node/256 KiB tree budgets, explicit truncation, pause verification, exact one-active-step state machine, shutdown cancellation, and stop requests.src/gdextension/expression_sandbox.cpp: tokenizer/policy, receiver-aware call validation, ClassDB-prebound native scalar property reads, context confinement, cooperative deadlines, and bounded Variant-to-JSON conversion.
The standalone router starts detached, then may auto-attach on first availability only to an unambiguous live canonical-project match: the sole session, or a unique editor among games. Preserve the tests that keep same-kind ambiguity detached, roll back failed handshakes, disable auto-selection after explicit attach/detach or quarantine, and make runtime_get_session revalidate the complete token-free identity. Keep local session management distinct from live engine calls. Never log or return descriptor tokens or full submitted expression source.
Routes are held per session, so several can be open at once. Two invariants keep that safe and both are covered by tests that fail without them. Every held route is released on shutdown, not only the selected one: the map holds the only reference to each route's ownership lock, so a route left in it is a session refused to every other Didi process until this one exits. And a request is only ever served on the session it named, which means new code that reaches for a route must take it by session id rather than by asking what the process is currently pointed at.
The release gate runs the complete native suite; the runner's reported total remains authoritative as cases evolve. Focused suites cover the existing session/routing/evaluation contracts plus search containment and lexical filtering, two-idle-frame reimport progress, exact diff arithmetic, cache eviction, public response completeness, and restoration guards. tests/run_godot_integration.ps1 creates disposable concurrent editor/game processes and verifies the complete live workflow on Godot 4.5.1, 4.6.2 and 4.7.2. Editor teardown first requests a normal window close, then uses PID-and-start-time-verified termination if a hidden Windows editor keeps an invisible native prompt alive; this fallback is limited to the disposable test process and cannot target a reused PID. Because forced exit cannot run the extension destructor, the harness removes a leftover descriptor only after its regular-file shape, session ID, PID, and process-start identity all match that editor instance.
Run from a clean worktree:
cmake --build build --config Release
.\build\Release\didi_tests.exe
.\tests\run_godot_integration.ps1 -GodotExecutable C:\Godot\Godot_v4.5.1-stable_win64_console.exe
.\tests\run_godot_integration.ps1 -GodotExecutable C:\Godot\Godot_v4.6.2-stable_win64_console.exe
.\tests\run_godot_integration.ps1 -GodotExecutable C:\Godot\Godot_v4.7.2-stable_win64_console.exeThe engine's own output is a result. The harness reads the editor's and
the game's --log-file after the run, prints every ERROR and WARNING line with
a count, and fails on any line that is not in $allowedEngineLines. Each
allowed entry names the request that causes the line on purpose, or the open
issue for a known defect. A new line means a request made the engine complain:
find the call (its answer carries the line under engine_diagnostics), then fix
the cause or, when the line is the point of the request, add an entry that says
which request and why. Never widen a pattern to make a run pass.
Lines the host causes -- a CI runner with no GPU or audio device, whose engine
says so while its drivers start -- are allowed by where in the engine they come
from (Where, matched against the at: line) as well as by what they say, so
the same message from anywhere else is still a finding. The engine's source
files move between lines: the audio fallback prints from
servers/audio_server.cpp on 4.5 and servers/audio/audio_server.cpp from
4.6, so read a new entry's at: line off all three engines before writing its
Where. A developer machine prints none of these, and the first CI run is the
first place they show.
Three things catch people out running the harness by hand on Windows.
Build every target, not just didi and didi_tests. The addon is its own
target, and the editor the harness starts loads whatever build of it was last
staged. A bridge change then looks absent rather than stale: the editor answers
with the contract it was built with, everything else looks normal, and the first
assertion to notice is whichever one covers your change, which reads exactly
like the change not working. The harness now fails on the first attach instead,
with the mismatch note Didi already reports in every session payload.
Run it on its own rather than piping it. Under Windows PowerShell 5.1,
redirecting a native command's stderr turns each line into an error record, and
Didi logs to stderr on startup, so .\tests\run_godot_integration.ps1 ... 2>&1 | Select-Object -Last 40 used to kill the run at its first exchange and report it
as didi.exe : [INFO ] Set working directory ..., which reads like Didi failed
to start. Every exchange now goes through one helper that relaxes the preference
for the duration of the call, so the run survives either way, but the output is
still easier to read unpiped. CI uses pwsh 7, where the redirection was always
harmless.
An IDE may connect to the harness's editor. The editor serves the GDScript
language server on port 6005, and an extension such as godot-tools connects to
whichever editor is there. On connect the editor re-parses every script, so the
fixture script that is broken on purpose prints Failed parse script from
reload_all_workspace_scripts. The allow list names that wording, so it shows
in the tally and passes. Any other line it brings is still a finding.
The native runner accepts --list to print every registered case and
--filter=<substring> to run a subset. The value goes after an = with no
space, and any other argument is refused by name with a non-zero exit: the space
form used to match nothing, leave the filter empty and run all of them, which
reads as one isolated test passing (#803):
.\build\Release\didi_tests.exe --list
.\build\Release\didi_tests.exe --filter=Tools.CaptureViewportWithIpcRun a test in isolation whenever it fails intermittently. The suite executes in a single process and shares the tool registry, the resource registry, and the working directory, so a case that does not register what it calls will pass only because an earlier case registered it, and will fail at a different assertion as the ordering changes. A test that passes in the full suite but fails alone is depending on a predecessor, not on the code under test.
For expression-policy changes, add a failing native scanner test and a real editor/game integration probe before changing implementation. A new accepted Node operation must prove it cannot dispatch script callbacks, traverse outside the active subtree, allocate unbounded data before a check, leak source/token text, or turn the cooperative timeout into a hard-preemption claim.
The CI MCP smoke must start Didi with an explicit fixture project. It verifies the live tools/list surface against the manifest emitted by didi --dump-tool-manifest from the same build, so counts are never written into the workflow, and it asserts every implemented flag rather than a sample. It also continues to assert offline-only search/deep-domain metadata, live-only reimport/diff/UI-hit-test metadata, strict Phase 4/5/6/7 schemas, local metadata for the four session tools, live metadata for routed runtime tools, cursor-shaped logs, implemented game input and profiler capabilities, and implemented: false only for physics_simulate_step, nav_bake_mesh, and runtime_get_call_stack. It is tests/test_mcp_wire_contract.py, so the Python suite runs it before a push, against the build tests/didi_binary.py picks.
Status: PARTIAL_DELIVERY
Canonical implementation: 117/120
Phase 7 registrations: 3/18 unimplemented
Feasibility: 15/18 implementation-feasible; 3/18 API-blocked
Phase 7 is PARTIAL_DELIVERY. The gate completed on 2026-08-29 against Godot 4.5.1 and 4.7.2: 15/18 names are implementation-feasible, while exactly 3/18 are API-blocked under their approved contracts: physics_simulate_step, nav_bake_mesh, and runtime_get_call_stack. For each blocker, no supported public API/semantics satisfying the exact approved contract was found on either tested version. Do not broaden that result into a permanent impossibility claim.
Feasibility is design evidence; a production trial is production behavior. All 15 feasible Phase 7 names are delivered, including the three TileMapLayer/GridMap tools. The 3 API-blocked names remain registered but unimplemented.
Governance authorized partial delivery, and all 15 feasible tools are now shipped, and the surface stands at 117/120 canonical implementations. Further work on physics_simulate_step, nav_bake_mesh, or runtime_get_call_stack requires new feasibility evidence on Godot 4.5.1 and 4.7.2 or an explicit contract amendment; do not weaken their contracts implicitly. Use PHASE_7_API_FEASIBILITY.md for reproducible evidence and PHASE_7_IMPLEMENTATION_PLAN.md for the approved executable plan.
src/tools/deep_domain_tools.cppandsrc/offline/deep_domain_support.cpp: bounded C#/shader diagnostics, public export-preset parsing, guarded export, deterministic MeshLibrary generation, and live UI hit-test registration/dispatch.include/didi/common/project_path.hpp: explicit project-root validation, canonical containment, and stable 16-hex project endpoint keys.src/runtime/session_lock.cpp: owner-only cross-platform OS locks and423exclusion for a second MCP owner.src/mcp/mutation_safety.cpp: mutation classification, schema decoration, handler-free previews, exact context binding, 120-second expiry, and single-use confirmation storage.tests/test_phase5.cpp,tests/test_phase6.cpp, thetests/test_phase7*.cppsuites, andtests/run_godot_integration.ps1: deep-domain contracts, project/lock/preview red-team cases, Phase 7 transport and validation contracts, and disposable Phases 1–7 Godot workflows.
When adding or reclassifying a mutation, update MutationSafety::isMutation, add dry_run schema coverage, and prove the dry-run never enters its handler. Add confirmation only for the documented high-risk set; changing that set is a public safety-contract change and requires updates to the Tool Reference, Capability Matrix, LLM instructions, and API specification.
Run the dependency-free documentation contract checks for any documentation, version, registration, capability, or release change:
python -m unittest tests.test_documentation_validator -v
python tools/validate_documentation.pyThe validator derives the release from CMakeLists.txt and checks the MCP server header, standalone version output, addon manifest, README, capability matrix, changelog, and security policy for alignment. It also locks the documented 120 canonical/10 legacy/130 total surface, the 117 implemented/3 unimplemented split, Phase 7's PARTIAL_DELIVERY status, 15/18 versus 3/18 feasibility result, exact three-tool blocker set, authoritative-record links, stale current-state prose, and all relative Markdown targets and anchors.
When the release changes, update these files in one change: CMakeLists.txt, include/didi/mcp/mcp_protocol.hpp, src/standalone/main.cpp, addons/didi/plugin.cfg, README.md, CHANGELOG.md, docs/CAPABILITIES.md, and SECURITY.md. When the tool surface or capability modes change, also update runtime discovery tests, docs/TOOL_REFERENCE.md, docs/ROADMAP.md, docs/LLM_INSTRUCTIONS.md, and the relevant quickstart/integration examples. Historical specs and plans record their original decisions and should not be rewritten as current release documentation.