Skip to content

Latest commit

 

History

History
558 lines (451 loc) · 38.4 KB

File metadata and controls

558 lines (451 loc) · 38.4 KB

Didi Developer & Extension Guide

This guide explains how to build, test, and extend Didi (godot-mcp-native).


🛠️ Build Environment Setup

Two Python floors, and they are different numbers. The C++ build needs 3.9 or newer, which is what a stock macOS ships through the Xcode Command Line Tools; CMakeLists.txt states it, so a generator step that slips past it fails at configure with a version rather than inside the generator. The Python test suite needs 3.10 or newer, because requirements-dev.txt pins jsonschema==4.26.0 and that release declares requires-python >= 3.10. On 3.9 the C++ build and its native suite are fine and pip install -r requirements-dev.txt fails with a resolution error that never names a Python version.

Install requirements-dev.txt into the Python environment used for tests. If it differs from the default interpreter, pass its absolute executable to CMake with -DPython3_EXECUTABLE=... so generators and CTest use that environment.

Windows (MSVC)

  • Visual Studio 2022 / Build Tools with C++20 support
  • CMake 3.20+
  • Python 3.9+ to build, 3.10+ to run the Python suite
  • Godot 4.5+
  • Windows PowerShell 5.1 or newer for the live integration harness (PowerShell 7 also works)
# Generate CMake solution
cmake -B build -S .

# Compile Release binaries
cmake --build build --config Release

# Run both registered suites
ctest --test-dir build -C Release --output-on-failure

Linux / macOS

cmake -B build -S . -DCMAKE_BUILD_TYPE=Release
cmake --build build -j$(nproc)
ctest --test-dir build --output-on-failure

Build options

Option Default Effect
DIDI_ELASTIC_INGRESS OFF Enables the explicitly requested safe-v1 profile on three read-only tools; strict calls stay strict. See Elastic Ingress for the contract and paired-build verification.
DIDI_BUILD_TESTS ON Builds the native suite, the Phase 7 signal bridge fixture and its probe, and registers both CTest entries. OFF produces only didi and didi_extension, and the configure fails if a test-only target is defined anyway.
DIDI_ENABLE_SANITIZERS OFF ASan and UBSan on GCC and Clang. MSVC has no supported combination here, so the configure refuses rather than building without the runtime. CI runs this on Ubuntu.

ctest runs the native suite and the Python suite, the latter discovered rather than listed, so the two cannot drift. Running didi_tests by path still works and is what the --list and --filter flags below are for.


📂 Source Code Layout

didi/
├── include/didi/
│   ├── common/           # Result<T>, Error, Logger, Base64/PNG, JSON, STB, IPC channels
│   ├── mcp/              # JSON-RPC 2.0, MCP server, tool/resource/prompt registries
│   ├── offline/          # GDScript diagnostics, resource indexer, test runner
│   └── gdextension/      # GDExtension interface, editor queue, Godot bridge, viewport renderer
├── src/
│   ├── common/           # Platform IPC (Win32 Named Pipes, POSIX sockets)
│   ├── mcp/              # MCP protocol handlers
│   ├── offline/          # AST analysis, file indexing, headless subprocess runner
│   ├── tools/            # Public tool handlers for the canonical and legacy surfaces
│   ├── gdextension/      # In-engine GDExtension module & renderer
│   └── standalone/       # main.cpp entry point for didi.exe
├── tests/                # Native suite plus real Godot smoke fixture/harness
├── addons/didi/          # Godot addon manifest, including the editor console's GDScript
│                         # and brand marks; the built addon is staged in build/addons/didi
└── demo/                 # Reference Godot 4 test project

🧪 Automated Test Suite

Test Inventory carries the current totals, derived from the suites themselves and regenerated by python tools/test_inventory.py; CI runs --check after the build and fails when the page and the suites disagree. The native runner's reported total is authoritative as the suite evolves. Coverage is organized by contract rather than duplicated here as a brittle test-name inventory:

  • JSON-RPC/MCP lifecycle, registration counts, capability metadata, output redaction, and structured errors.
  • IPC framing, bounded transport waits, editor-hook state transitions, and single-response ownership.
  • Descriptor validation, PID/start identity, opened-handle TOCTOU defenses, host publication, retirement, and cleanup races.
  • Native checkpoint-store tests cover snapshot bounds, manifests, corruption, unsafe paths/links, and restore staging; managed-process tests cover owned-child identity, launch, exit, and termination behavior.
  • Transactional attach, deterministic auto-selection, fresh identity handshakes, route supersession, kind-aware availability, deadlines, and quarantine.
  • Runtime log cursor/gap/filter behavior, UTF-8 and payload bounds, runtime-tree bounds, and expression-sandbox policy.
  • Tool/resource live and offline provenance, viewport/image encoding, GDScript diagnostics/patching/reflection, and resource indexing.
  • Blackboard path rejection, atomic patching, expiry, bounds, and the cross-process file lock; task claim exclusivity under concurrent claimers, lease expiry and reclaim, dependency readiness, cycle refusal, and lease ownership on update and complete.

Protocol guidance and test-fixture maintenance

kServerInstructions in src/mcp/mcp_server.cpp is the shared compiled guide for initialize and server/discover; docs/LLM_INSTRUCTIONS.md is the expanded reference, not a file loaded at runtime. When a documented route or boundary changes, update both and the API contract. Keep the guide compact and do not embed project paths, client data or live availability in it.

McpServer.InitializeInstructions checks the serialized field and resolves named routes against the registry. It also holds the guide to what discovery says about them: a tool the guide tells hosts not to call must still be unimplemented, every other tool it names must be implemented, and every parameter it spells as tool(param, ...) must be in that tool's inputSchema. Implementing or retiring a tool the guide names, or renaming a parameter it cites, fails the test until the guide is updated to match. tests/test_initialize_instructions.py drives the actual binary through malformed-handshake recovery, protocol fallback, pipelined/batched requests, concurrent clients and the guide's read workflow. Run it with python -m unittest tests.test_initialize_instructions -v. The normalization wire suite adds boundary and per-request opt-in tests; run tests.test_elastic_ingress with both compile-time profiles, setting DIDI_TEST_ELASTIC_INGRESS to 0 or 1 to verify the binary being tested.

New Python modules must participate in discovery and an explicit invocation in the existing CI workflow. Support both tests.<module> imports and discovery with -t tests. For simple test-owned stdio children, use tests/stdio_process.py to kill if needed, drain/close pipes and wait with a finite timeout. Managed editor fixtures keep their process-tree ownership cleanup. Avoid shared project/session state in new real-wire fixtures.

See the 2026-09-25 exploratory report for measured latency, cross-version live coverage, harness fixes and outstanding issues. Its local results do not substitute for cross-platform CI.

Contract snapshots

tests/contract_snapshots/ holds what a client is shown, so a change to it is a diff a reviewer reads rather than a surprise a client finds (Q3).

  • offline.json: initialize, tools/list, resources/list, resources/templates/list and prompts/list from a server with no engine. Entries are keyed by name, because the wire order is a hash map's and differs between standard libraries.
  • offline-core.json: what --tools core changes about those listings. A tool the profile leaves out is one <absent> line.
  • live-<line>.json, one for each engine line in CI's matrix: what attaching an editor changes in those listings, and the answers to the calls in calls.json. The calls run against a fresh copy of tests/contract_fixture/ in a headless editor with its own session directory and editor settings.
  • calls.json: the read-only calls, and every implemented read-only tool that is not called, with the reason. A call that should be refused carries "expect_error": true; any other error fails the recording.

Session ids, pids, the pipe endpoint, the build id, the server version, temporary paths, durations and timestamps become placeholders. The session id, endpoint, build id, version and paths are replaced by value, taken from the session Didi reports, so they are caught wherever they appear. The pid and the start time are short numbers that could match an unrelated value, so they are replaced by key, as durations are, and timestamps inside strings by pattern. A value none of these rules covers fails the two-recording check below; add the rule to tools/contract_snapshots.py. A text block that repeats structuredContent is stored as a marker, and any other JSON text, which includes every error, is stored parsed under text_json.

When a change to a schema or an answer is intended, regenerate in the same pull request and read the diff:

python tools/contract_snapshots.py --godot <Godot 4.5 exe> --godot <Godot 4.6 exe> --godot <Godot 4.7 exe>

Regenerating records everything twice, each time from a fresh editor, and writes nothing when the two disagree. The error names the JSON path that moved: a value the recorder does not yet normalise, which belongs in tools/contract_snapshots.py rather than in a snapshot. --check records once and prints a diff, and without --godot only offline.json is recorded. --binary names the server to record (the newest build by default, as the Python suites pick it), --out <dir> writes what a run recorded even when it failed, and --keep-work keeps each temporary project with its editor.log. The editor is given its own DIDI_SESSION_DIR, APPDATA and GODOT_BIN, so neither another editor on the machine nor your editor settings reach an answer.

CI checks offline.json in tests.test_contract_snapshots on all three build platforms, and each live-<line>.json in the Godot job for that line. A failed live check uploads what it recorded with the job's logs, so a contributor without Godot can read it or adopt it. A change to the fixture changes every live snapshot, so regenerate all three lines.

Tool list budgets and the core profile

didi --tools full|core chooses at startup which tools a session lists (Q4). tests/tool_list_budgets.json holds a byte budget for each profile's tools/list, measured by tests/test_tool_profiles.py from a server with no engine attached, and CI fails a listing over its budget. A new tool, or longer prose, can exceed one. Raising a budget takes its own pull request that says why, merged before the change that needs it, so growing what every session pays for is a decision rather than a side effect.

core is the union of the tools recorded field trials reached, in tools/field-trial/reached_tools.json, and every implemented tool the handshake guide names. The names are written out as coreProfileTools() in src/mcp/tool_registry.cpp, and the same test fails when that list drifts from either source: after a trial, or after the guide starts naming another tool. offline-core.json in the contract snapshots records what the profile changes about the full listing.

Response economy

A client that declares the didi/responseEconomy extension can decline the text copy of structuredContent and a session descriptor it already holds (Q5, API specification). The whole of it is economizeToolResult in src/mcp/response_economy.cpp, applied in one place: every tools/call answer in McpServer::handleRequest is encoded through its encode lambda, so a new return path there must use it too. Handlers never see the declaration and never need to. --session-descriptor once is the operator's way to the descriptor half for a client that declared nothing; it enters in McpServer::responseEconomyFor, beside the declarations.

Refusal remedies

Every refusal names what fixes it: retry_with, field, next_call, restart_with or retry_after_ms in error.data, or no_remedy saying why nothing can (Q6). Say it at the site when the site knows. Otherwise the error floor fills it from src/mcp/refusal_remedies.cpp, which is keyed by data.code and may look at the tool. A new code needs an entry there, or one in refusalsWithoutRemedy with the reason: tests/test_refusal_remedies.py scans the source for every code and fails on one the table does not cover, and the live harness fails a refusal that reaches a caller without a fix. Answer a failure with CallToolResult::errorJson, fromError or notConnected, never plain text, which has no error.data to carry a remedy. The constructor that makes plain text is private, so a new one does not compile, and the harness fails a plain-text failure it sees. RepeatedFailures, in src/mcp/repeated_failures.cpp, marks the second identical failure of a call on its way out of McpServer::handleRequest; it keys on the arguments as sent, so nothing a handler does changes what counts as the same call.

A mutation that leaves work undone names it under follow_up (Q6). src/mcp/follow_ups.cpp derives each step from a fact the answer already carries, such as scene_saved: false or requires_editor_restart: true, so publish the fact and the step follows. tests/follow_ups.json needs an entry for every mutating tool: the work it can leave and the fact behind it, why it leaves none, or an exemption with its issue. tests/test_follow_ups.py fails a mutating tool with no entry and a declared kind of work no rule produces, and the live harness fails a step a tool did not declare.

Bounded reads

Every read-only tool has an entry in tests/bounded_reads.json: bounded, naming what can cut its answer short, or unbounded, saying why nothing can (Q5). A new read with neither fails tests/test_bounded_reads.py on every build and the live harness's coverage check. A bounded tool puts a boolean truncated on every successful answer, on every path, true when a bound was reached; a limit that refuses the call instead makes a tool unbounded. The harness records every bounded answer it makes through Invoke-Didi and fails on one without the flag, and the Python test drives the offline paths and forces a bound to bite.

fields comes from ToolDefinition::sections: list the top-level keys a caller may choose between, never one the output schema requires, and registerTool publishes the argument and omitted_fields, while dispatchTool narrows the answer. It is for a large read made of separable sections; a read that is one long list takes a bound instead.

A client that declares nothing must get the same bytes. The contract snapshots are the check: they record as a client that declared nothing, so any change this makes to that client's answers shows as a diff. The live harness runs #776's seven-call arc as both kinds of client on every engine line, and fails unless the declared arc costs under half the bytes.

Opt-in live verification

The opt-in Python recovery suites launch real Godot through MCP stdio and cover owned-child crashes, saved-file persistence, the single restart, no replay, restore, source preservation, previews, and corrupt snapshots. Set DIDI_TEST_BINARY to the built host and DIDI_RECOVERY_GODOT to an absolute Godot executable:

$env:DIDI_TEST_BINARY = "D:/didi/build/Release/didi.exe"
$env:DIDI_RECOVERY_GODOT = "C:/Godot/Godot_v4.5.1-stable_win64.exe"
python -m unittest discover -s tests -t tests -p "test_managed_recovery*.py"

See Managed Recovery verification for the live and adversarial suite scope. These are opt-in engine tests; a skipped test is not live recovery evidence.

tests/test_editor_startup_live.py launches a fresh editor on a project with a main scene, opens another scene as soon as the session is listed, and reads the edited scene for ten seconds (#1069). Set DIDI_STARTUP_GODOT to a Godot executable beside DIDI_TEST_BINARY; CI runs it on every engine line after the recovery suites. It retries an open the route deadline cut short, and it skips, rather than fails, only when the editor died on the documentation-thread crash #285 tracks. It keeps the editor's output and prints it on any other failure.

$env:DIDI_STARTUP_GODOT = "C:/Godot/Godot_v4.6.2-stable_win64_console.exe"
python -m unittest tests.test_editor_startup_live -v

The Windows live integration harness copies the tracked fixture into build/ and starts real Godot processes. It preserves the Phase 1/2 sequence, adds Phase 3 concurrent editor/game routing, and now exercises Phase 4 bounded search, SVG reimport, reversible isolation, capture IDs, mutation diffs, exact undo restoration, and cleanup. The Phase 8 import block runs a menu scene whose music stops, as the control, then sets the loop options on an OGG and a WAV import with asset_configure_import after a dry run, checks seven refusals, and runs the scene again to see the music still playing. The earlier coverage still checks scripts, groups, autoloads, nested settings, InputEvent forms, persistence rollback, scene lifecycle, resource ownership, unsafe paths, and honest errors:

.\tests\run_godot_integration.ps1 `
  -GodotExecutable C:\Godot\Godot_v4.5.1-stable_win64_console.exe

The harness runs to completion on Windows PowerShell 5.1, including the persistence-rollback case that denies write rights through icacls. That case previously aborted the run: it selects its platform branch with $IsWindows, an automatic variable introduced in PowerShell 6, which is undefined on 5.1 and so took the POSIX branch and called chmod.

The Phase 1 substrate has also been run against Godot 4.6.2 and 4.7.2. The compatibility floor remains Godot 4.5.1; the live CI matrix runs the complete integration harness on 4.5.1, 4.6.2 and 4.7.2, which is every line the supported range covers, and bridge method hashes must remain valid on all three versions.


🎛️ Working on the editor console

addons/didi/ carries the plugin that draws Didi's main screen tab. It is GDScript, it compiles nothing, and it is the one part of Didi that is not behind the native test boundary, so it is worked on differently from the rest.

  • Adding a file to the addon means adding it to DIDI_ADDON_MANIFEST_FILES in CMakeLists.txt, to the expected list in the staged-addon step of .github/workflows/ci.yml (which compares against LC_ALL=C sort order), and to demo/addons/didi/. tests/test_editor_console.py fails the build when any of those disagree with the addon directory, so the list cannot go stale silently — but it will not add the file for you. Four more places list the files and no suite checks them: the expected strings in tools/localci/lane.sh and tools/localci/macos.sh, and the install trees in docs/QUICKSTART.md and docs/INTEGRATION_GUIDE.md. A new .gd ships with the .uid sidecar an engine mints for it; tools/vibe/addon_script_engines.py prints that uid and checks the script loads on every engine. A script the extension itself loads, as didi_await.gd and didi_import_watch.gd are, is also copied into the harness fixture by tests/run_godot_integration.ps1, because the fixture's addon is not the repository's.
  • The console must not become an MCP client. It reads the session descriptors Didi publishes and reports them. It calls no tool, speaks no part of the IPC protocol, and never reads the token out of a descriptor; a test enforces the last of those.
  • Verify it in a real editor, on the floor and the ceiling. Copy the addon into a scratch project and open it headless on 4.5.1 and on the newest engine in CI: Godot --headless --editor --path <project> --quit-after 6000. A GDScript parse error, a missing API on the 4.5 floor, and a tab that fails to build all show up there, and none of them show up in the native suite.
  • Brand marks are copies, not forks. The three SVGs in the addon are byte-identical to docs/brand/svg, held so by a test. Change the geometry in docs/brand/build.py and copy the regenerated source across.

➕ Adding a New MCP Tool

Do not register a success stub. A new name must either have a tested execution path or be classified as unimplemented and rejected.

A new name needs an accepted entry in Surface Amendments before it is registered. New capabilities normally arrive as an item in the Build Queue, whose section says what the tool has to do and when it is done, and whose Design Principles are the rules it has to keep.

To add a new tool (e.g. export_mesh_glb):

1. Register Tool Schema in src/mcp/tool_registry.cpp

ToolDefinition export_tool;
export_tool.name = "export_mesh_glb";
export_tool.description = "Exports a target 3D node to a GLB file.";
export_tool.inputSchema = {
    {"type", "object"},
    {"properties", {
        {"node_path", {{"type", "string"}, {"description", "Path to node"}}},
        {"output_path", {{"type", "string"}, {"description", "Target .glb path"}}}
    }},
    {"required", {"node_path", "output_path"}}
};
export_tool.handler = [this](const json& args) {
    return handleExportMesh(args, m_ipcClient);
};
registerTool(std::move(export_tool));

tools/call checks arguments against this schema before it dispatches, so what you declare here is enforced, not documentation. Declare required for every argument the handler cannot do without. A tool's arguments are closed by default, so a mistyped property name is a message naming it and listing what the tool does take; you do not have to remember "additionalProperties": false, and a new tool is covered on arrival. Declare "additionalProperties": true only if the tool genuinely takes arguments it does not publish. Nested objects are the other way round, because several of them are deliberately free-form maps: close one explicitly when it has a fixed set of keys.

type, enum, required, minimum, maximum, minLength, maxLength, minItems, maxItems, pattern and uniqueItems are all enforced. Do not publish a keyword outside that list expecting it to hold; a schema keyword the server ignores is worse than one it never advertised. Keep the handler's own checks for anything the schema cannot express.

2. Classify its execution modes

Update capabilityForTool in src/mcp/tool_registry.cpp. Choose only modes backed by tests: live, offline_fallback, both, or unimplemented.

If the handshake guide (kServerInstructions in src/mcp/mcp_server.cpp) names the tool, moving it between implemented and unimplemented also means updating the guide; see Protocol guidance.

3. Implement the Tool Handler in src/tools/

CallToolResult handleExportMesh(const json& args, std::shared_ptr<ipc::IIpcClient> ipc) {
    if (ipc && ipc->isConnected()) {
        auto res = ipc->sendRequest("mesh.export", args);
        if (res.isOk()) {
            return CallToolResult::successJson(res.value());
        }
        return CallToolResult::error(res.error().message);
    }
    return CallToolResult::error("Godot Editor is offline.");
}

4. Add live engine dispatch when applicable

Route the method from EditorHook::executeOnMainThread into a bounded implementation that performs real Godot calls on the main thread. Return a structured error whenever a required object, method bind, or engine operation is unavailable. Never return success metadata before the operation completes.

5. Test and document

  • Add native tests for capability metadata, offline behavior, and error propagation.
  • Add a real Godot integration case for live behavior and UndoRedo where relevant.
  • If the tool mutates, which is what gives it dry_run, give it an entry in tests/observed_post_state.json. Either list the answer fields it reads back after the write and add a case to tests/observed_post_state.ps1 that reads the same state through tests/godot_smoke/observed_witness.gd, or exempt it with the reason and the issue that tracks the gap. A field copied from the request or read before the write is not observed, whatever its name. tests/test_observed_post_state.py fails the build until the entry exists, and the live harness fails when an answer disagrees with the engine.
  • Regenerate the contract snapshots, because a new tool changes tools/list. It also grows the full profile's bytes, so it may need a budget raised first, in a pull request of its own. If the tool only reads, also give it a call in tests/contract_snapshots/calls.json, or an exclusion with the reason it cannot be recorded; tests/test_contract_snapshots.py fails the build until it has one.
  • Update Current Capability Matrix and Tool Reference.

Phase 3 and managed recovery implementation map

  • src/runtime/checkpoint_store.cpp: bounded saved-file inventory, manifest/hash validation, checkpoint publication, retention, and restore staging.
  • src/runtime/managed_process.cpp: owned-child launch, process identity, exit observation, logs, and termination.
  • src/runtime/managed_recovery.cpp: copied-project lifecycle, readiness, pre/post checkpoints, reconciliation latch, one automatic restart, and preserved-project restore.
  • src/runtime/session_client.cpp: descriptor discovery, opened-handle validation, cross-platform PID/process-start identity, transactional handshake on a finite deadline, token insertion, and local route state.
  • src/gdextension/session_host.cpp: bind-before-publish editor/game endpoint lifecycle, private descriptor generation, authentication stripping, and safe no-replace descriptor retirement.
  • src/gdextension/runtime_log.cpp: bounded 2,000-record ring, UTF-8-safe 16 KiB messages, 64 KiB details, cursor gaps, filtering, and logger sink mirroring.
  • src/gdextension/runtime_bridge.cpp: SceneTree resolution, UTF-8 field limits, 10,000-node/256 KiB tree budgets, explicit truncation, pause verification, exact one-active-step state machine, shutdown cancellation, and stop requests.
  • src/gdextension/expression_sandbox.cpp: tokenizer/policy, receiver-aware call validation, ClassDB-prebound native scalar property reads, context confinement, cooperative deadlines, and bounded Variant-to-JSON conversion.

The standalone router starts detached, then may auto-attach on first availability only to an unambiguous live canonical-project match: the sole session, or a unique editor among games. Preserve the tests that keep same-kind ambiguity detached, roll back failed handshakes, disable auto-selection after explicit attach/detach or quarantine, and make runtime_get_session revalidate the complete token-free identity. Keep local session management distinct from live engine calls. Never log or return descriptor tokens or full submitted expression source.

Routes are held per session, so several can be open at once. Two invariants keep that safe and both are covered by tests that fail without them. Every held route is released on shutdown, not only the selected one: the map holds the only reference to each route's ownership lock, so a route left in it is a session refused to every other Didi process until this one exits. And a request is only ever served on the session it named, which means new code that reaches for a route must take it by session id rather than by asking what the process is currently pointed at.

Phase 4 tests and release gate

The release gate runs the complete native suite; the runner's reported total remains authoritative as cases evolve. Focused suites cover the existing session/routing/evaluation contracts plus search containment and lexical filtering, two-idle-frame reimport progress, exact diff arithmetic, cache eviction, public response completeness, and restoration guards. tests/run_godot_integration.ps1 creates disposable concurrent editor/game processes and verifies the complete live workflow on Godot 4.5.1, 4.6.2 and 4.7.2. Editor teardown first requests a normal window close, then uses PID-and-start-time-verified termination if a hidden Windows editor keeps an invisible native prompt alive; this fallback is limited to the disposable test process and cannot target a reused PID. Because forced exit cannot run the extension destructor, the harness removes a leftover descriptor only after its regular-file shape, session ID, PID, and process-start identity all match that editor instance.

Run from a clean worktree:

cmake --build build --config Release
.\build\Release\didi_tests.exe
.\tests\run_godot_integration.ps1 -GodotExecutable C:\Godot\Godot_v4.5.1-stable_win64_console.exe
.\tests\run_godot_integration.ps1 -GodotExecutable C:\Godot\Godot_v4.6.2-stable_win64_console.exe
.\tests\run_godot_integration.ps1 -GodotExecutable C:\Godot\Godot_v4.7.2-stable_win64_console.exe

The engine's own output is a result. The harness reads the editor's and the game's --log-file after the run, prints every ERROR and WARNING line with a count, and fails on any line that is not in $allowedEngineLines. Each allowed entry names the request that causes the line on purpose, or the open issue for a known defect. A new line means a request made the engine complain: find the call (its answer carries the line under engine_diagnostics), then fix the cause or, when the line is the point of the request, add an entry that says which request and why. Never widen a pattern to make a run pass.

Lines the host causes -- a CI runner with no GPU or audio device, whose engine says so while its drivers start -- are allowed by where in the engine they come from (Where, matched against the at: line) as well as by what they say, so the same message from anywhere else is still a finding. The engine's source files move between lines: the audio fallback prints from servers/audio_server.cpp on 4.5 and servers/audio/audio_server.cpp from 4.6, so read a new entry's at: line off all three engines before writing its Where. A developer machine prints none of these, and the first CI run is the first place they show.

Three things catch people out running the harness by hand on Windows.

Build every target, not just didi and didi_tests. The addon is its own target, and the editor the harness starts loads whatever build of it was last staged. A bridge change then looks absent rather than stale: the editor answers with the contract it was built with, everything else looks normal, and the first assertion to notice is whichever one covers your change, which reads exactly like the change not working. The harness now fails on the first attach instead, with the mismatch note Didi already reports in every session payload.

Run it on its own rather than piping it. Under Windows PowerShell 5.1, redirecting a native command's stderr turns each line into an error record, and Didi logs to stderr on startup, so .\tests\run_godot_integration.ps1 ... 2>&1 | Select-Object -Last 40 used to kill the run at its first exchange and report it as didi.exe : [INFO ] Set working directory ..., which reads like Didi failed to start. Every exchange now goes through one helper that relaxes the preference for the duration of the call, so the run survives either way, but the output is still easier to read unpiped. CI uses pwsh 7, where the redirection was always harmless.

An IDE may connect to the harness's editor. The editor serves the GDScript language server on port 6005, and an extension such as godot-tools connects to whichever editor is there. On connect the editor re-parses every script, so the fixture script that is broken on purpose prints Failed parse script from reload_all_workspace_scripts. The allow list names that wording, so it shows in the tally and passes. Any other line it brings is still a finding.

The native runner accepts --list to print every registered case and --filter=<substring> to run a subset. The value goes after an = with no space, and any other argument is refused by name with a non-zero exit: the space form used to match nothing, leave the filter empty and run all of them, which reads as one isolated test passing (#803):

.\build\Release\didi_tests.exe --list
.\build\Release\didi_tests.exe --filter=Tools.CaptureViewportWithIpc

Run a test in isolation whenever it fails intermittently. The suite executes in a single process and shares the tool registry, the resource registry, and the working directory, so a case that does not register what it calls will pass only because an earlier case registered it, and will fail at a different assertion as the ordering changes. A test that passes in the full suite but fails alone is depending on a predecessor, not on the code under test.

For expression-policy changes, add a failing native scanner test and a real editor/game integration probe before changing implementation. A new accepted Node operation must prove it cannot dispatch script callbacks, traverse outside the active subtree, allocate unbounded data before a check, leak source/token text, or turn the cooperative timeout into a hard-preemption claim.

The CI MCP smoke must start Didi with an explicit fixture project. It verifies the live tools/list surface against the manifest emitted by didi --dump-tool-manifest from the same build, so counts are never written into the workflow, and it asserts every implemented flag rather than a sample. It also continues to assert offline-only search/deep-domain metadata, live-only reimport/diff/UI-hit-test metadata, strict Phase 4/5/6/7 schemas, local metadata for the four session tools, live metadata for routed runtime tools, cursor-shaped logs, implemented game input and profiler capabilities, and implemented: false only for physics_simulate_step, nav_bake_mesh, and runtime_get_call_stack. It is tests/test_mcp_wire_contract.py, so the Python suite runs it before a push, against the build tests/didi_binary.py picks.

Phase 7 feasibility gate

Status: PARTIAL_DELIVERY Canonical implementation: 117/120 Phase 7 registrations: 3/18 unimplemented Feasibility: 15/18 implementation-feasible; 3/18 API-blocked

Phase 7 is PARTIAL_DELIVERY. The gate completed on 2026-08-29 against Godot 4.5.1 and 4.7.2: 15/18 names are implementation-feasible, while exactly 3/18 are API-blocked under their approved contracts: physics_simulate_step, nav_bake_mesh, and runtime_get_call_stack. For each blocker, no supported public API/semantics satisfying the exact approved contract was found on either tested version. Do not broaden that result into a permanent impossibility claim.

Feasibility is design evidence; a production trial is production behavior. All 15 feasible Phase 7 names are delivered, including the three TileMapLayer/GridMap tools. The 3 API-blocked names remain registered but unimplemented.

Governance authorized partial delivery, and all 15 feasible tools are now shipped, and the surface stands at 117/120 canonical implementations. Further work on physics_simulate_step, nav_bake_mesh, or runtime_get_call_stack requires new feasibility evidence on Godot 4.5.1 and 4.7.2 or an explicit contract amendment; do not weaken their contracts implicitly. Use PHASE_7_API_FEASIBILITY.md for reproducible evidence and PHASE_7_IMPLEMENTATION_PLAN.md for the approved executable plan.

Phase 5 and Phase 6 implementation map

  • src/tools/deep_domain_tools.cpp and src/offline/deep_domain_support.cpp: bounded C#/shader diagnostics, public export-preset parsing, guarded export, deterministic MeshLibrary generation, and live UI hit-test registration/dispatch.
  • include/didi/common/project_path.hpp: explicit project-root validation, canonical containment, and stable 16-hex project endpoint keys.
  • src/runtime/session_lock.cpp: owner-only cross-platform OS locks and 423 exclusion for a second MCP owner.
  • src/mcp/mutation_safety.cpp: mutation classification, schema decoration, handler-free previews, exact context binding, 120-second expiry, and single-use confirmation storage.
  • tests/test_phase5.cpp, tests/test_phase6.cpp, the tests/test_phase7*.cpp suites, and tests/run_godot_integration.ps1: deep-domain contracts, project/lock/preview red-team cases, Phase 7 transport and validation contracts, and disposable Phases 1–7 Godot workflows.

When adding or reclassifying a mutation, update MutationSafety::isMutation, add dry_run schema coverage, and prove the dry-run never enters its handler. Add confirmation only for the documented high-risk set; changing that set is a public safety-contract change and requires updates to the Tool Reference, Capability Matrix, LLM instructions, and API specification.

Documentation and release gate

Run the dependency-free documentation contract checks for any documentation, version, registration, capability, or release change:

python -m unittest tests.test_documentation_validator -v
python tools/validate_documentation.py

The validator derives the release from CMakeLists.txt and checks the MCP server header, standalone version output, addon manifest, README, capability matrix, changelog, and security policy for alignment. It also locks the documented 120 canonical/10 legacy/130 total surface, the 117 implemented/3 unimplemented split, Phase 7's PARTIAL_DELIVERY status, 15/18 versus 3/18 feasibility result, exact three-tool blocker set, authoritative-record links, stale current-state prose, and all relative Markdown targets and anchors.

When the release changes, update these files in one change: CMakeLists.txt, include/didi/mcp/mcp_protocol.hpp, src/standalone/main.cpp, addons/didi/plugin.cfg, README.md, CHANGELOG.md, docs/CAPABILITIES.md, and SECURITY.md. When the tool surface or capability modes change, also update runtime discovery tests, docs/TOOL_REFERENCE.md, docs/ROADMAP.md, docs/LLM_INSTRUCTIONS.md, and the relevant quickstart/integration examples. Historical specs and plans record their original decisions and should not be rewritten as current release documentation.