|
| 1 | +# STRIX Reviewer Demo — `gps_denied_recon` (Software Replay) |
| 2 | + |
| 3 | +> Simulation-only. Zero hardware flights. No RF/sensor fidelity. No field validation. Classical-CBF is the only safety guarantee. |
| 4 | +
|
| 5 | +This is a **static, reviewer-facing** walkthrough of one deterministic software-only |
| 6 | +scenario replay. It shows exactly how to reproduce a single worked scenario and what |
| 7 | +the harness emits, so a reviewer can re-run it and obtain byte-identical output. |
| 8 | + |
| 9 | +- **Scope:** documentation-only. No source, manifest, or test was modified to produce this doc. |
| 10 | +- **Evidence kind:** `software_replay`, `fidelity = deterministic_kinematic_public_replay`. |
| 11 | +- **As-of commit:** `3a5d46de4ed69f4321c058ae55d84047e2ccad49` (`main`, working tree clean). |
| 12 | +- **Claim posture:** every figure below is a *prior measured software-replay result to be |
| 13 | + re-run on the exact submission commit*, never a final live fact, and never inferred |
| 14 | + hardware/RF/field readiness. See [`../CAPABILITY_BOUNDARY.md`](../CAPABILITY_BOUNDARY.md). |
| 15 | + |
| 16 | +The replay harness is a lightweight **kinematic** scenario player. It is **not** a |
| 17 | +field-physics, RF, or hardware-fidelity simulator. Its purpose is repeatable, |
| 18 | +inspectable public scenario behavior before any integration testing. |
| 19 | + |
| 20 | +--- |
| 21 | + |
| 22 | +## 1. The worked scenario |
| 23 | + |
| 24 | +`gps_denied_recon` — four aerial agents fly a reconnaissance pattern and transition to |
| 25 | +GPS-denied (degraded) navigation when GPS is lost at `t = 30 s`. The scenario's seed is |
| 26 | +fixed at **42001** inside the scenario file itself |
| 27 | +(`sim/scenarios/gps_denied_recon.yaml`), so the replay is deterministic by construction. |
| 28 | + |
| 29 | +``` |
| 30 | +Scenario: gps_denied_recon |
| 31 | +Seed: 42001 (defined in the scenario YAML; the harness reads it) |
| 32 | +Duration: 600 s Tick: 10 s Frames: 61 Agents: 4 |
| 33 | +``` |
| 34 | + |
| 35 | +--- |
| 36 | + |
| 37 | +## 2. Exact replay command |
| 38 | + |
| 39 | +The harness takes the scenario file as input and reads `seed: 42001` from it (there is no |
| 40 | +`--seed` flag — the seed is a scenario property). Run from the repository root: |
| 41 | + |
| 42 | +```bash |
| 43 | +python3 scripts/strix_sim_replay.py \ |
| 44 | + --scenario sim/scenarios/gps_denied_recon.yaml \ |
| 45 | + --output /tmp/a1_run.json \ |
| 46 | + --no-html |
| 47 | +``` |
| 48 | + |
| 49 | +`--no-html` writes only the JSON evidence. The process exits `0` when the scenario's |
| 50 | +pass-envelope is satisfied. |
| 51 | + |
| 52 | +--- |
| 53 | + |
| 54 | +## 3. Determinism proof (double run, byte-identical) |
| 55 | + |
| 56 | +The command was run **twice** in sequence into two separate files, then compared: |
| 57 | + |
| 58 | +```bash |
| 59 | +python3 scripts/strix_sim_replay.py --scenario sim/scenarios/gps_denied_recon.yaml --output /tmp/a1_run1.json --no-html |
| 60 | +python3 scripts/strix_sim_replay.py --scenario sim/scenarios/gps_denied_recon.yaml --output /tmp/a1_run2.json --no-html |
| 61 | +sha256sum /tmp/a1_run1.json /tmp/a1_run2.json |
| 62 | +``` |
| 63 | + |
| 64 | +Within a single checkout, the two runs are **byte-for-byte identical to each other** |
| 65 | +(`cmp -s /tmp/a1_run1.json /tmp/a1_run2.json` reports no difference): |
| 66 | + |
| 67 | +``` |
| 68 | +deeae065bbe543925cad81b3966a13c84c8b2cb834ad228f538b8dc5badbc957 /tmp/a1_run1.json |
| 69 | +deeae065bbe543925cad81b3966a13c84c8b2cb834ad228f538b8dc5badbc957 /tmp/a1_run2.json |
| 70 | +``` |
| 71 | + |
| 72 | +> **Do not expect this exact whole-file SHA on your machine.** The JSON embeds a |
| 73 | +> `repo` block (`commit`, `branch`, `working_tree_clean`) that reflects the git context |
| 74 | +> of *your* checkout, not the simulation — so the same seeded run yields a *different* |
| 75 | +> whole-file hash per checkout (e.g. a detached-`3a5d46d` worktree and a clean `main` |
| 76 | +> checkout produce different SHAs). The value above is only an **example** from one |
| 77 | +> clean `main` checkout. |
| 78 | +> |
| 79 | +> What **does** reproduce deterministically on any checkout is the seed-driven |
| 80 | +> simulation content — the `scenario`, `metrics`, and `envelope` blocks (Section 4). |
| 81 | +> Verify reproducibility by comparing **those blocks**, not the whole-file SHA. |
| 82 | +
|
| 83 | +--- |
| 84 | + |
| 85 | +## 4. Captured output excerpt |
| 86 | + |
| 87 | +A captured excerpt of the harness output for this run is stored alongside this doc at |
| 88 | +[`replay_output.txt`](replay_output.txt). The load-bearing parts are reproduced below. |
| 89 | + |
| 90 | +### 4.1 Scenario header — `replay["scenario"]` |
| 91 | + |
| 92 | +```json |
| 93 | +{ |
| 94 | + "config_hash": "829cc17c46b98d72", |
| 95 | + "duration_s": 600.0, |
| 96 | + "id": "gps_denied_recon", |
| 97 | + "name": "GPS-Denied Reconnaissance", |
| 98 | + "path": "sim/scenarios/gps_denied_recon.yaml", |
| 99 | + "seed": 42001, |
| 100 | + "tick_s": 10.0 |
| 101 | +} |
| 102 | +``` |
| 103 | + |
| 104 | +### 4.2 Metrics — `replay["metrics"]` |
| 105 | + |
| 106 | +```json |
| 107 | +{ |
| 108 | + "active_agents": 4, |
| 109 | + "area_coverage_pct": 100.0, |
| 110 | + "formation_coherence": 0.78, |
| 111 | + "frame_count": 61, |
| 112 | + "mean_energy_remaining_pct": 66.111, |
| 113 | + "min_constraint_clearance_m": 0.0, |
| 114 | + "offline_agents": 0, |
| 115 | + "position_error_rms_m": 4.5 |
| 116 | +} |
| 117 | +``` |
| 118 | + |
| 119 | +These are prior measured software-replay values for this seed/commit, to be re-run on the |
| 120 | +exact submission commit. They are not hardware, RF, or field measurements. |
| 121 | + |
| 122 | +### 4.3 Pass-envelope evaluation — `replay["envelope"]` |
| 123 | + |
| 124 | +```json |
| 125 | +{ |
| 126 | + "checks": [ |
| 127 | + { "metric": "area_coverage_pct", "observed": 100.0, "min": 80, "max": 100, "status": "passed" }, |
| 128 | + { "metric": "position_error_rms_m", "observed": 4.5, "min": 0.0, "max": 10.0, "status": "passed" }, |
| 129 | + { "metric": "formation_coherence", "observed": 0.78, "min": 0.6, "max": 1.0, "status": "passed" } |
| 130 | + ], |
| 131 | + "status": "passed" |
| 132 | +} |
| 133 | +``` |
| 134 | + |
| 135 | +### 4.4 Timeline excerpt — degraded-nav transition at GPS loss |
| 136 | + |
| 137 | +The per-agent `mode` field flips from `nominal` to `degraded_nav` exactly at the |
| 138 | +scheduled `t = 30 s` `gps_loss` event, and scenario events surface at their scheduled |
| 139 | +times. This is the documented degraded-mode / EW behavior under GPS loss. |
| 140 | + |
| 141 | +``` |
| 142 | + t_s agent_1.mode agent_1 (x, y, z) events |
| 143 | +----- ---------------- ------------------------------ ------------------------ |
| 144 | + 0 nominal ( 0.00, 200.00, -50.00) - |
| 145 | + 10 nominal ( -51.89, 218.94, -50.00) - |
| 146 | + 20 nominal ( -100.98, 201.07, -50.00) - |
| 147 | + 30 degraded_nav ( -144.39, 171.75, -50.00) 30.0s:gps_loss |
| 148 | + 40 degraded_nav ( -180.79, 134.60, -50.00) - |
| 149 | + 120 degraded_nav ( -76.23, -212.21, -50.00) 120.0s:wind_gust |
| 150 | + 300 degraded_nav ( -144.40, 172.08, -50.00) 300.0s:sensor_degradation |
| 151 | + 600 degraded_nav ( -222.22, 39.13, -50.00) - |
| 152 | +``` |
| 153 | + |
| 154 | +This scenario defines no threat constraints, so `min_constraint_clearance_m` is `0.0` |
| 155 | +and no constraint-avoidance gating fires in this run. The shipped safety story rests on |
| 156 | +classical control-barrier-function gating plus ROE plus traces plus the simulator; this |
| 157 | +particular scenario exercises the GPS-denied degraded-navigation path, not CBF gating. |
| 158 | + |
| 159 | +--- |
| 160 | + |
| 161 | +## 5. Diagrams (labels only) |
| 162 | + |
| 163 | +### 5.1 Replay data flow |
| 164 | + |
| 165 | +```mermaid |
| 166 | +flowchart LR |
| 167 | + Y["gps_denied_recon.yaml<br/>seed=42001"] --> R["strix_sim_replay.py<br/>kinematic replay"] |
| 168 | + R --> J["replay JSON<br/>scenario / metrics / envelope / frames"] |
| 169 | + J --> E["pass-envelope check"] |
| 170 | + J --> T["timeline / modes"] |
| 171 | +``` |
| 172 | + |
| 173 | +### 5.2 Navigation-mode state (this scenario) |
| 174 | + |
| 175 | +``` |
| 176 | + t < 30s t >= 30s (gps_loss) |
| 177 | + +----------+ +---------------+ |
| 178 | + | nominal | -------> | degraded_nav | |
| 179 | + +----------+ +---------------+ |
| 180 | +``` |
| 181 | + |
| 182 | +--- |
| 183 | + |
| 184 | +## 6. Reproduce checklist |
| 185 | + |
| 186 | +1. Check out commit `3a5d46de4ed69f4321c058ae55d84047e2ccad49`, clean working tree. |
| 187 | +2. From the repo root, run the command in Section 2. |
| 188 | +3. Run it a second time to a different output path. |
| 189 | +4. `cmp -s` the two outputs — within your checkout they are byte-identical to each |
| 190 | + other. (The whole-file `sha256sum` in Section 3 is an example only: it embeds git |
| 191 | + metadata and varies per checkout, so do not expect that exact value.) |
| 192 | +5. Compare your `scenario` / `metrics` / `envelope` blocks to Section 4 — these are the |
| 193 | + seed-deterministic, checkout-independent reproducibility check. |
| 194 | + |
| 195 | +--- |
| 196 | + |
| 197 | +## 7. Boundary reminder |
| 198 | + |
| 199 | +- Software replay only. No hardware flights, no RF/sensor fidelity, no field validation. |
| 200 | +- Numbers are prior measured software-replay results, re-run on the submission commit. |
| 201 | +- No fielded deployment, no delivered ROS2/MAVLink hardware integration, no defence |
| 202 | + validation/accreditation, no trained-neural safety guarantee is claimed. |
| 203 | +- Authoritative claim map: [`../CAPABILITY_BOUNDARY.md`](../CAPABILITY_BOUNDARY.md). |
0 commit comments