Skip to content

Commit c84a45c

Browse files
RMANOVclaude
andcommitted
docs(demo): A1 submission-safe gps_denied_recon demo doc (sim-only)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
1 parent 3a5d46d commit c84a45c

2 files changed

Lines changed: 285 additions & 0 deletions

File tree

‎Project_Docs/demo/DEMO.md‎

Lines changed: 203 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,203 @@
1+
# STRIX Reviewer Demo — `gps_denied_recon` (Software Replay)
2+
3+
> Simulation-only. Zero hardware flights. No RF/sensor fidelity. No field validation. Classical-CBF is the only safety guarantee.
4+
5+
This is a **static, reviewer-facing** walkthrough of one deterministic software-only
6+
scenario replay. It shows exactly how to reproduce a single worked scenario and what
7+
the harness emits, so a reviewer can re-run it and obtain byte-identical output.
8+
9+
- **Scope:** documentation-only. No source, manifest, or test was modified to produce this doc.
10+
- **Evidence kind:** `software_replay`, `fidelity = deterministic_kinematic_public_replay`.
11+
- **As-of commit:** `3a5d46de4ed69f4321c058ae55d84047e2ccad49` (`main`, working tree clean).
12+
- **Claim posture:** every figure below is a *prior measured software-replay result to be
13+
re-run on the exact submission commit*, never a final live fact, and never inferred
14+
hardware/RF/field readiness. See [`../CAPABILITY_BOUNDARY.md`](../CAPABILITY_BOUNDARY.md).
15+
16+
The replay harness is a lightweight **kinematic** scenario player. It is **not** a
17+
field-physics, RF, or hardware-fidelity simulator. Its purpose is repeatable,
18+
inspectable public scenario behavior before any integration testing.
19+
20+
---
21+
22+
## 1. The worked scenario
23+
24+
`gps_denied_recon` — four aerial agents fly a reconnaissance pattern and transition to
25+
GPS-denied (degraded) navigation when GPS is lost at `t = 30 s`. The scenario's seed is
26+
fixed at **42001** inside the scenario file itself
27+
(`sim/scenarios/gps_denied_recon.yaml`), so the replay is deterministic by construction.
28+
29+
```
30+
Scenario: gps_denied_recon
31+
Seed: 42001 (defined in the scenario YAML; the harness reads it)
32+
Duration: 600 s Tick: 10 s Frames: 61 Agents: 4
33+
```
34+
35+
---
36+
37+
## 2. Exact replay command
38+
39+
The harness takes the scenario file as input and reads `seed: 42001` from it (there is no
40+
`--seed` flag — the seed is a scenario property). Run from the repository root:
41+
42+
```bash
43+
python3 scripts/strix_sim_replay.py \
44+
--scenario sim/scenarios/gps_denied_recon.yaml \
45+
--output /tmp/a1_run.json \
46+
--no-html
47+
```
48+
49+
`--no-html` writes only the JSON evidence. The process exits `0` when the scenario's
50+
pass-envelope is satisfied.
51+
52+
---
53+
54+
## 3. Determinism proof (double run, byte-identical)
55+
56+
The command was run **twice** in sequence into two separate files, then compared:
57+
58+
```bash
59+
python3 scripts/strix_sim_replay.py --scenario sim/scenarios/gps_denied_recon.yaml --output /tmp/a1_run1.json --no-html
60+
python3 scripts/strix_sim_replay.py --scenario sim/scenarios/gps_denied_recon.yaml --output /tmp/a1_run2.json --no-html
61+
sha256sum /tmp/a1_run1.json /tmp/a1_run2.json
62+
```
63+
64+
Within a single checkout, the two runs are **byte-for-byte identical to each other**
65+
(`cmp -s /tmp/a1_run1.json /tmp/a1_run2.json` reports no difference):
66+
67+
```
68+
deeae065bbe543925cad81b3966a13c84c8b2cb834ad228f538b8dc5badbc957 /tmp/a1_run1.json
69+
deeae065bbe543925cad81b3966a13c84c8b2cb834ad228f538b8dc5badbc957 /tmp/a1_run2.json
70+
```
71+
72+
> **Do not expect this exact whole-file SHA on your machine.** The JSON embeds a
73+
> `repo` block (`commit`, `branch`, `working_tree_clean`) that reflects the git context
74+
> of *your* checkout, not the simulation — so the same seeded run yields a *different*
75+
> whole-file hash per checkout (e.g. a detached-`3a5d46d` worktree and a clean `main`
76+
> checkout produce different SHAs). The value above is only an **example** from one
77+
> clean `main` checkout.
78+
>
79+
> What **does** reproduce deterministically on any checkout is the seed-driven
80+
> simulation content — the `scenario`, `metrics`, and `envelope` blocks (Section 4).
81+
> Verify reproducibility by comparing **those blocks**, not the whole-file SHA.
82+
83+
---
84+
85+
## 4. Captured output excerpt
86+
87+
A captured excerpt of the harness output for this run is stored alongside this doc at
88+
[`replay_output.txt`](replay_output.txt). The load-bearing parts are reproduced below.
89+
90+
### 4.1 Scenario header — `replay["scenario"]`
91+
92+
```json
93+
{
94+
"config_hash": "829cc17c46b98d72",
95+
"duration_s": 600.0,
96+
"id": "gps_denied_recon",
97+
"name": "GPS-Denied Reconnaissance",
98+
"path": "sim/scenarios/gps_denied_recon.yaml",
99+
"seed": 42001,
100+
"tick_s": 10.0
101+
}
102+
```
103+
104+
### 4.2 Metrics — `replay["metrics"]`
105+
106+
```json
107+
{
108+
"active_agents": 4,
109+
"area_coverage_pct": 100.0,
110+
"formation_coherence": 0.78,
111+
"frame_count": 61,
112+
"mean_energy_remaining_pct": 66.111,
113+
"min_constraint_clearance_m": 0.0,
114+
"offline_agents": 0,
115+
"position_error_rms_m": 4.5
116+
}
117+
```
118+
119+
These are prior measured software-replay values for this seed/commit, to be re-run on the
120+
exact submission commit. They are not hardware, RF, or field measurements.
121+
122+
### 4.3 Pass-envelope evaluation — `replay["envelope"]`
123+
124+
```json
125+
{
126+
"checks": [
127+
{ "metric": "area_coverage_pct", "observed": 100.0, "min": 80, "max": 100, "status": "passed" },
128+
{ "metric": "position_error_rms_m", "observed": 4.5, "min": 0.0, "max": 10.0, "status": "passed" },
129+
{ "metric": "formation_coherence", "observed": 0.78, "min": 0.6, "max": 1.0, "status": "passed" }
130+
],
131+
"status": "passed"
132+
}
133+
```
134+
135+
### 4.4 Timeline excerpt — degraded-nav transition at GPS loss
136+
137+
The per-agent `mode` field flips from `nominal` to `degraded_nav` exactly at the
138+
scheduled `t = 30 s` `gps_loss` event, and scenario events surface at their scheduled
139+
times. This is the documented degraded-mode / EW behavior under GPS loss.
140+
141+
```
142+
t_s agent_1.mode agent_1 (x, y, z) events
143+
----- ---------------- ------------------------------ ------------------------
144+
0 nominal ( 0.00, 200.00, -50.00) -
145+
10 nominal ( -51.89, 218.94, -50.00) -
146+
20 nominal ( -100.98, 201.07, -50.00) -
147+
30 degraded_nav ( -144.39, 171.75, -50.00) 30.0s:gps_loss
148+
40 degraded_nav ( -180.79, 134.60, -50.00) -
149+
120 degraded_nav ( -76.23, -212.21, -50.00) 120.0s:wind_gust
150+
300 degraded_nav ( -144.40, 172.08, -50.00) 300.0s:sensor_degradation
151+
600 degraded_nav ( -222.22, 39.13, -50.00) -
152+
```
153+
154+
This scenario defines no threat constraints, so `min_constraint_clearance_m` is `0.0`
155+
and no constraint-avoidance gating fires in this run. The shipped safety story rests on
156+
classical control-barrier-function gating plus ROE plus traces plus the simulator; this
157+
particular scenario exercises the GPS-denied degraded-navigation path, not CBF gating.
158+
159+
---
160+
161+
## 5. Diagrams (labels only)
162+
163+
### 5.1 Replay data flow
164+
165+
```mermaid
166+
flowchart LR
167+
Y["gps_denied_recon.yaml<br/>seed=42001"] --> R["strix_sim_replay.py<br/>kinematic replay"]
168+
R --> J["replay JSON<br/>scenario / metrics / envelope / frames"]
169+
J --> E["pass-envelope check"]
170+
J --> T["timeline / modes"]
171+
```
172+
173+
### 5.2 Navigation-mode state (this scenario)
174+
175+
```
176+
t < 30s t >= 30s (gps_loss)
177+
+----------+ +---------------+
178+
| nominal | -------> | degraded_nav |
179+
+----------+ +---------------+
180+
```
181+
182+
---
183+
184+
## 6. Reproduce checklist
185+
186+
1. Check out commit `3a5d46de4ed69f4321c058ae55d84047e2ccad49`, clean working tree.
187+
2. From the repo root, run the command in Section 2.
188+
3. Run it a second time to a different output path.
189+
4. `cmp -s` the two outputs — within your checkout they are byte-identical to each
190+
other. (The whole-file `sha256sum` in Section 3 is an example only: it embeds git
191+
metadata and varies per checkout, so do not expect that exact value.)
192+
5. Compare your `scenario` / `metrics` / `envelope` blocks to Section 4 — these are the
193+
seed-deterministic, checkout-independent reproducibility check.
194+
195+
---
196+
197+
## 7. Boundary reminder
198+
199+
- Software replay only. No hardware flights, no RF/sensor fidelity, no field validation.
200+
- Numbers are prior measured software-replay results, re-run on the submission commit.
201+
- No fielded deployment, no delivered ROS2/MAVLink hardware integration, no defence
202+
validation/accreditation, no trained-neural safety guarantee is claimed.
203+
- Authoritative claim map: [`../CAPABILITY_BOUNDARY.md`](../CAPABILITY_BOUNDARY.md).
Lines changed: 82 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,82 @@
1+
STRIX software-replay -- captured deterministic output excerpt
2+
Scenario: gps_denied_recon Seed: 42001
3+
Source: scripts/strix_sim_replay.py (kind=software_replay, fidelity=deterministic_kinematic_public_replay)
4+
As-of commit: 3a5d46de4ed69f4321c058ae55d84047e2ccad49 (main, working tree clean)
5+
6+
This file is a trimmed excerpt from the JSON output emitted by the replay harness. Two consecutive
7+
runs in the same checkout produce a byte-identical file. The whole-file SHA256 below is
8+
an EXAMPLE from one clean main checkout only: the JSON embeds a git repo block, so the
9+
whole-file hash varies per checkout. The seed-driven scenario/metrics/envelope content
10+
is what reproduces deterministically across checkouts.
11+
Example replay output SHA256 (clean main checkout): deeae065bbe543925cad81b3966a13c84c8b2cb834ad228f538b8dc5badbc957
12+
13+
--- scenario header (replay['scenario']) ---
14+
{
15+
"config_hash": "829cc17c46b98d72",
16+
"duration_s": 600.0,
17+
"id": "gps_denied_recon",
18+
"name": "GPS-Denied Reconnaissance",
19+
"path": "sim/scenarios/gps_denied_recon.yaml",
20+
"seed": 42001,
21+
"tick_s": 10.0
22+
}
23+
24+
--- metrics (replay['metrics']) ---
25+
{
26+
"active_agents": 4,
27+
"area_coverage_pct": 100.0,
28+
"formation_coherence": 0.78,
29+
"frame_count": 61,
30+
"mean_energy_remaining_pct": 66.111,
31+
"min_constraint_clearance_m": 0.0,
32+
"offline_agents": 0,
33+
"position_error_rms_m": 4.5
34+
}
35+
36+
--- pass-envelope evaluation (replay['envelope']) ---
37+
{
38+
"checks": [
39+
{
40+
"max": 100,
41+
"metric": "area_coverage_pct",
42+
"min": 80,
43+
"observed": 100.0,
44+
"status": "passed"
45+
},
46+
{
47+
"max": 10.0,
48+
"metric": "position_error_rms_m",
49+
"min": 0.0,
50+
"observed": 4.5,
51+
"status": "passed"
52+
},
53+
{
54+
"max": 1.0,
55+
"metric": "formation_coherence",
56+
"min": 0.6,
57+
"observed": 0.78,
58+
"status": "passed"
59+
}
60+
],
61+
"status": "passed"
62+
}
63+
64+
--- timeline excerpt: agent_1 navigation mode at key timestamps ---
65+
(degraded_nav engages exactly at the t=30s scheduled gps_loss event)
66+
67+
t_s agent_1.mode agent_1 (x, y, z) events
68+
----- ---------------- ------------------------------ ------------------------
69+
0 nominal ( 0.00, 200.00, -50.00) -
70+
10 nominal ( -51.89, 218.94, -50.00) -
71+
20 nominal ( -100.98, 201.07, -50.00) -
72+
30 degraded_nav ( -144.39, 171.75, -50.00) 30.0s:gps_loss
73+
40 degraded_nav ( -180.79, 134.60, -50.00) -
74+
120 degraded_nav ( -76.23, -212.21, -50.00) 120.0s:wind_gust
75+
300 degraded_nav ( -144.40, 172.08, -50.00) 300.0s:sensor_degradation
76+
600 degraded_nav ( -222.22, 39.13, -50.00) -
77+
78+
Total frames: 61 Agents: 4 Constraints: 0
79+
80+
Caveat: simulation-only deterministic kinematic replay. No RF/sensor fidelity,
81+
no hardware, no field validation. Classical-CBF is the only safety guarantee.
82+

0 commit comments

Comments
 (0)