Skip to content

Commit 3cce7d1

Browse files
jssblckclaude
andauthored
Make Jev semantic checks best effort (#75)
Jev checks now surface anything they cannot evaluate as warnings that name the skipped comment or file and why, instead of failing. `nudge check` exits 1 only for error findings and deterministic violations; exit 2 is removed. Uncertain judgments are warnings for both warn and block rules. - Retry 429, 5xx, and connection failures up to three times with jittered exponential backoff (200/400/800 ms), bounded by the deadline. - Send batches through an 8-worker pool; a failed request skips only its own batch, and a malformed answer skips only its own comment. - Judge comments without following code against the preceding code or the enclosing block instead of failing the whole file; skip comments with no code in scope. - Bump tree-sitter-rust to 0.24.2, which parses `&raw` borrows of a variable named `raw`. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
1 parent 1c16b15 commit 3cce7d1

25 files changed

Lines changed: 719 additions & 345 deletions

File tree

‎.agents/skills/nudge/references/ci.md‎

Lines changed: 3 additions & 3 deletions
Original file line numberDiff line numberDiff line change
@@ -17,7 +17,6 @@ Exit behavior:
1717

1818
- `0`: no checkable violations were found, or no file-based rules exist.
1919
- `1`: one or more checkable violations were found.
20-
- `2`: uncertain or incomplete semantic checks; takes precedence over findings.
2120
- Other non-zero exits can indicate configuration, argument, or runtime errors.
2221

2322
With explicit operands, every path or glob must resolve to at least one file.
@@ -48,8 +47,9 @@ It does not evaluate live-hook surfaces:
4847
If the user wants CI coverage, prefer file-content rules for conventions that
4948
must be enforceable outside a live agent.
5049

51-
Semantic warnings exit 0; error/block findings exit 1. Uncertain or incomplete
52-
semantic scans exit 2, taking precedence over errors.
50+
Semantic error/block findings exit 1. Warnings, uncertain judgments, and skipped
51+
semantic checks exit 0; skipped checks print a warning naming what was not
52+
evaluated.
5353

5454
## GitHub Actions Example
5555

‎.agents/skills/nudge/references/hook-responses.md‎

Lines changed: 3 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -80,7 +80,8 @@ For noisy, silent, or surprising rules, read
8080

8181
Opt-in Rust-comment rules choose `action: warn` or `action: block`. Completed
8282
findings at or above `thresholds.violation` warn or block according to that action.
83-
Uncertain judgments and incomplete checks always allow the tool with model-visible
83+
Uncertain judgments and skipped checks always allow the tool with model-visible
8484
context. Treat uncertainty as a question to review, not an established violation.
85-
Missing credentials or service errors do not establish that code is clean.
85+
A skipped check names the comment or file that was not evaluated and why; it does
86+
not establish that the code is clean.
8687
Deterministic blocks still take priority, and warnings preserve substitutions.

‎.agents/skills/nudge/references/rule-writing.md‎

Lines changed: 28 additions & 19 deletions
Original file line numberDiff line numberDiff line change
@@ -523,10 +523,12 @@ rules:
523523
Add an equivalent `tool: Edit` matcher to check edits. On Edit, Nudge evaluates
524524
comments whose text or adjacent code changed in the resulting file. It does not
525525
interpret a replacement fragment as a complete file. Ambiguous replacements,
526-
missing source files, invalid Rust syntax, or comments without associated code
527-
produce incomplete-check warnings. Consecutive comments are grouped; doc comments
528-
and comment-like strings are excluded. Rust fences in Markdown are supported with
529-
`target: { kind: MarkdownCodeBlock, language: rust }`.
526+
missing source files, or invalid Rust syntax skip the file with a warning.
527+
Consecutive comments are grouped; doc comments and comment-like strings are
528+
excluded. A comment is judged with the code after it, the code before it when it
529+
closes a block, or its enclosing block when that block has no other code.
530+
Comments with no code in scope are not evaluated. Rust fences in Markdown are
531+
supported with `target: { kind: MarkdownCodeBlock, language: rust }`.
530532

531533
Optional `content` (Write) or `new_content` (Edit) matchers act as deterministic
532534
preconditions on each selected target. With semantic Edit rules these conditions
@@ -535,28 +537,35 @@ apply to the resulting file or code block, not just the replacement string.
535537
Choose `action: warn` for warning-level findings or `action: block` for
536538
error-level findings that prevent the hook operation. Values at or below `clear`
537539
pass; values at or above `violation` trigger the chosen action. Values between
538-
them report uncertainty and allow the operation. API failures and incomplete
539-
checks also allow the operation with a warning, even for `action: block`.
540+
them are uncertain and reported as warnings, even for `action: block`.
540541

541542
`thresholds.violation` is the probability that the violation statement is true,
542543
not a separate confidence score. For example, `violation: 0.95` triggers at 95%
543544
or higher. The thresholds must satisfy `0 <= clear < violation <= 1`. Choose
544545
thresholds for each rule based on your code and tolerance for false positives.
545546
Semantic judgments are probabilistic; selecting `block` does not make them facts.
546547

547-
The client pins `jev-1.13.0`, batches independent judgments, and uses a one-second
548-
semantic deadline for a hook invocation. It does not retry failures or follow
549-
HTTP redirects. Missing credentials, rate limiting, timeout, oversized context,
550-
or malformed responses produce explicit incomplete warnings. Check mode has a
551-
30-second network-evaluation budget for the scan. Source selection and Jev's
552-
network call are separate from deterministic checks; existing block rules still
553-
take priority in hooks.
554-
555-
`nudge check` exits 0 for a complete scan with no errors, including warning-only
556-
findings; 1 for error-level semantic findings or deterministic violations; and 2
557-
for uncertain or incomplete semantic checks. Incomplete or uncertain takes
558-
precedence over errors, and never prints an all-clear result. Warnings are printed
559-
but do not fail CI.
548+
Semantic checks are best effort. The client pins `jev-1.13.0`, batches
549+
independent judgments, and sends up to eight batches concurrently. It retries
550+
rate limits (429), server errors including overload (5xx), and connection
551+
failures up to three times with jittered exponential backoff (about 200, 400,
552+
and 800 ms), never past the deadline. It does not retry other failures or follow
553+
HTTP redirects. Hooks have a one-second semantic deadline; `nudge check` has 30
554+
seconds for the whole scan.
555+
556+
When a check cannot run, Nudge reports a skipped-check warning naming the file,
557+
line, and rule it did not evaluate, and why: for example, missing credentials,
558+
exhausted retries, timeout, oversized context, or a malformed answer. Failures
559+
are isolated. A failed request skips only its own batch, and a malformed answer
560+
skips only its own comment. Skipped checks never block a hook or fail CI; they
561+
mean Nudge cannot vouch for that code, not that it is clean.
562+
563+
Source selection and Jev's network call are separate from deterministic checks;
564+
existing block rules still take priority in hooks.
565+
566+
`nudge check` exits 1 for error-level semantic findings or deterministic
567+
violations and 0 otherwise. Warnings, uncertain judgments, and skipped checks
568+
are printed, with a count of skipped checks, but do not fail CI.
560569

561570
Validate configuration offline with `nudge validate`. For a sample evaluation:
562571

‎.claude/skills/nudge/references/ci.md‎

Lines changed: 3 additions & 3 deletions
Original file line numberDiff line numberDiff line change
@@ -17,7 +17,6 @@ Exit behavior:
1717

1818
- `0`: no checkable violations were found, or no file-based rules exist.
1919
- `1`: one or more checkable violations were found.
20-
- `2`: uncertain or incomplete semantic checks; takes precedence over findings.
2120
- Other non-zero exits can indicate configuration, argument, or runtime errors.
2221

2322
With explicit operands, every path or glob must resolve to at least one file.
@@ -48,8 +47,9 @@ It does not evaluate live-hook surfaces:
4847
If the user wants CI coverage, prefer file-content rules for conventions that
4948
must be enforceable outside a live agent.
5049

51-
Semantic warnings exit 0; error/block findings exit 1. Uncertain or incomplete
52-
semantic scans exit 2, taking precedence over errors.
50+
Semantic error/block findings exit 1. Warnings, uncertain judgments, and skipped
51+
semantic checks exit 0; skipped checks print a warning naming what was not
52+
evaluated.
5353

5454
## GitHub Actions Example
5555

‎.claude/skills/nudge/references/hook-responses.md‎

Lines changed: 3 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -80,7 +80,8 @@ For noisy, silent, or surprising rules, read
8080

8181
Opt-in Rust-comment rules choose `action: warn` or `action: block`. Completed
8282
findings at or above `thresholds.violation` warn or block according to that action.
83-
Uncertain judgments and incomplete checks always allow the tool with model-visible
83+
Uncertain judgments and skipped checks always allow the tool with model-visible
8484
context. Treat uncertainty as a question to review, not an established violation.
85-
Missing credentials or service errors do not establish that code is clean.
85+
A skipped check names the comment or file that was not evaluated and why; it does
86+
not establish that the code is clean.
8687
Deterministic blocks still take priority, and warnings preserve substitutions.

‎.claude/skills/nudge/references/rule-writing.md‎

Lines changed: 28 additions & 19 deletions
Original file line numberDiff line numberDiff line change
@@ -523,10 +523,12 @@ rules:
523523
Add an equivalent `tool: Edit` matcher to check edits. On Edit, Nudge evaluates
524524
comments whose text or adjacent code changed in the resulting file. It does not
525525
interpret a replacement fragment as a complete file. Ambiguous replacements,
526-
missing source files, invalid Rust syntax, or comments without associated code
527-
produce incomplete-check warnings. Consecutive comments are grouped; doc comments
528-
and comment-like strings are excluded. Rust fences in Markdown are supported with
529-
`target: { kind: MarkdownCodeBlock, language: rust }`.
526+
missing source files, or invalid Rust syntax skip the file with a warning.
527+
Consecutive comments are grouped; doc comments and comment-like strings are
528+
excluded. A comment is judged with the code after it, the code before it when it
529+
closes a block, or its enclosing block when that block has no other code.
530+
Comments with no code in scope are not evaluated. Rust fences in Markdown are
531+
supported with `target: { kind: MarkdownCodeBlock, language: rust }`.
530532

531533
Optional `content` (Write) or `new_content` (Edit) matchers act as deterministic
532534
preconditions on each selected target. With semantic Edit rules these conditions
@@ -535,28 +537,35 @@ apply to the resulting file or code block, not just the replacement string.
535537
Choose `action: warn` for warning-level findings or `action: block` for
536538
error-level findings that prevent the hook operation. Values at or below `clear`
537539
pass; values at or above `violation` trigger the chosen action. Values between
538-
them report uncertainty and allow the operation. API failures and incomplete
539-
checks also allow the operation with a warning, even for `action: block`.
540+
them are uncertain and reported as warnings, even for `action: block`.
540541

541542
`thresholds.violation` is the probability that the violation statement is true,
542543
not a separate confidence score. For example, `violation: 0.95` triggers at 95%
543544
or higher. The thresholds must satisfy `0 <= clear < violation <= 1`. Choose
544545
thresholds for each rule based on your code and tolerance for false positives.
545546
Semantic judgments are probabilistic; selecting `block` does not make them facts.
546547

547-
The client pins `jev-1.13.0`, batches independent judgments, and uses a one-second
548-
semantic deadline for a hook invocation. It does not retry failures or follow
549-
HTTP redirects. Missing credentials, rate limiting, timeout, oversized context,
550-
or malformed responses produce explicit incomplete warnings. Check mode has a
551-
30-second network-evaluation budget for the scan. Source selection and Jev's
552-
network call are separate from deterministic checks; existing block rules still
553-
take priority in hooks.
554-
555-
`nudge check` exits 0 for a complete scan with no errors, including warning-only
556-
findings; 1 for error-level semantic findings or deterministic violations; and 2
557-
for uncertain or incomplete semantic checks. Incomplete or uncertain takes
558-
precedence over errors, and never prints an all-clear result. Warnings are printed
559-
but do not fail CI.
548+
Semantic checks are best effort. The client pins `jev-1.13.0`, batches
549+
independent judgments, and sends up to eight batches concurrently. It retries
550+
rate limits (429), server errors including overload (5xx), and connection
551+
failures up to three times with jittered exponential backoff (about 200, 400,
552+
and 800 ms), never past the deadline. It does not retry other failures or follow
553+
HTTP redirects. Hooks have a one-second semantic deadline; `nudge check` has 30
554+
seconds for the whole scan.
555+
556+
When a check cannot run, Nudge reports a skipped-check warning naming the file,
557+
line, and rule it did not evaluate, and why: for example, missing credentials,
558+
exhausted retries, timeout, oversized context, or a malformed answer. Failures
559+
are isolated. A failed request skips only its own batch, and a malformed answer
560+
skips only its own comment. Skipped checks never block a hook or fail CI; they
561+
mean Nudge cannot vouch for that code, not that it is clean.
562+
563+
Source selection and Jev's network call are separate from deterministic checks;
564+
existing block rules still take priority in hooks.
565+
566+
`nudge check` exits 1 for error-level semantic findings or deterministic
567+
violations and 0 otherwise. Warnings, uncertain judgments, and skipped checks
568+
are printed, with a count of skipped checks, but do not fail CI.
560569

561570
Validate configuration offline with `nudge validate`. For a sample evaluation:
562571

‎AGENTS.md‎

Lines changed: 5 additions & 5 deletions
Original file line numberDiff line numberDiff line change
@@ -138,12 +138,12 @@ When Nudge has something to share, it responds in one of several ways:
138138
- **Continue**: For UserPromptSubmit hooks, Nudge injects context as plain text
139139
- **Learned context**: For UserPromptSubmit hooks, Nudge searches `.nudge/learned/*.md` with BM25, or hybrid BM25 plus local embeddings when `learn.embeddings.enabled` is set in `.nudge.yaml` or `.nudge.yml`, and injects the most relevant incident notes when the prompt resembles a known scenario. For supported PreToolUse command surfaces, learned context can be surfaced as an allow-with-context warning.
140140
- **Interrupt**: For PreToolUse hooks, Nudge blocks the operation and explains what to fix
141-
- **Warning**: For uninspectable provider inputs and opt-in Jev semantic findings, uncertainty, or incomplete checks, Nudge allows the operation and returns model-visible context.
141+
- **Warning**: For uninspectable provider inputs and opt-in Jev semantic findings, uncertainty, or skipped checks, Nudge allows the operation and returns model-visible context.
142142
- **Substitute**: For deterministic PreToolUse Bash rules, Nudge rewrites the command and lets it proceed
143143

144144
The response type is determined by the hook type:
145145
- `PreToolUse` block rules **interrupt** (block provider-supported Write/Edit/WebFetch/Bash operations)
146-
- `PreToolUse` semantic rules choose `warn` (**allow with context**) or `block` (**deny completed findings**). Uncertain or incomplete checks always allow with context.
146+
- `PreToolUse` semantic rules choose `warn` (**allow with context**) or `block` (**deny completed findings**). Uncertain judgments and skipped checks always allow with context.
147147
- `PreToolUse` substitute rules **allow with updated input** (Claude Code, Codex CLI, Grok Build, and Cursor Bash commands)
148148
- `UserPromptSubmit` rules always **continue** (inject guidance into the conversation)
149149
- `PermissionRequest` is parsed but always **passes through** until Nudge has a permission-specific rule surface
@@ -192,7 +192,7 @@ client reads `TYPESAFE_API_KEY`, falling back to user-level `credentials.json`
192192
saved by `nudge login typesafe.ai`, and uses a fixed TypeSafe endpoint.
193193
`src/semantic.rs` owns planning and policy; `src/semantic/` contains configuration,
194194
selection, exact edit snapshots, and the HTTP client. Tests inject a transport,
195-
without reading credentials or contacting Jev. Hooks allow uncertain/incomplete
196-
checks with warnings. Check mode exits 0 for warning-only findings, 1 for errors,
197-
and 2 for uncertain or incomplete semantic scans.
195+
without reading credentials or contacting Jev. Checks are best effort: retryable
196+
failures back off and retry, and uncertain judgments or checks that cannot run are
197+
warnings naming what was skipped. Only error findings block hooks or fail `nudge check`.
198198
See [the semantic rules guide](docs/semantic-rules.md) for the full contract.

‎CLAUDE.md‎

Lines changed: 5 additions & 5 deletions
Original file line numberDiff line numberDiff line change
@@ -126,12 +126,12 @@ When Nudge has something to share, it responds in one of several ways:
126126
- **Continue**: For UserPromptSubmit hooks, Nudge injects context as plain text
127127
- **Learned context**: For UserPromptSubmit hooks, Nudge searches `.nudge/learned/*.md` with BM25, or hybrid BM25 plus local embeddings when `learn.embeddings.enabled` is set in `.nudge.yaml` or `.nudge.yml`, and injects the most relevant incident notes when the prompt resembles a known scenario. For supported PreToolUse command surfaces, learned context can be surfaced as an allow-with-context warning.
128128
- **Interrupt**: For PreToolUse hooks, Nudge blocks the operation and explains what to fix
129-
- **Warning**: For uninspectable provider inputs and opt-in Jev semantic findings, uncertainty, or incomplete checks, Nudge allows the operation and returns model-visible context.
129+
- **Warning**: For uninspectable provider inputs and opt-in Jev semantic findings, uncertainty, or skipped checks, Nudge allows the operation and returns model-visible context.
130130
- **Substitute**: For deterministic PreToolUse Bash rules, Nudge rewrites the command and lets it proceed
131131

132132
The response type is determined by the hook type:
133133
- `PreToolUse` block rules **interrupt** (block provider-supported Write/Edit/WebFetch/Bash operations)
134-
- `PreToolUse` semantic rules choose `warn` (**allow with context**) or `block` (**deny completed findings**). Uncertain or incomplete checks always allow with context.
134+
- `PreToolUse` semantic rules choose `warn` (**allow with context**) or `block` (**deny completed findings**). Uncertain judgments and skipped checks always allow with context.
135135
- `PreToolUse` substitute rules **allow with updated input** (Claude Code, Codex CLI, Grok Build, and Cursor Bash commands)
136136
- `UserPromptSubmit` rules always **continue** (inject guidance into the conversation)
137137
- `PermissionRequest` is parsed but always **passes through** until Nudge has a permission-specific rule surface
@@ -180,7 +180,7 @@ client reads `TYPESAFE_API_KEY`, falling back to user-level `credentials.json`
180180
saved by `nudge login typesafe.ai`, and uses a fixed TypeSafe endpoint.
181181
`src/semantic.rs` owns planning and policy; `src/semantic/` contains configuration,
182182
selection, exact edit snapshots, and the HTTP client. Tests inject a transport,
183-
without reading credentials or contacting Jev. Hooks allow uncertain/incomplete
184-
checks with warnings. Check mode exits 0 for warning-only findings, 1 for errors,
185-
and 2 for uncertain or incomplete semantic scans.
183+
without reading credentials or contacting Jev. Checks are best effort: retryable
184+
failures back off and retry, and uncertain judgments or checks that cannot run are
185+
warnings naming what was skipped. Only error findings block hooks or fail `nudge check`.
186186
See [the semantic rules guide](docs/semantic-rules.md) for the full contract.

‎Cargo.lock‎

Lines changed: 3 additions & 2 deletions
Some generated files are not rendered by default. Learn more about customizing how changed files appear on GitHub.

‎docs/ci.md‎

Lines changed: 6 additions & 7 deletions
Original file line numberDiff line numberDiff line change
@@ -21,9 +21,8 @@ nudge check "**/*.rs"
2121

2222
Exit behavior:
2323

24-
- `0`: no errors were found (semantic warnings are allowed), or no file-based rules exist.
24+
- `0`: no errors were found (semantic warnings, uncertain judgments, and skipped checks are allowed), or no file-based rules exist.
2525
- `1`: one or more deterministic violations or semantic errors were found.
26-
- `2`: a semantic check was uncertain or incomplete; this takes precedence over errors.
2726
- Other non-zero exits: configuration, argument, or runtime errors.
2827

2928
With no path arguments, `nudge check` scans the whole project from the current
@@ -49,11 +48,11 @@ Checked 25 files against 6 rules
4948

5049
Semantic Rust-comment rules with `action: warn` or `action: block` also run in check
5150
mode. Use a saved credential from `nudge login typesafe.ai`, or `TYPESAFE_API_KEY`
52-
for CI. Selected source context is sent to TypeSafe. Warning findings exit 0;
53-
error/block findings exit 1; uncertain or incomplete semantic evaluation exits 2
54-
and takes precedence over other findings. Check mode never prints an all-clear
55-
result for those states. See [semantic rules](semantic-rules.md) for thresholds,
56-
budgets, and setup. Static `nudge validate` remains offline.
51+
for CI. Selected source context is sent to TypeSafe. Semantic checks are best
52+
effort: error/block findings exit 1, while warnings, uncertain judgments, and
53+
checks skipped for missing credentials or service failures exit 0 with a warning
54+
that names what was not evaluated. See [semantic rules](semantic-rules.md) for
55+
thresholds, retries, budgets, and setup. Static `nudge validate` remains offline.
5756

5857
`nudge check` loads the same rule files as hook mode:
5958

0 commit comments

Comments
 (0)