Skip to content

feat(source): add ‘source clean’ command - #261

Merged
teng-lin merged 7 commits into
teng-lin:mainfrom
Flosters:feat/source-clean
May 13, 2026
Merged

teng-lin merged 7 commits into
teng-lin:mainfrom
Flosters:feat/source-clean

Conversation

@Flosters

@Flosters Flosters commented Apr 9, 2026 •

Copy link
Copy Markdown
Contributor

Summary

After failed deep-research runs, bulk imports, or bot-blocked crawls, notebooks can accumulate junk sources: errored ones, gateway/anti-bot pages (Cloudflare, 403, 404), and duplicates. There is currently no way to clean these up in bulk — users must delete them one by one.

This PR adds notebooklm source clean, which automatically identifies and removes junk sources in one command.

What it removes

Category Criteria
Error/unknown status is error or unknown
Gateway / anti-bot Title matches patterns: 403, 404, Forbidden, Access Denied, Just a Moment, Attention Required, Security Check, CAPTCHA
Duplicates Same URL already seen (normalized: scheme+host+path, no query/fragment); keeps the oldest copy

Flags

notebooklm source clean             # interactive confirmation
notebooklm source clean --dry-run   # preview without deleting
notebooklm source clean -y          # skip confirmation
notebooklm source clean -n <id>     # target specific notebook

Implementation notes

  • Sources are sorted oldest-first before deduplication so the earliest copy is always kept
  • Deletions are chunked in batches of 10 with a 0.5s inter-chunk delay to avoid rate limiting on large notebooks
  • Uses existing asyncio.gather pattern consistent with other bulk operations in the codebase

Test plan

  • --dry-run prints count and exits without deleting
  • Sources with status=error are removed
  • Duplicate URLs: only the newest copy is deleted
  • Gateway-titled sources (e.g. title "403 Forbidden") are removed
  • Clean notebook prints "Notebook is already clean"
  • -y skips confirmation prompt
  • Deletions are batched (no 429 on 50+ junk sources)

Summary by CodeRabbit

  • New Features

    • Added notebooklm source clean CLI to find and optionally remove junk sources (status errors, gateway/security blocks, duplicate URLs). Keeps oldest copy per URL, preserves query strings, strips fragments, and prints a summary table. Supports --dry-run (preview) and --yes (skip confirmation); performs batched deletions with brief pauses and reports successes or top failures.
  • Tests

    • Added unit and end-to-end tests for classification, dedupe rules, dry-run, confirmation flow, and deletion error reporting.

Review Change Stack

agustinsilvazambrano added 2 commits April 8, 2026 21:27
Adds `notebooklm source clean` to automatically remove junk sources from
a notebook in bulk. Useful after failed deep-research runs, bot-blocked
crawls, or bulk imports that left behind error or duplicate sources.

Removes sources that match any of:
- Status is 'error' or 'unknown'
- Title matches gateway/anti-bot patterns (403, 404, Access Denied,
  Cloudflare 'Just a Moment', CAPTCHA, etc.)
- URL is a duplicate of an already-seen source (keeps oldest)

Flags:
  --dry-run   Preview what would be deleted without deleting
  -y/--yes    Skip confirmation prompt
  -n          Target a specific notebook

Deletions are batched in chunks of 10 with a 0.5s delay to avoid
hitting rate limits on large notebooks.
@coderabbitai

coderabbitai Bot commented Apr 9, 2026 •

Copy link
Copy Markdown
Contributor

Note

Reviews paused

It looks like this branch is under active development. To avoid overwhelming you with review comments due to an influx of new commits, CodeRabbit has automatically paused this review. You can configure this behavior by changing the reviews.auto_review.auto_pause_after_reviewed_commits setting.

Use the following commands to manage reviews:

  • @coderabbitai resume to resume automatic reviews.
  • @coderabbitai review to trigger a single review.

Use the checkboxes below for quick actions:

  • ▶️ Resume reviews
  • 🔍 Trigger review

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro

Run ID: 9aa3d84e-f690-4455-a3c6-1d59211bd515

📥 Commits

Reviewing files that changed from the base of the PR and between 489d9a3 and 2601936.

📒 Files selected for processing (2)
  • src/notebooklm/cli/source.py
  • tests/unit/cli/test_source.py
🚧 Files skipped from review as they are similar to previous changes (2)
  • tests/unit/cli/test_source.py
  • src/notebooklm/cli/source.py

📝 Walkthrough

Walkthrough

Adds a source clean CLI subcommand that resolves a notebook, lists sources, classifies junk (error status, gateway/anti-bot titles, normalized-URL duplicates), shows a Rich table, and optionally deletes candidates asynchronously in batches (dry-run and confirmation flags supported).

Changes

New Source Clean Command

Layer / File(s) Summary
Classification and URL normalization
src/notebooklm/cli/source.py
Implements gateway/title regex and eligible-junk status set; URL normalization that strips only fragments while preserving queries; deterministic sort-by-created_at and per-normalized-URL duplicate detection that marks later copies duplicate_of:<short-id>; candidate table printer.
CLI command and deletion orchestration
src/notebooklm/cli/source.py
Adds source clean Click subcommand: resolves notebook, lists sources, computes candidates, supports --dry-run and --yes/-y, prompts when necessary, deletes in async batches of 10 with 0.5s delay, captures exceptions, and reports successes/failures.
Tests for classification and CLI
tests/unit/cli/test_source.py
Adds _src fixture, unit tests for _classify_junk_sources (status/title/deduplication invariants and edge cases), Click invocation tests for dry-run/confirmation/deletion/partial-failure reporting, and _CleanPatch to mock client and auth token fetching.

Sequence Diagram

sequenceDiagram
    actor User
    participant CLI as "CLI"
    participant Resolver as "resolve_notebook_id"
    participant SourceSvc as "sources.list / sources.delete"
    participant Selector as "identify_candidates"
    participant Normalizer as "normalize_url"
    participant Deleter as "chunked_deleter"

    User->>CLI: source clean --notebook X [--dry-run] [--yes]
    CLI->>Resolver: resolve_notebook_id(X)
    Resolver-->>CLI: notebook_id
    CLI->>SourceSvc: list_sources(notebook_id)
    SourceSvc-->>CLI: sources[]
    CLI->>Selector: identify_candidates(sources)
    Selector->>Normalizer: normalize_url(src.url)
    Normalizer-->>Selector: normalized_url (fragment stripped)
    Selector-->>CLI: candidate_ids[] (status/title/duplicate reasons)
    alt dry-run
        CLI->>User: report candidate count & table
    else proceed
        CLI->>User: prompt confirmation (unless --yes)
        User-->>CLI: confirm
        CLI->>Deleter: delete_in_batches(candidate_ids, size=10)
        loop per batch
            Deleter->>SourceSvc: delete_source(id) (concurrent via asyncio.gather)
            SourceSvc-->>Deleter: success / exception
            Deleter->>Deleter: sleep(0.5s)
        end
        Deleter-->>CLI: successes & failures
        CLI->>User: final report (errors truncated to first 5)
    end
Loading

Estimated code review effort

🎯 4 (Complex) | ⏱️ ~45 minutes

Poem

🐰 I hopped through links both broken and spare,
Found gateways, duplicates, and errors laid bare.
Dry-run I pondered, then confirmed with a thump,
In tidy batches I cleared each cluttered stump.
Notebooks gleam anew — a rabbit's small jump.

🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 25.71% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title 'feat(source): add 'source clean' command' clearly and specifically describes the main change: a new CLI subcommand is being added.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Inline comments:
In `@src/notebooklm/cli/source.py`:
- Around line 982-992: The deletion loop currently swallows all errors by
calling asyncio.gather(..., return_exceptions=True) and always prints success;
update the loop around client.sources.delete / asyncio.gather to inspect the
gathered results, count actual successful deletions vs exceptions, and log or
print any failures (include exception messages and the corresponding sid) as
they occur; after the loop, change the console.print call to report the real
number of successful deletions (and optionally failed count) using the counted
successes so users aren’t shown a false “Successfully cleaned” message for
delete_list, chunk_size, delete_tasks, and nb_id_resolved.
- Around line 945-950: The code incorrectly converts the enum s.status to a
string with str(...).lower(), so comparisons never match; replace that
conversion with the helper source_status_to_str(s.status) (handle None if
needed) and use that result to check if status is "error" or "unknown", then add
s.id to to_delete as before (look at variables s, to_delete, and function
source_status_to_str for where to change).
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro

Run ID: 4ae7d06c-d332-4bb1-9a3a-f0412794bf36

📥 Commits

Reviewing files that changed from the base of the PR and between a997718 and da0b8bd.

📒 Files selected for processing (1)
  • src/notebooklm/cli/source.py

Comment thread src/notebooklm/cli/source.py Outdated
Comment thread src/notebooklm/cli/source.py Outdated

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces a new source clean command to the CLI, designed to automatically remove duplicate, error, and access-blocked sources from a NotebookLM notebook. The command identifies sources based on their status, title, and normalized URL, and includes options for dry-run and confirmation. A critical issue was identified in the status checking logic: the s.status enum is incorrectly converted to a string, which prevents sources in an error or unknown state from being properly detected and cleaned. The suggested fix is to use the source_status_to_str() helper function.

Comment thread src/notebooklm/cli/source.py Outdated
Agustín Silva Zambrano and others added 2 commits April 14, 2026 19:32
- Use source_status_to_str() instead of str().lower() so error/unknown
  status sources are correctly identified and removed
- Track per-result success/failure from asyncio.gather so the output
  reflects what actually happened instead of always reporting success

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🧹 Nitpick comments (1)
src/notebooklm/cli/source.py (1)

957-965: URL dedup misses case-only differences and trailing slashes.

urlparse preserves the scheme/host casing and raw path, so HTTPS://Example.com/a and https://example.com/a/ will not dedupe against https://example.com/a, even though they point at the same resource. Lowercasing scheme + netloc and stripping a trailing / on the path gives better duplicate coverage without affecting correctness.

♻️ Proposed tweak
-                    parsed = urlparse(url)
-                    normalized = urlunparse((parsed.scheme, parsed.netloc, parsed.path, "", "", ""))
+                    parsed = urlparse(url)
+                    path = parsed.path.rstrip("/") or "/"
+                    normalized = urlunparse(
+                        (parsed.scheme.lower(), parsed.netloc.lower(), path, "", "", "")
+                    )
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@src/notebooklm/cli/source.py` around lines 957 - 965, The current URL
normalization misses case differences and trailing slashes; update the
normalization logic around urlparse/urlunparse so scheme and netloc are
lowercased and the path has a single canonical form (strip any trailing slash
except preserve root as "/" or empty path as ""), then use that normalized
string for deduping. Concretely: take parsed = urlparse(url), set scheme =
parsed.scheme.lower(), netloc = parsed.netloc.lower(), compute path_norm =
parsed.path.rstrip('/') (and if path_norm == '' and parsed.path == '/' keep '/'
or otherwise use '' per your desired canonicalization), then call
urlunparse((scheme, netloc, path_norm, "", "", "")) and use that for
seen_urls/to_delete comparisons (same variables: normalized, seen_urls,
to_delete, s.id).
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Inline comments:
In `@src/notebooklm/cli/source.py`:
- Around line 952-955: The current junk-filtering deletes any source whose title
starts with "http://" or "https://", which is too broad; update the logic around
GATEWAY_PATTERNS and the title.startswith(("http://", "https://")) check so it
only drops URL-title sources when the fetch clearly failed—i.e., gate that
branch behind the source status check (use status in ["error","unknown"]) or
remove the URL-prefix heuristic entirely; modify the block that adds s.id to
to_delete (referencing GATEWAY_PATTERNS, the title.startswith check, and the
status variable) so in-progress/ready items are not silently purged.

---

Nitpick comments:
In `@src/notebooklm/cli/source.py`:
- Around line 957-965: The current URL normalization misses case differences and
trailing slashes; update the normalization logic around urlparse/urlunparse so
scheme and netloc are lowercased and the path has a single canonical form (strip
any trailing slash except preserve root as "/" or empty path as ""), then use
that normalized string for deduping. Concretely: take parsed = urlparse(url),
set scheme = parsed.scheme.lower(), netloc = parsed.netloc.lower(), compute
path_norm = parsed.path.rstrip('/') (and if path_norm == '' and parsed.path ==
'/' keep '/' or otherwise use '' per your desired canonicalization), then call
urlunparse((scheme, netloc, path_norm, "", "", "")) and use that for
seen_urls/to_delete comparisons (same variables: normalized, seen_urls,
to_delete, s.id).
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro

Run ID: 2da1f3e1-768a-42a6-ad0b-62fb7e078798

📥 Commits

Reviewing files that changed from the base of the PR and between 1f47282 and e19eac6.

📒 Files selected for processing (1)
  • src/notebooklm/cli/source.py

Comment thread src/notebooklm/cli/source.py Outdated
@teng-lin teng-lin added the enhancement New feature or request label May 3, 2026
@teng-lin

teng-lin commented May 3, 2026

Copy link
Copy Markdown
Owner

This is a destructive command, so it needs more care before we can merge: (1) make --dry-run the default behavior (require an explicit --yes or --force to actually delete); (2) document the heuristics being used to flag a source as cleanable; (3) add unit tests covering both the dry-run and destructive paths, including one test that verifies sources matching exclusion criteria are NOT deleted. Once safer defaults and tests are in place this is reviewable.

teng-lin added 2 commits May 12, 2026 20:38
Polish pass on the new `source clean` command after a 3-way review
(Claude/Gemini/Codex) flagged data-loss risks and a missing test suite.

Safety fixes
- Drop the `title.startswith("http://"...)` heuristic. Source.title is
  documented as "may be URL if not yet processed", so the previous code
  would silently delete legitimate sources whose metadata had not yet
  resolved. The gateway-title regex still catches the obvious
  Cloudflare/CAPTCHA/403 blocking pages.
- Remove "unknown" from `_JUNK_STATUSES`. Only explicit `error` is
  cleaned; future NotebookLM status codes we don't recognize and
  missing-status payloads are now preserved instead of wiped.
- Preserve query string in URL normalization. Stripping it collapsed
  distinct YouTube videos (`?v=AAA` vs `?v=BBB`), Google Docs IDs, and
  arXiv versions. Only the `#fragment` is stripped now.
- Push undated sources to the end of the sort (was: epoch-0, which made
  them the "oldest" anchor during dedup so real dated copies got
  deleted).

UX
- Dry-run and confirmation now print a Rich table of id/title/status/
  reason instead of just a count, so the user can audit what the
  heuristics matched before authorizing a destructive operation.
- Partial-failure output lists up to 5 failing source IDs with their
  exception messages so the user can retry or investigate.

Refactor
- Extract pure `_classify_junk_sources()` helper so the classification
  logic is unit-testable without spinning up the Click command.
- Add `_print_clean_candidates()` for the table rendering.

Tests
- TestSourceCleanClassify: 15 unit tests covering each branch (error
  status, gateway titles, dedup with oldest-first ordering, dedup when
  oldest is error, query-string preservation, fragment stripping,
  undated-source ordering, no-URL no-dedup, unknown/zero/processing
  statuses not flagged).
- TestSourceCleanCommand: 5 end-to-end Click tests covering
  already-clean, dry-run, --yes, declined confirmation, and partial
  delete failures.

All 88 source tests + 2635 full-suite tests pass; ruff/mypy clean.
@teng-lin

Copy link
Copy Markdown
Owner

Taking over this PR for finalization. Pushed polish commit 489d9a3 on top of the original 4 commits.

Polish pass (after a 3-way review by Claude / Gemini / Codex):

Issue raised by reviewers Fix in 489d9a3
title.startswith("http") deletes in-flight legitimate sources (Source.title is documented as "may be URL if not yet processed") Removed entirely. Gateway regex still catches CAPTCHA / Cloudflare / 4xx pages.
"unknown" in _JUNK_STATUSES wipes future / unrecognized status codes Restricted to {"error"} only
URL normalization strips query string → collapses YT ?v=AAA vs ?v=BBB, Google Docs IDs, arXiv versions Only #fragment is stripped now; query is preserved
created_at=None sorted to position 0 — undated source becomes "oldest" and is kept while real dated copies get deleted Use float("inf") to push undated sources to the end
Dry-run / confirm show count only — defeats the purpose of audit Print a Rich table of id / title / status / reason so the user can see exactly what will be deleted
Partial failures swallowed — user sees "success" even if every delete failed Collect (id, exception) tuples; show up to 5 failing IDs with their error messages
Zero unit tests for a destructive bulk-delete command 20 new tests: 15 TestSourceCleanClassify (pure-function branches) + 5 TestSourceCleanCommand (E2E Click invocations)

Live validation on a real notebook: created a notebook with 6 sources (1 errored + 3 example.com/ duplicates + 1 example.com/?ref=tracker + 1 distinct path + 1 Wikipedia). source clean --dry-run correctly flagged exactly 2 (the errored one and 1 duplicate of the most-duplicated URL), preserved the ?ref=tracker variant as distinct (proving the query-string fix), then source clean -y deleted them and a re-run reported "Notebook is already clean."

Local checks: ruff format --check . && ruff check . && mypy src/notebooklm --ignore-missing-imports && pytest — all 2635 tests pass.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🧹 Nitpick comments (1)
tests/unit/cli/test_source.py (1)

1800-1826: 💤 Low value

Consider contextlib.ExitStack for the patch stack.

The hand-rolled __enter__/__exit__ pair leaks state if a later patch(...).__enter__() raises: previously-entered patches won't be exited because the current __enter__ doesn't wrap subsequent enters in try/except. contextlib.ExitStack is the idiomatic Python primitive for this and gives you safe rollback for free.

♻️ Sketch
-import importlib
+import importlib
+from contextlib import ExitStack
 ...
 class _CleanPatch:
     def __init__(self, sources):
         self._sources = sources
-        self._stack: list = []
+        self._stack = ExitStack()

     def __enter__(self):
-        mock_client_cls_ctx = patch_client_for_module("source")
-        mock_client_cls = mock_client_cls_ctx.__enter__()
-        self._stack.append(mock_client_cls_ctx)
-
-        mock_client = create_mock_client()
-        mock_client.sources.list = AsyncMock(return_value=self._sources)
-        mock_client.sources.delete = AsyncMock(return_value=None)
-        mock_client_cls.return_value = mock_client
-        self.sources = mock_client.sources
-
-        fetch_ctx = patch("notebooklm.auth.fetch_tokens_with_domains", new_callable=AsyncMock)
-        fetch_mock = fetch_ctx.__enter__()
-        fetch_mock.return_value = ("csrf", "session")
-        self._stack.append(fetch_ctx)
-        return self
+        self._stack.__enter__()
+        try:
+            mock_client_cls = self._stack.enter_context(patch_client_for_module("source"))
+            mock_client = create_mock_client()
+            mock_client.sources.list = AsyncMock(return_value=self._sources)
+            mock_client.sources.delete = AsyncMock(return_value=None)
+            mock_client_cls.return_value = mock_client
+            self.sources = mock_client.sources
+
+            fetch_mock = self._stack.enter_context(
+                patch("notebooklm.auth.fetch_tokens_with_domains", new_callable=AsyncMock)
+            )
+            fetch_mock.return_value = ("csrf", "session")
+        except BaseException:
+            self._stack.__exit__(*sys.exc_info())
+            raise
+        return self

-    def __exit__(self, *exc):
-        for ctx in reversed(self._stack):
-            ctx.__exit__(*exc)
+    def __exit__(self, *exc):
+        return self._stack.__exit__(*exc)
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@tests/unit/cli/test_source.py` around lines 1800 - 1826, Replace the manual
enter/exit stack in class _CleanPatch with contextlib.ExitStack: initialize
self._stack = ExitStack() in __init__, replace calls to
patch_client_for_module(...).__enter__() and patch(...).__enter__() with
self._stack.enter_context(patch_client_for_module("source")) and
self._stack.enter_context(patch("notebooklm.auth.fetch_tokens_with_domains",
new_callable=AsyncMock)), and use self._stack.close() (or rely on
ExitStack.__exit__) in __exit__ so any failure during enter will automatically
roll back previously-entered contexts; keep references to mock_client_cls and
fetch_mock from the enter_context returns and retain self.sources assignment as
before.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@src/notebooklm/cli/source.py`:
- Around line 58-66: _normalize_url_for_dedup currently preserves original
casing of scheme and netloc so equivalent URLs with different casing (e.g.
HTTPS://Example.com) won't dedupe; update the function to normalize scheme and
host to lowercase before urlunparse (lowercase parsed.scheme and parsed.netloc
or use parsed.hostname with port handling) while still preserving path, params,
and query and stripping fragment, and add a unit test in TestSourceCleanClassify
that supplies mixed-case schemes/hosts (e.g. "HTTPS://Example.com/a" vs
"https://example.com/a") to assert they produce the same normalized key.

---

Nitpick comments:
In `@tests/unit/cli/test_source.py`:
- Around line 1800-1826: Replace the manual enter/exit stack in class
_CleanPatch with contextlib.ExitStack: initialize self._stack = ExitStack() in
__init__, replace calls to patch_client_for_module(...).__enter__() and
patch(...).__enter__() with
self._stack.enter_context(patch_client_for_module("source")) and
self._stack.enter_context(patch("notebooklm.auth.fetch_tokens_with_domains",
new_callable=AsyncMock)), and use self._stack.close() (or rely on
ExitStack.__exit__) in __exit__ so any failure during enter will automatically
roll back previously-entered contexts; keep references to mock_client_cls and
fetch_mock from the enter_context returns and retain self.sources assignment as
before.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro

Run ID: 028e233f-fab3-4c57-87f1-217c66c574ba

📥 Commits

Reviewing files that changed from the base of the PR and between e19eac6 and 489d9a3.

📒 Files selected for processing (2)
  • src/notebooklm/cli/source.py
  • tests/unit/cli/test_source.py

Comment thread src/notebooklm/cli/source.py Outdated
Follow-up from CodeRabbit review on commit 489d9a3:

- `_normalize_url_for_dedup` now lowercases scheme and netloc per RFC 3986
  (both are case-insensitive), so `https://Example.COM/a` dedups against
  `https://example.com/a`. Path/params/query/fragment handling unchanged.
- `_CleanPatch` test helper uses `contextlib.ExitStack` so a mid-setup
  exception in any patch correctly unwinds the patches that already
  entered, instead of leaking them into the next test.

Adds `test_dedup_is_case_insensitive_on_scheme_and_host` to lock in the
behaviour.
@teng-lin
teng-lin merged commit 3286709 into teng-lin:main May 13, 2026
19 checks passed
teng-lin pushed a commit that referenced this pull request May 21, 2026
Restructures the [0.5.0] - UNRELEASED section to lead with what users
need to act on, drop internal-refactoring prose, and stay accurate
against the actual commits since v0.4.1.

Structure (Keep a Changelog + upfront Breaking summary):
  Breaking changes (13 items) → Added (Auth/Chat/CLI/Python API)
  → Changed → Deprecated → Removed → Fixed → Security

Notable additions surfaced by the triple-model review (claude + codex
+ agy) and verified against gh pr view:
- Breaking: sources.add_* kw-only (#756), server_error_max_retries=3
  default (#629), max_concurrent_rpcs=16 default (#630), --storage
  context-isolation script-impact (#467).
- Added: NOTEBOOKLM_BASE_URL enterprise (#402), source clean (#261),
  create --use (#220/#413), Chromium profile selectors (#648), public
  rpc_call (#646), observability+drain (#643), correlation IDs +
  categorized logging (#430/#431), upload timeouts (#618), ChatRef
  enhancements (#686), citation hover save (#675).
- Changed: server-assigned conversation_id (#659/#667), cross-event-
  loop fail-fast (#633).
- Security: bundled #746 + #803 + #903 into one consolidated entry.

No source-code changes — CHANGELOG.md only.
teng-lin pushed a commit that referenced this pull request May 21, 2026
Restructures the [0.5.0] - UNRELEASED section to lead with what users
need to act on, drop internal-refactoring prose, and stay accurate
against the actual commits since v0.4.1.

Structure (Keep a Changelog + upfront Breaking summary):
  Breaking changes (13 items) → Added (Auth/Chat/CLI/Python API)
  → Changed → Deprecated → Removed → Fixed → Security

Notable additions surfaced by the triple-model review (claude + codex
+ agy) and verified against gh pr view:
- Breaking: sources.add_* kw-only (#756), server_error_max_retries=3
  default (#629), max_concurrent_rpcs=16 default (#630), --storage
  context-isolation script-impact (#467).
- Added: NOTEBOOKLM_BASE_URL enterprise (#402), source clean (#261),
  create --use (#220/#413), Chromium profile selectors (#648), public
  rpc_call (#646), observability+drain (#643), correlation IDs +
  categorized logging (#430/#431), upload timeouts (#618), ChatRef
  enhancements (#686), citation hover save (#675).
- Changed: server-assigned conversation_id (#659/#667), cross-event-
  loop fail-fast (#633).
- Security: bundled #746 + #803 + #903 into one consolidated entry.

No source-code changes — CHANGELOG.md only.
teng-lin added a commit that referenced this pull request May 21, 2026
)

Restructures the [0.5.0] - UNRELEASED section to lead with what users
need to act on, drop internal-refactoring prose, and stay accurate
against the actual commits since v0.4.1.

Structure (Keep a Changelog + upfront Breaking summary):
  Breaking changes (13 items) → Added (Auth/Chat/CLI/Python API)
  → Changed → Deprecated → Removed → Fixed → Security

Notable additions surfaced by the triple-model review (claude + codex
+ agy) and verified against gh pr view:
- Breaking: sources.add_* kw-only (#756), server_error_max_retries=3
  default (#629), max_concurrent_rpcs=16 default (#630), --storage
  context-isolation script-impact (#467).
- Added: NOTEBOOKLM_BASE_URL enterprise (#402), source clean (#261),
  create --use (#220/#413), Chromium profile selectors (#648), public
  rpc_call (#646), observability+drain (#643), correlation IDs +
  categorized logging (#430/#431), upload timeouts (#618), ChatRef
  enhancements (#686), citation hover save (#675).
- Changed: server-assigned conversation_id (#659/#667), cross-event-
  loop fail-fast (#633).
- Security: bundled #746 + #803 + #903 into one consolidated entry.

No source-code changes — CHANGELOG.md only.

Co-authored-by: Claude <claude@zfs.local>
zeekay pushed a commit to Dream-AI-4444/notebooklm-py that referenced this pull request Sep 9, 2026
Adds `notebooklm source clean` — bulk-remove junk sources (errored, gateway/anti-bot-blocked, duplicates) from a notebook.

## What it removes

- Sources with `status=error`
- Sources whose title matches a known gateway / anti-bot pattern (Cloudflare "Just a Moment", Access Denied, 403, 404, CAPTCHA, etc.)
- Duplicate URLs — fragment stripped, query preserved, scheme/host lowercased per RFC 3986; oldest copy of each URL is kept

## Flags

```
notebooklm source clean             # interactive confirmation
notebooklm source clean --dry-run   # preview as a Rich table without deleting
notebooklm source clean -y          # skip confirmation
notebooklm source clean -n <id>     # target specific notebook
```

## Implementation notes

- Dry-run and confirmation print a Rich table of id / title / status / reason so the user can audit before authorizing a destructive operation.
- Deletions are chunked (10 per batch with a 0.5s sleep between chunks) using the existing `asyncio.gather` pattern.
- Partial failures collect `(id, exception)` tuples and show up to 5 failing source IDs with their messages.
- Pure `_classify_junk_sources()` helper is unit-testable without spinning up the Click command — 16 classify tests + 5 end-to-end Click tests covering every branch.

Co-authored-by: Flosters <flosters@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

enhancement New feature or request

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants