How to contribute to the Dependency Scout codebase — adding checks, ecosystems, detection patterns, fixing bugs, or improving the classifier.
To build on top of the Scout without modifying this repo (custom ecosystems, classifiers, or check plugins), see extending.md instead.
This is the lowest-barrier contribution: a two-line YAML edit, no Python required.
Detection patterns (regex strings that flag suspicious code in package diffs) live in checks/signatures/:
| File | What it covers |
|---|---|
net_calls.yaml |
Outbound network calls in library code, by language extension |
obfuscation.yaml |
Encoded payloads, zero-width Unicode tricks |
persistence.yaml |
OS persistence mechanisms, npm worm propagation |
file_types.yaml |
Suspicious filenames, dangerous binary extensions |
Add a pattern — for example, a new HTTP client library that a recent attack used:
# checks/signatures/net_calls.yaml, under the .py block:
- pattern: 'evil_http\.fetch\b'
desc: EvilHTTP library used in SupplyChainCorp May 2026 attackUse single-quoted YAML strings for regex — backslashes work as-is without doubling.
Run uv run pytest tests/test_signatures.py -v to verify your pattern compiles. Then run the full suite (uv run pytest -x -q).
Use the /add-detection Claude Code skill for a guided walkthrough in Claude Code.
The plugin architecture makes this straightforward. Adding a new ecosystem is approximately 150 lines in one file.
Also see ecosystems/README.md for the coverage table and required-method reference.
Step 1 — Create ecosystems/{name}.py
Implement the EcosystemProvider Protocol by inheriting from EcosystemProviderBase. Copy an existing provider (e.g. rubygems.py) as a starting point:
from ecosystems import EcosystemProviderBase, is_major, parse_upload_time, validate_archive_url
from models import AttestationChecks, MaintainerChecks, MetadataChecks, ReleaseAgeChecks
class ComposerProvider(EcosystemProviderBase):
async def fetch_metadata(self, package, old_version, new_version) -> MetadataChecks: ...
async def fetch_release_age(self, package, new_version) -> ReleaseAgeChecks: ...
async def fetch_maintainer(self, package, old_version, new_version) -> MaintainerChecks: ...
async def get_archive_url(self, client, package, version) -> tuple[str, str, str] | None: ...
def extract_archive(self, archive_bytes, filename, dest) -> None: ...
async def fetch_attestations(self, package, old_version, new_version) -> AttestationChecks: ...
async def fetch_release(self, package, old_version, version) -> ReleaseChecks: ...get_archive_url returns (url, filename, integrity_string). Call validate_archive_url(url) before returning — this enforces the CDN allowlist. Add your registry's CDN host to ALLOWED_CDN_HOSTS in ecosystems/__init__.py.
Step 2 — Register metadata in ecosystems/_registry.py
Add an EcosystemMeta entry to BUILTIN_ECOSYSTEMS. This is the single place that controls the ecosystem name, Dependabot slug, OSV ecosystem string, and package-name validation regex:
EcosystemMeta(
name="composer",
dependabot_slug="composer",
osv_name="Packagist",
name_re=re.compile(r"^[A-Za-z0-9][A-Za-z0-9._-]{0,99}/[A-Za-z0-9][A-Za-z0-9._-]{0,99}$"),
module="ecosystems.composer",
class_name="ComposerProvider",
),This replaces what previously required edits in three separate files (pr_parser.py, api/webhook.py, and the class attributes). get_provider(), get_dependabot_slug_map(), and get_name_re() all read from here.
Step 3 — Update the type model
In models/__init__.py, add the new ecosystem name to the Literal[...] types for ecosystem names.
Step 4 — Write tests
Add a test file under tests/ following the existing patterns (e.g. tests/test_pip_*.py or tests/test_npm_*.py). Use respx for HTTP mocking and ActivityEnvironment for the activity harness. Each method needs at minimum: success case, 404/not-found case, and (for attestations) a no-attestation case. Aim to keep overall coverage above 95%.
Step 5 — Regenerate replay fixtures (if needed)
If you changed any workflow code (unlikely for a pure ecosystem add, but possible):
uv run python tests/generate_fixtures.pyEach check is a parallel activity that gathers one category of supply chain data and returns a typed sub-model. All checks run at the same time; a failing check gets degraded defaults rather than crashing the workflow.
Step 1 — Define the sub-model in models/
class MyNewChecks(BaseModel):
some_flag: bool = False
some_score: int | None = NoneAll fields need defaults so the model can be instantiated as a zero/degraded state when the activity fails. Then add it as a nested field in PackageChecks:
class PackageChecks(BaseModel):
...
my_new: MyNewChecks = Field(default_factory=MyNewChecks)Step 2 — Create checks/my_new_check.py
from temporalio import activity
from models import MyNewChecks
@activity.defn(name="activities.my_new_check.check")
async def check(ecosystem: str, package: str, old_version: str, new_version: str) -> MyNewChecks:
try:
# fetch data from an external API
return MyNewChecks(some_flag=True, some_score=42)
except Exception:
return MyNewChecks() # degraded defaults — never raise from a check activityThe activity name string is what the workflow references. It must match exactly.
Caching — add an ActivityCache to avoid redundant network calls when multiple repos bump the same package simultaneously:
from helpers.cache import ActivityCache, INDEFINITE
# Pick a TTL that matches how often the data can legitimately change:
_cache = ActivityCache() # immutable (archive contents, provenance, upload timestamps)
_cache = ActivityCache(ttl_seconds=3600) # can change, but rarely within an hour (CVEs, scores)
_cache = ActivityCache(ttl_seconds=86400) # changes slowly (deprecation, repo health)
@activity.defn(name="activities.my_new_check.check")
async def check(ecosystem, package, old_version, new_version):
key = (ecosystem, package, new_version) # omit old_version if it doesn't affect the result
if (hit := _cache.get(key)) is not None:
activity.logger.debug("my_new_check cache hit: %s %s", package, new_version)
return hit
result = await _fetch(...)
_cache.set(key, result) # only cache successful results — don't cache degraded defaults
return resultTests are isolated automatically — a conftest.py fixture clears all caches before and after each test.
Step 3 — Add to _CHECK_REGISTRY in workflows/package_triage_workflow.py
Append a row to _CHECK_REGISTRY — do not insert mid-list, as this changes Temporal's command sequence and breaks replay of existing histories:
_CHECK_REGISTRY: list[tuple[str, str, type, bool]] = [
...
("my_new", "activities.my_new_check.check", MyNewChecks, False),
# ^activity name string ^result type ^True = 2-min timeout
]The first element ("my_new") must match the field name you added to PackageChecks.
Step 4 — Use the check in classifiers/
Add logic to _rule_based using signals.my_new.some_flag, and/or let the LLM see it — it already appears in the trusted JSON block via PackageChecks.model_dump(). Forgetting this step will cause a test failure in tests/test_check_wiring.py.
Step 5 — Regenerate replay fixtures
Any change to the workflow's gather list changes its Temporal command sequence. Regenerate the determinism fixtures:
# Temporal must be running — the generator executes real workflows
temporal server start-dev # in a separate terminal, if not already running
uv run python tests/generate_fixtures.pyThe script prints one line per fixture as it runs. If it hangs, Temporal isn't reachable. Commit the updated files in tests/fixtures/.
make bootstrap # uv sync + installs pre-commit hooks
uv run ruff format . # format
uv run ruff check . # lint
uv run mypy . # type check
uv run pytest # testsOr with Docker (no local Python required):
cp .env.example .env # fill in what you have
docker compose upThe Temporal UI will be at http://localhost:8233.
If you change workflows/package_triage_workflow.py or workflows/pr_action_workflow.py, regenerate the replay fixtures:
temporal server start-dev # must be running — generator executes real workflows
uv run python tests/generate_fixtures.pyCommit the updated files in tests/fixtures/. The CI pytest run will catch any determinism regression.
See architecture.md for the design rationale.
The built-in classifiers (AnthropicClassifier, RuleBasedClassifier) are selected automatically based on whether ANTHROPIC_API_KEY is set. You can override either choice or supply your own engine.
Force a built-in by name (useful to pin rule-based even when a key is present):
# .env
CLASSIFIER=rule_based # or: claudeAdd a new built-in classifier to this repo:
- Create
classifiers/{name}.pywith a class implementingasync def classify(self, signals: PackageChecks) -> Verdict - In
classifiers/__init__.py, add an entry to_BUILTIN_CLASSIFIERS - Write tests following the patterns in
tests/test_classifier.py
Register a third-party classifier in a plugin package — see extending.md.
- Graceful degradation — missing API keys or upstream errors produce degraded check defaults, not a crash. This is enforced at two levels: each activity catches its own errors and returns a zero-state model; the workflow's
asyncio.gatherusesreturn_exceptions=Trueso a single failing activity never discards the other ten results. - Attacker-controlled data stays sandboxed — package descriptions, socket alert strings, release notes, and diff content go into clearly-labelled XML tags in the LLM prompt and are explicitly named in the system prompt as untrusted.
- No silent fallbacks — use
ApplicationError(..., non_retryable=True)for permanent errors (404, auth failure) so Temporal doesn't retry endlessly. - Archive URLs are validated before any HTTP request — add new CDN hosts to
ALLOWED_CDN_HOSTS, never skip thevalidate_archive_url()call. - Checks are nested, not flat —
PackageChecksholds typed sub-models (signals.age.release_age_hours, notsignals.release_age_hours). Field name collisions between checks are structurally impossible. - The registry is the source of truth —
_CHECK_REGISTRYinpackage_triage_workflow.pydrives the gather order, theCHECK_ACTIVITY_NAMESconstant, and the worker registration test. When adding a check, the registry row is the one required workflow-layer edit. - Worker registration is automatic —
worker.pyscanschecks/*.pyandplatform/*.pyat startup for@activity.defn-decorated functions. Creating a new activity file is sufficient; no manual entry inACTIVITIESis needed. - Ecosystem registry is explicit and metadata-first —
ecosystems/_registry.pyholds all built-in metadata (slug, OSV name, name regex, module path) as stdlib-only data.get_provider()lazy-imports exactly one provider module on first use. Plugin providers register anEcosystemMetainstance via thedependency_scout.ecosystemsentry point — no module scan, no class-attribute sniffing.
- pip/uv, npm, RubyGems, Maven/Gradle (Java/JVM), Composer (PHP), NuGet (.NET), Cargo (Rust), Go Modules, GitHub Actions, Mix/Hex (Elixir), Pub (Dart/Flutter), Elm, Docker, Terraform, Swift
- Eleven parallel check sources (OSV, Socket.dev, diff, release age, maintainer, SLSA/Sigstore, OpenSSF Scorecard, deps.dev deprecation, version staleness, PR file audit, metadata)
- LLM classifier with rule-based fallback
- GitHub and GitLab support
- FastAPI webhook receiver
- Per-repo config via
.github/dependency-scout.yml - Observe-only safe default (comment-only with no config file)
- Replay test fixtures (workflow determinism guarantee)
- Ecosystem plugin architecture — entry points +
RemoteEcosystemProviderHTTP bridge for non-Python stacks - Pluggable classifier — Claude, OpenAI, Ollama, or any
dependency_scout.classifiersplugin - Custom check plugin architecture — third-party checks via
dependency_scout.checksentry points, no Temporal internals required, surfaced to LLM automatically - Temporal Cloud support — TLS credentials in
.env, no code changes needed vs local dev - Renovate full support — title variants with/without
dependencykeyword, arrow and from/to body extraction, pre-release versions, false-positive prevention