Skip to content

feat: vault entry type auto-detection (local heuristics + AI-assisted with redaction) #19

Description

@f3rdy

Problem

Currently, vault entry types (usernamePassword, sshKey, certificate, secretText) must be set manually via the type field inside structured dict entries. For existing vaults with many entries — especially those migrated from plain key-value vaults — there is no automated way to detect or suggest the correct type.

The detect_entry_type() function in types.py only reads the explicit type field from dict entries or falls back to secretText. It does not infer types from key names, field patterns, or value structure.

Goal: Provide tooling to automatically detect or suggest entry types for untyped vault entries, enabling bulk migration of legacy vaults to the structured type system.

Critical constraint: Secret values must NEVER leave the local machine or be exposed to external services.


Proposed Approaches

Approach 1: Local Heuristic Detection (AI-free)

A built-in command (vaultctl detect-types or vaultctl analyze) that runs entirely locally using rule-based heuristics.

Detection signals:

Signal Example Inferred Type
Key name contains _user, _pass, _password, _login db_user, db_password usernamePassword
Key name contains _key, _ssh, _privkey deploy_ssh_key sshKey
Key name contains _cert, _certificate, _pem, _crt tls_cert certificate
Dict entry has username + password fields {username: ..., password: ...} usernamePassword
Dict entry has private_key or key field {private_key: ..., public_key: ...} sshKey
Dict entry has certificate or chain field {certificate: ..., chain: ...} certificate
Value starts with -----BEGIN (PEM header) -----BEGIN RSA PRIVATE KEY----- sshKey or certificate
Value starts with ssh-rsa, ssh-ed25519 ssh-ed25519 AAAA... sshKey

Heuristic priority: Structure (dict fields) > Value patterns (PEM headers) > Key name patterns.

Output modes:

  • --dry-run (default): Show suggestions without modifying anything
  • --apply: Write detected types into vault entries and/or vault-keys.yml
  • --json: Machine-readable output for scripting
  • --confidence: Show confidence level (high/medium/low) for each suggestion

Advantages:

  • Zero external dependencies
  • Works offline / air-gapped
  • Deterministic and auditable
  • Fast execution
  • No security concerns — everything stays local

Disadvantages:

  • Limited to known patterns — cannot handle unusual naming conventions
  • Requires maintenance of heuristic rules
  • May produce false positives for ambiguous entries

Approach 2: AI-Assisted Detection (with Redaction)

An optional command (vaultctl detect-types --ai) that sends redacted vault structure to an LLM for analysis.

Redaction strategy:

All secret values are replaced BEFORE any external communication. Only structural metadata is transmitted.

# ORIGINAL (never sent)
db_credentials:
  type: usernamePassword
  username: admin
  password: s3cr3t-p4ssw0rd!

# REDACTED (what gets sent)
db_credentials:
  type: usernamePassword
  username: "***REDACTED***"
  password: "***REDACTED***"

For plain string entries:

# ORIGINAL
api_token: ghp_1234567890abcdef

# REDACTED
api_token: "***REDACTED***"

What IS transmitted:

  • Key names (e.g., db_credentials, api_token)
  • Dict field names (e.g., username, password, private_key)
  • Explicit type fields (non-sensitive metadata)
  • Value type indicators (string, multiline, dict)

What is NEVER transmitted:

  • Actual secret values
  • Passwords, tokens, keys, certificates
  • File contents referenced by secrets

Safety controls:

  1. --ai flag must be explicitly provided (never default behavior)
  2. Interactive confirmation showing exactly what will be sent
  3. --show-payload to inspect the redacted payload without sending
  4. Audit log written to .vaultctl-ai-audit.log with timestamp, hash of payload, and endpoint
  5. Configurable endpoint in .vaultctl.yml (for self-hosted LLMs)

Advantages:

  • Can handle unusual naming conventions and edge cases
  • More nuanced suggestions with reasoning
  • Can suggest types for ambiguous entries where heuristics fail

Disadvantages:

  • Requires network access and API credentials
  • Key names themselves may be sensitive in some environments
  • Additional dependency on external service availability
  • Cost per API call
  • Requires user trust in redaction completeness

Implementation Plan

Phase 1: Local Heuristics (MVP)

  1. Add redact module (src/vaultctl/redact.py) — deterministic value redaction utility
  2. Add heuristic detection engine to types.py (or new detect.py module)
  3. Add vaultctl detect-types CLI command with --dry-run, --apply, --json, --confidence
  4. Unit tests for all heuristic rules
  5. Integration test with sample vault containing mixed entry types

Phase 2: AI-Assisted (Optional Extension)

  1. Add --ai flag to detect-types command
  2. Implement redaction pipeline with --show-payload preview
  3. Add interactive consent prompt
  4. Add audit logging
  5. Add .vaultctl.yml configuration for AI endpoint
  6. Tests for redaction completeness (fuzz testing for value leakage)

Security Requirements

These are non-negotiable constraints for both approaches.

  • No secret exfiltration: Actual secret values MUST NEVER be sent to any external service
  • Redaction is deterministic: The same input always produces the same redacted output — no randomness that could encode information
  • Redaction is complete: Every leaf value in the vault data structure must be replaced; nested dicts and lists must be recursively processed
  • Explicit opt-in: AI-assisted mode requires --ai flag; it is never the default
  • User consent: Before any network request, the user must see and confirm the exact payload
  • Audit trail: Every AI request is logged locally with timestamp, payload hash (SHA-256), and target endpoint
  • Configurable endpoint: Users can point to self-hosted LLMs (Ollama, vLLM) to avoid cloud services entirely
  • Key name sensitivity warning: Before transmitting, warn that key names themselves may reveal information (e.g., prod_db_password reveals infrastructure details)
  • No caching of AI responses with vault context: AI suggestions must not be persisted in a way that links them back to the original vault

Acceptance Criteria

Phase 1 (Local Heuristics)

  • vaultctl detect-types lists all vault entries with suggested types and confidence levels
  • --dry-run (default) shows suggestions without modifying files
  • --apply writes detected types into vault entries (dict entries get type field) and updates vault-keys.yml metadata
  • --json outputs machine-readable results
  • Entries that already have an explicit type field are skipped (or shown as "confirmed")
  • Heuristics cover all KNOWN_TYPES: secretText, usernamePassword, sshKey, certificate
  • Confidence levels: high (multiple signals match), medium (single signal), low (name-only guess)
  • Unit tests for each heuristic rule
  • Integration test with a vault containing at least one entry per type
  • Coverage >= 70% for new code

Phase 2 (AI-Assisted)

  • --ai flag enables AI-assisted detection
  • --show-payload prints the redacted payload and exits (no network request)
  • Interactive confirmation before sending (bypassable with --yes)
  • Audit log entry for every AI request
  • Configurable endpoint via .vaultctl.yml
  • Fuzz test: no secret value survives redaction across 1000+ random vault structures
  • Graceful fallback: if AI endpoint is unavailable, fall back to local heuristics with a warning

Open Questions

  1. Should detect-types also handle _previous backup keys (e.g., db_password_previous)? Probably skip them.
  2. Should we support custom type definitions beyond KNOWN_TYPES? Extensibility via config?
  3. For Phase 2: Which LLM providers should be supported out of the box? OpenAI, Anthropic, Ollama?
  4. Should the redaction module be usable standalone (e.g., vaultctl redact for debugging/export)?

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions