Skip to content

brightstaff: add output_filter_mode: buffered for model-listener output filters - #1040

Open
ninadphalak wants to merge 1 commit into
katanemo:mainfrom
ninadphalak:fix/1036-buffered-output-filters
Open

ninadphalak wants to merge 1 commit into
katanemo:mainfrom
ninadphalak:fix/1036-buffered-output-filters

Conversation

@ninadphalak

Copy link
Copy Markdown

Closes #1036.

What changes

A model listener gets an optional output_filter_mode:

listeners:
  - type: model
    name: llm_gateway
    port: 12000
    output_filters:
      - output_redactor
    output_filter_mode: buffered   # default: streaming
  • streaming (default) is today's behaviour, unchanged: each upstream chunk goes through the filter chain on its own.
  • buffered collects the whole upstream body, calls the filter chain once with it, and returns the result.

This addresses both findings in #1036:

  1. Split values. In streaming mode a filter never sees text that spans a chunk boundary, so SECRET_ + TOKEN. reaches the client unredacted. In buffered mode the filter gets The token is SECRET_TOKEN. in one call.
  2. Filter errors. In streaming mode a failed filter call forwards the original chunk. In buffered mode a filter error, an upstream stream error, or a body over 64 MiB withholds the body (the client gets the upstream status and headers with an empty body, and the error is logged and recorded on the span/metrics). This also covers the stream: false case from the issue, where the filter received gzip data in two pieces and could not decompress either.

Why opt-in

Buffering costs streaming: the client sees nothing until the provider finishes. Filters that do not need whole values (logging, metrics, per-event transforms) should not pay that, so the default stays streaming. A carry-over buffer would keep streaming, but the gateway does not know how long a filter's match can be, so that needs a filter-side protocol (the connection model in #834). This PR is the small, safe option in the meantime.

Other changes

  • Content-Length from the upstream response is no longer copied when output filters are configured, since a filter can change the body length. This applies in both modes.
  • config/plano_config_schema.yaml accepts output_filter_mode (streaming | buffered); otherwise planoai rejects the key.
  • Reference config (plano_config_full_reference*.yaml) and the model-listener demo README document the mode.

Tests

  • streaming::output_filter_mode_tests (mockito filter, two-chunk upstream The token is SECRET_ + TOKEN.):
    • buffered: the filter is called once with the joined body and the client gets The token is [REDACTED].
    • buffered: the filter returns 500, the client body is empty
    • streaming: the filter is called twice, never with the whole value (documents the default)
  • configuration::test::test_listener_output_filter_mode_deserialize: buffered, absent (defaults to streaming), unknown value rejected.
  • CLI: valid_listener_output_filter_mode_buffered schema case (fails without the schema change).

Run locally: cargo fmt --all -- --check, cargo test -p brightstaff (250 passed, 2 ignored), cargo test -p common, pytest cli/test/test_config_generator.py (27 passed). cargo clippy --all-targets --all-features -- -D warnings on Rust 1.99 reports only clippy::double_must_use at state/mod.rs:64 and session_cache/mod.rs:175, identical on main; with that lint allowed it is clean. I did not rebuild the Docker image to rerun the #1036 repro end to end.

…ut filters

Output filters receive each upstream chunk on its own, so a value split
across two chunks is never seen whole, and a filter error forwards the
original chunk (katanemo#1036).

output_filter_mode: buffered collects the whole upstream response, runs
the filter chain once and returns the result. A stream error, a filter
error or a body over 64 MiB withholds the body. The default stays
streaming. Content-Length from the upstream is dropped when output
filters run, since a filter can change the body length.

Closes katanemo#1036

@fairozfouz-max fairozfouz-max left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Output filters see each streamed chunk alone, so a value split across two SSE events passes output_redactor

2 participants