Skip to content

missing cost fixes - #137

Merged
adambalogh merged 7 commits into
mainfrom
ani/missing-cost-fixes
Aug 5, 2026
Merged

adambalogh merged 7 commits into
mainfrom
ani/missing-cost-fixes

Conversation

@dixitaniket

@dixitaniket dixitaniket commented Jul 27, 2026 •

Copy link
Copy Markdown
Collaborator
  • provider tests for chat and image models
  • image url retry for providers (glm image)

@adambalogh
adambalogh marked this pull request as ready for review July 27, 2026 15:25
@adambalogh
adambalogh requested a review from Copilot July 28, 2026 14:26

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR fixes cases where successful inferences were missing token-usage metadata (and therefore missing OpenGradient cost/settlement data), with a particular focus on xAI’s LangChain response shapes and streaming behavior.

Changes:

  • Expand extract_usage() to normalize usage from both LangChain’s usage_metadata and raw provider payloads under response_metadata (e.g., token_usage / usage).
  • Ensure xAI non-streaming requests don’t accidentally use a streaming-configured cached model, and ensure xAI streaming requests explicitly request usage when applicable.
  • Add targeted regression tests covering xAI streaming/non-streaming and the new extract_usage() fallbacks.

Reviewed changes

Copilot reviewed 4 out of 4 changed files in this pull request and generated no comments.

File Description
tests/test_opengradient_field.py Adds regression tests ensuring xAI streaming/non-streaming paths propagate usage/cost behavior correctly.
tee_gateway/test/test_tee_core.py Adds unit tests for llm_backend.extract_usage() fallback behavior when LangChain omits usage_metadata.
tee_gateway/llm_backend.py Extends extract_usage() to support additional metadata shapes and normalize reasoning tokens.
tee_gateway/controllers/chat_controller.py Adjusts xAI streaming/non-streaming invocation kwargs and uses extract_usage() for streaming usage accumulation; logs when usage is missing.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

This comment was marked as outdated.

…nt stream kwargs (#142)

The completions endpoint invoked models without forcing stream=False, so
xAI requests hit the same summed-cumulative-usage inflation the chat
endpoint was just fixed for. Hoist the workaround into
non_streaming_invoke_kwargs() and use it from both controllers.

Also remove the per-call stream_usage=True on model.stream(): every
provider is already constructed with stream_usage=True, and the
web_search gate implied the deprecated flag still selected the Responses
API. Replace the two tests pinning that kwarg with one covering the
keep-latest cumulative snapshot behavior, and add a completions-endpoint
test for the stream=False kwarg.


Claude-Session: https://claude.ai/code/session_013bcpZQZ8pAgHL6z4gBZrDM

Co-authored-by: Claude <noreply@anthropic.com>

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Copilot reviewed 9 out of 9 changed files in this pull request and generated no new comments.

Suppressed comments (2)

tee_gateway/controllers/chat_controller.py:571

  • This comment claims every provider is constructed with stream_usage=True, but Google models (ChatGoogleGenerativeAI) are constructed without that flag (see tee_gateway/llm_backend.py around the provider == "google" branch). Updating the comment avoids misleading future readers about why usage is available on stream chunks.
                    # Terminal usage needs no per-call opt-in: every provider
                    # is constructed with stream_usage=True (see
                    # get_chat_model_cached).

tee_gateway/test/test_provider_usage_integration.py:58

  • In this live integration test list, glm-5.2 is labeled as provider Z.ai and tied to ZAI_API_KEY, but the model registry routes glm-5.2 through the ByteDance/ModelArk client (provider bytedance; see tee_gateway/model_registry.py:591-596). This mismatch is misleading and makes it harder to reason about which key is actually required for the test case.
    _ProviderCase("Z.ai", "glm-5.2", "ZAI_API_KEY"),

@adambalogh
adambalogh merged commit 3eead2e into main Aug 5, 2026
10 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants