Repository navigation
missing cost fixes - #137
Merged
Merged
Conversation
Contributor
There was a problem hiding this comment.
Pull request overview
This PR fixes cases where successful inferences were missing token-usage metadata (and therefore missing OpenGradient cost/settlement data), with a particular focus on xAI’s LangChain response shapes and streaming behavior.
Changes:
- Expand
extract_usage()to normalize usage from both LangChain’susage_metadataand raw provider payloads underresponse_metadata(e.g.,token_usage/usage). - Ensure xAI non-streaming requests don’t accidentally use a streaming-configured cached model, and ensure xAI streaming requests explicitly request usage when applicable.
- Add targeted regression tests covering xAI streaming/non-streaming and the new
extract_usage()fallbacks.
Reviewed changes
Copilot reviewed 4 out of 4 changed files in this pull request and generated no comments.
| File | Description |
|---|---|
| tests/test_opengradient_field.py | Adds regression tests ensuring xAI streaming/non-streaming paths propagate usage/cost behavior correctly. |
| tee_gateway/test/test_tee_core.py | Adds unit tests for llm_backend.extract_usage() fallback behavior when LangChain omits usage_metadata. |
| tee_gateway/llm_backend.py | Extends extract_usage() to support additional metadata shapes and normalize reasoning tokens. |
| tee_gateway/controllers/chat_controller.py | Adjusts xAI streaming/non-streaming invocation kwargs and uses extract_usage() for streaming usage accumulation; logs when usage is missing. |
💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.
…ing-cost-fixes # Conflicts: # tee_gateway/image_generation.py
…nt stream kwargs (#142) The completions endpoint invoked models without forcing stream=False, so xAI requests hit the same summed-cumulative-usage inflation the chat endpoint was just fixed for. Hoist the workaround into non_streaming_invoke_kwargs() and use it from both controllers. Also remove the per-call stream_usage=True on model.stream(): every provider is already constructed with stream_usage=True, and the web_search gate implied the deprecated flag still selected the Responses API. Replace the two tests pinning that kwarg with one covering the keep-latest cumulative snapshot behavior, and add a completions-endpoint test for the stream=False kwarg. Claude-Session: https://claude.ai/code/session_013bcpZQZ8pAgHL6z4gBZrDM Co-authored-by: Claude <noreply@anthropic.com>
Contributor
There was a problem hiding this comment.
Pull request overview
Copilot reviewed 9 out of 9 changed files in this pull request and generated no new comments.
Suppressed comments (2)
tee_gateway/controllers/chat_controller.py:571
- This comment claims every provider is constructed with
stream_usage=True, but Google models (ChatGoogleGenerativeAI) are constructed without that flag (see tee_gateway/llm_backend.py around theprovider == "google"branch). Updating the comment avoids misleading future readers about why usage is available on stream chunks.
# Terminal usage needs no per-call opt-in: every provider
# is constructed with stream_usage=True (see
# get_chat_model_cached).
tee_gateway/test/test_provider_usage_integration.py:58
- In this live integration test list,
glm-5.2is labeled as providerZ.aiand tied toZAI_API_KEY, but the model registry routesglm-5.2through the ByteDance/ModelArk client (providerbytedance; see tee_gateway/model_registry.py:591-596). This mismatch is misleading and makes it harder to reason about which key is actually required for the test case.
_ProviderCase("Z.ai", "glm-5.2", "ZAI_API_KEY"),
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Uh oh!
There was an error while loading. Please reload this page.