Skip to content

[Feature]: Model Failover Policy #3356

Description

@Thenujan-Nagaratnam

Please select the area the issue is related to

Gateway

Please select the aspect the issue is related to

Aspect/API (API backends, definitions, contracts, interfaces, OpenAPI)

Suggested Feature

An application built on top of an LLM proxy API has no way to stay up when the specific model or provider it's calling degrades or goes down — a 500-class response from the primary model is simply handed back to the app as-is. Today the app developer has to write and maintain their own retry-and-fallback logic client-side: catch the failure, pick a fallback model, potentially re-authenticate and re-shape the request for a completely different provider's API, and retry — for every app that calls the gateway.

We need a policy for this and automatically fallback to the configured models when sth goes wrong with a model.

Related Issues

No response

Steps to Verify

  • Design Document — A detailed design document has been created and reviewed, covering architecture, data flow, and edge cases.
  • Design Mail — A design summary email has been sent to relevant stakeholders for awareness and feedback.
  • Code Review — All code changes have been peer-reviewed and approved according to the project's review standards.
  • Testing Complete — Adequate unit, integration, and/or end-to-end tests have been written and are passing.
  • Documentation Review — User-facing and/or developer documentation has been updated to reflect the new feature and reviewed.
  • Feature Complete — The feature is fully implemented, all checklist items above are done, and it is ready for release.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Labels

Aspect/APIAPI definitions, contracts, OpenAPI, interfaces

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions