Skip to content

Return rate limit error instead of quota error from LanguageModelRateLimitingPlugin - #1912

Open
waldekmastykarz wants to merge 2 commits into
dotnet:mainfrom
waldekmastykarz:llm-rate-limit-error-body
Open

waldekmastykarz wants to merge 2 commits into
dotnet:mainfrom
waldekmastykarz:llm-rate-limit-error-body

Conversation

@waldekmastykarz

Copy link
Copy Markdown
Collaborator

Closes #1911

When throttling with whenLimitExceeded: "Throttle", LanguageModelRateLimitingPlugin returned OpenAI's insufficient_quota billing error, which tells clients to stop retrying. It now returns OpenAI's tokens-per-minute rate limit error so clients back off and retry:

{
  "error": {
    "message": "Rate limit reached for gpt-4o on tokens per min (TPM): Limit 10, Used 10. Please try again in 299s.",
    "type": "tokens",
    "param": null,
    "code": "rate_limit_exceeded"
  }
}
  • The message includes the model (when present in the request), the exhausted limit, tokens used, and seconds until reset (same value as the retry-after header).
  • whenLimitExceeded: "Custom" is unchanged, for anyone who wants to simulate a quota error on purpose.
  • Updated the skill docs example and added an integration test.

…eModelRateLimitingPlugin. Closes dotnet#1911

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot AI balanced review requested due to automatic review settings October 3, 2026 18:18
@waldekmastykarz
waldekmastykarz requested a review from a team as a code owner October 3, 2026 18:18

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

The new error message underreports consumed tokens when usage exceeds the configured limit.

Review effort: Balanced
Findings: 1 Medium severity

Open (1)
What changed in this PR

Changes the default language-model throttle response to a retryable OpenAI rate-limit error, addressing #1911.

Changes:

  • Reports the model, token limit, usage, and retry delay.
  • Adds integration coverage for the throttle response.
  • Updates documentation while preserving custom responses.
File Description
skills/​dev-proxy/​references/​test-llm-apps.md Updates error examples and explains custom quota errors.
DevProxy.Plugins/​Behavior/​LanguageModelRateLimitingPlugin.cs Returns rate-limit errors with token and retry details.
DevProxy.Integration.Tests/​BehaviorPluginsIntegrationTests.cs Tests the default throttle response.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread DevProxy.Plugins/Behavior/LanguageModelRateLimitingPlugin.cs Outdated
…ttle message

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

LanguageModelRateLimitingPlugin returns a billing error (insufficient_quota) instead of a rate limit error

2 participants