Bundle
dsh-llm-rate-limit
LLM API rate limiting, concurrency control, queuing, and adaptive cooldown for DeepSeek Harness
Asong68241
2 results
LLM API rate limiting, concurrency control, queuing, and adaptive cooldown for DeepSeek Harness
Patient auto-retry for DeepSeek Harness: honor the upstream gateway's retry_after_seconds on a 429 capacity cooldown instead of giving up after two fast retries, and wait out the mislabeled modality 400 a saturated gateway replays from its client-error circuit. Optional floating countdown badge included.