guarded-hotshard

Tenant-aware request scheduling for LLM inference. Cuts premium-tenant p99 latency 60-70% on a contested GPU at <5% extra cost.

Installation

In a virtualenv (see these instructions if you need to create one):

pip3 install guarded-hotshard

Dependencies

Releases

Version Released Bullseye
Python 3.9
Bookworm
Python 3.11
Trixie
Python 3.13
Files
0.2.0 2026-05-02      
0.1.0 2026-05-02      

Issues with this package?

Page last updated 2026-05-17 03:13:53 UTC