server · Issues· 895 open
Open on GitHubLocally synced open issues (discussions stay on GitHub)
- #8979
k8s-onprem autoscaling metric avg_time_queue_us under-reports queue time by roughly the number of loaded models
Updated Sep 18, 2026 - #8978
python_backend: error-string shm freed before parent reads it in is_ready() readiness check -> allocator corruption -> permanent hang
Updated Sep 18, 2026 - #8969
TensorRT backend leaks GPU resources after model initialization failure; unload does not recover memory
Updated Sep 11, 2026 - #8947
Random latency regression on L40S nodes after new pod creation — node becomes persistently "bad"
performanceUpdated Sep 11, 2026 - #8301
[Server] Support for numpy2.x
Updated Sep 8, 2026 - #8882
Intermittent server crash (`boost::interprocess::lock_exception` / `std::bad_alloc` / SIGSEGV) in Python backend under concurrent BLS load — 25.02 (2.55.0)
Updated Sep 6, 2026 - #8945
Triton TRT-LLM Multimodal guide outdated
Updated Aug 31, 2026 - #8853
Inaccuracies in PyTorch backend documentation
documentation issueDocumentationUpdated Aug 30, 2026 - #8117
[Question] Reloading models and latency spikes
Updated Aug 19, 2026 - #8922
InferenceRequest.async_exec(decoupled=True) intermittently raises `RuntimeError: Invalid argument` under concurrent in-flight BLS to a decoupled model
Updated Aug 15, 2026 - #8924
vLLM backend passes deprecated raw prompts to InputProcessor
Updated Aug 10, 2026 - #8841
Python stub aborted with glibc heap corruption, leading to 0 RPS on gRPC serving thread
Updated Aug 9, 2026 - #8919
test: L0_lifecycle uses blind time.sleep(5) for model load/unload, causing flaky CI failures
Updated Aug 6, 2026 - #8894
[vLLM backend][Feature Request] Support vLLM-Omni for native audio/video input
Updated Aug 4, 2026 - #8375
[Question] Different performance of vLLM server and Triton Inference Server
Updated Jul 30, 2026