[Serve] [SGLang] [Draft] POC PD disaggregation - #63741
Conversation
There was a problem hiding this comment.
Code Review
This pull request introduces SGLang Prefill-Decode (PD) disaggregated LLM serving by adding SGLangPDPrefillServer and SGLangPDDecodeServer deployments, and updating the SGLang engine to pass bootstrap coordination fields. Feedback on these changes highlights critical issues: first, setting attributes directly on Pydantic v2 models will raise runtime errors, so object.__setattr__ should be used instead; second, the prefill generator must be consumed in a background task to prevent premature garbage collection and engine hangs; third, asyncio needs to be imported to support this background task; and finally, the new servers should be added to the __all__ list in deployment.py to be properly exposed in the public API.
|
can you add a target in |
|
Added the target in release_tests.yaml. Thanks |
jeffreywang88
left a comment
There was a problem hiding this comment.
Left some design questions for you to consider while revising the RFC.
| to the right place. The decode server generates the bootstrap_room upfront and | ||
| dispatches both sides simultaneously — it does not wait for a prefill response | ||
| before starting decode. |
There was a problem hiding this comment.
Looks similar to the parallel handoff pattern (https://github.com/ray-project/ray/pull/63950/changes#diff-ccc2365e3ea6c5f4fd309223ecbaaa15d44deb1d510f0acf985604060a7f1747R56). Could you assess the feasibility of leveraging the established concurrent handoff mechanism?
| # LLMServer.get_deployment_options calls get_engine_config(), which | ||
| # unconditionally imports vLLM. The SGLang byod image uninstalls vLLM, |
There was a problem hiding this comment.
We will need to somehow decouple vLLM from the common paths. It might require a refactor. Could you explore some options?
| carrying the same prefill bootstrap host/port/room — the decode | ||
| KVReceiver connects to that prefill bootstrap server and blocks | ||
| internally waiting for the KV cache to arrive. | ||
| 5. Streams the decode response back to the client. |
There was a problem hiding this comment.
This is exactly what we need, nice!
| For each chat/completions request it: | ||
| 1. Reads the PREFILL node's bootstrap_host and bootstrap_port, fetched | ||
| from the prefill deployment at init (the bootstrap server lives there). | ||
| 2. Generates a unique bootstrap_room integer. |
There was a problem hiding this comment.
In the design doc, it'd be great to clarify what bootstrap_room does and where it lives.
| # TODO: Users currently need to set disaggregation_mode manually in engine_kwargs. | ||
| # The builder should set this automatically since it already knows which | ||
| # config is prefill and which is decode. |
There was a problem hiding this comment.
We should emulate ray serve LLM's PD API for vLLM here. The user API should be similar.
|
|
||
| Unlike the vLLM flow, we do not wait for a prefill response before | ||
| starting decode — the bootstrap_room is established upfront and both | ||
| sides coordinate directly via SGLang's bootstrap server. |
There was a problem hiding this comment.
Is this bootstrap server strictly required?
| pass | ||
|
|
||
|
|
||
| @PublicAPI(stability="beta") |
There was a problem hiding this comment.
We should start from alpha.
| - UCX_TLS=all | ||
| - UCX_NET_DEVICES=all |
There was a problem hiding this comment.
Could you help me understand why we need these?
There was a problem hiding this comment.
When we use SGLang with NIXL, it uses a networking library called UCX to move data directly from one GPU to another.
By default, our testing environment (the CI box) has strict or limited network paths. Without these two settings, UCX gets confused, picks a network path that doesn't actually have access to the GPUs, and the data transfer crashes.
Here is what these two lines specifically tell UCX to do:
-
UCX_TLS=all: Tells it, 'You are allowed to use any available transport method (like shared memory, TCP, or CUDA-IPC) to move this data.'
-
UCX_NET_DEVICES=all: Tells it, 'You are allowed to look at all network devices on this machine to find a working path.'
Basically, it forces UCX to stop being picky and use whatever path works on our test machines so the test doesn't fail. Once we know the absolute bare minimum network settings needed for this specific CI box, we can narrow this down."
Signed-off-by: Limark Dcunha <limarkdcunha@gmail.com>
|
This pull request has been automatically marked as stale because it has not had You can always ask for help on our discussion forum or Ray's public slack channel. If you'd like to keep this open, just leave any comment, and the stale label will be removed. |
|
The work on this PR has been paused due to unavailability of GPU's I am working on getting access to some resources, commenting it here so this PR is not marked as stale. |
Signed-off-by: Limark Dcunha <limarkdcunha@gmail.com>
Signed-off-by: Limark Dcunha <limarkdcunha@gmail.com>
Signed-off-by: Limark Dcunha <limarkdcunha@gmail.com>
Signed-off-by: Limark Dcunha <limarkdcunha@gmail.com>
Signed-off-by: Limark Dcunha <limarkdcunha@gmail.com>
Signed-off-by: Limark Dcunha <limarkdcunha@gmail.com>
Signed-off-by: Limark Dcunha <limarkdcunha@gmail.com>
Signed-off-by: Limark Dcunha <limarkdcunha@gmail.com>
Signed-off-by: Limark Dcunha <limarkdcunha@gmail.com>
Signed-off-by: Limark Dcunha <limarkdcunha@gmail.com>
There was a problem hiding this comment.
Cursor Bugbot has reviewed your changes using default effort and found 1 potential issue.
Reviewed by Cursor Bugbot for commit 9cc692b. Configure here.
Signed-off-by: Limark Dcunha <limarkdcunha@gmail.com>
Launching 8 SGLang processes simultaneously spikes the container's cgroup PID/thread count past its limit for a moment (each process spins up its own tokenizer thread pool, scheduler, detokenizer), crashing whichever process happens to hit the ceiling with a rayon ThreadPoolBuildError. A few seconds between launches avoids it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Signed-off-by: Limark Dcunha <limarkdcunha@gmail.com>
Signed-off-by: Limark Dcunha <limarkdcunha@gmail.com>

Description
A POC PR for the my proposal (#63257) related to Ray Serve SGLang PD disaggregation support.
Related issues - #62792 #63257