[rollout][vllm] Garbled multi-language output after weight sync when `free_cache_engine=true` (sleep/resume corrupts rollout weights)
Environment
- verl 0.8.0 (pip), vllm 0.28.0, flashinfer 0.6.16.post3, torch 2.13 / CUDA 13, single NVIDIA A100-80GB
- Model: Qwen3-8B (dense, bf16), loaded from HF-merged checkpoint (no
lora_adapter_path) - LoRA r=64 trained on top of the merged base,
actor_rollout_ref.model.lora.merge=true - GRPO,
n=4,total_training_steps=20, multi-turn agent loop (2-stage label/utterance generation, 12-turn cap) actor_rollout_ref.rollout.free_cache_enginedefault (true)
Symptom
Starting from the second update_weights (i.e., after the first non-trivial weight sync), rollout outputs degrade into multi-language token soup (mixed CJK / Cyrillic / Arabic fragments of real vocabulary tokens). Per-step attribution logged in the rollout trajectories:
| global_step | n | script-corrupted | CJK |
|---|---|---|---|
| 0 (initial val, greedy) | 60 | 0 | 0 |
| 1 (train rollouts) | 32 | 0 | 1 |
| 2 (train rollouts + final val) | 52 | 32 | 32 |
The final greedy validation pass scored 0.0 (all trajectories E=A=0) because every response was corrupted. The user simulator even replies "I think your message got a bit garbled there. Could you say that again?" — and the next response is still garbage, so it is not a one-off sampling artifact.
What we ruled out (full evidence chain)
- Weights are correct at the sender: instrumented
get_per_tensor_param(merge branch) and compared the merged dict against a manual HF merge (base + B@A·alpha/r): 396 tensors, 0 exceed 1e-3 max abs diff. - Weights are correct at the receiver: instrumented
_update_weightsinverl/workers/rollout/vllm_rollout/utils.pyand dumped the received weights; they match the sender dump exactly (max diff 0.0 on common keys). - vllm's weight-loading paths are clean in isolation: a standalone AsyncLLMEngine replay of the exact same config — disk
reload_weights×4, in-memorymodel.load_weights×2, and 8-concurrent two-stage generation with a hot reload between batches — produced zero corruption in 96 utterances. - The RL update itself is innocent: the trained LoRA delta after 2 steps is ~1e-6 in magnitude; HF-merging it and generating locally is clean.
- The single remaining variable:
free_cache_engine. Re-running the exact same training withfree_cache_engine=falseproduces zero corrupted trajectories at gs=0/1/2 (script 0/0/0, CJK 0/1/0) and a healthy reward signal (E/A non-zero 78%).
The corruption therefore occurs somewhere in the sleep/resume cycle that free_cache_engine=true triggers around each weight sync (sleep(level=1) on the vllm engine followed by resume), where the rollout engine wakes with weights that are not what was synced — consistent with the GRPO sleep/wake gibberish fix in modelscope/ms-swift#7017.
Impact
Silent corruption: no error is raised; rollouts simply score zero and the training signal collapses (this masked the bug as "GRPO reward collapse" for several debugging rounds). It is the default configuration (free_cache_engine: bool = True).
Cross-version note
verl 0.9.0's vllm_async_server.py sleep/resume implementation is byte-identical to 0.8's (both call await self.engine.sleep(level=1) in colocated mode, same resume_kv_cache), and the default remains true, so 0.9 with the same vllm 0.28 is expected to be affected the same way (we could not complete an experimental confirmation on 0.9 because the pip-installed legacy runner hits an unrelated FSDP assertion in update_weights with lora.merge=true).
Workaround
Set actor_rollout_ref.rollout.free_cache_engine=false (verified clean).
Questions for maintainers
- Is
engine.sleep(level=1)/ wake-up expected to fully restore weights on vllm 0.28, or does this belong on the vllm side? - Would a warning on
free_cache_engine=true(or forcing a full weight re-sync after wake, cf. ms-swift #7017) be an acceptable fix?
Information
- The official example scripts
- My own modified scripts
Tasks
- An officially supported task in the
examplesfolder (such as GLUE/SQuAD, ...) - My own task or dataset (give details below)
Reproduction
Reproduction
Launch command (single A100-80GB; free_cache_engine is the default true):
CUDA_VISIBLE_DEVICES=0 python -m verl.trainer.main_ppo \
--config-name ppo_trainer \
algorithm.adv_estimator=grpo \
data.train_files=data/mcppo_esconv/train.parquet \
data.val_files=data/mcppo_esconv/val.parquet \
+data.apply_chat_template_kwargs.enable_thinking=false \
data.train_batch_size=8 \
actor_rollout_ref.actor.optim.lr=3e-6 \
data.max_prompt_length=1024 \
data.max_response_length=4096 \
actor_rollout_ref.model.path=/path/to/hf-merged-checkpoint \
actor_rollout_ref.model.lora_rank=64 \
actor_rollout_ref.model.lora.merge=true \
actor_rollout_ref.rollout.name=vllm \
actor_rollout_ref.rollout.gpu_memory_utilization=0.40 \
actor_rollout_ref.rollout.enforce_eager=true \
actor_rollout_ref.rollout.free_cache_engine=true \
actor_rollout_ref.rollout.temperature=0.2 \
actor_rollout_ref.rollout.n=4 \
actor_rollout_ref.rollout.agent.num_workers=4 \
actor_rollout_ref.rollout.agent.agent_loop_config_path=configs/agent_loop.yaml \
actor_rollout_ref.rollout.agent.default_agent_loop=cstpo_agent \
actor_rollout_ref.rollout.multi_turn.enable=true \
actor_rollout_ref.rollout.multi_turn.max_assistant_turns=12 \
trainer.total_training_steps=2 \
trainer.test_freq=2 \
custom_reward_function.path=custom_reward.py \
custom_reward_function.name=compute_score
The custom agent loop is a standard multi-turn loop: per turn it generates a strategy label under a structured_outputs choice constraint, then a free-form utterance, steps the environment, and appends the user reply (user tokens masked). Its only notable divergence from the built-in loop is logging global_step/is_validate into extra_fields so each trajectory can be attributed to a training step. The full source is available on request.
There is no error or stack trace — this is silent corruption. The first visible symptom is the reward collapse; the actual evidence is the rollout text itself. Example trajectory after the second weight sync (step-2 rollout):
{"seed_id": "esconv_04", "global_step": 2, "is_validate": false,
"dialogue": [{"role": "assistant", "text": "Hi there"},
{"role": "user", "text": "Hello"},
{"role": "assistant", "text": "If-away Extend Hardware.transitionsauc-exc Osborneةحفاكم_FindModelErrorundy comfy計 colspanQ ريبكلا join Norrispositive涨停んでいる喙uated SN中式验收읐 considerationculator Fold_blocked Fou yak ganz팡سلجم profesional书面senhaơ/Observable"},
{"role": "user", "text": "Sorry, I think your message got a bit garbled there. Could you say that again?"},
{"role": "assistant", "text": "Short晚上-await TSRMLS礴arlסרופ|RF蚌埠_ACTION[op讲究湍 wordKeywords Mons AIM HtmlWebpackPluginAudit拿到了 Crownsubmitted early一直没有wers dortTranslateartment qualifiesclassified octaveckeditor ةميق╠campaign℉˹-{지는nett.drawRectcoon ver"}]}
Thought for 9s
|---|---|---|
| 0 (initial val, greedy) | 60 | 0 |
| 1 (train rollouts) | 32 | 1 |
| 2 (train rollouts + final val) | 52 | 32 |
**Control**: the exact same command with `actor_rollout_ref.rollout.free_cache_engine=false` produces 0 corrupted trajectories at every step and a healthy reward signal.Per-step corruption counts across the whole run (script-corrupted = mixed CJK/Cyrillic/Arabic fragments):
| global_step | n | corrupted |
|---|---|---|
| 0 (initial v |
Expected behavior
Rollout generations should stay coherent after every update_weights. Since the synced weights are verified correct on both sender and receiver sides (see Reproduction), the model served by the vllm rollout engine should keep producing normal text after sleep/resume cycles — identical in quality to the free_cache_engine=false run. free_cache_engine should only trade GPU memory, never model behavior.
Source: verl-project/verl