trl · Issues· 284 open
Open on GitHubLocally synced open issues (discussions stay on GitHub)
- #7270
[Bug] GMPOTrainer with use_liger_kernel=True trains on the GRPO objective
Updated Sep 18, 2026 - #7268
SFTTrainer's QLoRA fp16 downcast also hits QDoRA magnitude vector, causing silent no-op updates
Updated Sep 18, 2026 - #6808
GMPO: use_liger_kernel=True silently bypasses the GMPO loss; MoE aux loss and entropy bonus dropped (gaps not covered by #6095)
Updated Sep 18, 2026 - #6789
GRPO: vLLM importance-sampling ratio is biased when top_p/top_k/min_p truncate sampling
Updated Sep 18, 2026 - #7244
Fused Triton kernels for the loss head
Updated Sep 17, 2026 - #7063
[Tracking] Move off liger-kernel and make the memory-efficient loss the default
Updated Sep 17, 2026 - #7222
router_aux_loss_coef is silently ignored on MoE configs without output_router_logits
Updated Sep 17, 2026 - #7202
Audit the trl-internal-testing repos that have no generation script
Updated Sep 17, 2026 - #7137
Align tiny test model configs with their reference models
Updated Sep 17, 2026 - #7182
[RFC] Criteria for removing code from `trl.experimental`
Updated Sep 17, 2026 - #7245
DPO and KTO skip DDP gradient synchronization with use_liger_kernel=True
Updated Sep 16, 2026 - #7015
get_repetition_penalty_reward accepts non-positive ngram_size
Updated Sep 16, 2026 - #7220
Apple Silicon (MPS): create_model_from_path can segfault while loading bf16 checkpoints
Updated Sep 15, 2026 - #6981
[Bug] get_dataset raises ZeroDivisionError for zero-sum fractions
Updated Sep 15, 2026 - #6151
Support router-based multi-teacher (MOPD) distillation in DistillationTrainer
Updated Sep 15, 2026