feat(sft): support logits selection with Ulysses sequence parallelism - #10191
Open
Excelius-Wang wants to merge 1 commit into
Open
Excelius-Wang wants to merge 1 commit into
Excelius-Wang wants to merge 1 commit into
Conversation
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Sequence-parallel SFT currently rejects explicit
use_logits_to_keep=true. Add opt-in selection after SP has shifted and sharded labels, retaining a sentinel on prompt-only ranks. Restore scalar token losses and integer predictions to their local positions, then reuse existing gathering, weighting, normalization and accuracy. Evaluation keeps full logits; shared DPO/KTO/GKD preparation is unchanged.Support covers native tensor-selection text causal LMs and text-only Qwen3.5/3.6 MoE inputs, with Ulysses, per-device batch one and no packing/padding-free. Media inputs, other multimodal architectures, Ring, custom loss, label smoothing, Unsloth and Liger remain excluded. Accept the all-zero modality IDs emitted by the Qwen text collator while rejecting nonzero media IDs.
Related #9765. This implements the reporter's model-family text SP path, but does not establish that their full 35B/128K, eight-H20, ZeRO3 workload fits in memory.
Validation:
CPU integration uses torch 2.9.1 and transformers 5.12.1. No ZeRO/FSDP, full pretrained long-context, image/video or GPU performance claim; full upstream suite not run.