forked from NVIDIA/Megatron-LM
-
Notifications
You must be signed in to change notification settings - Fork 0
Pull requests: lmcafee-nvidia/Megatron-LM
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
Honor per-request logprob policy in async scheduling
#46
opened Aug 11, 2026 by
lmcafee-nvidia
Owner
•
Draft
Preserve async request metadata during compaction
#44
opened Aug 11, 2026 by
lmcafee-nvidia
Owner
•
Draft
Clear pending async logits when inference state resets
#43
opened Aug 11, 2026 by
lmcafee-nvidia
Owner
•
Draft
Invalidate prefix caches across generation changes
#42
opened Aug 10, 2026 by
lmcafee-nvidia
Owner
•
Draft
Fix CUDA graph lifecycle across inference suspend and resume
#40
opened Aug 10, 2026 by
lmcafee-nvidia
Owner
•
Draft
Fix checkpointed prefix-cache result reconstruction
#39
opened Aug 10, 2026 by
lmcafee-nvidia
Owner
•
Draft
Publish Mamba state at aligned chunk endpoints
#38
opened Aug 10, 2026 by
lmcafee-nvidia
Owner
Loading…
Add prefix cache stress coverage
Run functional tests
#35
opened Jul 30, 2026 by
lmcafee-nvidia
Owner
•
Draft
2 of 6 tasks
Support routing replay with async scheduling
#34
opened Jul 24, 2026 by
lmcafee-nvidia
Owner
•
Draft
Add cumulative async scheduling feature support
#33
opened Jul 24, 2026 by
lmcafee-nvidia
Owner
•
Draft
Support prefix caching and chunked prefill with async scheduling
#31
opened Jul 24, 2026 by
lmcafee-nvidia
Owner
•
Draft
Support log probabilities with async scheduling
#29
opened Jul 21, 2026 by
lmcafee-nvidia
Owner
•
Draft
Support sampling modes with async scheduling
#28
opened Jul 20, 2026 by
lmcafee-nvidia
Owner
•
Draft
ProTip!
Follow long discussions with comments:>50.