A scalable, agentic-first, and HuggingFace-native RL framework for research (9k lines).
-
Updated
Oct 4, 2026 - Python
A scalable, agentic-first, and HuggingFace-native RL framework for research (9k lines).
SC'25 UltraAttn: Efficiently Parallelizing Attention through Hierarchical Context-Tiling
Deterministic sparse-index localization for Apple silicon (MLX/Metal) and CUDA/Triton, with stable ordering, a Python oracle, and reproducible evidence.
Hardware barrier-free distributed computing plane (FNG V3) fusing Burgers' viscous dissipation and moment correction directly into XLA registers to pre-rectify tensor skewness & minimize latency spikes inside distributed LLM attention rails.
Distributed training framework for DeepSeek-V3 (Multi-Head Latent Attention, DeepSeekMoE, auxiliary-loss-free load balancing) with composable DP/FSDP/HSDP/TP/PP/CP/EP parallelism, torchtitan-inspired.
To associate your repository with the context-parallelism topic, visit your repo's landing page and select "manage topics."