High-performance RLHF/GRPO pipeline scaling Gemma 3 on GKE Ray Clusters (B200/H200) using NVIDIA NeMo-RL. Includes native FSDP checkpoint merging and zero-shot vLLM benchmarking.
-
Updated
Mar 27, 2026 - Shell
High-performance RLHF/GRPO pipeline scaling Gemma 3 on GKE Ray Clusters (B200/H200) using NVIDIA NeMo-RL. Includes native FSDP checkpoint merging and zero-shot vLLM benchmarking.
Deterministic mathematical and symbolic verification environments for Reinforcement Learning with Verifiable Rewards (RLVR).
A minimal reasoning pipeline for Qwen3-0.6B where one verifier plays three roles: evaluation, reward, validation.
To associate your repository with the math-500 topic, visit your repo's landing page and select "manage topics."