Cycle-level performance and power of real Hopper LLM kernels with Accel-Sim 2.0 + AccelWattch, run on a Mac; traced Qwen2.5-0.5B layer on H100; split-K deadlock report (accel-sim#561)
-
Updated
Sep 19, 2026 - Shell
Cycle-level performance and power of real Hopper LLM kernels with Accel-Sim 2.0 + AccelWattch, run on a Mac; traced Qwen2.5-0.5B layer on H100; split-K deadlock report (accel-sim#561)
Fork of GPGPU-Sim/Accel-Sim adding a Distributed Shared Memory (thread-block cluster) model ported from ClusterSim, with a timing-validated SM-to-SM interconnect and a measured comparison against conventional shared memory.
To associate your repository with the accel-sim topic, visit your repo's landing page and select "manage topics."