Currently, there is no proper batching infrastructure, so the data cannot really be streamed in a GPU when there is smaller VRAM. For example: in the below screenshot when running with OPR and mixed states on a 32GB GPU, the memory is insufficient very often. This needs to be taken care of to make the same code GPU agnostic.

Currently, there is no proper batching infrastructure, so the data cannot really be streamed in a GPU when there is smaller VRAM. For example: in the below screenshot when running with OPR and mixed states on a 32GB GPU, the memory is insufficient very often. This needs to be taken care of to make the same code GPU agnostic.