ggml-org/gemma-4-12B-GGUF
Any-to-Any • 12B • Updated • 1.32k • 2
That's almost exactly what I'm estimating based on Deepseek R1 pulling about 1 or 2 TPS. Offloading a few layers to some A6000s didnt really help much either.
I have enough DDR4 to put this on a 64 core threadripper, but I'm worried about how brutally slow it will be :)
btw we got it using 2 nvlinked RTX 6000 GPUs and its even better.