50 points | by Maverick617 6 hours ago ago
6 comments
I didn't see any benchmarks against vllm, sglang, exllama, etc
Since this is basically a wrapper around libllama.so, I would assume that the performance is roughly the same as llama.cpp upstream.
From what I've seen, Vulkan adds a lot of overhead on Intel hardware.
> llama.cpp via Vulkan (AMD / Intel / NVIDIA) or CPU fallback
I got excited about someone paying attention to intel. Oh well.
What hardware do you have? I’ve been playing with a 258V and OpenVINO has come a longggggg way.
llama.cpp sycl and vllm xmx work is pretty incredible right now - you just gotta build it with some extra flags
I didn't see any benchmarks against vllm, sglang, exllama, etc
Since this is basically a wrapper around libllama.so, I would assume that the performance is roughly the same as llama.cpp upstream.
From what I've seen, Vulkan adds a lot of overhead on Intel hardware.
> llama.cpp via Vulkan (AMD / Intel / NVIDIA) or CPU fallback
I got excited about someone paying attention to intel. Oh well.
What hardware do you have? I’ve been playing with a 258V and OpenVINO has come a longggggg way.
llama.cpp sycl and vllm xmx work is pretty incredible right now - you just gotta build it with some extra flags