Tested hardware & providers
We ran our tests on the following hardware:- NVIDIA GeForce RTX 3060 (mobile)
- NVIDIA GeForce RTX 3070 (Scaleway GPU-3070-S)
- NVIDIA A10 (Lambda Cloud gpu_1x_a10)
- NVIDIA A10G (AWS g5.xlarge)
- NVIDIA L4 (Scaleway L4-1-24G)
Tested LLMs
The results are available for the following LLMs (cf. Ollama hub):- Deepseek Coder 6.7b - instruct (Ollama, HuggingFace)
- OpenCodeInterpreter 6.7b (Ollama, HuggingFace, paper)
- Dolphin Mistral 7b (Ollama, HuggingFace, paper)
- CodeQwen 1.5 7b (Ollama, HuggingFace, blog)
- LLaMA 3 7b (Ollama, HuggingFace, blog)
- Phi 3 3.8b (Ollama, HuggingFace, paper)
- Coming soon: StarChat v2 (HuggingFace, paper)
q3_K_M, q4_K_M, q5_K_M.
Throughput benchmark
NVIDIA GeForce RTX 3060 (mobile)
NVIDIA GeForce RTX 3070 (Scaleway GPU-3070-S)
NVIDIA A10 (Lambda Cloud gpu_1x_a10)
NVIDIA A10G (AWS g5.xlarge)
NVIDIA L4 (Scaleway L4-1-24G)
If you’re looking for the latest benchmark results, head over here

