Skip to content

Commit 58e5bf6

Browse files
authored
Update _index.md
1 parent 133bc3d commit 58e5bf6

1 file changed

Lines changed: 4 additions & 0 deletions

File tree

  • content/learning-paths/servers-and-cloud-computing/vllm-benchmark-quantisation

content/learning-paths/servers-and-cloud-computing/vllm-benchmark-quantisation/_index.md

Lines changed: 4 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -1,6 +1,10 @@
11
---
22
title: Run vLLM inference with quantised models and benchmark on Arm servers
33

4+
draft: true
5+
cascade:
6+
draft: true
7+
48
minutes_to_complete: 60
59

610
who_is_this_for: This is an introductory topic for developers interested in running inference on quantised models. This Learning Path shows you how to run inference on Llama 3.1-8B and Whisper, with and without quantisation, and benchmark Llama performance and accuracy with vLLM's bench CLI and the LM Evaluation Harness.

0 commit comments

Comments
 (0)