Records · kv:2610.00030v1

Reproduction of vllm-project/vllm: offline throughput of Qwen2.5-1.5B-Instruct

K-Veritas Team

Submitted by K-Veritas Team

Published 2026-10-08

Independent reproduction · Artifact evaluation · Systems and performance

PaperReproduction of: original code

Record PDF (9 KB)

Abstract

Ran vLLM's offline throughput benchmark with the released vLLM 0.31.0 package (PyTorch 2.13 for CUDA 13.0, FlashInfer kernels compiled for sm_120 at startup): Qwen2.5-1.5B-Instruct on 2,000 random prompts of 512 input and 256 output tokens, default engine settings. The bundle holds the v0.31.0 tag's benchmark code and project files; the rest of the source tree was set aside because the installed package, recorded in the signed environment, is what ran. One NVIDIA RTX 5060 Ti.

Sealed report

Qwen2.5-1.5B throughput, 2000 prompts

Sealed 2026-10-08 · K-Veritas server

Data hash bc31cd8c61d010a64596faef5b0e1343ad41bb2c43fb04c836c027879dd8b50f

Cite this record

@misc{kveritas261000030v1,
  title={Reproduction of vllm-project/vllm: offline throughput of Qwen2.5-1.5B-Instruct},
  author={K-Veritas Team},
  year={2026},
  howpublished={K-Veritas Records},
  note={kv:2610.00030v1},
  url={https://kveritas.org/records/2610.00030v1},
}

References

  1. [1] Mamadou K. Keita and Christopher Homan. Computer Science Conferences Should Require Nonrepudiable Experimental Results. NeurIPS 2026 Position Paper Track. arXiv:2605.08586.

Sealed with K-Veritas [1]. Listed, not endorsed. A record shows these results came from the sealed code, unchanged. Whether the work is sound is for the reader to judge.