Records · kv:2610.00015v1

Reproduction of ggml-org/llama.cpp: Qwen2.5-7B Q4_K_M throughput and WikiText-2 perplexity

K-Veritas Team

Submitted by K-Veritas Team

Published 2026-10-07

Independent reproduction · Benchmark or leaderboard entry · Systems and performance

Reproduction of: original code

Record PDF (10 KB)

Abstract

Built llama.cpp at the pinned commit with CUDA for the RTX 5060 Ti (sm_120) and ran the repository's own tools on the official Qwen2.5-7B-Instruct Q4_K_M GGUF: llama-bench with its default tests (512-token prompt processing and 128-token generation, 5 repetitions, all layers on the GPU), then llama-perplexity on the WikiText-2 raw test set with default settings. One NVIDIA RTX 5060 Ti.

Sealed report

Qwen2.5-7B Q4_K_M bench and perplexity

Sealed 2026-10-07 · K-Veritas server

Data hash b984556d40a3984d6689ad4afa528098bf4626b887cada3a059b5bfa54876972

Cite this record

@misc{kveritas261000015v1,
  title={Reproduction of ggml-org/llama.cpp: Qwen2.5-7B Q4_K_M throughput and WikiText-2 perplexity},
  author={K-Veritas Team},
  year={2026},
  howpublished={K-Veritas Records},
  note={kv:2610.00015v1},
  url={https://kveritas.org/records/2610.00015v1},
}

References

  1. [1] Mamadou K. Keita and Christopher Homan. Computer Science Conferences Should Require Nonrepudiable Experimental Results. NeurIPS 2026 Position Paper Track. arXiv:2605.08586.

Sealed with K-Veritas [1]. Listed, not endorsed. A record shows these results came from the sealed code, unchanged. Whether the work is sound is for the reader to judge.