Records · kv:2610.00025v1

Reproduction of huggingface/lighteval: SmolLM2-1.7B-Instruct on ARC-Challenge and HellaSwag

K-Veritas Team

Submitted by K-Veritas Team

Published 2026-10-08

Independent reproduction · Benchmark or leaderboard entry · Natural language processing

Reproduction of: original code

Record PDF (9 KB)

Abstract

Ran the README evaluation with lighteval at the pinned commit (accelerate backend): SmolLM2-1.7B-Instruct on the full ARC-Challenge test set with 25-shot prompts and the full HellaSwag validation set with 10-shot prompts, batch size 8. At this commit HellaSwag is scored generatively (exact match) rather than by log-likelihood accuracy. A fresh cache directory was used so every sample was computed in the sealed run. One NVIDIA RTX 5060 Ti.

Sealed report

SmolLM2-1.7B ARC-C and HellaSwag

Sealed 2026-10-08 · K-Veritas server

Data hash 89ce7eee4602d9d203d3f3986b0360b2646900cf6a05df02b9b18608839cc520

Cite this record

@misc{kveritas261000025v1,
  title={Reproduction of huggingface/lighteval: SmolLM2-1.7B-Instruct on ARC-Challenge and HellaSwag},
  author={K-Veritas Team},
  year={2026},
  howpublished={K-Veritas Records},
  note={kv:2610.00025v1},
  url={https://kveritas.org/records/2610.00025v1},
}

References

  1. [1] Mamadou K. Keita and Christopher Homan. Computer Science Conferences Should Require Nonrepudiable Experimental Results. NeurIPS 2026 Position Paper Track. arXiv:2605.08586.

Sealed with K-Veritas [1]. Listed, not endorsed. A record shows these results came from the sealed code, unchanged. Whether the work is sound is for the reader to judge.