Records · kv:2610.00036v1
Reproduction of Lightning-AI/lightning-thunder: Hugging Face LLM generation, eager vs Thunder
K-Veritas Team
Submitted by K-Veritas Team
Published 2026-10-11
Independent reproduction · Benchmark or leaderboard entry · Programming languages and compilers
Reproduction of: original code
Record PDF (9 KB)Abstract
Ran the quickstart examples/quickstart/hf_llm.py at the pinned commit: greedy generation of 100 tokens with a static cache, timed with Thunder's benchmark_n over two runs, once in PyTorch eager mode and once compiled with Thunder's hf-transformers recipe. The script's default model, meta-llama/Llama-3.2-1B, is gated, so the first public model from its own list was used instead (deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B, bf16). PyTorch 2.8 with CUDA 12.8 to match the pinned nvFuser build. One NVIDIA RTX 5060 Ti.
Sealed report
Eager vs Thunder generation
Sealed 2026-10-10 · K-Veritas server
Data hash d5f0b916cca0952bbef89c86fb15264e33d54a4fbc9b0a4b3e48186de56ffeff
Cite this record
@misc{kveritas261000036v1,
title={Reproduction of Lightning-AI/lightning-thunder: Hugging Face LLM generation, eager vs Thunder},
author={K-Veritas Team},
year={2026},
howpublished={K-Veritas Records},
note={kv:2610.00036v1},
url={https://kveritas.org/records/2610.00036v1},
}References
- [1] Mamadou K. Keita and Christopher Homan. Computer Science Conferences Should Require Nonrepudiable Experimental Results. NeurIPS 2026 Position Paper Track. arXiv:2605.08586.
Sealed with K-Veritas [1]. Listed, not endorsed. A record shows these results came from the sealed code, unchanged. Whether the work is sound is for the reader to judge.