Records · kv:2610.00035v1

Reproduction of linkedin/Liger-Kernel: fused linear cross entropy, RMSNorm and SwiGLU benchmarks

K-Veritas Team

Submitted by K-Veritas Team

Published 2026-10-11

Independent reproduction · Artifact evaluation · Systems and performance

PaperReproduction of: original code

Record PDF (13 KB)

Abstract

Ran three of the repository's kernel benchmarks at the pinned commit (fused linear cross entropy, RMSNorm, SwiGLU) on Triton 3.6 with PyTorch 2.11, comparing Liger's kernels with the baseline each script uses (Hugging Face modules, or plain PyTorch for fused linear cross entropy) in speed and peak memory, forward, backward and full passes. The report holds the geometric-mean ratio of baseline to Liger over each script's sweep; values above 1 favour Liger. On this GPU the fused linear cross entropy kernel uses less memory but is slower than the PyTorch baseline. One NVIDIA RTX 5060 Ti (Blackwell, sm_120).

Sealed report

FLCE, RMSNorm, SwiGLU vs baseline

Sealed 2026-10-10 · K-Veritas server

Data hash 8e2bf2961fc0da8e721b6f9c0eec4508e92617114d33f4893e7cd6d64efeeadc

Cite this record

@misc{kveritas261000035v1,
  title={Reproduction of linkedin/Liger-Kernel: fused linear cross entropy, RMSNorm and SwiGLU benchmarks},
  author={K-Veritas Team},
  year={2026},
  howpublished={K-Veritas Records},
  note={kv:2610.00035v1},
  url={https://kveritas.org/records/2610.00035v1},
}

References

  1. [1] Mamadou K. Keita and Christopher Homan. Computer Science Conferences Should Require Nonrepudiable Experimental Results. NeurIPS 2026 Position Paper Track. arXiv:2605.08586.

Sealed with K-Veritas [1]. Listed, not endorsed. A record shows these results came from the sealed code, unchanged. Whether the work is sound is for the reader to judge.