Compute cost & hardware
Bind the declared workload to physical hardware evidence, measured for the run's own process, not the whole machine.
Compute-cost attestation
Declare a model card so K-Veritas knows the scale of the work:
print("KVERITAS_MODEL params=25600000 arch=resnet50 precision=fp16")
print("KVERITAS_WORKLOAD dataset_size=1281167 epochs=90 batch_size=256")At seal time a certificate checks the declared FLOPs against the hardware evidence with three physical bounds. Anyone can recompute them from the numbers on the report.
Declared FLOPs cannot exceed what the hardware could deliver: GPU peak times GPU-active seconds plus CPU peak times CPU core-seconds.
Declared FLOPs cannot exceed measured GPU joules divided by the minimum energy per FLOP.
Declared weights should fit the observed GPU memory.
A wall-clock check catches fast fakes: the declared FLOPs need a minimum number of seconds on any machine, so a script that declares a large model and exits instantly fails even with no telemetry. Verdicts: PASS, REVIEW, FABRICATION-IMPOSSIBLE, N/A. It proves the work ran at the declared scale, not that the result is correct.
Execution coherence (HMCA)
Metric-blind: HMCA never looks at the result. It asks whether the run's telemetry channels move as one process. A real computation drives every channel (CPU, memory, context switches, page faults, CPU frequency, I/O, and GPU utilization, memory, power, and temperature) from one activity, so they move together. A faked, replayed, or spliced trace does not.
HMCA scores that coupling from 0 to 1, with a verdict of coherent, weak, incoherent, or inconclusive. A light run is judged on the activity it has, not penalized for being light. The web verifier graphs the channels over time.
Per-process measurement
K-Veritas measures only the process tree that kveritas run launched: CPU and memory from the tree, GPU memory and utilization filtered to its PIDs, GPU power scaled by its share. Other jobs running at the same time do not inflate the numbers.
/proc and nvidia-smi). On macOS and Windows the sampler falls back to system-wide readings until per-process support arrives there.