Energy characterisation of LLM inference across heterogeneous CPU islands, measured with RAPL counters. Finding: the efficiency cores cost 2.3× the joules per token — the penalty is duration, not power.
benchmarking quantization rapl energy-efficiency heterogeneous-computing llm-inference llm-inference-efficiency llm-inference-optimization llm-inference-kv-cache
-
Updated
Aug 2, 2026 - Python