Heterogeneous Structured Pruning Under Resource Constraints for Hardware-Efficient Large Language Models
1 School of Informatics, University of Edinburgh
2 Leiden Institute of Advanced Computer Science, Leiden University
3 School of Engineering, University of Edinburgh
∗ Equal contribution
Pruning code: Coming soon · Sparse kernel code: Coming soon
Overview
Pruning a language model reduces its number of weights, but the resulting latency and memory use depend on how those weights are stored and processed on the target hardware.
Heterogeneous Structured Pruning (HSP) chooses a sparsity structure and density for each sublayer under a latency or memory budget. It combines dense, unstructured, 2:4, block-sparse, head-pruned, and channel-pruned options in a single training-free search, using importance scores to minimise proxy pruning loss.

Selected results


Funding acknowledgements
This work was supported by the APRIL AI Hub. Cameron Barker is supported by the UKRI EPSRC Centre for Doctoral Training in Machine Learning Systems (EP/Y03516X/1).
This publication is part of the project LESSEN with project number NWA.1389.20.183 of the research program NWA ORC 2020/21 which is (partly) financed by the Dutch Research Council (NWO).
Henry Gouk was supported by the Royal Academy of Engineering under the Research Fellowship programme.
Social posts
Links to posts and discussions will appear here.