Skip to main navigation Skip to search Skip to main content

Compressing Pre-trained Language Models by Matrix Decomposition

  • Intel
  • The Allen Institute for Artificial Intelligence

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

62 Scopus citations

Abstract

Large pre-trained language models reach state-of-the-art results on many different NLP tasks when fine-tuned individually; They also come with a significant memory and computational requirements, calling for methods to reduce model sizes (green AI). We propose a two-stage model-compression method to reduce a model's inference time cost. We first decompose the matrices in the model into smaller matrices and then perform feature distillation on the internal representation to recover from the decomposition. This approach has the benefit of reducing the number of parameters while preserving much of the information within the model. We experimented on BERT-base model with the GLUE benchmark dataset and show that we can reduce the number of parameters by a factor of 0.4x, and increase inference speed by a factor of 1.45x, while maintaining a minimal loss in metric performance.

Original languageEnglish
Title of host publicationProceedings of the 1st Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 10th International Joint Conference on Natural Language Processing, AACL-IJCNLP 2020
EditorsKam-Fai Wong, Kevin Knight, Hua Wu
PublisherAssociation for Computational Linguistics (ACL)
Pages884-889
Number of pages6
ISBN (Electronic)9781952148910
DOIs
StatePublished - 2020
Event1st Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 10th International Joint Conference on Natural Language Processing, AACL-IJCNLP 2020 - Virtual, Online, China
Duration: 4 Dec 20207 Dec 2020

Publication series

NameProceedings of the 1st Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 10th International Joint Conference on Natural Language Processing, AACL-IJCNLP 2020

Conference

Conference1st Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 10th International Joint Conference on Natural Language Processing, AACL-IJCNLP 2020
Country/TerritoryChina
CityVirtual, Online
Period4/12/207/12/20

Bibliographical note

Publisher Copyright:
© 2020 Association for Computational Linguistics.

Fingerprint

Dive into the research topics of 'Compressing Pre-trained Language Models by Matrix Decomposition'. Together they form a unique fingerprint.

Cite this