— Original model: MiniCPM-1B-sft-bf16 — Model creator and fine-tuned by: ModelBest, OpenBMB, and THUNLP — Paper: link The utilization of activation sparsity, namely the existence of considerable weakly-contributed elements among activation outputs, is a promising method for inference acceleration of large language models (LLMs) (Liu et al., 2023; Song et al., 2023). Concretely, acceleration methods based on activation sparsity usually achieve higher inference speed by making wiser resource allocation and computation policies to avoid resource waste on these weakly-contributed parameters. Adopting ReLU as the activation function is a straightforward method to achieve activation sparsity. However, most recent mainstream LLMs adopt activation functions without intrinsic sparsity (e.g., GELU and Swish). Some efforts (Zhang et al., 2022; Mirzadeh et al., 2023; Zhang et al., 2024) introduce ReLU or its variants as the substitutive activation function to help non-ReLU LLMs achieve activation sparsity and inference acceleration, but few can concurrently obtain high sparsity and comparable task-specific performance. In this work, we introduce a simple and effective sparsification method…
Модальности:
Генерация текста
Задача: Генерация текста
Автор: SparseLLM
Теги: MiniCPM, ModelBest, THUNLP, custom_code, en, zh
Лайков: 3 | Загрузок: 599
Описание основано на материалах HuggingFace. Перевод выполнен автоматически.