mysticbeing/Llama-3.1-Nemotron-70B-Instruct-HF-FP8-DYNAMIC - Каталог нейросетей
Генерация текста

mysticbeing/Llama-3.1-Nemotron-70B-Instruct-HF-FP8-DYNAMIC

Добавлено:
mysticbeing/Llama-3.1-Nemotron-70B-Instruct-HF-FP8-DYNAMIC

— Model Architecture: Llama-3.1-Nemotron — Input: Text — Output: Text — Model Optimizations: — Weight quantization: FP8 — Activation quantization: FP8 — Intended Use Cases: Intended for commercial and research use in multiple languages. Similarly to Llama-3.1-Nemotron-70B-Instruct-HF, this model is intended for chat between a user and AI assistant. — Out-of-scope: Use in any manner that violates applicable laws or regulations (including trade compliance laws). Use in languages other than English. — Release Date: 10/31/2024 — Version: 1.0 — License(s): llama3.1 — Model Developers: mysticbeing — Method used to quantize the weights (quantmethod) compressed-tensors — Weights format float-quantized — Architecture LlamaForCausalLM — Attention heads 64 — KV heads 8 — Hidden Activation** Sigmoid Linear Unit (SiLU) By accessing this model, you are agreeing to the LLama 3.1 terms and conditions of the license, acceptable use policy and Meta’s privacy policy Quantized version of It achieves an average score of [TBD] on the OpenLLM benchmark (version 1), whereas the unquantized model achieves 86.79. FP8 quantized models require…

Модальности:
Генерация текста

Области применения:
Диалог / чат Следование инструкциям


Задача: Генерация текста
Автор: mysticbeing
Теги: llama, fp8, vllm, conversational, en, de, fr, it
Лайков: 3  |  Загрузок: 17

Открыть на HuggingFace →

Описание основано на материалах HuggingFace. Перевод выполнен автоматически.