Llama-3.1-Nemotron-70B-Instruct-HF-FP8-DYNAMIC
— Model Architecture: Llama-3.1-Nemotron — Input: Text — Output: Text — Model Optimizations: — Weight quantization: FP8 —...
— Model Architecture: Llama-3.1-Nemotron — Input: Text — Output: Text — Model Optimizations: — Weight quantization: FP8 —...
— Model Architecture: Meta-Llama-3.2 — Input: Text — Output: Text — Model Optimizations: — Weight quantization: FP8 —...
This model was obtained by quantizing the weights and activations of Bielik-11B-v.2.2-Instruct to FP8 data type, ready for...
— Model Architecture: Meta-Llama-3.1 — Input: Text — Output: Text — Model Optimizations: — Weight quantization: FP8 —...
Преобразованная и квантованная контрольная точка nvidia/Nemotron-4-340B-Instruct. В частности, он был получен из контрольной точки v1.0 .nemo на NGC....
FP8 квантованная контрольная точка nvidia/Minitron-8B-Base для использования с vLLM. Модальности:Генерация текста Задача: Генерация текста Автор: mgoin Теги: nemotron,...
FP8 квантованная контрольная точка nvidia/Minitron-4B-Base для использования с vLLM. Модальности:Генерация текста Задача: Генерация текста Автор: mgoin Теги: nemotron,...
— Model Architecture: Mistral-7B-Instruct-v0.3 — Input: Text — Output: Text — Model Optimizations: — Weight quantization: FP8 —...
Meta-Llama-3-70B-Instruct quantized to FP8 weights and activations using per-tensor quantization, ready for inference with vLLM >= 0.5.0. This...
— Model Architecture: Qwen2 — Input: Text — Output: Text — Model Optimizations: — Weight quantization: FP8 —...