Chat & support: my new Discord server Want to contribute? TheBloke’s Patreon page This is GGML format quantised 4-bit, 5-bit and 8-bit GGML models of Medalpaca 13B. This repo is the result of quantising to 4-bit, 5-bit and 8-bit GGML for CPU (+CUDA) inference using llama.cpp. 4-bit GPTQ models for GPU inference. 4-bit, 5-bit 8-bit GGML models for llama.cpp CPU (+CUDA) inference. * medalpaca’s float32 HF format repo for GPU inference and further conversions. llama.cpp recently made another breaking change to its quantisation methods — https://github.com/ggerganov/llama.cpp/pull/1508 I have quantised the GGML files in this repo with the latest version. Therefore you will require llama.cpp compiled on May 19th or later (commit 2d5db48 or later) to use them. For files compatible with the previous version of llama.cpp, please see branch previousllamaggmlv2. medalpaca-13B.ggmlv3.q40.bin | q40 | 4bit | 8.14GB | 10.5GB | 4-bit. | medalpaca-13B.ggmlv3.q41.bin | q41 | 4bit | 8.14GB | 10.5GB | 4-bit. Higher accuracy than q40 but not as high as q50. However has quicker inference than q5 models. | medalpaca-13B.ggmlv3.q50.bin | q50 | 5bit | 8.95GB | 11.0GB | 5-bit. Higher accuracy, higher…
Модальности:
Генерация текста
Области применения:
Медицина
Задача: Генерация текста
Автор: jayantdocplix
Теги: llama, medical, en, text-generation-inference
Лайков: 3 | Загрузок: 9
Описание основано на материалах HuggingFace. Перевод выполнен автоматически.