Chat & support: TheBloke’s Discord server Want to contribute? TheBloke’s Patreon page TheBloke’s LLM work is generously supported by a grant from andreessen horowitz (a16z) These files are GPTQ model files for Jan Philipp Harries’ Vicuna 13B v1.3 German. Multiple GPTQ parameter permutations are provided; see Provided Files below for details of the options provided, their parameters, and the software used to create them. GPTQ models for GPU inference, with multiple quantisation parameter options. 2, 3, 4, 5, 6 and 8-bit GGML models for CPU+GPU inference * Original unquantised fp16 model in pytorch format, for GPU inference and for further conversions Multiple quantisation parameters are provided, to allow you to choose the best one for your hardware and requirements. Each separate quant is in a different branch. See below for instructions on fetching from different branches. — In text-generation-webui, you can add :branch to the end of the download name, eg TheBloke/Vicuna-13B-v1.3-German-GPTQ:gptq-4bit-32g-actorderTrue` — With Git, you can clone a branch with: — In Python Transformers code, the branch is the revision parameter; see below. Please make sure you’re using the…
Модальности:
Генерация текста
Задача: Генерация текста
Автор: TheBloke
Теги: llama, de, en, text-generation-inference, 4-bit, gptq
Лайков: 3 | Загрузок: 10
Описание основано на материалах HuggingFace. Перевод выполнен автоматически.