Efficient-Large-Model/VILA1.5-40b-AWQ - Каталог нейросетей
Генерация текста

Efficient-Large-Model/VILA1.5-40b-AWQ

Добавлено:
Efficient-Large-Model/VILA1.5-40b-AWQ

Model type: VILA is a visual language model (VLM) pretrained with interleaved image-text data at scale, enabling multi-image VLM. VILA is deployable on the edge, including Jetson Orin and laptop by AWQ 4bit quantization through TinyChat framework. We find: (1) image-text pairs are not enough, interleaved image-text is essential; (2) unfreezing LLM during interleaved image-text pre-training enables in-context learning; (3)re-blending text-only instruction data is crucial to boost both VLM and text-only performance. VILA unveils appealing capabilities, including: multi-image reasoning, in-context learning, visual chain-of-thought, and better world knowledge. Paper or resources for more information: https://github.com/NVLabs/VILA — The code is released under the Apache 2.0 license as found in the LICENSE file. — The pretrained weights are released under the CC-BY-NC-SA-4.0 license. — The service is a research preview intended for non-commercial use only, and is subject to the following licenses and terms: — Model License of LLaMA — Terms of Use of the data generated by OpenAI — Dataset Licenses for each one used during training. Where to send questions or comments about the model:…

Модальности:
Генерация текста


Задача: Генерация текста
Автор: Efficient-Large-Model
Теги: llava_llama, VILA, VLM, endpoints_compatible
Лайков: 3  |  Загрузок: 23

Открыть на HuggingFace →

Описание основано на материалах HuggingFace. Перевод выполнен автоматически.