We have a Qwen 2.5 (all model sizes) free Google Colab Tesla T4 notebook-Alpaca.ipynb) All notebooks are beginner friendly! Add your dataset, click «Run All», and you’ll get a 2x faster finetuned model which can be exported to GGUF, vLLM or uploaded to Hugging Face. — This Llama 3.2 conversational notebook-Conversational.ipynb) is useful for ShareGPT ChatML / Vicuna templates. — This text completion notebook-TextCompletion.ipynb) is for raw text. This DPO notebook replicates Zephyr. — Kaggle has 2x T4s, but we use 1. Due to overhead, 1x T4 is 5x faster. Qwen2.5-1M is the long-context version of the Qwen2.5 series models, supporting a context length of up to 1M tokens. Compared to the Qwen2.5 128K version, Qwen2.5-1M demonstrates significantly improved performance in handling long-context tasks while maintaining its capability in short tasks. The model has the following features: — Type: Causal Language Models — Training Stage: Pretraining & Post-training — Architecture: transformers with RoPE, SwiGLU, RMSNorm, and Attention QKV bias — Number of Parameters: 7.61B — Number of Paramaters (Non-Embedding): 6.53B — Number of Layers: 28 — Number of Attention Heads (GQA): 28 for Q and 4…
Модальности:
Генерация текста
Области применения:
Диалог / чат Следование инструкциям
Задача: Генерация текста
Автор: unsloth
Теги: qwen2, unsloth, qwen, conversational, en, text-generation-inference, endpoints_compatible, 4-bit
Лайков: 3 | Загрузок: 95
Описание основано на материалах HuggingFace. Перевод выполнен автоматически.