litert-community/Qwen2.5-Coder-3B-Instruct - Каталог нейросетей
Генерация текста

litert-community/Qwen2.5-Coder-3B-Instruct

Добавлено:
litert-community/Qwen2.5-Coder-3B-Instruct

This repository contains LiteRT-LM variant of Qwen/Qwen2.5-Coder-3B-Instruct optimized for on-device text generation. Ready to integrate this into your product? Get started in the LiteRT-LM documentation. Measured with the LiteRT-LM CLI: litert-lm benchmark -p 256 -d 256 —runs 3 —cache no (litert-lm 0.15.0) on an idle Apple M4 Max (macOS); 256 prefill / 256 decode tokens, 3 iterations averaged by the tool. A desktop reference point — phone-side figures vary by SoC and backend. Measured on a physical Samsung Galaxy S26 (SM-S942Q, Snapdragon 8 Elite Gen 5 / SM8850, Android 16) with litertlmadvancedmain from the litert-lm v0.16.0 release; the GPU backend is OpenCL (LITERTCL). One fixed 205-token prompt text (223 tokens under this tokenizer), —benchmark. Two runs per backend taken back-to-back — cells show the range. Peak RSS is the process VmHWM. Before quoting, the same file was run on each backend with a real prompt: both backends produced a correct text answer. The GPU takes the whole graph — decode 1603/1603 ops and prefill 1452/1452 on LITERTCL`. It wins prefill 2.2–2.7× and decode 1.2–1.4×, and peaks 4.6× lower (864 against 3934 MB) — for a 3.4 GB int8 bundle the RSS…

Модальности:
Генерация текста

Области применения:
Генерация кода Следование инструкциям


Задача: Генерация текста
Автор: litert-community
Теги: litert-lm, litertlm, qwen, Qwen2.5, en
Лайков: 4  |  Загрузок: 2,351

Открыть на HuggingFace →

Описание основано на материалах HuggingFace. Перевод выполнен автоматически.