litert-community/Qwen3.5-2B - Каталог нейросетей
Генерация текста

litert-community/Qwen3.5-2B

Добавлено:
litert-community/Qwen3.5-2B

LiteRT is Google’s on-device runtime, the new name for TensorFlow Lite (Android: com.google.ai.edge.litert:litert), and litert-torch, the renamed ai-edge-torch, is its PyTorch converter: a PyTorch model converted unmodified with literttorch.convert` matched the original to 4e-7 on a Galaxy S26 (measured, LiteRT 2.2.0, Android 16, 2026-09-05). Measured on device (edge-compat): Galaxy S26 · LiteRT-LM 0.16.0 · GPU · decode 17.5 tok/s · prefill 450 tok/s · TTFT 520 ms · all 25703 ops delegated (2026-09-05); Galaxy S26 · LiteRT-LM 0.16.0 · CPU · decode 19.0 tok/s · prefill 181 tok/s · TTFT 1.20 s (2026-09-05); Raspberry Pi 5 · LiteRT-LM 0.16.1 · CPU, 4 threads · decode 4.2 tok/s · prefill 56 tok/s · TTFT 5.04 s (2026-09-01). Record: https://github.com/john-rocky/edge-compat/blob/main/cards/qwen35-2b-vl-int8/CARD.md Qwen/Qwen3.5-2B converted to the LiteRT-LM (.litertlm) format for on-device inference with Google’s LiteRT-LM runtime. Requires litert-lm ≥ 0.15 (both backends gated on 0.15.0 and 0.16.0). Same conversion rail as our Qwen3.5-0.8B and Qwen3.5-4B — the GPU-delegable rank-≤4 gated-delta kernel. Qwen3.5 is Alibaba’s hybrid architecture: GatedDeltaNet (gated delta rule linear…

Модальности:
Генерация текста


Задача: Генерация текста
Автор: litert-community
Теги: litert-lm, litert, litertlm, on-device, edge, hybrid, gated-deltanet, qwen3_5
Лайков: 4  |  Загрузок: 1,887

Открыть на HuggingFace →

Описание основано на материалах HuggingFace. Перевод выполнен автоматически.