litert-community/Qwen3-4B-Thinking-2507 - Каталог нейросетей
Генерация текста

litert-community/Qwen3-4B-Thinking-2507

Добавлено:
litert-community/Qwen3-4B-Thinking-2507

LiteRT is Google’s on-device runtime, the new name for TensorFlow Lite (Android: com.google.ai.edge.litert:litert), and litert-torch, the renamed ai-edge-torch, is its PyTorch converter: a PyTorch model converted unmodified with literttorch.convert` matched the original to 4e-7 on a Galaxy S26 (measured, LiteRT 2.2.0, Android 16, 2026-09-05). Qwen/Qwen3-4B-Thinking-2507 converted to the LiteRT-LM (.litertlm) format for on-device inference with Google’s LiteRT-LM runtime (the engine behind the official litert-community/ models). Qwen3-4B-Thinking-2507 is a dense 4B reasoning model (Qwen3ForCausalLM, 36 layers) that operates exclusively in thinking mode — it emits a … chain before its answer — so it rides the existing Qwen3 converter and runtime directly. Qwen34bthinkingdynamicwi4b32afp32.litertlm is a dynamic INT4 variant (block-32 weights, FP32 activations). It was converted through the LiteRT Torch (litert-torch`) path and quantized with AI Edge Quantizer. This artifact incorporates LiteRT-LM GPU graph optimizations, including composite ops for RoPE, fused QKV, and fused Gate/Up projections, and is configured with static prefill memory allocation. > Update (2026-08-31) — thought…

Модальности:
Генерация текста

Области применения:
Логика и рассуждение


Задача: Генерация текста
Автор: litert-community
Теги: litert-lm, litert, litertlm, on-device, edge, qwen3, reasoning, thinking
Лайков: 4  |  Загрузок: 1,159

Открыть на HuggingFace →

Описание основано на материалах HuggingFace. Перевод выполнен автоматически.