This repository contains a CoreML conversion of the DeepSeek-R1-Distill-Llama-8B model optimized for Apple Silicon devices. This conversion features stateful key-value caching for efficient text generation. DeepSeek-R1-Distill-Llama-8B is a distilled 8 billion parameter language model from the DeepSeek-AI team. The model is built on the Llama architecture and has been distilled to maintain performance while reducing the parameter count. This CoreML conversion provides: — Full compatibility with Apple Silicon devices (M1, M2, M3 series) — Stateful inference with KV-caching for efficient text generation — Optimized performance for on-device deployment — Base Model: deepseek-ai/DeepSeek-R1-Distill-Llama-8B — Parameters: 8 billion — Context Length: Configurable (default: 64, expandable based on memory constraints) — Quantization: FP16 — File Format: .mlpackage — Deployment Target: macOS 15+ — Architecture: State -…
Модальности:
Генерация текста
Области применения:
Диалог / чат Логика и рассуждение
Задача: Генерация текста
Автор: anthonymikinka
Теги: coreml, llama, conversational, endpoints_compatible
Лайков: 3 | Загрузок: 2
Описание основано на материалах HuggingFace. Перевод выполнен автоматически.