phixtral-3x2_8
phixtral-3x2_8 is the first Mixure of Experts (MoE) made with two microsoft/phi-2 models, inspired by the mistralai/Mixtral-8x7B-v0.1 architecture....
phixtral-3x2_8 is the first Mixure of Experts (MoE) made with two microsoft/phi-2 models, inspired by the mistralai/Mixtral-8x7B-v0.1 architecture....
phixtral-2x2_8 is the first Mixure of Experts (MoE) made with two microsoft/phi-2 models, inspired by the mistralai/Mixtral-8x7B-v0.1 architecture....
— Model creator: Universal-NER — Original model: UniNER-7B-all This repo contains GGUF format model files for UniNER-7B-all. GGUF...
Note that phi-2 here is only used as an instruct model, instead of a chat model. Модальности:Генерация текста...
Training data contains 143,587 Japanese lyrics which are collected from uta-net by lyric_download Модальности:Генерация текста Задача: Генерация текста...
このモデルはrinna/japanese-gpt2-meduimを教師として蒸留したものです。 蒸留には、HuggigFace Transformersのコードをベースとし、りんなの訓練コードと組み合わせてデータ扱うよう改造したものを使っています。 学習に当たり、Google Startup Programにて提供されたクレジットを用いました。 a2-highgpu-4インスタンス(A100 x 4)を使って4か月程度、何度かのresumeを挟んで訓練させました。 Wikipediaをコーパスとし、perplexity 40 程度となります。 rinna/japanese-gpt2-meduim を直接使った場合、27 程度なので、そこまで及びません。 何度か複数のパラメータで訓練の再開を試みたものの、かえって損失が上昇してしまう状態となってしまったので、現状のものを公開しています。 This model...
GPT-Neo-vi-small is a transformer model designed using EleutherAI’s replication of the GPT-3 architecture. GPT-Neo-vi-smal was trained on the...
A fine-tuned version of numind/NuExtract-tiny-v1.5 (Qwen2.5-0.5B backbone) specialised for resume / CV structured extraction. Given raw resume text...
4-bit quantized GGUF version of distil-qwen3-4b-text2sql for efficient local inference. Only 2.5GB — runs on most laptops and...
This model is a fine-tuned version of Qwen/Qwen3-0.6B specifically designed for parsing cultural heritage person names into structured...