Fett-uccine_Mini_3B
This is a merge of pre-trained language models created using mergekit. This model was merged using the task...
This is a merge of pre-trained language models created using mergekit. This model was merged using the task...
This model is a fine-tuned version of microsoft/phi-2 on the WhiteRabbit Cybersecurity dataset. It achieves the following results...
phixtral-3x2_8 is the first Mixure of Experts (MoE) made with two microsoft/phi-2 models, inspired by the mistralai/Mixtral-8x7B-v0.1 architecture....
— Просто и понятно: МИНИМАЛЬНЫЕ изменения кода и архитектуры. Введен только один слой проекции вверх-вниз, без необходимости кэширования...
— Просто и понятно: МИНИМАЛЬНЫЕ изменения кода и архитектуры. Введен только один слой проекции вверх-вниз, без необходимости кэширования...
Grafted WhiteRabbitNeo-13B-v1 and NexusRaven-V2-13B with mergekit. Use the WhiteRabbitNeo template for regular code, and the NR template for...
DeciLM-7B-instruct is a model for short-form instruction following. It is built by LoRA fine-tuning on the SlimOrca dataset....
DeciLM-7B-instruct is a model for short-form instruction following. It is built by LoRA fine-tuning on the SlimOrca dataset....
[Github] | [Colab Demo] | [Huggingface] | [Discord] | [Twitter] | [Blog] OpenMoE is a project aimed at...
Experimental model reccomended by this user: https://huggingface.co/VivyAI It turned out quite interesting, I will be iterating on this...