360Zhinao-7B-Chat-32K-Int4
360Zhinao (360智脑) 🤗 HuggingFace   |    🤖 ModelScope   |    💬 WeChat (微信)   Feel free to visit 360Zhinao’s...
360Zhinao (360智脑) 🤗 HuggingFace   |    🤖 ModelScope   |    💬 WeChat (微信)   Feel free to visit 360Zhinao’s...
Phi-Bode é um modelo de linguagem ajustado para o idioma português, desenvolvido a partir do modelo base Phi-2B...
Если вы используете text-generator-webui Select Transformers — Compute d-type: bfloat16 — Quantization Type : nf4 — Load in...
~~Mistral 7b v0.2 with attention_dropout=0.6, for training purposes~~ 1. Download original weights from https://models.mistralcdn.com/mistral-7b-v0-2/mistral-7B-v0.2.tar 2. Convert with https://github.com/huggingface/transformers/blob/main/src/transformers/models/mistral/convertmistralweightstohf.py...
OpenCSG stands for Converged resources, Software refinement, and Generative LM. The ‘C’ represents Converged resources, indicating the integration...
This is a DenseFormer implementation of Mistral-7B-v0.1. The details about DenseFormer are in the paper. You will need...
❤️ This repo contains the model LLaMA-8x265M-MoE(970M totally), which activates 2 out of 8 experts (332M parameters). This...
This model is created using MoE (Mixture of Experts) through mergekit based on Qwen/Qwen1.5-7B-Chat and abacusai/Liberated-Qwen1.5-7B without further...
GEMMA IS NOW FIXED WITHIN TRANSFORMERS — DISREGARD THIS REPO This model card corresponds to the 7B base...
> Note: If you wish to use GemMoE while it is in beta, you must add the flag...